Engineering methodology for deploying AI systems with operational discipline.
A system that ships predictably beats one that demos brilliantly.
Flagship: Animus — AI operating environment. Local-first, multi-agent, evidence-graded maturity. 17,000+ tests, six-layer architecture, autonomous improvement via Citizens (Architect, Conversation Designer, Knowledge Curator, Session Steward, Test Oracle).
Shipped tools:
- ai-spend —
htopfor AI spend. Cross-provider cost aggregation (Anthropic, OpenAI, OpenRouter).pip install ai-spend - mcp-manager — MCP server lifecycle manager across Claude Code, Cursor, Windsurf.
pip install arete-mcp - ai-skills — Production-ready skills for agent orchestration and Claude Code personas
Methodology: the-human-stack — Evidence-graded engineering reference (E0–E5), benchmark methodology, failure taxonomy
Infrastructure:
- arete-evals — LLM eval suites with bootstrap A/B comparison and rubric-based scoring
- agent-lint — Workflow YAML cost estimator + anti-pattern linter for agent orchestration
- context-hygiene — Heuristic context-window analyzer with position-decay scoring
How the pieces fit. Solid lines = active integration. Dashed lines = methodology guidance or historical extraction.
flowchart TB
subgraph Flagship["FLAGSHIP"]
Animus[Animus<br/>AI Operating Environment]
end
subgraph Ring1["CORE TOOLS — orbit Animus"]
direction LR
AISPEND[ai-spend<br/>Cost Observability]
MCPMGR[mcp-manager<br/>MCP Lifecycle]
EVALS[arete-evals<br/>Eval Pipeline]
LINT[agent-lint<br/>Workflow Governance]
HYGIENE[context-hygiene<br/>Context Quality]
SKILLS[ai-skills<br/>Agent Capabilities]
end
subgraph Ring2["METHODOLOGY — governs all"]
direction LR
THS[the-human-stack<br/>Evidence-graded reference]
NOTES[notes<br/>Cross-session memory]
CI[ci-templates<br/>Reusable CI]
end
subgraph Ring3["HISTORICAL — extracted into Animus"]
direction LR
MEMBOOT[memboot<br/>→ AST chunking]
ANCHOR[anchormd<br/>→ Pattern extraction]
end
AISPEND -->|"spend data"| Animus
MCPMGR -->|"server registry"| Animus
EVALS -->|"benchmark results"| Animus
LINT -->|"workflow checks"| Animus
Animus -->|"orchestrates"| SKILLS
THS -.->|"guides"| Animus
THS -.->|"guides"| AISPEND
THS -.->|"guides"| LINT
MEMBOOT -.->|"upstreamed"| Animus
ANCHOR -.->|"patterns archived"| NOTES
Most AI projects ship as demos and die in production. The failure modes are boringly consistent: costs spiral because no one tracks token spend, pipelines break halfway through and restart from scratch wasting compute, and multi-agent systems bottleneck on a single supervisor that burns tokens just watching. These aren't capability problems — they're operational problems.
The Human Stack is my response: a living engineering methodology that treats AI systems as operational systems. Every claim is evidence-graded (E0–E5). Every architectural choice is append-only, dated, and cross-referenced in an ADL. Every tool carries adversarial tests before it ships. The methodology is public in the-human-stack.
The Arete Stack is its implementation — the tools I actually build and maintain. Each repo maps to a layer of the methodology: skills for session discipline, cost observability for spend anomalies, governance linting for workflow anti-patterns, orchestration for autonomous improvement, and the methodology reference itself. This profile is arranged as an onion — start at the outer layer and peel inward.
Before software I spent 17 years in manufacturing and logistics (IBM, Toyota Production System). That discipline shows up in every codebase: standard work, visual management, error-proofing, continuous improvement.
Substack · LinkedIn · jamesyng79@gmail.com
Every post is grounded in a repo. Every repo is explained in a post.
- I Oversaw Scaling a Portland Ice Cream Line to 4,800 Pints Per Hour. AI Engineers Keep Making the Same Mistakes. — Manufacturing discipline applied to production AI: waste visibility, checkpoint/resume, and error-proofing before error-catching. (Methodology: the-human-stack · Implementation: Animus Session Controller)
- The Eval Is the Product — Why evaluation infrastructure matters more than model choice. (Tool: arete-evals · Integration: Animus Forge)
- Local-First AI Stack — Zero-cost inference with Ollama on consumer hardware. (Stack: RX 7900 XTX + 4-model tier · Budget tool: ai-spend)
- Decision Logging as Operational Memory — The ADL methodology: append-only, evidence-graded, revisit-conditioned. (Practice: notes/decisions)
These are the repos I maintain to production quality — evidence, tests, and documentation first.
| Project | Stars | Language | What It Is |
|---|---|---|---|
| animus | ⭐1 | Python | Personal AI operating environment — bitemporal memory, evidence-graded maturity, autonomous improvement |
| mcp-manager | ⭐1 | Python | MCP server lifecycle manager across Claude Code, Cursor, Windsurf — now with permission-prompt audit |
| ai-spend | ⭐0 | Python | Cross-provider AI cost aggregation with provider retry, config encryption, and health checks |
| agent-lint | ⭐2 | Python | Workflow YAML cost estimator + anti-pattern detection for agent orchestration |
| the-human-stack | ⭐0 | Markdown | Living engineering reference — evidence-graded chapters, benchmark methodology, failure taxonomy |
| ai-skills | ⭐3 | Python | Production-ready skills for Claude Code personas and orchestrated workflows |
| arete-evals | ⭐1 | Python | LLM eval suites and run records — eval engineering vs eval practice |
Hobby projects (EVE Online lore, personal context servers) are kept private. What you see here is what I ship.
The fastest way to understand the methodology is to use it. These repos are arranged as an onion — start at the outer layer and peel inward.
| Layer | Repo | What It Is | Entry Point |
|---|---|---|---|
| Skills | ai-skills | Production-ready skills for Claude Code, agent orchestration, and workflows | ./tools/install.sh --bundle arete-studio-ops |
| Eval | arete-evals | LLM eval suites and run records — eval engineering vs eval practice | pip install -e . |
| Cost | ai-spend | htop for AI spend — cross-provider cost aggregation |
pip install ai-spend |
| Governance | agent-lint | Workflow YAML cost estimator + anti-pattern linter | pip install agentlinter |
| Orchestration | animus | Personal AI operating environment with local-first control, evidence-graded maturity, session persistence, autonomous improvement | git clone + docker compose up |
| Methodology | the-human-stack | The engineering reference itself — evidence-graded chapters, benchmark methodology, failure taxonomy | Read VISION.md |
Animus is the primary reference implementation — a personal AI operating environment with local-first control and evidence-graded maturity tracking.
| Subsystem | What It Does | Tests |
|---|---|---|
| Kernel | Bitemporal memory core, checkpoint/resume, session persistence | 488 green |
| Head | Quality gates, fallback routing, natural-language interface | 96 green |
| Citizens | Architect (autonomous improvement proposals), Conversation Designer, Knowledge Curator | 71 green |
| Forge | Eval pipeline, benchmark execution, rubric-based scoring | Integrated |
| Session Controller | Token-budgeted session lifecycle, graceful wrap, auto-restart | 22 green |
Key design decisions (visible in the code):
- Evidence Framework — Six-stage maturity model (Concept → Self-improving) with Coverage KPI. Makes "documented but unverified" into "inspectable evidence."
- ADL-governed architecture — Every major decision is append-only, dated, and cross-referenced. No tribal knowledge.
- Quality gates before merge — Deterministic scoring (tool/completeness/structure) with adversarial test harness. Features don't ship without gate passage.
| Tool | What It Does | Install |
|---|---|---|
| ai-spend | htop for AI spend — cross-provider cost aggregation with OpenRouter, Anthropic, OpenAI |
pip install ai-spend |
| mcp-manager | MCP server manager across agentic IDEs (Claude Code, Cursor, Windsurf) | pip install arete-mcp |
| agent-lint | Workflow YAML cost estimator + anti-pattern linter | pip install agentlinter |
| context-hygiene | Heuristic context-window analyzer — position-decay scoring + regex patterns | pip install context-hygiene |
| arete-evals | LLM eval suites with benchmark scoring, A/B comparison with bootstrap CIs | pip install -e . |
| ci-templates | Reusable GitHub Actions workflows, lint configs, release automation | Copy + adapt |
| RedOPS | Red-team playbook for AI infrastructure — incident response, supply-chain audit | Read + run |
| overwatch | Multi-service health monitoring with structured logging and alerting | pip install -e . |
- Decision logs, not memory. Every architectural choice is an ADL entry — rationale, rejected options, tradeoffs, kill criteria.
- Tests are deliverables. Not an afterthought. Adversarial tests prevent regression of quality-gate contracts.
- Local-first where possible. RX 7900 XTX inference stack (Ollama, 4-model tiered) for zero-cost eval calibration.
- Ship, measure, error-proof, repeat. Kaizen learned from scaling an ice-cream line 6× — applied to software.
I develop primarily with Claude Code (Anthropic's CLI for Claude). Think of it as pair programming with a fast, literal pair — not autonomous AI.
What that means:
- Every commit is human-reviewed. The
claudecontributor in git history is the application, not a person. - Every test is human-verified. Coverage numbers reflect assertions I wrote and ran.
- Every architectural decision is logged in the ADL with rationale, rejected options, and tradeoffs.
- AI accelerates implementation; humans own validation, integration, and shipping.
Why this matters: AI-assisted development is becoming standard. What differentiates engineering discipline from cargo-culting is the verification layer: adversarial tests, evidence-graded maturity, and append-only decision logs. The code works because I test it. The methodology holds because I revisit it.
Direct entry points for "what does the code actually look like":
- Animus Kernel — bitemporal memory + adversarial tests — v2.3 scaffold: valid-time / transaction-time axes, quality-gate contracts enforced before any feature ships. Architect Citizen produces ImprovementProposals from codebase observation.
- ai-spend — OpenRouter provider —
GET /api/v1/generationspagination withnative_costfallback to token-based estimates. Pattern: provider registry ABC with side-effect registration. - Animus Session Controller — Token-budgeted session lifecycle (96% utilization trigger), graceful finalization with model-generated summary, checkpoint + auto-restart. 22 tests, 96 existing head core tests green.
- Animus P5 Discovery — MCP scanner, OpenAPI ingestion, annotated script discovery, 4-dimension schema validator, hash deduplication. 41 tests, 194 total P0–P5 tests green.
17 years in manufacturing and logistics operations (IBM, Toyota Production System). Scaled an ice-cream production line from 740 pints/day to 4,800/hour using Kaizen. Now apply the same discipline to AI infrastructure: standard work, visual management, error-proofing, continuous improvement.
See it through. Do it better. Leave something real.




