Notes
Working notes on agentic engineering, AI, and quantitative finance. Less polished than a paper, more honest than a tweet.
-
claude
Claude Agent SDK vs Claude Code CLI: Choosing Per Use Case
A deep-dive explainer on Claude Agent SDK vs Claude Code CLI: Choosing Per Use Case: methodology, historical context, worked examples with real numbers, and com
-
ai
Ox Alpha: Fact-Checking the Free Stealth Model Everyone Is Adopting
A mystery lab shipped Ox Alpha: free, 1M-token context, reasoning-first. The timeline, tokenizer theories, and the 80% DeepSWE rumor, fact-checked.
-
agents
Agent Memory That Beats Grep: How I Built Heimdall
Why agents redo work, how I built Heimdall — a verified, self-healing knowledge layer for AI coding agents — and how it compares to graphify and Graft.
-
claude
Claude Code Hook Cookbook: Patterns From Production
A deep-dive explainer on Claude Code Hook Cookbook: Patterns From Production: methodology, historical context, worked examples with real numbers, and common pit
-
token-economics
The AI Cost Reality for Indie Developers
DeepSeek's peak pricing, Codex and Claude usage caps, and the Muse Spark contributor tier — what the end of the subsidy era actually costs.
-
memory
Memory vs Context vs Plans: Three Persistence Layers Engineers Confuse
A deep-dive explainer on Memory vs Context vs Plans: Three Persistence Layers Engineers Confuse: methodology, historical context, worked examples with real numb
-
deepseek
DeepSeek V4 Pro 0813: Frontier Agentic Coding at Commodity Price
DeepSeek's GA release closes the agentic coding gap with Claude and GPT — Terminal-Bench 2.1 at 87.9, DeepSWE up 5x — for roughly 1/12th the price.
-
foundation-models
DeepSeek V4 Flash 0731: Loop Economics and the Price Cascade
V4 Flash 0731 sits one point behind GPT-5.6 Luna at a fraction of the cost. Loop math, the best harnesses, and the price cascade is a gift to indie developers.
-
deterministic
Deterministic Tool Rewriting: Why Hooks Beat Model Discipline
A deep-dive explainer on Deterministic Tool Rewriting: Why Hooks Beat Model Discipline: methodology, historical context, worked examples with real numbers, and
-
routing
Routing Through Free-Tier LLM Providers Without Hitting a Wall
A deep-dive explainer on Routing Through Free-Tier LLM Providers Without Hitting a Wall: methodology, historical context, worked examples with real numbers, and
-
foundation-models
DeepSeek V4 Flash 0731 and GPT-5.6 Luna: The Routing Table Just Moved
DeepSeek's official V4 Flash 0731 beta and OpenAI's 80% Luna price cut redraw the practical price-performance frontier for production agent fleets.
-
github-trending
Alibaba’s Open Code Review pairs deterministic checks with an LLM code
A factual overview of Alibaba’s open-source code-review project, which presents a hybrid deterministic-pipeline and LLM-agent approach to line-level review.
-
making
Making LLM Code Review Actually Catch Things
A deep-dive explainer on Making LLM Code Review Actually Catch Things: methodology, historical context, worked examples with real numbers, and common pitfalls w
-
developer-infrastructure
GPT-5.6 Luna Is the Value Tier. Terra Is Not Useless.
GPT-5.6 Luna is the rational default for bounded agent work, while Terra remains the escalation tier for ambiguity, mixed context, and harder coding loops.
-
open-source-ai
Open Source Frontier: Giant Models Meet Lower-Cost Coding Access
Kimi K3, Qwen3.8-Max-Preview, GLM-5.2, OpenCode Go, Qoder, and ClinePass show the frontier shifting from model announcements to usable agent economics.
-
model
Model Release Roundup: What Actually Changed
A deep-dive explainer on Model Release Roundup: What Actually Changed: methodology, historical context, worked examples with real numbers, and common pitfalls w
-
tool-use
Tool-Use Loops: The Failure Modes Nobody Documents
A deep-dive explainer on Tool-Use Loops: The Failure Modes Nobody Documents: methodology, historical context, worked examples with real numbers, and common pitf
-
prompt
Prompt Caching in Practice: The 5-Minute Cache and Workflow Design
A deep-dive explainer on Prompt Caching in Practice: The 5-Minute Cache and Workflow Design: methodology, historical context, worked examples with real numbers,
-
subagent
Subagent Orchestration: Fanout, Pipeline, and Where Both Break
A deep-dive explainer on Subagent Orchestration: Fanout, Pipeline, and Where Both Break: methodology, historical context, worked examples with real numbers, and
-
trending
Trending AI Repos Worth Cloning This Week
A deep-dive explainer on Trending AI Repos Worth Cloning This Week: methodology, historical context, worked examples with real numbers, and common pitfalls when
-
this
This Week in Claude Code: Features Worth Trying
A deep-dive explainer on This Week in Claude Code: Features Worth Trying: methodology, historical context, worked examples with real numbers, and common pitfall
-
agent
Agent Harness Comparison: Claude Code, Aider, Cursor Agent, Codex CLI
A deep-dive explainer on Agent Harness Comparison: Claude Code, Aider, Cursor Agent, Codex CLI: methodology, historical context, worked examples with real numbe
-
context
Context Window Management: Tactics That Survive Real Sessions
A deep-dive explainer on Context Window Management: Tactics That Survive Real Sessions: methodology, historical context, worked examples with real numbers, and
-
model
Model Context Protocol Server Design Patterns That Actually Hold Up
A deep-dive explainer on Model Context Protocol Server Design Patterns That Actually Hold Up: methodology, historical context, worked examples with real numbers
-
hooks
Hooks vs Skills vs Subagents: Picking the Right Claude Code Primitive
A deep-dive explainer on Hooks vs Skills vs Subagents: Picking the Right Claude Code Primitive: methodology, historical context, worked examples with real numbe
-
token
Token Economics of Long-Running Agent Loops
A deep-dive explainer on Token Economics of Long-Running Agent Loops: methodology, historical context, worked examples with real numbers, and common pitfalls wh
-
how
How Claude Code's Skills System Actually Works
A deep-dive explainer on How Claude Code's Skills System Actually Works: methodology, historical context, worked examples with real numbers, and common pitfalls
-
agentic-engineering
Agentic engineering patterns that survive contact with production
Field notes on the patterns that hold up when you put coding agents on real work. Context budgets, tool design, planner-executor splits, evaluation loops.
-
ai
Frontier AI in 2026, what actually changed and what did not
A working note on the shifts in the AI frontier through mid 2026. Long context, agentic capability, open weights, and the parts the headlines got wrong.
-
finance
LLM agents in quantitative finance, where they actually pay off
A working note on where coding agents earn their keep in equity research, forensic accounting, and DCF construction, and where they still do not.