Building an AI-Native Development Workflow with Pi
Eleven extensions, six skills, eighteen agents and a token budget: how I rebuilt my whole working day inside a terminal-first coding agent, from morning briefing to verified pull request.
Most developers bolt an AI assistant onto their existing toolchain — a chat sidebar here, a copilot tab-completion there. I went the other way. Over the past few months I’ve rebuilt my entire development workflow inside pi.dev, a terminal-first coding agent, turning it into a unified system that handles everything from morning standup prep to evening journal reconciliation, with a full agentic SDLC pipeline in between.
This post walks through my setup: 11 custom extensions, 6 skills, 18+ specialised agents, 4 MCP servers, and a token economy that keeps the whole thing affordable. Everything lives in ~/.config/pi/agent/ — version-controlled, portable, and composable.
The Stack at a Glance
| Setting | Value |
|---|---|
| Provider | GitHub Copilot |
| Model (interactive) | Claude Opus 4.6 (high thinking) |
| Model (agents) | Configurable via subagent-config.json — 3 tiers |
| Model (none) | Native REST tools bypass the LLM entirely |
| Thinking level | High |
| Theme | Dracula (custom JSON) |
| Output compression | Ultra (token-pruner) |
| Editor | Vi-mode (custom modal editor) |
| MCP servers | GitHub, Filesystem, Atlassian (remote), Google Workspace (local) |
| Packages | pi-mcp-adapter |
The core philosophy: right-size every token. Five tiers:
- Opus 4.6 (interactive session) — where I’m reasoning through problems, making architectural decisions, and orchestrating the pipeline. High thinking budget. Worth every token.
- Reasoning tier (5 agents: spec, plan, adversarial test review, security review, hypothesis generation) — tasks that need deep analysis, multi-step reasoning, and creative problem-solving. Sonnet 4.6 with high thinking by default.
- Workhorse tier (12 agents: code, test, gather, explore, verify, code review, design review, etc.) — standard execution. Sonnet 4.6 with medium thinking.
- Lightweight tier (1 agent:
phase-gapsfor ticket creation) — mechanical, structured output. Haiku 4.5 with low thinking. ~25× cheaper than Opus. - No model at all — the
tools/extension makes native REST calls to Jira, Slack, Gmail, and Calendar directly.defineTool()wires API calls with auth gating and MCP fallback, returning structured data without spending a single LLM token. When the morning skill fetches 3 JQL queries, 3 Gmail searches, calendar events, and Slack history in parallel — that’s ~15 API calls that never touch a model.
Tiers are abstracted via subagent-config.json — a single config file that maps logical tier names to concrete models and thinking levels:
{
"tiers": {
"reasoning": { "model": "github-copilot/claude-sonnet-4.6", "thinking": "high" },
"workhorse": { "model": "github-copilot/claude-sonnet-4.6", "thinking": "medium" },
"lightweight": { "model": "github-copilot/claude-haiku-4.5", "thinking": "low" }
}
}
Agents declare a tier instead of a hardcoded model:
# Before (hardcoded)
model: github-copilot/claude-sonnet-4.6
# After (tier-based)
tier: reasoning
Change one line in subagent-config.json and every reasoning agent switches to a different model or thinking level. Individual agents can override the tier’s thinking level with an explicit thinking: field in their frontmatter — useful for agents that need more or less thinking than their tier default.
Then compress aggressively on both sides: input tokens truncated at source, output compressed to terse fragments, old turns stubbed to one-liners. Maximum intelligence where it counts, zero waste everywhere else.
Extensions — The Real Power
Pi auto-discovers extensions from ~/.config/pi/agent/extensions/. Each is a TypeScript file or directory with an index.ts barrel that receives the ExtensionAPI at startup. Here are the eleven I run daily.
🔴 Dangerous Command Gate
Every bash tool call passes through a risk classifier before execution. Commands are matched against pattern rules and sorted into two tiers:
- High risk (🔴):
rm -rf,sudo,mkfs, system control — always prompts, blocked entirely in non-interactive mode. - Medium risk (🟡):
chmod,git push --force,docker rm, package uninstalls — prompts interactively, allowed with warning otherwise.
When a dangerous command is caught, an interactive prompt lets me allow once, allow the whole category for the session, or block. The gate also injects system-prompt guidance nudging the LLM toward safer alternatives: trash over rm -rf, dry-run flags before destructive operations, narrow permission scopes over chmod -R 777.
Toggle with Ctrl+Alt+G or /gate on|off. It’s the seatbelt I never knew I needed — especially when the agent is autonomously running shell commands during a code phase.
🏗️ Dev Harness
This is the backbone. The dev-harness extension provides the runtime for a full agentic SDLC — from ticket to PR. It operates through transparent event hooks that fire at phase boundaries, so agents stay focused on their task while the extension handles validation, verification, cost tracking, and gating.
The design principle is shift left — catch every category of defect at the earliest possible stage, where it’s cheapest to fix. A missing acceptance criterion caught during the spec interview costs one follow-up question. The same gap caught during code review costs a respawn. Caught in production, it costs an incident.
The standard pipeline:
/dev PROJ-1234 add refresh tokens
→ phase-gather (context collection + enrichment from Confluence, Slack, Figma, etc.)
→ phase-spec (multi-turn interview → spec JSON → review-spec self-review)
→ phase-plan (multi-turn planning → plan JSON → review-plan self-review)
→ phase-test × N (per task: write tests [RED] in clean context → review-test [advisory pre-impl])
→ phase-code × N (per task: implement [GREEN] against tests it didn't write → extension type_check + test)
→ phase-verify-acceptance (against spec ACs, with extension-triggered preflight)
→ done (or phase-gaps → follow-up tickets)
Shift-Left: Defect Prevention by Phase
Each phase has built-in quality gates that prevent defects from propagating downstream:
| Phase | What’s caught | Mechanism |
|---|---|---|
| Spec | Missing ACs, untestable requirements, scope gaps, secret exposure | 16-topic interview sequence + review-spec self-review (structural consistency, AC→TC coverage, infra/security/DX completeness) |
| Plan | Infeasible decomposition, missing AC coverage, stub-covered MUSTs, security discipline violations | Finalisation checks + review-plan self-review |
| Test authoring | Confirmation bias (tests shaped to pass), trivial assertions, missing ACs | Separate phase-test agent with clean context (no impl knowledge) + adversarial review-test (devises wrong impls that pass) |
| Implementation | Type errors, test failures, regressions, misunderstood tests | Separate phase-code agent with clean context (reads tests from disk) + test_issue escalation for bad tests + extension type_check (hard gate) + test (advisory) |
| Verify | Spec drift, acceptance gaps, stale artifacts | Staleness check (spec/plan mtime vs verify mtime) + full verification_commands preflight |
The spec alone covers 16 mandatory topics — from problem statement through acceptance criteria, API contracts, constraints, risks, deployment strategy, rejected alternatives, infra/tooling, security topology, dev ergonomics, and test topology. Every MUST acceptance criterion must have a matching test case before the spec is finalised. The review-spec agent then validates structural consistency, and critical findings trigger a fix-and-re-review cycle (up to 2 rounds) before the spec ever reaches the planner.
TDD with Context-Isolated Test Authoring
The harness enforces strict TDD — but with a crucial twist: the agent that writes the tests is not the agent that writes the code. They run in separate subprocesses with clean context windows, eliminating confirmation bias.
Within a traditional single-agent TDD loop, the same LLM writes tests and then immediately implements against them. It “knows” how it intends to implement, so it writes tests shaped to pass — trivially-true assertions, missing edge cases, tests that exercise a code path without verifying the outcome. The harness fixes this by splitting the loop across three agents:
per task:
phase-test (clean context: spec ACs + plan guidance → write tests → confirm RED)
↓
review-test (advisory: pre-impl mode → audit test quality)
↓
phase-code (clean context: spec ACs + plan guidance → read tests from disk → implement → GREEN)
↓
extension (automatic: type_check hard gate + test advisory)
Here’s the detailed flow:
-
Write tests —
phase-test(RED) — Receives the spec’s acceptance criteria, test cases, plan guidance, and constraints. Does NOT receive any production code context or implementation hints. Writes tests purely from the spec’s observable behaviour contract — what the system should do, not how. Runs the test command and confirms tests fail (RED). A passing test before implementation is a deviation event: the behaviour already exists, or the test is trivially true. -
Review tests —
review-testadversarial analysis (advisory) — Spawned after everyphase-testcompletion. This agent doesn’t just audit tests — it actively tries to break them. For each acceptance criterion, it devises adversarial implementations: the simplest wrong code that would make every test pass while violating the spec. A hardcoded return value. A counter that resets on every request. A function that returns the right shape with wrong semantics. If an adversarial implementation exists, the tests are insufficient. Specific checks:- Adversarial implementation analysis — for each AC, construct a wrong implementation that passes all tests. If one exists → critical finding with the exploit and the missing assertion.
- AC negation — negate each criterion (“rate limiting is NOT enforced”) and walk through every test. If no test fails → the tests don’t actually depend on the correct behaviour.
- Side-effect coverage — for ACs implying side effects (DB writes, events, audit logs), check whether any test asserts the side effect. An adversarial impl can skip untested side effects entirely.
- Assertion specificity — every trivial assertion (
toBeDefined(),toBeTruthy()) is paired with the adversarial impl it permits and the specific assertion that would block it.
On critical findings, the orchestrator respawns
phase-testwith the adversarial analysis — up to 2 respawns. The test author now knows exactly how a wrong implementation could sneak through. -
Implement —
phase-code(GREEN) — Receives the spec ACs and plan guidance, but NOT the test source code. Reads test files from disk in its own context — approaching them fresh, as a contract to satisfy rather than code it authored. Implements production code to make the tests pass. If it finds a test that appears incorrect or untestable, it emits atest_issueevent back to the orchestrator (with a recommendation:fix_test,clarify_spec, oracceptable) and continues implementing. -
Extension verification — After
phase-codecompletes, the dev-harness extension automatically runstype_check(hard gate — failure respawns the agent) andtest(advisory — results injected into the orchestrator’s context).
The test_issue event is the escape valve. When phase-code encounters a test it believes is wrong, it doesn’t modify the test — it flags the issue and keeps working. The orchestrator sees the test_issue events after the code phase and can decide whether to respawn phase-test with the feedback, escalate to the engineer, or accept the test as-is.
This creates a four-layer verification stack:
| Layer | Agent | What it catches |
|---|---|---|
| 1. Test authoring | phase-test |
Spec → executable contract (clean context, no impl bias) |
| 2. Test review | review-test |
Adversarial analysis: devises wrong impls that pass tests, AC negation, side-effect gaps |
| 3. Implementation | phase-code |
Test issues surfaced by a fresh reader; impl bugs caught by RED→GREEN |
| 4. Extension | dev-harness hooks | Type errors (hard gate), test failures (advisory), schema violations |
Each layer catches a different category of defect, and the context isolation between layers 1 and 3 prevents the most insidious class of bug: tests that are green but don’t actually validate behaviour.
Extension-Driven Verification
The key design choice is that agents never decide when to verify — the extension’s event hooks handle that transparently:
tool_resulthook on subagent completion: If the agent requires verification (registered in the agent registry), and it didn’t returnstuck/blocked, the extension runstype_check+testand appends formatted results to the tool result. Type check failure injects a hard-gate message.tool_callhook on verify-phase dispatch: Beforephase-verify-*agents launch, the extension checks artifact staleness (spec/plan modified after the last verify?) and runs the fullverification_commandsarray as a preflight. All-fail blocks the launch; partial results are injected into the agent’s brief.tool_resulthook on file writes: Every write to.tasks/ordocs/is validated against JSON schemas. Invalid shapes are rejected before they hit disk.
Other Key Features
- Multi-pattern support:
implement/standard(full ceremony),implement/express(gather → code → verify),implement/orchestrated(parallel worktrees),analyse(investigation, no code),debug(hypothesise → test → verify). - Auto-config via
/dev init: Scans forgo.mod,Cargo.toml,package.json,pyproject.toml, etc., infers test runners, type checkers, linters from actual project config, fills gaps from built-in lang profiles, and writes a## Dev HarnessJSON block toAGENTS.md. Supports Go, Rust, TypeScript, Python, Ruby, Java, and C#/.NET out of the box. - Usage tracking: Every subagent result is costed and logged to JSONL. The orchestrator sees cumulative cost per work item.
- Config injection: Every subagent brief is automatically enriched with the project stack config (language, test command, type checker, linter) and relevant JSON schemas — agents never read config files themselves.
📋 Plan Mode
Sometimes I want the agent to analyse code and formulate a plan without touching anything. /plan (or Ctrl+Alt+P) flips into read-only mode: only read, bash (allowlisted commands like grep, git log, ls), and find are available. The agent produces a numbered plan under a Plan: header, then I choose whether to execute it. During execution, steps are tracked with [DONE:n] markers and a progress widget shows completion.
It’s the difference between “think first, then act” and “act and hope for the best.”
✂️ Token Pruner
The single most impactful extension for cost control. It attacks token usage from both sides:
Input pruning (60–80% savings):
- Strategy A — truncate at source: When a tool result exceeds configured byte limits (bash: 10KB, read: 20KB, grep/find/ls: 5KB), it’s truncated before entering the context window. A notice tells the LLM what was trimmed and how to re-read with
offset/limit. - Strategy B — stub old turns: Before each LLM call, tool results older than 3 turns are replaced with one-line stubs like
[bash output — 150 lines, pruned from older turn]. The stored session is unchanged — this runs on a deep copy.
Output pruning (~75% savings):
A system-prompt injection instructs the LLM to maximise information density. At ultra level: abbreviations, arrows for causality, fragments OK. "Inline obj prop → new ref → re-render. useMemo." Code, file paths, error messages, and config values are never compressed — only natural-language explanations.
Safety guardrails auto-expand to full verbosity for security warnings, irreversible-action confirmations, or when I appear confused.
🌳 Git Worktrees
Work on multiple things concurrently without branch-switching. /wt create auth-refactor creates a worktree at .worktrees/auth-refactor with branch wt/auth-refactor. The extension provides tools for status, merge, PR, and cleanup — and crucially, integrates with the subagent system for parallel code execution:
subagent({ tasks: [
{ agent: "phase-code", task: "Task A brief", cwd: ".worktrees/task-a" },
{ agent: "phase-code", task: "Task B brief", cwd: ".worktrees/task-b" },
]})
The implement/orchestrated pipeline uses this for monorepo-scale parallelism.
🔧 Tool Integrations
A declarative framework for native API tools. Each tool is ~25 lines via defineTool():
export default defineTool({
name: "jira-search",
params: { jql: "string", maxResults: "number" },
async execute(p) { /* native REST call */ },
mcp: { server: "atlassian", tool: "searchJiraIssuesUsingJql", ... },
format(result) { return { text: `Found ${result.length} issues` }; },
});
The framework generates TypeBox schemas from params, wires a provider chain (native REST first, MCP fallback), and calls format() once regardless of which provider succeeded. Currently covers:
| Service | Tools |
|---|---|
| Jira | search, get (native → Atlassian MCP fallback) |
| Gmail | search, get, thread (native OAuth → Google Workspace MCP fallback) |
| Calendar | list events (native OAuth → MCP fallback) |
| Slack | search, unread, DM history, channel history, user info, conversations (native REST, no MCP) |
Per-service setup via /jira-setup, /slack-setup, /google-setup.
📓 Journal
Domain-specific tools for my PARA-style work journal (work-journal.md). Three tools: journal-read-entry (structured parse of a date’s sections, frontmatter, guard check), journal-carry-over (daily or weekly open-item computation with dedup and recency badges), and journal-write-entry (create/update daily entries, merge sections, update frontmatter). These tools power the morning and evening skills described below.
🤖 Subagent
Spawn isolated pi subprocesses with separate context windows. Discovers agents from ~/.config/pi/agent/agents/ (user-level) and .pi/agents/ (project-level), with project agents overriding user agents by name. Each subagent runs in its own process with its own token budget — the orchestrator only sees a structured return packet.
⌨️ Vi Mode
A custom ModalEditor that intercepts keystrokes and implements vim-like modal editing directly in pi’s input area. Escape switches to NORMAL mode; motions (h j k l w b e 0 $ ^ gg G), operators (d c with w b $ 0 targets, dd cc), edits (x X D C S J u p), and mode switches (i a I A o O) all work. A mode indicator (NORMAL / INSERT + pending operator) renders in the bottom-right of the input area.
It’s not full vim — no count prefixes, no visual mode, no registers — but it covers 90% of my editing muscle memory. For the rest, Ctrl+G opens $EDITOR (neovim).
📊 Statusline
A p10k-lean-style three-line footer:
~/.config/pi/agent main
| claude-opus-4.6 [▓▓▓░░░░░░░] 72% 45.2k/200.0k tok 🌳 2 worktrees $0.3842 14:32:07
| ✂ 15.2KB truncated, 5 stubs 🔥 terse ULTRA
- Line 1: Directory (with
~shortening) + git branch + mode tag (plan mode, etc.) - Line 2: Model name + context bar (green ≥50%, amber <50%, red <20%) + token count + extension statuses + cumulative cost + clock
- Line 3: Token-pruner stats (truncation + stub counts, output compression level)
The context bar turns red and shows a 🔥 overage warning when tokens exceed the 200k standard window (where providers charge a 1.25× multiplier). It’s the financial dashboard I check reflexively.
🔔 Notify
Sends an OSC 9 terminal notification when the agent finishes and awaits input. Handles tmux passthrough via DCS escape + direct client tty write. Simple, essential — I often switch to another tmux pane while the agent works, and the bell brings me back when it’s ready.
Skills — Natural Language Commands
Skills are prompt-driven workflows that pi dispatches based on intent matching. Mine cover the full day.
/dev — Agentic Harness Orchestrator
The command-line entry point for all harness-driven work. Parses subcommands (help, init, tasks, review, resume, spec, plan, verify, gaps, deviations, spec-gaps, amend-spec) or classifies free-text intent into a pipeline pattern.
Key design choices:
- Minimal context: The orchestrator holds only the work item JSON (~1KB) + agent summaries. Source files, spec bodies, and test output are never loaded into the orchestrator’s context — always delegated to agents.
- Disk-persistent state: All state lives in
.tasks/(runtime) anddocs/{specs,plans,verify}/(committed). You can/clearbetween rounds and resume with/dev resume PROJ-1234. - Deterministic tools:
dev_bootstrap,dev_transition,dev_code_brief,dev_promote_events, etc. handle all file I/O and validation. The orchestrator never parses JSON manually.
morning — Daily Briefing
Runs on invocation, pulling from four sources in parallel:
- Jira: Three JQL queries — active issues, recently updated, backlog. Deduped, routed to Active Work or Backlog sections.
- Gmail + Google Chat: Unread email, unread chat, GDPR-specific threads. Thread-level analysis to surface only actionable items.
- Calendar: Today’s events with embedded document link extraction for pre-meeting reading.
- Slack: @mentions, DMs, monitored channel activity (7 channels). Top 5 actionable items per search.
Computes daily carry-over (open items from the previous weekday) and weekly carry-over (Mondays — all open items from the prior week with carried over badges). Writes a structured journal entry via journal-write-entry, commits and pushes. A guard check prevents duplicate runs — if the entry already exists with Active Work + Backlog populated, it only refreshes meetings.
evening — Journal Reconciliation
Cross-references today’s journal against what actually happened:
- Jira: Issues updated since last reconciliation → mark completed items
[x], add missing work - Gmail sent: Replies sent → mark “Needs response” items
[x] - Calendar: Attended events (accepted/tentative, time passed) →
[x]; declined →[-] - Slack: Messages I posted → evidence of action on items
- Google Docs: Documents I edited → mark “Docs to review” items
[x]
Appends concise notes to items with partial progress. Updates reconciled_at in frontmatter so the next evening run only queries the delta. Five checkbox states: [ ] open, [x] done, [~] partial, [-] skipped, [>] deferred.
commit — Context-Aware Commits
Calls git_commit_context to gather status, diff, log, branch, secrets scan, artifacts (spec/plan/verify files), and ticket ID in one shot. Drafts a structured commit message:
[PROJ-10107] Add anonymous order preview endpoint
Context:
Patients need pricing before creating an account.
Decisions:
- invoice_items over subscription_details (no active subscription yet)
- Phase mapped at service boundary to keep API contract clean
Test areas:
- Anonymous preview returns correct line items per phase
- Authenticated preview path unchanged (regression)
Always waits for confirmation before executing. Warns on detected secrets. Enriches from dev-harness artifacts when present.
pr — Push & PR
Calls gh_pr_context for existing PR, branch, commits, diffstat, spec, and verify report. If a PR exists, pushes (update flow). If not, drafts title + body with sections: Summary, Acceptance Criteria (from spec), Test Plan, Verification Evidence (from verify report), Risks/Notes. Never force-pushes. Reports push rejections cleanly.
review — Standalone Code Review
Spawns review-code and review-test agents in parallel against the current diff (staged → unstaged → branch diff fallback). review-test only runs if test files changed. Synthesises a single report with Critical / Warnings / Suggestions / Test Quality sections. Works without a spec or plan — useful for quick fixes and pre-commit sanity checks.
Agent Fleet
The agents/ directory contains 18+ markdown files, each defining a specialised agent with frontmatter (name, description, model, tools) and a system prompt body. They run as isolated subagent processes — the orchestrator never shares their context window.
Phase Agents (Implementation Pipeline)
| Agent | Role |
|---|---|
phase-gather |
Collect context from codebase, tickets, docs, Slack, Confluence |
phase-spec |
Multi-turn interview → docs/specs/<ID>-spec.json |
phase-plan |
Multi-turn planning → docs/plans/<ID>-plan.json with ordered tasks |
phase-code |
Implement production code against tests it didn’t write — clean context, reads tests from disk |
phase-test |
Write tests in a clean context from spec ACs — no impl knowledge |
phase-verify-acceptance |
Verify implementation against spec ACs. Receives preflight results from the extension |
phase-verify-infra |
Infrastructure-specific verification (Terraform plans, Helm diffs, etc.) |
phase-explore |
Open-ended codebase exploration for investigations |
phase-hypothesise |
Generate and rank hypotheses for debugging |
phase-gaps |
Identify gaps from verify reports → create follow-up tickets |
Review Agents (Quality Gates)
| Agent | Role |
|---|---|
review-code |
Code quality, simplification, drift from spec |
review-test |
Adversarial test review: devises wrong implementations that would pass the tests |
review-design |
Architecture, coupling, API design |
review-security |
Vulnerability scanning, auth patterns, data exposure |
review-spec |
Spec completeness, ambiguity, testability |
review-plan |
Plan feasibility, ordering, risk |
review-deviation |
Flag and resolve implementation deviations from plan |
review-investigation |
Evaluate investigation quality and completeness |
Tracker Providers & Enrichment Sources
The agents directory also includes pluggable tracker providers (jira.md, github.md, gitlab.md, plain-text.md) that teach phase-gather how to fetch ticket context from different issue trackers, and enrichment sources (confluence.md, figma.md, github-discussions.md, github-pr.md, google-docs.md, slack.md) that add domain-specific context during the gather phase.
MCP Servers
Four MCP servers provide external data access:
| Server | Transport | Purpose |
|---|---|---|
| GitHub | npx @modelcontextprotocol/server-github |
Repository operations, PR management, issues |
| Filesystem | npx @modelcontextprotocol/server-filesystem |
Broad filesystem access beyond the working directory |
| Atlassian | Remote MCP (https://mcp.atlassian.com/v1/mcp) |
Jira + Confluence (direct tools enabled) |
| Google Workspace | Local node server (workspace-server/dist/index.js) |
Gmail, Calendar, Drive |
The native tool integrations in extensions/tools/ call APIs directly when credentials are available, falling back to MCP servers when they’re not. This means the fast path (native REST) handles most requests, with MCP as a resilient fallback. The Atlassian MCP server is the only one with directTools: true, making its tools available to the LLM without the native wrapper.
Token Economy
The tiered model strategy handles the macro economics — reasoning tier for deep analysis, workhorse for standard execution, lightweight for mechanical tasks, native REST for data fetching. All configurable from a single subagent-config.json. The token-pruner extension handles the micro economics — squeezing waste out of every individual call.
The Tools Extension: Zero-Token API Calls
The biggest cost saving isn’t in the pruner — it’s in the tools/ extension. Every Jira search, Gmail query, Slack history fetch, and Calendar lookup is a native HTTP call that returns structured data directly to the agent’s context. No LLM interprets the API response; no MCP server translates it. The defineTool() framework wires auth gating, a provider chain (native first, MCP fallback), and a format() function that runs once regardless of which provider succeeded.
The morning skill alone makes ~15 parallel API calls across four services. At Opus token prices, having the LLM make those calls via function-calling would cost significantly more than having the tools extension handle them as plain HTTP requests. The LLM receives pre-formatted results and reasons over them — it never sees raw API payloads.
Input Budget
Per-tool byte limits control what enters the context:
| Tool | Limit | Direction | Rationale |
|---|---|---|---|
bash |
10 KB | Tail (keep end) | Exit codes and errors appear last |
read |
20 KB | Head (keep start) | Imports and declarations appear first |
grep |
5 KB | Head | First matches are usually most relevant |
find |
5 KB | Head | First results suffice for discovery |
ls |
5 KB | Head | Directory listing is front-loaded |
The 3-turn recency window means that by the time the agent is 4 interactions deep, early tool results are compressed to stubs. Combined with truncation at source, a typical session uses 60–80% fewer input tokens than it would without pruning.
Error results are never truncated — the LLM needs full diagnostics to self-correct.
Output Budget
At ultra compression, the system-prompt injection produces responses like:
Inline obj prop → new ref → re-render. useMemo.
Instead of:
“Sure! The issue here is that you’re creating a new object reference on every render by defining the prop inline. This causes the child component to re-render unnecessarily. You can fix this by wrapping the value in
useMemo.”
The savings are dramatic. An ultra-compressed explanation is typically 3-4× shorter than a normal one, with no information loss for an experienced developer.
The Statusline as Financial Dashboard
The context bar in the statusline provides real-time cost awareness:
- Green (≥50% remaining): Plenty of context budget. Work freely.
- Amber (<50% remaining): Getting expensive. Consider
/clearor starting a new session. - Red (<20% remaining): In the overage zone. Tokens past 200k incur a 1.25× multiplier. The
🔥indicator shows exactly how much overage.
Cumulative session cost is always visible (e.g. $0.3842). Between this and the pruner, I can run Opus 4.6 all day without cost anxiety.
The Dracula Theme
A full semantic theme covering every pi surface: syntax highlighting (pink keywords, green functions, yellow strings, purple numbers, cyan types), markdown rendering (orange headings, cyan links, green code), diff colors, tool result backgrounds (green tint for success, red tint for errors, purple tint for custom messages), and — my favourite detail — thinking-level gradients:
| Level | Color |
|---|---|
| Off | Dim grey (#545978) |
| Minimal | Comment blue (#6272a4) |
| Low | Cyan (#8be9fd) |
| Medium | Purple (#bd93f9) |
| High | Pink (#ff79c6) |
| X-High | Red (#ff5555) |
When the model is thinking hard, the indicator glows pink. When it’s barely thinking, it fades to grey. It’s a small thing, but it gives an intuitive sense of how much cognitive budget the model is spending on each response.
A Day in the Life
Here’s what a typical workday looks like with this setup.
08:30 — Morning Briefing
> morning
The skill fires four parallel API queries (Jira, Gmail, Calendar, Slack), computes carry-over from the previous weekday, writes a structured journal entry with sections for GDPR, Needs Response, Active Work, Meetings, Docs to Review, and Backlog. Commits and pushes. Total time: ~15 seconds. I scan the terminal output, mentally prioritise, and start.
09:00 — Feature Work
> /dev PROJ-1234 add refresh token rotation
The harness classifies this as implement/standard, creates a work item, and launches phase-gather to collect context from the codebase, Jira ticket, linked Confluence docs, and relevant Slack threads. Then phase-spec runs a multi-turn interview — asking about edge cases, failure modes, backwards compatibility. I answer 3-4 rounds of questions, the spec is finalised as JSON with acceptance criteria.
phase-plan breaks the spec into ordered tasks with file-level guidance. Then for each task, phase-test writes tests in a clean context — working purely from the spec’s acceptance criteria, with no knowledge of how the implementation will work. The adversarial review-test then tries to devise wrong implementations that would pass those tests — if it finds one, phase-test is respawned to strengthen the assertions. Then phase-code receives the spec and plan guidance (but not the test source code), reads the tests from disk in its own clean context, and implements production code to make them pass. The extension runs type_check (hard gate) and test (advisory) automatically after each code phase.
After all tasks complete, phase-verify-acceptance checks each AC against the implementation. Gaps become follow-up tickets via phase-gaps.
11:00 — Code Review
> review
Before committing a quick fix, I run the standalone review skill. It spawns review-code and review-test in parallel against my staged diff, then synthesises a single report with Critical/Warnings/Suggestions sections. Quick, focused, spec-free.
14:00 — Safe Exploration
/plan
> How is authentication handled in the order service?
Plan mode restricts tools to read-only. The agent explores the codebase, traces auth middleware, and produces a numbered plan for how I might refactor it. I review the plan, toggle /plan off, and decide whether to execute.
Meanwhile, the dangerous-command gate has already caught two rm -rf attempts from the agent during the morning’s code phase — both safely gated behind interactive prompts.
16:30 — Commit & PR
> commit
The commit skill gathers diff, status, branch, ticket ID, and any dev-harness artifacts. It drafts a structured message with Context (why), Decisions (trade-offs), and Test Areas (what could break). I confirm, it commits.
> pr
The PR skill pushes, detects no existing PR, and drafts one with Summary, Acceptance Criteria (copied from the spec), Test Plan, and Verification Evidence (extracted from the verify report). One command, full PR.
17:30 — Evening Reconciliation
> evening
The skill reads my journal entry, queries all five activity sources for what happened since the last reconciliation, and updates checkboxes: meetings attended → [x], emails replied to → [x], Jira issues completed → [x]. Adds items I worked on that weren’t in the morning briefing. Appends concise notes to items with partial progress. Updates reconciled_at so tomorrow’s run only queries the delta.
How It All Fits Together
┌─────────────────────────────────────────────────────────────────────┐
│ ~/.config/pi/agent/ │
│ │
│ settings.json ─── provider, model, thinking, theme, packages │
│ │
│ extensions/ (auto-loaded on startup) │
│ ├── dangerous-command-gate/ ── bash hook ── risk classify ── gate │
│ ├── dev-harness/ ── 6 event hooks ── validation, verification, │
│ │ usage tracking, config guard, brief injection │
│ ├── plan-mode/ ── tool restriction ── progress tracking │
│ ├── token-pruner/ ── tool_result hook ── context hook ── syspr. │
│ ├── worktree/ ── git worktree tools ── parallel execution │
│ ├── tools/ ── Jira, Slack, Gmail, Calendar ── defineTool() │
│ ├── journal/ ── read, write, carry-over ── PARA journal │
│ ├── subagent/ ── process spawning ── agent discovery │
│ ├── git-tools/ ── commit, diff, PR tools │
│ ├── vi-mode.ts ── modal editing (NORMAL/INSERT) │
│ ├── statusline.ts ── 3-line p10k footer │
│ └── notify.ts ── OSC 9 terminal bell │
│ │
│ skills/ (dispatched by intent) │
│ ├── dev/ ── agentic SDLC orchestrator │
│ ├── morning/ ── daily briefing → journal │
│ ├── evening/ ── journal reconciliation │
│ ├── commit/ ── structured commits │
│ ├── pr/ ── push + PR open/update │
│ └── review/ ── standalone diff review │
│ │
│ agents/ (spawned as isolated subprocesses) │
│ ├── phase-{gather,spec,plan,test,code,verify-*,explore,...}.md │
│ ├── review-{code,test,design,security,spec,plan,deviation,...}.md │
│ └── providers/ ── tracker + enrichment source agents │
│ │
│ mcp.json ─── GitHub, Filesystem, Atlassian, Google Workspace │
│ subagent-config.json ─── reasoning/workhorse/lightweight → models │
│ themes/dracula.json ─── full semantic Dracula theme │
│ token-pruner.json ─── per-tool limits, output level, stubs │
└─────────────────────────────────────────────────────────────────────┘
Data flow:
User → skill (intent match) → orchestrator (subcommand dispatch)
→ subagent (isolated process) → agent prompt + tools
→ native API tools (Jira, Slack, Gmail, Calendar)
→ MCP fallback if native unavailable
→ dev-harness hooks fire on every tool_result + tool_call
→ schema validation, verification gates, usage tracking
→ return packet → orchestrator promotes events, transitions phase
→ statusline updates (tokens, cost, context %, extension statuses)
Key Takeaways
Pi’s extension system is composable enough to build an entire workflow. What started as “add vi-mode” became a full SDLC pipeline, a daily productivity system, and a cost management layer — all from the same ExtensionAPI surface.
Right-size every token. Opus for orchestration, three configurable agent tiers (reasoning/workhorse/lightweight), native REST for data fetching — all tunable from a single subagent-config.json. Layer on input truncation, turn stubbing, and output compression, and the whole fleet runs for a few dollars a day.
Transparent hooks beat explicit orchestration. The dev-harness doesn’t ask the LLM to run tests — it runs them automatically at phase boundaries and injects results. The dangerous-command gate doesn’t ask the LLM to be careful — it intercepts dangerous commands before they execute. The best guardrails are the ones the agent doesn’t know about.
State on disk enables /clear without fear. Every piece of harness state lives in .tasks/ and docs/. I can /clear the context, close the terminal, come back tomorrow, and /dev resume PROJ-1234 picks up exactly where I left off. The orchestrator’s working context is just the work item JSON (~1KB) plus agent summaries.
The terminal is the IDE. With vi-mode, a p10k statusline, Dracula theme, native tool integrations, and tmux notifications — pi isn’t a chat assistant bolted onto my workflow. It is my workflow.
Everything described here is in ~/.config/pi/agent/, version-controlled as part of my dotfiles. The extensions are TypeScript, the agents are markdown, and the whole thing reloads with /reload.