Skip to main content

Command Palette

Search for a command to run...

How I Route Every Claude Code Prompt to the Right Model

Updated
•14 min read•View as Markdown
S
Senior Software Engineer with 8+ years of experience designing scalable backend systems and high-traffic platforms. This blog documents my journey exploring system design, backend architecture, cloud infrastructure, and AI — sharing practical engineering lessons from real-world systems.

A small Python hook, three subagents and a CLAUDE.md file. Full code below, plus the parts that still don't work well.

TL;DR

I used to run one model for everything. That meant I either overpaid for simple lookups or got shallow answers on hard problems. Now a small Python hook scores each prompt and sends it to one of three Claude Code subagents: quick (Haiku), standard (Sonnet) or deep (Opus).

  • A UserPromptSubmit hook adds a one-line [task-router] hint to every prompt. Rules in CLAUDE.md tell the main session to follow it.

  • When the scores are close, it picks the stronger model. A stronger model rarely makes an answer worse. A weaker one often does.

  • The hint is a suggestion, not a hard rule, and the keyword matching gets some prompts wrong. I list those at the end.


The problem

Here's a normal day of my Claude Code prompts:

  • "Where is the login handler defined?"

  • "Run the tests."

  • "Add a /health endpoint and write a test for it."

  • "Our web socket workers hang intermittently under load. Find the root cause."

If I use Opus for all of these, the first two cost far more than they should. If I use Haiku, the last one gets a fast, confident and shallow answer, and I lose an hour going down the wrong path. Sonnet for everything is a reasonable middle, but I'm still overpaying on one end and underpowered on the other.

What I wanted was simple. Pick the cheapest model that still gives a good answer, and do it automatically. When it's not clear, go with the stronger one.


How it fits together

 You type a prompt
        |
        v
+------------------------------+
| UserPromptSubmit hook        |   route_task.py
| - keyword + length scoring   |   (runs locally, ~ms, no LLM call)
| - softmax -> percentages     |
| - <50% confidence: escalate  |
+------------------------------+
        |  additionalContext:
        |  "[task-router] Delegate this prompt to the `standard`
        |   subagent ... Scores: quick 22%, standard 61%, deep 17%"
        v
+------------------------------+
| Main session = dispatcher    |   follows routing rules in CLAUDE.md
| (cheap model, e.g. Sonnet)   |   may go one tier higher, never lower
+------------------------------+
        |  Agent tool
   +----+----------------+---------------------+
   v                     v                     v
+-----------+     +--------------+     +--------------+
| quick     |     | standard     |     | deep         |
| Haiku/low |     | Sonnet/medium|     | Opus/high    |
+-----------+     +--------------+     +--------------+
        |
        v
 Reply starts with "Model: Sonnet, medium effort (standard subagent)"

There are four pieces:

Piece What it does Where it lives
Subagents Set the model and effort level for a task ~/.claude/agents/*.md
Hook Suggests which subagent should take the prompt ~/.claude/hooks/route_task.py
CLAUDE.md Tells the main session to follow the suggestion ~/.claude/CLAUDE.md
settings.json Registers the hook and sets the main session model ~/.claude/settings.json

Note: This is user-level config, so it applies to every project. Each Claude Code profile has its own config directory, set with CLAUDE_CONFIG_DIR (for example CLAUDE_CONFIG_DIR=~/.claude-work claude). Wherever I write ~/.claude, use your profile's directory.


Now let's build it

Step 1: Directory layout

~/.claude/
├── CLAUDE.md              # routing rules for the main session
├── settings.json          # hook registration + main-session model
├── agents/
│   ├── quick.md           # Haiku, low effort
│   ├── standard.md        # Sonnet, medium effort
│   └── deep.md            # Opus, high effort
└── hooks/
    └── route_task.py      # the prompt scorer

bash

mkdir -p ~/.claude/agents ~/.claude/hooks

Step 2: Create the three subagents

Each subagent is a Markdown file with YAML frontmatter. The main session reads the description to decide whether a subagent fits the task, so be specific there.

~/.claude/agents/quick.md

---
name: quick
description: Fast, low-cost worker for simple tasks: file lookups, finding where something is defined, small single-file edits, renames, formatting, one-line answers, running a command and reporting the result. Use when the task needs little reasoning and has an obvious correct answer.
model: haiku
effort: low
---

You handle small, well-defined tasks quickly. Do exactly what was asked, nothing more.
Keep the change minimal and match the surrounding code style. Report what you did in a few lines.
If the task turns out to be harder than it looked (unclear root cause, several files, design choices), stop and say so instead of guessing.

~/.claude/agents/standard.md

---
name: standard
description: Balanced worker for everyday engineering: implementing a feature or endpoint, fixing a bug with a clear cause, writing tests, explaining code, multi-file edits that follow existing patterns. Use for most normal coding tasks.
model: sonnet
effort: medium
---

You implement everyday features and fixes in the current codebase.
Read the relevant code first and follow its existing patterns. Keep changes scoped to the task.
Verify your work (run it, or run the tests) when practical, and report what changed and how you checked it.
If the problem needs architectural decisions or the cause is unclear after investigation, say so.

~/.claude/agents/deep.md

---
name: deep
description: Most capable worker for hard problems: system or architecture design, debugging with an unclear root cause, concurrency or performance issues, security reviews, large refactors or migrations, RAG/LLM pipeline design, anything needing careful multi-step reasoning. Use when getting it wrong is costly.
model: opus
effort: high
---

You handle the hardest tasks: design, deep debugging, security and large refactors.
Investigate before concluding: read the relevant code, form hypotheses, and verify them.
Weigh the trade-offs explicitly and recommend one approach.
Report your findings, the decision and why, what you changed, and what you verified.

Step 3: Add the routing rules to CLAUDE.md

The hook only makes a suggestion. CLAUDE.md is what tells the main session to act as a dispatcher and actually follow it.

~/.claude/CLAUDE.md

# Task routing (model + effort)

The main session is a dispatcher. Do the work by delegating to the subagent most likely to give the best answer at the lowest cost:

| Subagent   | Model / effort        | Use for |
|------------|-----------------------|---------|
| `quick`    | Haiku, low effort     | Lookups, "where is X", small single-file edits, renames, running a command, one-line answers |
| `standard` | Sonnet, medium effort | Normal features, endpoints, bug fixes with a clear cause, tests, code explanations |
| `deep`     | Opus, high effort     | Architecture, code reviews, unclear or intermittent bugs, concurrency/performance, security, large refactors/migrations, RAG/LLM design |

Rules:
- A `[task-router]` line is attached to each prompt. Delegate to the tier it names. You may go one tier higher if your own read of the task says so, never lower.
- Delegate every task except trivial chit-chat and questions about this conversation, even when the codebase is small or you could do it faster yourself. Don't do the work inline.
- If you're unsure between two tiers, choose the higher one. Quality beats saving a few tokens.
- Split mixed tasks: e.g. use `quick` to locate the code, then `deep` to design the fix.
- If a subagent reports that the task was harder than expected, re-run it on the next tier up.
- Start every reply with one line naming the model that did the work, e.g. `Model: Opus, high effort (deep subagent)` or `Model: main session (<your model>), answered directly`.

Most of this is housekeeping. Three of the rules carry the weight:

  1. One tier higher, never lower. A regex can't tell what you mean, but the model can. So the model is allowed to overrule the hint, but only upward. It can't quietly downgrade a task.

  2. If unsure, pick the higher tier. The hook follows the same rule.

  3. The Model: line. This is how I check what really happened. Without it, I had no idea whether routing was working at all.

Step 4: Write the hook script

A UserPromptSubmit hook runs before Claude sees your prompt. It gets JSON on stdin with a prompt field. Whatever it writes to hookSpecificOutput.additionalContext goes into the model's context. The systemMessage field shows up for you as a banner.

I kept the scoring basic on purpose. It runs in milliseconds:

  1. Base scores. Every prompt starts at quick 0.8, standard 1.0, deep 0.0.

  2. Keyword weights. Each regex match adds points to its tier. Deep keywords (debug, race condition, security, refactor, RAG and so on) add 1.3 each. Review words ("review", "audit", "go through") add 2.5, because a review is deep work even when the request is short. Standard keywords add 0.8 and quick keywords add 1.0.

  3. Prompt length. Over 150 words adds 1.5 to deep. 41 to 150 words adds 0.8 to standard and 0.5 to deep. Under 8 words adds 0.8 to quick.

  4. Code. A fenced code block or a Traceback adds 1.0 to deep.

  5. Softmax. The scores become percentages that add up to 100%.

  6. Escalation. If the top tier is under 50%, the hint moves up one tier.

  7. Skips. Empty prompts, slash commands like /compact and @agent-... mentions get no hint. Questions about the conversation itself ("which model did you use?") stay with the main session.

~/.claude/hooks/route_task.py (the full script, standard library only):


#!/usr/bin/env python3
"""
UserPromptSubmit hook: scores each prompt and tells Claude which subagent
(quick / standard / deep) is most likely to give the best answer.

The scores become probabilities with a softmax. If the top tier is below
ESCALATE_BELOW, the hint moves up one tier: a stronger model rarely hurts the
answer, a weaker one can.
"""
import json
import math
import re
import sys

ESCALATE_BELOW = 0.5
TIERS = ("quick", "standard", "deep")
PROFILE = {"quick": "Haiku, low effort", "standard": "Sonnet, medium effort", "deep": "Opus, high effort"}

DEEP = [
    r"\b(architect\w*|design (a|the)|system design|trade-?offs?|scal\w+)\b",
    r"\b(debug\w*|root cause|race condition|deadlock|memory leak|flaky|intermittent\w*|hangs?)\b",
    r"\b(security|vulnerab\w+|auth\w*|injection|threat model)\b",
    r"\b(refactor\w*|migrat\w+|rewrite the|overhaul)\b",
    r"\b(performance|optimi[sz]\w+|latency|bottleneck|concurren\w+)\b",
    r"\b(rag|embedding\w*|vector (db|store|search)|retrieval|agent\w*)\b",
    r"\b(think hard|in depth|in-depth|thorough\w*|carefully|step by step)\b",
]
# Reviews are always deep work, even when the prompt is short.
REVIEW = [r"\b(review\w*|audit\w*|go through)\b"]
STANDARD = [
    r"\b(implement|add|build|create|write|update|change|fix|test\w*)\b",
    r"\b(endpoint|route|model|schema|migration|service|feature|function|class)\b",
    r"\b(explain|why|how (do|does|can|should))\b",
]
QUICK = [
    r"^\s*(hi|hello|thanks|thank you|ok|okay|yes|no)\b",
    r"\b(where is|which file|find (the|a) file|show me|what does|rename|typo|format)\b",
    r"\b(run|install|commit|status)\b",
]
# Questions about the conversation itself: answered by the main session, no delegation.
META = [
    r"\b(which|what) (model|tier|effort)\b",
    r"\b(did|do) you (use|delegate)\b",
    r"\byour (last|previous) (answer|reply|response)\b",
]


def hits(patterns, text):
    return sum(len(re.findall(p, text, re.IGNORECASE)) for p in patterns)


def probabilities(prompt):
    words = len(prompt.split())
    scores = {"quick": 0.8, "standard": 1.0, "deep": 0.0}
    scores["deep"] += 1.3 * hits(DEEP, prompt) + 2.5 * hits(REVIEW, prompt)
    scores["standard"] += 0.8 * hits(STANDARD, prompt)
    scores["quick"] += 1.0 * hits(QUICK, prompt)
    if words > 150:
        scores["deep"] += 1.5
    elif words > 40:
        scores["standard"] += 0.8
        scores["deep"] += 0.5
    elif words < 8:
        scores["quick"] += 0.8
    if "```" in prompt or "Traceback" in prompt:
        scores["deep"] += 1.0
    top = max(scores.values())
    exps = {t: math.exp(s - top) for t, s in scores.items()}
    total = sum(exps.values())
    return {t: exps[t] / total for t in TIERS}


def main():
    try:
        prompt = json.load(sys.stdin).get("prompt", "")
    except (json.JSONDecodeError, OSError):
        return
    # Slash commands and explicit agent mentions already say what to run.
    if not prompt.strip() or prompt.lstrip().startswith("/") or "@agent-" in prompt:
        return

    if hits(META, prompt):
        context = (
            "[task-router] Meta question about this conversation: answer it directly in the "
            "main session, don't delegate."
        )
        banner = "[task-router] main session (meta question, not delegated)"
    else:
        probs = probabilities(prompt)
        tier = max(probs, key=probs.get)
        confidence = probs[tier]
        note = ""
        if confidence < ESCALATE_BELOW and tier != "deep":
            tier = TIERS[TIERS.index(tier) + 1]
            note = f" (escalated: top score was only {confidence:.0%})"

        dist = ", ".join(f"{t} {p:.0%}" for t, p in probs.items())
        context = (
            f"[task-router] Delegate this prompt to the `{tier}` subagent ({PROFILE[tier]}) "
            f"via the Agent tool{note}. Scores: {dist}. Only answer directly if it is pure "
            f"chit-chat. You may pick a higher tier, never a lower one."
        )
        banner = f"[task-router] -> {tier} ({PROFILE[tier]}){note}"

    print(json.dumps({
        "systemMessage": banner,
        "hookSpecificOutput": {"hookEventName": "UserPromptSubmit", "additionalContext": context},
    }))


if __name__ == "__main__":
    main()
chmod +x ~/.claude/hooks/route_task.py

I made two choices on purpose here. First, if anything goes wrong, the hook does nothing. Bad JSON or an empty prompt means no output, and Claude Code carries on as normal. A routing hook should never block you from working. Second, I use softmax so the raw scores turn into percentages. That gives me a clean 50% cutoff and a hint I can actually read.

Step 5: Register the hook in settings.json

{
  "model": "sonnet",
  "hooks": {
    "UserPromptSubmit": [
      {
        "hooks": [
          {
            "type": "command",
            "command": "/usr/bin/python3 /Users/you/.claude/hooks/route_task.py"
          }
        ]
      }
    ]
  }
}

The "model" key sets the model for the main session, which is the dispatcher in this setup. More on why I picked Sonnet in Step 7.

Step 6: Test the hook with sample JSON

You don't need to start Claude Code to test the scoring. Just pipe JSON into the script:

echo '{"prompt":"fix the flaky test"}' | /usr/bin/python3 ~/.claude/hooks/route_task.py | python3 -m json.tool
{
    "systemMessage": "[task-router] -> standard (Sonnet, medium effort)",
    "hookSpecificOutput": {
        "hookEventName": "UserPromptSubmit",
        "additionalContext": "[task-router] Delegate this prompt to the `standard` subagent (Sonnet, medium effort) via the Agent tool. Scores: quick 22%, standard 61%, deep 17%. Only answer directly if it is pure chit-chat. You may pick a higher tier, never a lower one."
    }
}

For testing a batch of prompts, I use this loop:

for p in "where is the login handler defined?" \
         "add a /health endpoint to the FastAPI app and write a test for it" \
         "review the auth middleware for security issues" \
         "can you look at the payment retry logic" \
         "which model did you use for the last answer?" \
         "/compact"; do
  printf '%-70s ' "$p"
  python3 -c 'import json,sys; print(json.dumps({"prompt": sys.argv[1]}))' "$p" \
    | /usr/bin/python3 ~/.claude/hooks/route_task.py \
    | python3 -c 'import json,sys; d=sys.stdin.read(); print(json.loads(d)["systemMessage"] if d.strip() else "(skipped)")'
done
where is the login handler defined?                                    [task-router] -> quick (Haiku, low effort)
add a /health endpoint to the FastAPI app and write a test for it      [task-router] -> standard (Sonnet, medium effort)
review the auth middleware for security issues                         [task-router] -> deep (Opus, high effort)
can you look at the payment retry logic                                [task-router] -> deep (Opus, high effort) (escalated: top score was only 46%)
which model did you use for the last answer?                           [task-router] main session (meta question, not delegated)
/compact                                                               (skipped)

Once that looks right, start a new Claude Code session. Hooks load at startup. You should now see the [task-router] banner under every prompt and a Model: line at the top of each reply.


What it looks like in practice

These are real outputs from the script above:

Prompt Routed to Scores (q / s / d)
"where is the login handler defined?" quick 78% / 16% / 6%
"fix the typo in README" quick 66% / 29% / 5%
"rename this to async and run it" quick 91% / 7% / 2%
"add a /health endpoint to the FastAPI app and write a test for it" standard 3% / 95% / 1%
"explain this function" standard 26% / 69% / 5%
"find the file that handles retries and update it" standard (escalated from quick) 46% / 46% / 8%
"review the auth middleware for security issues" deep 3% / 2% / 96%
"our websocket workers hang intermittently under load, find the root cause" deep 4% / 5% / 91%
"which model did you use for the last answer?" main session n/a (meta)

And in an actual session:

> review the auth middleware for security issues
  [task-router] -> deep (Opus, high effort)

Model: Opus, high effort (deep subagent)
Found 3 issues in middleware/auth.py ...

The money saved on any single prompt is small. What I like is the default. Simple work goes to the cheap model without me thinking about it, and Opus is still one keyword or one escalation away.


Wrapping up

The whole thing is four small files and an afternoon of tweaking weights. It doesn't always pick the right model. But I went from "same model for everything" to "a sensible model per prompt, leaning stronger when unsure", and the Model: line tells me what actually ran.

If you try it, I'd like to know which prompt your router gets most wrong. Send me the prompt, the tier it picked and the tier you expected.

K

Biasing to the stronger model on close scores is the right asymmetry, since over-spending on Opus just costs money while under-routing a hard debug costs you a wrong answer you trust. The part I would instrument is exactly the misroutes you list at the end: keyword matching drifts, and without logging routing decisions against outcomes you cannot tell if the router is actually saving anything. Are you tracking how often the deep tier gets a prompt Haiku would have nailed?

S

Agreed on the asymmetry. That's why close calls go up a tier: overspending costs money, but a confidently wrong debug costs much more.

To answer your question: no, not yet. Right now I only know which tier a prompt was sent to, not whether a cheaper model would have handled it. So I can't yet say how much the router saves.

My plan is to log every routing decision (prompt hash, scores, tier, escalated or not) to a JSONL file. Then I'll take a sample of prompts that went to Opus, re-run them on Haiku and Sonnet, and compare the answers. That shows how often Opus was overkill, and turns "it probably saves money" into an actual number.

I

Routing by prompt score is smart, but the costly failures are misroutes in one direction: a hard refactor sent to the cheap tier comes back confidently shallow and costs more in rework than the big model would have. Do you log the cases where you escalated manually after a cheap answer, and feed those back into the scoring rules?

iin1005h23

S

Yes, that's the failure that matters most. A shallow answer you trust costs more in rework than Opus would have.

Today there are two guards. Claude can move a task up a tier, and Haiku and Sonnet are told to stop and hand off if the work turns out harder than it looked. But I don't log manual escalations yet, so nothing feeds back into the scoring.

That's the next step. Every escalation will be logged with the original prompt and both tiers: automatic hand-offs, Claude moving a task up, and me re-running a task on a stronger model. Each one becomes a labelled example ("this prompt needed deep"). They go into a regression test set, and I'll tune the rules against that set rather than changing weights on gut feel.