SaratogaDeveloper guide · October 2026
For Saratoga developers · AI in the SDLC

Route each task to the right model.

One default model for everything is either overpaying or underpowered. Match the task to the model on reasoning depth, token cost and speed, and escalate only when a cheaper model falls short.

OpenAI models Anthropic models
01 · OpenAI

OpenAI models

From frontier horsepower down to cheap, narrow decisions.

OpenAI

GPT-6 Astra

Frontier intelligence and high technical complexity

Use when

Long-running, tireless work that needs extreme intelligence and predictable execution with custom skills.

Trade-off

Reserve it for jobs that truly need it. Ultra-fast mode burns plan allowance quickly, and without strong steering it makes expensive mistakes fast.

FrontierBurns allowanceNeeds steering
OpenAI

GPT-6.1 Sol

Heavy data synthesis and high-token workloads

Use when

Big projects with massive context: piles of data, code or documents, or computer use. Persistent without draining weekly limits.

Trade-off

Approaches Astra on some benchmarks at about one-fifth of the token price. Shade task complexity slightly down.

≈ ⅕ Astra's priceHuge context
OpenAI · Decisions API

Luna

Fast, cheap classification

Use when

Constrained "Type 1" decisions: tagging incoming email, sorting customer records, classifying changes by pricing or release date.

Trade-off

Built for narrow decision structures, not open-ended writing. That focus is what makes it so fast and cheap.

Very fastLow costNarrow only
02 · Anthropic

Anthropic models

Insight, writing quality and dependable structured execution.

Anthropic

Fable 5.1

Intuitive synthesis and pattern discovery

Use when

Deep qualitative analysis across messy, massive archives, like months of transcripts and long documents, where people can't see the forest for the trees.

Strength

Finds patterns you didn't know to look for, pulls new insight straight from sources and explains it in clear, well-written prose.

Pattern findingMessy archives
Anthropic

Opus 5.5

High-quality, natural writing

Use when

Producing a large volume of clean, high-quality written work, and you can give the model time to work.

Strength

Very responsive to steering away from stiff, mannered prose. Slower to generate than Astra.

Writing qualitySteerable toneNot the fastest
Anthropic

Sonnet 5.5

Efficient, structured execution

Use when

Well-defined, structured tasks with moderate to high complexity.

Strength

Priced the same as Sol. A dependable, highly competent workhorse that reliably gets through the work.

WorkhorsePriced like Sol
03 · Comparison

Side by side

Each model rated 1–5. Cost uses published API prices; speed uses measured throughput where it exists.

🧠 Brainpower: hardest work it handles well $ Cost: output price per 1M tokens 🏃 Speed: how quickly it responds and writes est no public figure yet
Model 🧠Brainpower $CostAPI price, in / out per 1M 🏃Speed Best at

Context size isn't a differentiator. All six models now take roughly 1 million tokens of input and return up to 128K. Choose on brainpower, cost and speed instead. On OpenAI models, a single request over 272K input tokens is billed at double the input rate.

04 · API specs

The numbers you'll budget with

Prices in US$ per 1M tokens. Context shown as max input / max output.

ModelInputCached inputOutputContextThroughputWatch out for
GPT-6 Astra
$10.00$1.00$50.001.05M / 128KNot publishedFast mode: 2× speed at 2× price. Ultrafast: $60 / $300.
GPT-6.1 Sol
$2.00$0.10$10.001.05M / 128KNot publishedAbout one-fifth of Astra per task on coding benchmarks.
Luna
$0.10$0.01$0.501.05M / 128KNot published100× cheaper than Astra. Built for narrow classification.
Fable 5.1
$10.00$0.25$50.001M / 128K≈ 67 tok/sThinking is always on. First token can take minutes at max effort.
Opus 5.5
$4.00See pricing page$20.001M / 128K≈ 106 tok/sHighest Artificial Analysis score (58). Anthropic’s recommended default.
Sonnet 5.5
$2.00See pricing page$10.001M / 128K≈ 116 tok/sUses ~60% more output tokens than Opus at max effort. Measure cost per task.

OpenAI long-context surcharge: if a single request sends more than 272K input tokens, the whole request is billed at 2× input and 1.5× output. Max input per OpenAI request is about 922K tokens even though the window is 1.05M.

05 · Routing guide

Who gets the job?

Typical developer tasks, with a prompt to copy and adapt.

Hard, cross-cutting bugs and architecture work that cheaper models have already failed on→Astra
Try askingAttached are the order-processing modules and logs from three production incidents. Find the race condition causing duplicate orders, explain the exact sequence that triggers it, and propose a fix with tests. Don't change public interfaces. Ask before reading files outside these modules.
Repo-wide sweeps: upgrades, deprecations, consistency checks, long agentic runs→Sol
Try askingFind every use of the deprecated auth library in this solution. List call sites by file and line, group them by usage pattern, and draft the replacement for each pattern. Run the test suite after each group and stop if anything fails.
CI and log triage, PR labelling, routing tickets at volume→Luna
Try askingClassify each failed CI test as flaky, environment, regression or unknown. Return JSON only: [{"test": "", "label": "", "reason": ""}] with a one-line reason.
Recovering intent in legacy systems from commits, tickets and old docs→Fable 5.1
Try askingRead the commit history, Jira tickets and wiki pages for the billing module since 2019. Explain why the rounding rules work the way they do, which decisions were deliberate, and where the code now contradicts the documented rules. Cite commits and tickets.
Design docs, ADRs, code review of risky changes, technical writing for clients→Opus 5.5
Try askingWrite an architecture decision record for moving our notification service from polling to an event queue. Use the ADR template attached, cover the options we considered and their trade-offs, and keep it short and plain.
Day-to-day feature work, unit tests, refactors with clear acceptance criteria→Sonnet 5.5
Try askingImplement the endpoint described in the ticket below. Follow the patterns in OrdersController, add unit tests for each listed edge case, and keep the change under 300 lines. List any assumptions at the end.
06 · Escalation

Start cheap, escalate on evidence

Move up a level only when the model below has clearly failed the task.

Luna triage→ Sonnet / Sol build→ Opus reason & review→ Astra / Fable hardest 5%
07 · Practice

Keep cost and risk down

1

Cache the stable prefix

Put the system prompt, repo map and style rules first and reuse them. Cached input is 10× cheaper on Astra ($1 vs $10) and 40× cheaper on Fable.

2

Stay under 272K per OpenAI call

Above that the whole request is billed at 2× input. Send the relevant files, not the whole repo.

3

Measure cost per task

Per-token price misleads. Sonnet is half Opus’s rate but can use more tokens to finish. Log tokens per ticket and compare.

4

Turn effort down for routine work

High effort adds latency and thinking tokens. Save max effort for genuinely hard problems.

5

Ask for a plan first

On Astra and long agent runs, approve a short plan before it starts editing. A wrong direction at speed is the expensive failure.

6

Check before you paste

Client code, secrets and personal data only go to approved enterprise or API accounts that exclude training. Ask if unsure.