1Cache the stable prefix
Put the system prompt, repo map and style rules first and reuse them. Cached input is 10× cheaper on Astra ($1 vs $10) and 40× cheaper on Fable.
2Stay under 272K per OpenAI call
Above that the whole request is billed at 2× input. Send the relevant files, not the whole repo.
3Measure cost per task
Per-token price misleads. Sonnet is half Opus’s rate but can use more tokens to finish. Log tokens per ticket and compare.
4Turn effort down for routine work
High effort adds latency and thinking tokens. Save max effort for genuinely hard problems.
5Ask for a plan first
On Astra and long agent runs, approve a short plan before it starts editing. A wrong direction at speed is the expensive failure.
6Check before you paste
Client code, secrets and personal data only go to approved enterprise or API accounts that exclude training. Ask if unsure.