Claude model tiers ਦੀ ਤੇਜ਼, ਅਮਲੀ ਗਾਈਡ. Routing blueprints, prompting patterns ਅਤੇ token-budget playbook ਵਾਲੇ ਪੂਰੇ deep-dive ਲਈ ਪੜ੍ਹੋ ਵੱਡਾ article.
claude-haiku-4-5 ਤੇਜ਼ ਅਤੇ ਸਸਤੇ intent ਕੰਮ ਲਈ, claude-sonnet-5 ਸਭ ਤੋਂ ਵਧੀਆ speed/quality balance ਲਈ, claude-opus-4-8 ਗੁੰਝਲਦਾਰ agentic coding ਅਤੇ enterprise ਕੰਮ ਲਈ, ਅਤੇ ਰਾਖਵਾਂ ਰੱਖੋ claude-fable-5 (ਅਤੇ gated claude-mythos-5) ਅਸਲ ਵਿੱਚ ਔਖੀਆਂ, long-horizon jobs ਲਈ ਜੋ ਘੰਟੇ ਜਾਂ ਦਿਨ ਲੈਂਦੀਆਂ ਹਨ.
ਇੱਕ ਨਜ਼ਰ ਵਿੱਚ model lineup
| ਮਾਡਲ | ਸਭ ਤੋਂ ਵਧੀਆ | Context window | Max output | Pricing (input / output) | Thinking behavior |
|---|---|---|---|---|---|
Claude Fable 5 (claude-fable-5) |
Long-running agents ਅਤੇ ਸਭ ਤੋਂ ਔਖੀਆਂ unresolved ਸਮੱਸਿਆਵਾਂ | 1M tokens | 128K tokens | $10 / $50 per MTok | Adaptive thinking always on |
Claude Opus 4.8 (claude-opus-4-8) |
ਗੁੰਝਲਦਾਰ agentic coding ਅਤੇ enterprise ਕੰਮ | 1M tokens | 128K tokens | $5 / $25 per MTok | Adaptive thinking always on |
Claude Sonnet 5 (claude-sonnet-5) |
ਸਭ ਤੋਂ ਵਧੀਆ speed/quality balance | 1M tokens | 128K tokens | $3 / $15 per MTok | Adaptive thinking always on |
Claude Haiku 4.5 (claude-haiku-4-5-20251001) |
ਸਸਤੇ, high-volume tasks ਲਈ ਸਭ ਤੋਂ ਤੇਜ਼ model | 200K tokens | 64K tokens | $1 / $5 per MTok | Extended thinking available (adaptive off) |
Claude Mythos 5 (claude-mythos-5) ਇੱਕ gated model ਹੈ ਜੋ Fable 5 ਵਰਗੀਆਂ specs ਅਤੇ pricing ਸਾਂਝੀਆਂ ਕਰਦਾ ਹੈ, ਪਰ Fable 5 ਵਾਲੇ safety classifiers ਸ਼ਾਮਲ ਨਹੀਂ ਕਰਦਾ. ਉਪਲਬਧਤਾ Project Glasswing ਰਾਹੀਂ ਮਨਜ਼ੂਰ partners ਤੱਕ ਸੀਮਿਤ ਹੈ.
ਸਹੀ model ਚੁਣੋ (ਤੇਜ਼ ਨਿਯਮ)
- Haiku 4.5: routing, classification, extraction, ਅਤੇ ਕੋਈ ਵੀ workflow ਜਿੱਥੇ ਤੇਜ਼ ਜਵਾਬ ਚਾਹੀਦੇ ਹੋਣ ਅਤੇ ਉੱਚ output-token budgets ਨਾ ਚੁੱਕੇ ਜਾ ਸਕਣ.
- Sonnet 5: ਰੋਜ਼ਾਨਾ coding ਅਤੇ document ਕੰਮ ਜਿੱਥੇ Opus/Fable ਕੀਮਤ ਤੋਂ ਬਿਨਾਂ ਮਜ਼ਬੂਤ ਗੁਣਵੱਤਾ ਚਾਹੀਦੀ ਹੋਵੇ.
- Opus 4.8: ਗੁੰਝਲਦਾਰ agentic coding, enterprise analysis, ਅਤੇ ਔਖੀ debugging ਜਿੱਥੇ ਲੰਬਾ, ਵਧੇਰੇ autonomous ਕੰਮ ਚਾਹੀਦਾ ਹੋਵੇ.
- Fable 5: ਉਹ ਔਖੀਆਂ, long-horizon ਸਮੱਸਿਆਵਾਂ ਜੋ ਘੰਟੇ ਜਾਂ ਦਿਨ ਲੈਂਦੀਆਂ ਹਨ, ਖਾਸ ਕਰਕੇ ਜਦੋਂ model ਨੂੰ ਟੀਚਾ ਰੱਖਣਾ, sub-tasks ਸੌਂਪਣਾ ਅਤੇ ਇਕਸਾਰ ਰਹਿਣਾ ਹੋਵੇ.
Effort ਅਤੇ thinking: ਲਾਗਤ ਉੱਤੇ ਅਸਰ
API ਉੱਤੇ, effort Fable 5 ਅਤੇ Mythos 5 ਉੱਤੇ intelligence, latency ਅਤੇ cost ਵਿਚਕਾਰ trade-off ਦਾ ਮੁੱਖ ਨਿਯੰਤਰਣ ਹੈ. Opus 4.8 ਅਤੇ Sonnet 5 ਲਈ ਤੁਸੀਂ effort ਸਪਸ਼ਟ ਰੂਪ ਵਿੱਚ ਵੀ ਸੈੱਟ ਕਰ ਸਕਦੇ ਹੋ ਜਦੋਂ ਵੱਖਰਾ cost/latency profile ਚਾਹੀਦਾ ਹੋਵੇ.
ਅਮਲੀ ਅਰਥ: ਹਰ request ਨੂੰ maximum effort ਉੱਤੇ ਡਿਫਾਲਟ ਕਰਨਾ ਅਕਸਰ output tokens ਉੱਤੇ ਬਜਟ ਖਰਚਣ ਦਾ ਸਭ ਤੋਂ ਤੇਜ਼ ਤਰੀਕਾ ਹੈ.
Token strategy ਜੋ ਸਾਰੇ tiers ਉੱਤੇ ਕੰਮ ਕਰੇ
Claude output tokens ਮਹਿੰਗਾ ਹਿੱਸਾ ਹਨ. ਇੱਕ ਸਧਾਰਨ ਨਿਯਮ ਹੈ: output length ਸੀਮਿਤ ਕਰੋ, TLDR-first ਲਈ prompt, ਅਤੇ ਵਰਤੋ prompt caching ਤਾਂ ਜੋ stable prefixes (system instructions ਅਤੇ tool schemas) ਦੁਬਾਰਾ ਵਰਤੇ ਜਾਣ.
- ਸਭ ਤੋਂ ਸਸਤਾ model ਵਰਤੋ ਜੋ ਪੂਰਾ ਕਰ ਸਕੇ. ਇੱਕ router (Haiku -> Sonnet/Opus -> Fable) ਅਕਸਰ ਗੁਣਵੱਤਾ ਉੱਚੀ ਰੱਖਦਿਆਂ ਲਾਗਤ ਨਾਟਕੀ ਤੌਰ ਤੇ ਘਟਾਉਂਦਾ ਹੈ.
- Output limits ਸੈੱਟ ਕਰੋ. Production ਵਿੱਚ ਹਮੇਸ਼ਾ max output / token cap ਸੈੱਟ ਕਰੋ. ਬੇਅੰਤ ਜਵਾਬ budgeting accident ਹਨ.
- Summaries ਰਣਨੀਤਕ ਤੌਰ ਤੇ ਵਰਤੋ. ਹਰ turn ਉੱਤੇ ਪੂਰੇ transcripts ਭੇਜਣ ਦੀ ਬਜਾਏ, ਇੱਕ compact “lessons” memory ਰੱਖੋ ਅਤੇ ਸਿਰਫ਼ ਉਸ ਨੂੰ ਅੱਪਡੇਟ ਕਰੋ.
ਇੱਕ ਸਧਾਰਨ router blueprint
- Intent ਅਤੇ ਅਨੁਮਾਨਿਤ complexity classify ਕਰੋ
claude-haiku-4-5. - “Normal hard” ਕੰਮ ਭੇਜੋ
claude-sonnet-5(ਜਾਂclaude-opus-4-8ਜਦੋਂ ਤੁਸੀਂ heavier agentic coding ਦੀ ਉਮੀਦ ਕਰੋ). - Escalate ਕਰੋ
claude-fable-5ਸਿਰਫ਼ ਜਦੋਂ task ਸਪਸ਼ਟ ਤੌਰ ਤੇ long-horizon ਹੋਵੇ ਜਾਂ ਹੇਠਲੇ tiers ਉੱਤੇ fail ਹੋ ਚੁੱਕਿਆ ਹੋਵੇ. - ਜੇ Fable 5 safety classifiers ਰਾਹੀਂ ਇਨਕਾਰ ਕਰੇ, Opus 4.8 ਉੱਤੇ ਵਾਪਸ ਜਾਓ (API configured ਹੋਣ ਤੇ fallbacks ਨਾਲ refiring ਸਹਾਰਦਾ ਹੈ).