Skip to the board
DIDRESET
UTC

Why does GPT-6 Astra drain my Codex 5-hour limit so quickly?

by vortwangUpdated

Because GPT-6 Astra costs the most per message of any model on the shared allowance. OpenAI's own table estimates 5–45 Astra local messages per five-hour period on Plus, against 10–100 for GPT-5.6 Sol and 250–2,000 for GPT-5.6 Luna; on the credit rate card Astra is 250 credits per million input tokens versus 100 for Sol. Higher reasoning effort, Fast mode (a 2.5x multiplier) and long multi-step tasks add to it. Switch with /model, check /status.

What the official docs say

How it is counted

The published estimates make the gap explicit. The help center's table of 'estimated local messages per five-hour period' gives GPT-6 Astra 5–45 on Plus, 25–225 on Pro 5x and 100–900 on Pro 20x, while GPT-5.6 Sol gets 10–100 / 50–500 / 200–2,000 and GPT-5.6 Luna 250–2,000 / 1,250–10,000 / 5,000–40,000 on the same plans. 'These are not fixed message limits. Actual usage varies by task, model and settings, and weekly limits may also apply.' On the token rate card, Astra is 250 credits per million input tokens, 25 cached and 1,250 output, against 100 / 10 / 500 for Sol and 5 / 0.5 / 30 for Luna.

Settings multiply the model cost. 'Different models can use different amounts of your allowance for the same task. Larger inputs and outputs, higher reasoning settings, Fast mode and tasks with multiple steps can also increase usage.' 'Higher effort can use more of your allowance and does not always produce a better result', and 'Fast mode applies a 2.5x multiplier to Astra's Standard rate.' The pricing page adds that 'larger projects, long-running tasks, or extended sessions that require the agent to hold more context will use significantly more per message.'

What to change, in the CLI: /model to 'Choose the active model (and reasoning effort, when available)'; /fast to toggle the Fast tier off; /compact to 'Summarize the visible chat to free tokens' or /new for a fresh chat; /status to read the '5h limit' bar before and after. OpenAI's advice on effort: 'Astra at Low effort can outperform Sol at High effort. If you've been getting good results with Sol at High, try Astra at Low or Medium as a starting point.' And a warning that switching alone is not a refill: 'Switching models does not restore allowance in a shared usage pool.' Issue #42987 (2026-09-05) is a user report of a Plus 5-hour allowance going 100% → 42% → 0% in two Astra Medium turns; treat it as one observation, not a rule.

How it is counted
ModelPlus: est. local messages / 5 hPro 5xPro 20xCredits per 1M tokens (input / cached / output)Source
GPT-6 Astra5–4525–225100–900250 / 25 / 1,250[1] [2]
GPT-5.6 Sol10–10050–500200–2,000100 / 10 / 500[1] [2]
GPT-5.6 Terra25–200125–1,000500–4,00050 / 5 / 300[1] [2]
GPT-5.6 Luna250–2,0001,250–10,0005,000–40,0005 / 0.5 / 30[1] [2]
Any model, Fast mode2.5x the Standard rate (Astra)2.5x2.5xUses more of your included allowance[1] [2]

FAQ

Does lowering reasoning effort on Astra save allowance?

Usually, but it is not a fixed amount: 'A reasoning level does not set a fixed amount of usage for a task or guarantee a better result.' OpenAI suggests starting Astra at Low or Medium, noting that 'Astra at Low effort can outperform Sol at High effort.' Set it with /model in the CLI.

If I switch from Astra to Luna mid-task, do I get my allowance back?

No. 'Switching models does not restore allowance in a shared usage pool.' The allowance you already spent is gone; a cheaper model only slows the rate at which the remainder is used. Check /status to see what is left in the 5h limit bar.

Is it normal for two Astra turns to use up a Plus 5-hour window?

OpenAI publishes a range, not a guarantee: 5–45 Astra local messages per five hours on Plus, and usage rises with context, effort, Fast mode and multi-step tasks. Issue #42987 reports exactly two Medium-effort turns emptying a Plus window; that is a user report. If your numbers look wrong, the help center says to contact Support with the model, reasoning level, Fast mode setting and a screenshot.

Does a long session cost more per message?

Yes. The pricing page: 'larger projects, long-running tasks, or extended sessions that require the agent to hold more context will use significantly more per message.' Use /compact to summarize the chat or /new to start fresh before a new task.

Sources

  1. OpenAI Help Center: Managing usage with GPT-6 Astra in Work and Codex (article 20001516, read via Wayback snapshot dated 2026-09-16)verified 2026-09-18
  2. Codex docs: Pricing (learn.chatgpt.com): token rate card, Fast mode multiplier, /statusverified 2026-09-18
  3. Codex CLI docs: Slash commands (learn.chatgpt.com): /model, /fast, /compact, /new, /status, /usageverified 2026-09-18
  4. openai/codex issue #42987 (2026-09-05, user report): GPT-6 Astra Medium depleted 100% of a Plus 5-hour quota in two short turnsverified 2026-09-18

Want the next Codex reset on your phone?

Sponsors