Back to blog

AI Coding / Agent Architecture

Claude Code vs. Codex: What Changed in My Agent Runtime

My Claude Code vs. Codex experience: agent-runtime refactoring, Astra and Sol, quota usage, current prices and what coding benchmarks actually show.

10 min readHeiner Giehl
Claude Code vs. Codex: What Changed in My Agent Runtime cover image

My agent runtime still wasn't behaving reliably. OpenAI Codex had produced another round of changes, the code was getting more complicated, and my usage allowance was disappearing. The architecture problem I had asked it to solve was still there.

After repeated attempts with GPT-6 Astra and Sol, I moved the work to Claude Code. Its first refactoring attempt gave me a runtime that worked noticeably better. The difference in speed and effort was large enough that I started planning more ambitious changes to the product.

This is my experience of Claude Code vs. Codex on a real agent-runtime refactor, with current pricing and selected coding benchmarks for context. I care most about how much time, usage and complexity it takes to reach a working result.

Product information checked on October 2, 2026. The project observations below are my own; these sessions were not a controlled test with identical starting conditions. Codex and Claude Code are coding tools, while Astra, Sol and Opus are models.

The project: building a reliable agent loop

I'm building a software product with its own chatbot runtime. Its agent loop receives user input, decides what to do next, calls tools where appropriate and processes their results before responding.

The difficult part is how those steps behave together. A user might give an incomplete instruction, change direction or write something that makes little sense. A tool might fail or return an unexpected result. The runtime needs to keep track of what has happened and respond sensibly.

I wanted help with that architecture. Adding individual functions wasn't enough: the system needed a clear way to manage the conversation, tool execution and run state. A coding agent had to understand why the current design was unreliable before proposing changes.

For the broader operational side of this work, I've written about running a Laravel AI support bot in production. This refactor concerned the foundation beneath those features.

GPT-6 Astra: high usage, more code, the same unresolved problem

I used Astra mostly with high reasoning effort, and sometimes on Medium. My expectation was that a capable model would be useful for a difficult architecture problem.

In these sessions, it repeatedly proposed more changes and added more code. The solution grew, but the underlying problem remained. Explanations could sound convincing even when the runtime still didn't behave reliably afterward.

At times, its reasoning appeared to rely on assumptions about the code or execution flow that I couldn't substantiate. I experienced that as hallucination, although I didn't measure a hallucination rate. Unsupported confidence made the work harder because I had to review both the proposed implementation and the assumptions behind it.

What disappointed me most was that Astra didn't identify and resolve the important weaknesses in my agent loop in those attempts. I ended up with more proposed code to understand and a problem that still needed solving.

The expensive part was a sequence of changes that consumed my allowance without producing a clean solution.

During intensive work, I felt I could burn through much of the usable capacity on a $100 or $200 plan in less than a day. That's an observation about the subscription allowance, not a claim that I incurred a $100 or $200 API bill that day.

The workaround: GitHub research, review and implementation slices

I changed the workflow when direct implementation became too costly and frustrating. I connected ChatGPT to my remote GitHub repository and asked for a detailed research and review report.

I wanted it to investigate how projects such as Codex and OpenCode approached similar problems: handling user input, invoking tools and maintaining sensible behavior through difficult conversation flows. The aim was to find architecture ideas that could be useful for my runtime.

I then handed the report to Codex and asked Astra, on High, to review it against the code. Sometimes that produced improvements. I wanted this additional check because I wasn't certain how completely the analysis through the GitHub integration had captured the project state.

Next, I asked for an implementation plan split into reasonably sized slices. Each slice needed a coherent purpose and should be possible to implement in a separate thread when that made sense. I was trying to avoid one conversation accumulating all the research, failed attempts and subsequent changes.

I still think that approach can help. A useful slice should have a clear scope and an outcome you can verify. A new thread also needs the relevant decisions and current code state; starting with an empty conversation doesn't automatically make the task cheaper.

Sol cost less, but implementation was slow in my sessions

I usually asked Sol to implement the plan: GPT-6 Sol initially, then GPT-6.1 Sol. Using Astra for analysis and review, with Sol handling implementation, seemed like a reasonable division of work.

I do give Sol credit for its lower cost. The implementation work in my sessions was very slow, though, and further problems appeared along the way. Research, review and a sliced plan still didn't get me to the runtime I wanted.

A well-organized workflow has limits. Working through a plan step by step doesn't help enough if the plan or its implementation fails to address the cause of the problem.

I haven't run a fresh, comparable speed test to establish whether GPT-6.1 Sol has improved since those sessions. My doubts about its suitability for this architecture task come from this project; they aren't a verdict on every task the model can perform.

Moving to Claude Code changed the result

I eventually switched to Claude Code and started with the Pro plan at roughly $20 per month. I was surprised by how quickly the work progressed and how much I could accomplish with that allowance.

The weekly limit wasn't my main problem at first. The five-hour session window interrupted concentrated work on larger changes, so I upgraded to the $100 plan.

For my workflow, the capacity then felt generous both within a session and across the week. More importantly, Claude Code refactored and rebuilt parts of my chatbot runtime so that its first attempt worked better than the results of my previous Astra and Sol attempts.

That didn't establish that every possible failure case had been eliminated. It did give me an immediately better foundation to work with, which was what I needed.

I observed roughly five to six percent usage in the quota display for that refactor. This is an approximate observation, not a token invoice. I haven't assigned it retrospectively to a weekly or session budget, and a percentage point in Claude's display isn't directly comparable to one in Codex's.

The practical difference was still substantial. With fewer interruptions and a better result, I started taking on larger refactorings in the product. The amount of review and rework a tool leaves behind affects what I'm willing to attempt next.

Current pricing: subscriptions, API tokens and Codex credits

A monthly subscription price, the included usage allowance and the API token rate are separate things. Keeping them separate helps explain why a model can be economical for a project without having the lowest token price.

Subscriptions and session limits

Claude Pro currently costs $20 with monthly billing. Max starts at $100, with options offering five or twenty times Pro's per-session usage. Max still has a five-hour session window and a weekly limit. Displayed prices exclude applicable taxes. Sources: Claude pricing and Max usage limits.

OpenAI currently lists Plus at $20 and Pro plans at $100, $200 or $500 per month. Its current documentation says Pro plans have no five-hour limit; other usage limits can apply. Model choice, context, reasoning and tool use affect consumption. Source: Codex and ChatGPT Work pricing.

Standard API pricing

These are USD list prices per million tokens, excluding tool charges, cache writes and special processing modes. The OpenAI figures use the short-context tier, up to 272,000 input tokens.

Standard API list prices, checked October 2, 2026
ModelInputCached inputOutput
GPT-6 Astra$10.00$1.00$50.00
GPT-6.1 Sol$2.00$0.10$10.00
Claude Opus 5.5$4.00$0.20$20.00

Sources: OpenAI API pricing and Claude Opus 5.5 model and pricing information. Opus 5.5 is an Anthropic model; OpenAI's GPT-5.5 is a different model. I include Opus here as a current pricing reference, separately from my session observations.

Astra costs five times as much as Sol at these input and output rates. Opus 5.5 sits between them and costs more per uncached token than Sol. My positive cost experience with Claude Code therefore can't be explained simply by the lowest token price: reaching a useful result sooner mattered.

Codex's Standard credit rates also show the gap. Astra is listed at 250 credits per million input tokens and 1,250 per million output tokens; GPT-6.1 Sol at 50 and 250. Those rates don't imply an identical consumption multiplier for the included subscription allowance. Source: Codex token credit rates.

What coding benchmarks do — and don't — establish

Claude Code was more effective on my project. That experience alone doesn't establish which model is the most intelligent overall.

For additional context, Anthropic publishes the following results in its Opus 5.5 announcement. This is a manufacturer-compiled comparison.

Selected results from Anthropic's Opus 5.5 announcement
BenchmarkOpus 5.5GPT-6 Astra
Terminal-Bench 4.066.4%57.9%
FrontierCode v1.1 (Main)54.4%53.3%
AutomationBench40.0%41.4%

Opus leads on these two coding benchmarks; Astra leads on AutomationBench. Settings aren't identical throughout: Terminal-Bench lists Opus at xhigh and Astra at high, while other results use different settings. The table contains no current GPT-6.1 Sol column, so it cannot rank all the models in my workflow. Source and methodology: Anthropic's Opus 5.5 announcement.

Anthropic also reports over 30% faster output generation than Opus 5 and roughly 40% lower costs on typical workloads in its own tests. Those comparisons concern its predecessor; they aren't speed or cost improvements I measured against Astra or Sol.

OpenAI positions GPT-6.1 Sol as delivering near-Astra performance at a lower cost for complex work. My disappointing runtime refactor doesn't rule out its usefulness elsewhere. It does give me a reason to test that claim against my own tasks.

How I'll judge the next refactor

I want to see evidence that a coding agent understands the cause of a failure earlier in the process. A detailed plan or a long list of changed files isn't sufficient evidence.

For an agent runtime, I'd assess an implementation plan against these questions:

  • Is the failure concrete? A problematic user input or tool response should be describable as a reproducible sequence.
  • Are responsibilities clear? It should be understandable which part manages run state, executes tools and delivers responses.
  • Are difficult paths checked? These include new input during an active step, tool failures and repeated actions.
  • Is each change reviewable? A slice should solve a recognizable part of the problem and have an outcome that can be verified.
  • Does the design become easier to reason about? Additional special cases need a specific justification.

I'd also make future comparisons more controlled: the same starting commit, the same failure case, explicit acceptance criteria and recorded time and usage. That would help separate the model's contribution from the coding tool and the knowledge already gained. Claude Code inherited a different situation from my earliest Astra attempts.

My conclusion: measure the cost of a working change

Astra disappointed me on this project. I invested substantial time and allowance without getting the agent-loop architecture resolved cleanly. Sol was cheaper, but the sessions I ran were still slow and left too many implementation problems.

Claude Code exceeded my expectations. It produced a noticeably better runtime quickly, and the $100 plan gave me enough room to keep working. That combination encouraged me to tackle larger refactorings.

For this project, switching was the right decision. The question I now ask is straightforward: How much time, usage and added complexity does it take before the change actually works?

You can read more about the product behind this work on my AI Agent Workflow Builder page. Projects like this are where I decide whether a coding agent earns its price.

Browse more guides

Continue through the blog index or explore the product pages.

Browse blog