Cursor vs Claude Code vs Codex (2026): Pick by Work Shape

Cursor vs Claude Code vs Codex in October 2026: features, pricing, value for money and a simple rule for choosing by the shape of your work.

· By Guilherme Salgueiro

Put the Cursor, Claude Code and Codex websites side by side this autumn and you will get the uneasy feeling of marking three students' homework after someone left the answer sheet on the desk. All three run agents in the cloud while your laptop sleeps, all three split big jobs across swarms of subagents, all three ship a command-line tool, and all three park a review bot on your pull requests. On 18 September 2026 the last difference on paper vanished, when Claude Code version 2.1.277 started reading AGENTS.md, the instruction file the other two already read. You can spend a whole weekend on comparison tables and finish exactly where you started, only more tired, because the tables measure the one thing that changes every month.

Tools will come and tools will go

In May 2025 Rick Rubin, the producer behind records for Jay-Z, Kanye West, Adele and Justin Timberlake, and the man Justin Bieber went into the studio with when he wanted to sound grown up, published The Way of Code with Anthropic: Lao Tzu's Tao Te Ching rewritten, chapter by chapter, for people who build software by talking to machines. Most of it is gentle and strange, full of valleys, water and uncarved wood, and then chapter 28 lands a line that reads like a verdict on every tool war since: "The Vibe Coder knows the tools, yet keeps to the block. Thus, he can use all things. Tools will come and tools will go. Only The Vibe Coder remains." The block is the uncarved wood that every tool is cut from, which in plain terms means your judgement, your taste and your sense of what the thing should become, and the advice is to know the tools well while keeping your weight on the part of you that doesn't depreciate.

Rubin is the right person to say it, because he has lived the argument for forty years. When Anderson Cooper asked him on 60 Minutes in January 2023 whether he played instruments, he said "barely", and when asked whether he could work a soundboard he went further: "I have no technical ability. And I know nothing about music." What he does know, he explained, is what he likes and what he doesn't, and he is decisive about it, which happens to be the exact job description of someone directing an AI agent. The studios he worked in moved from tape to Pro Tools to laptops, the engineers and the instruments changed with them, and none of it changed what artists were paying him for: "the confidence that I have in my taste." In The Creative Act, the book that interview was promoting, he puts it even more simply: "No matter what tools you use to create, the true instrument is you."

That is the lens for this guide. You still have to pick a tool on Monday morning, and the differences between Cursor, Claude Code and Codex are real enough to cost you weeks if you pick badly, so the sections that follow compare them with dated facts and without mercy. But the comparison is arranged around the part that stays, because a vibe coder who knows what good looks like will get good work out of any of the three, while someone without that ear will get confident nonsense out of all of them, only at different speeds and prices.

Seen from the producer's chair, the three tools stop looking like rival products and start looking like three different ways of making a record. Cursor is playing the instrument yourself with a gifted session player at your elbow, who finishes your phrase before you've found it and lets you keep or erase every note, so your hands never leave the keys. Claude Code is producing a band in the room, the way Rubin does it from the couch: you describe the feeling you're after, the band plays it, and you stop the take, ask for it "again, but slower and sadder", and listen once more. Codex is sending your tracks to a mastering house with a precise brief, then getting back a finished master and a spec sheet proving it met the brief, without ever hearing the work in progress. Each is a legitimate way to make a great record, and each produces disasters when you choose it for the wrong song, because you don't send a half-written demo to mastering and you don't book a full band to fix one wrong note.

So the questions worth asking are about you rather than the tool. What shape is the work in front of you, and how much listening can you honestly do before your ears get tired? The second question matters more than any benchmark, because in Stack Overflow's 2025 Developer Survey 66% of developers named AI answers that are "almost right, but not quite" as their biggest frustration, which makes your attention the scarcest instrument in the room. And before we walk through the three studios, it's worth seeing how fast chapter 28 is already coming true, because in the last three months alone one of these tools was sold, another was folded into a bigger product, and the third swapped its engine.

Feature by feature: what Cursor, Claude Code and Codex share, and where they still differ

Before taste, the facts. The table below compares the three tools as they stand on 8 October 2026, checked against each vendor's own docs and pricing pages, and it is worth reading slowly, because about two thirds of it is the same in every column. That shared ground is the good news for anyone choosing: whichever you pick, you get an agent that edits across files, runs commands, works in the cloud, reviews pull requests and reads an open instruction file. The rows where the columns diverge are the ones that should drive your choice.

| | Cursor | Claude Code | Codex | |---|---|---|---| | Started life as | An AI code editor (a VS Code fork) | A terminal agent | A cloud agent plus an open-source CLI | | Where you use it | Editor and Agents Window, CLI (agent), web, iOS and iPad, Slack, JetBrains | Terminal, VS Code and JetBrains, desktop app, web, mobile, Slack | ChatGPT desktop app, CLI, IDE extension, web, mobile, Slack, Linear | | Models | Composer 2.5 and Grok (latest: Grok 4.7, 21 Sep 2026) in the cheaper in-house pool; Claude and Gemini at API rates; OpenAI models due to leave (proposed shutoff 12 Nov 2026) | Claude only: Opus 5.5 (default), Sonnet 5.5, Fable 5.1, Haiku 5.5 (since 7 Oct 2026) | OpenAI only: GPT-6.1 Sol (recommended default), GPT-6 Astra, GPT-6 Luna | | Reads AGENTS.md | Yes, including nested folders | Yes since 18 Sep 2026, but by default only when there is no CLAUDE.md | Yes, natively (32 KiB combined cap) | | Reads CLAUDE.md | Yes, always applied | Yes, natively | Only if you add it as a fallback name | | Agents in the cloud | Cloud Agents on their own VMs | Cloud sessions on the web, scheduled routines | Codex cloud tasks, scheduled tasks | | Parallel work | Cloud subagents, Projects coordinator (beta) | Subagents, worktrees, workflows, agent teams (experimental) | Subagents, worktrees, Ultra reasoning that splits work across agents | | Pull request review | Bugbot (usage-billed on Pro, included on Teams) | Code Review and the GitHub Action | @codex and automatic PR review | | Browser and computer use | Built-in browser; cloud agents can use a desktop | Desktop app browser, Chrome extension, computer use on macOS and Windows (CLI: macOS only) | In-app browser, computer use on macOS and Windows | | Extend it with | Rules, skills, hooks, MCP, plugins (also loads Claude Code skills and hooks) | Skills, hooks, plugins, mods, MCP | Skills, plugins, hooks, MCP, /import from Claude Code and Cursor | | Safety by default | Sandbox toggle, review every diff in the editor | Permission modes; auto mode is the default since 28 Sep 2026 | OS-level sandbox; workspace-write with approvals on request | | Entry plan | Pro, $20/month | Pro, $20/month | Plus, $20/month (Free and Go get Luna only) | | Heavy-use plan | Ultra, $200/month | Max 20x, $200/month | Pro, $100, $200 or $500/month |

Three rows explain most of the real-world differences. The models row matters most, because each tool is now married to a model family: Claude Code runs only Claude, Codex runs only OpenAI, and Cursor, once the neutral ground where you could switch between everyone, is tilting towards its new owner's models. The safety row tells you how each tool expects to be supervised, with Codex fencing the agent in at the operating-system level, Claude Code trusting a permission system and a classifier, and Cursor assuming your eyes are on the diff. The browser and computer-use row looks close on paper, since both Claude Code and Codex can now drive a browser and a desktop, but it is where Codex's polished app and plugin ecosystem make the bigger difference for work beyond code, which matters if half your week is research, spreadsheets and admin rather than features.

What changed this summer

Most comparisons still ranking on Google were written before three changes that reshaped the table above, which is chapter 28 playing out in a single season.

None of this makes any of the three a bad choice, but it does change what is worth investing in. Hours spent memorising one tool's quirks or one vendor's billing unit depreciate on somebody else's schedule, while an instruction file in an open format, a habit of reviewing what comes back and a clear sense of what good looks like move with you from one studio to the next. With the table in hand, the next three sections take each tool in turn: what it does best, where it lets you down, and which work I give it.

Cursor: the best choice when you want your hands on the code

Cursor is the session player at your elbow. It began as a fork of VS Code, the most popular code editor in the world, and its whole design assumes that you are reading and touching the code yourself while the AI suggests, completes and edits alongside you. That is no longer the whole story, though, because 2026 has been the year Cursor learned to work when your hands are off the keys, and most of what it shipped since April is about agents running without you watching.

Where Cursor is strongest

Where it falls short

My take

I don't use Cursor, and the reason says more about the reader than about the tool. Cursor's main selling point used to be access to every model in one place, but I already pay for Claude Code and Codex, so I get Anthropic's and OpenAI's best models directly, and the in-house alternatives, Composer 2.5 and Grok, were noticeably less reliable in my work. Setting up parallel worktrees also took more fiddling than in Claude Code or Codex, where it works out of the box. As a surface I like it less than Claude Code and much less than Codex. If you are a developer who wants to stay inside the code and collaborate with the AI line by line, Cursor may well be the most natural home. For a non-developer like me, though, an editor built around reading code adds nothing that the other two desktop apps don't already do better. The one exception is Grok Bot, the persistent cloud assistant that now comes with paid Cursor plans, which in my experience is probably the best product in its category. In studio terms, the session player is wonderful if you play the instrument, and an expensive piano in the corner if you don't.

Claude Code: the best choice when you want to direct and still hear every take

Claude Code is the band in the room, with you on the couch. It started in February 2025 as an agent that lived in the terminal, and its design still assumes that you describe what you want in plain words, let the agent do the playing, and stop the take whenever it drifts. The terminal is now only one of its doors, because the same agent runs in VS Code and JetBrains, in a desktop app, on the web and on your phone. What changed this autumn is bigger than another door, though. In the space of three weeks the band got a new lead singer, a manager who runs several sessions at once, and a set designer, and Claude Code quietly stopped being a coding tool and became a studio for making almost anything.

Where Claude Code is strongest

Where it falls short

My take

Claude Code was the first coding tool I picked up after Lovable, and I started using it the moment it came out. For many months it was the only tool I used, and almost everything I built until about four months ago was built in it. Like every model family, though, Anthropic's has had its hits and misses, and some of its releases were disappointing: in the months before this autumn Codex was ahead more than once, and I started splitting my coding and AI work between the two. Opus 5.5 is what won me back. It is the strongest model I have used, it feels faster in my hands than Codex on GPT-6.1 Sol, because it writes nearly twice as fast even if Sol tends to finish whole tasks sooner, and it talks to me in plain English instead of filling its reports with "contracts" and "receipts" that I have to translate before I can judge them. That plain speech matters more than it sounds, because trust comes from understanding what happened, and you can only direct a band whose playing you can follow. Today it is where I go for anything that needs design sense or creativity; I run Projects to keep several pieces of work moving at once, and I reach for Design when an idea needs a face. The shared limits are the price of admission. Tools will come and tools will go, and I have watched this one fall behind and come back, but right now, when the work needs taste, this is the band I want in the room. If your work is mostly design, writing or anything where you judge the result by feel rather than by a test, that is the case for starting here too.

Codex: the best choice when the work goes beyond code

Codex is the best work surface of the three, and the only one that mixes coding and knowledge work so well that you stop noticing the switch. It still does what made its name, taking a precise brief and coming back with a finished change and proof that it met the brief, but it now lives inside the ChatGPT desktop app, next to Chat and ChatGPT Work, and that app is the easiest and the most beautiful place to work with an agent today. It rules at computer use and browser use, and its MCP connections simply hold: where Claude Code still drops a server and asks you to reconnect, Codex keeps working. If the mastering house was where you sent finished tracks, Codex has become the whole studio complex, with the best equipment in every room.

Where Codex is strongest

Where it falls short

My take

Codex is the better tool, and Claude Code is the better collaborator, which is why I pay for both. When Codex pulled ahead of Claude this summer I started splitting my work, and that split has stuck: Codex is where I go for knowledge work, for anything that needs a browser, a desktop or a long chain of plugins and MCP servers, and it has the nicest interface of any agent I have used. The models are the catch. GPT-6 was a disaster for me, GPT-6.1 Sol is a real improvement on 5.6, but even at its best it feels slower in my hands, wordier and less clear than Opus 5.5. The numbers explain the feeling: Opus 5.5 writes about 96 tokens a second against Sol's 54, so it reads faster while you watch, even though Sol finishes the whole task sooner because it takes fewer, shorter steps. And Sol still likes to add things that are not critical to what I asked. In studio terms, Codex has the best building in town and the best equipment in every room, and I book it for the sessions where the equipment matters most. If your work is more operations than craft, more inbox, browser and spreadsheet than design and prose, that is exactly the case for making Codex your first studio rather than your second.

One instructions file can serve all three tools

Every session musician who walks into a studio gets the same thing on the music stand: a lead sheet with the chords, the key and a few notes on feel, written so that any good player can pick it up and play. AGENTS.md is that lead sheet for coding agents, and since 18 September 2026 it is the first one all three tools can read, which is the most practical gift this autumn has given anyone who works with more than one of them. You write down once how your project works, what to avoid and how to check the work, and the same file briefs Cursor, Claude Code and Codex alike, so you stop maintaining three slightly different versions of the truth and wondering which one each agent believed.

The catch is that each tool reads the sheet slightly differently, and one of those differences can quietly switch it off:

The setup that works in all three is a single AGENTS.md as the source of truth and, only if you need Claude-specific extras, a short CLAUDE.md whose first line imports it:

Anthropic's own docs recommend exactly this pattern, and say that keeping the import never makes Claude read AGENTS.md twice. Because Cursor reads CLAUDE.md as well, keep that file to the import and a few Claude-only lines, and never copy the shared rules into it. Keep AGENTS.md itself short, too: an ETH Zurich study from February 2026 found that context files did not generally make coding agents more successful but raised their running cost by over 20%, whether an AI or a developer wrote them, and that repository overviews in particular did not help. Where the files earned their place was in spelling out the unusual rules of a project, the things an agent could not guess on its own. I have written about the file itself in Start with CLAUDE.md. The short version is that the lead sheet is the one part of your setup that moves with you when the tools change, which makes it the best hour you will spend on any of them.

What Cursor, Claude Code and Codex cost in October 2026

Studio time has always been sold in two ways: by the hour, where you pay for exactly what you use and wince at every overrun, or by the block, where you pay up front and the clock stops mattering until you hit the end of the booking. All three tools now sell a block with an hourly meter hidden behind it, and the sticker prices look almost identical, which is precisely why they are misleading.

| | Cursor | Claude Code | Codex | |---|---|---|---| | Free tier | Hobby: limited agent requests | No Claude Code on Free | Free and Go ($8): GPT-6 Luna only | | Entry plan | Pro, $20/month | Pro, $20/month ($17 billed yearly) | Plus, $20/month | | Middle plan | Pro+, $60/month | Max 5x, $100/month | Pro, $100/month | | Heavy-use plan | Ultra, $200/month | Max 20x, $200/month | Pro, $200 or $500/month | | Team seat | Teams, $40/user/month | Team Standard, $25/user/month ($125 Premium) | Business, $25/user/month ($125 Premium) | | What the plan buys | A usage allowance, generous on Composer and Grok, at API rates for Claude and Gemini | A five-hour window plus a weekly cap, shared with all of Claude; Fable up to half the weekly cap on Max | Usage shared with ChatGPT Work, measured in messages per five hours and weekly limits | | When you run out | On-demand usage, billed afterwards | Usage credits at API rates, if you turn them on | ChatGPT credits |

Prices are monthly in US dollars before tax, checked on each vendor's pricing page on 9 October 2026.

The $20 tier in all three is a test drive rather than a working plan: enough for a few focused sessions a week, and Cursor's own docs admit that daily agent users typically spend $60 to $100 a month and power users $200 or more. The real comparison starts at $100 to $200, and there the difference is less the price than the meter. Codex publishes the clearest numbers, a range of messages per five hours for each model, although OpenAI has reset and rebalanced them several times this year. Claude Code publishes only multiples, Max being "5x or 20x" Pro, and its pool is shared with everything else you do in Claude. Cursor's limits are the hardest to see, because it stopped publishing what each plan includes in August.

Value for money is where my own view has moved most this year. For a long stretch Claude Code burned through its limits fast, and Codex was the better value for money, comfortably. Since Opus 5.5 arrived on 22 September, that has changed. Anthropic says it uses fewer tokens per task than Opus 5 and raised five-hour limits on the same day, and in normal, reasonable use I no longer hit mine, and today I would call the two balanced. Both give good value for what they cost, so the choice between them comes down to the shape of your work rather than the price. Cursor is the odd one out. Its allowance resets monthly rather than weekly, and on its own models, Composer and Grok, it lasts a very long time; switch to Claude or GPT and even the top plan burns fast, because those run at API rates. The shame is that the models that last are not the ones you most want doing the work.

What no price table shows is how much of the booking you will actually use. Theo Browne made the bluntest version of this point in a video on 1 October 2026: once you have paid, "your ability to work and put out even bigger things is only limited by your token usage." His test is how many agent threads you have running right now, and his bar is five. I would keep the idea and drop the number. A plan earns its price when it is busy on work you had already decided was worth doing, ideally while you sleep, and an expensive plan sitting idle is the worst deal of the three. Five threads of the wrong song are still the wrong song, though, which is why the cheapest hour of studio time is the one you don't waste on a take you never needed, and where the next section comes in.

Which one to pick: start from the work, not the tool

A good song can still be ruined by the wrong room. An intimate ballad drowns in a hall built for an orchestra, and a loud band sounds boxed in inside a vocal booth, even when the musicians and the gear are exactly the same. The same is true here, which is why "which tool is best?" is the wrong first question. The better ones are about the work itself: what shape is it, and how much of what comes back can you honestly listen to before it ships?

Shape comes first, and it is easier to read than it sounds:

| If the work is mostly… | Start in | Because | |---|---|---| | Code you want to read, shape and own line by line | Cursor | You stay at the keyboard, with an agent one keystroke away | | Judged by feel: design, writing, product decisions, a blank page | Claude Code | The strongest model right now, and the clearest collaborator while the brief is still forming | | Well specified, long, or spread across apps, browsers and documents | Codex | The best work surface, computer use that works, and jobs that keep running while you're away | | A bit of everything, and you are not a developer | Claude Code or Codex | Cursor's edge lives in the editor, the one place you won't spend your day |

The second question is the one no pricing page asks. Every agent thread hands you something to listen back to before it ships, and your ears do not scale the way the threads do. This is why I keep Theo's idea and drop his number: the right number of threads is however many you can review properly in the time you actually have, and for most people on most days that is fewer than their plan allows. A thread that finishes and then sits unheard until tomorrow was never saving you time. It was only queueing it, which is the same lesson I found in running several agents without losing judgement.

Do you need two studios? I pay for both, Claude Max and ChatGPT Pro at $200 a month each, because my work splits between them, and each does its half better than the other would. If yours doesn't split, one is enough, and the saving is the smallest part of it: one set of instructions, one set of habits, one place to look when something breaks. If you do add a second, give each one a job rather than asking both the same question, which I've written about in Stop Being Loyal to One AI Stack.

Whichever you pick, you will feel sure about it within a week, and that feeling is exactly what the last section is about.

Only the vibe coder remains, so measure like one

The line this guide opened with began as a joke. In February 2025 Andrej Karpathy coined "vibe coding" for building software by describing what you want and letting the AI write it, and within weeks the internet had chosen its face: a photo of Rick Rubin, eyes closed, headphones on, one hand resting on a mouse. Rubin had never seen the picture and assumed it was AI-generated, but it was real, taken at a hi-fi show in Germany where the mouse was only turning the volume up. The week after that, he said, "a company called Cursor, which I don't really know what they do," published fifteen rules of vibe coding with his photo at the top. So he did what a good producer does with an accident and leaned into it. He posted the first joke of his life, "Tools will come and tools will go. Only the vibe coder remains," watched it pass a million views when his daily posts usually drew tens of thousands, and turned the joke into a book, The Way of Code, with the gloriously deadpan subtitle The Timeless Art of Vibe Coding, for a term that was, in his words, "maybe ten weeks old."

He wasn't only joking, though. Before vibe coding, he said, you had to be a virtuoso coder to make something great, "and now everybody can do it." That is what these three tools really are under the pricing tables: they hand the designer with an app in a sketchbook, the writer with a product in their head and the founder who never learned to code everything they need to put an idea live. I'm not an engineer either, and the site you're reading this on exists because describing it turned out to be enough.

That freedom has a catch. Making something is now easy, but easy doesn't tell you whether it is any good, and the honeymoon with a new tool makes that harder to hear: the first big job lands in minutes, and from then on every good result confirms the switch while every bad one gets filed under bad luck. That is outcome bias, the habit this blog is named after, judging a decision by one result rather than by whether it was sound. Don't expect the researchers to settle it for you either. METR, the lab that caught developers working slower with AI in 2025 while they believed they were faster, tried to run its study again and said in February 2026 that the new data was "unreliable", partly because so many developers now refuse to work without AI, and partly because nobody could time a task cleanly once several agents were running at once.

So listen to your own work the way Rubin listens to a take. When he describes how he makes things with AI, it is never a single prompt and a verdict: he asks for different versions, compares them, narrows them to two he likes, often combines the best of both, and stays "open to being wrong." For two weeks, do the same with your tools. Give each one the kind of work the table in the last section sends its way, and keep a one-line note per task: what you asked for, how long until you had something you could use, how many minutes you spent checking it, and whether any of it had to be redone. Ten tasks a tool is enough to see a pattern, and what you are looking for is not the fastest take but the lowest total cost of reaching something you would ship, because that is where the honeymoon hides its bill.

By this time next year at least one of these three studios will look nothing like it does today, and every comparison table, this one included, will need rewriting. Your notes won't, because what they train is the thing Rubin has been selling for forty years without playing a note: knowing what you like, knowing what you don't, and checking the one against the other. Pick a studio, write down what happens, and make the thing only you would make. Tools will come and tools will go. The joke was right.

Frequently asked questions

Which is better, Cursor, Claude Code or Codex?

None wins outright in October 2026. Cursor is best when a developer wants to stay in the code, Claude Code when the work is judged by feel (design, writing, product decisions), and Codex for long, well-specified work across apps, browsers and documents.

Is Claude Code or Codex better value for money?

Since Opus 5.5 launched on 22 September 2026, the two are roughly balanced. Codex used to give more usage per dollar, but Anthropic cut Opus 5.5's token use and raised five-hour limits, so at $100 to $200 a month both give good value.

Can one instructions file work for Cursor, Claude Code and Codex?

Yes. Since 18 September 2026 all three read AGENTS.md. Keep it as the single source of truth and, if you need Claude-only settings, add a short CLAUDE.md whose first line is @AGENTS.md.

Do I need to pay for more than one AI coding tool?

Only if your work splits cleanly, for example craft in Claude Code and operations in Codex. If it doesn't, one tool is enough, with one set of instructions and habits.

How do I know if an AI coding tool actually makes me faster?

Log two weeks of real tasks: what you asked for, time to a usable result, minutes spent reviewing, and whether anything had to be redone. Even METR said in February 2026 that agent work has become hard to time, so your own log beats the feeling of speed.