Reflections

Why I’m Leaving OpenAI and Lessons

After about a year and a half at OpenAI, I’m leaving to build something of my own. During my time there, I worked on the core Codex team. Watching it grow from an experiment into such a huge part of OpenAI, with more than 30 million weekly active users today, has been a once-in-a-lifetime experience.

My friends and family keep telling me I’m terrible at documenting what I learn, so I’m using this article to finally change that a little. I want to share what it was actually like to build on that team, how the experience changed the way I think and work, and some of the lessons I’m taking with me.

Specifically:

I’m writing this as much for myself as for anyone else trying to build a company in the AI era.

Welcome to Codex

I joined the Codex team two weeks before the Codex Web launch in May 2025.

I expected a gradual onboarding ramp. Instead, my first assignment was, roughly, to help get the product ready to ship. Reliability was below a single nine. I partnered with Tom Wiltzius, one of the best engineers I’ve ever worked with, to fix anything and everything while the rest of the team kept adding features.

I definitely felt thrown into the deep end without much context, but there wasn’t really time to complain because:

  • Failures showed up at every layer: repository setup could fail, GitHub connections could be flaky, branches could change underneath a running task, and consolidation could fail in its own ways. Then there was the question of recovery: sometimes we retried too little, sometimes far too much. You know, typical distributed systems stuff.
  • Then there were the agent problems. The agent could call the wrong tool, a prompt could send it in the wrong direction, or the eval itself could be broken. Getting a task to finish meant figuring out which of these things had actually gone wrong.

Being dropped straight into the work forced me to use Codex Web and Cursor to learn the codebase quickly, with the help of the very product I was trying to fix. I also hadn’t written frontend code in about a decade, beyond the odd toy example. Suddenly, I was making frontend changes to Codex Web.

I would later onboard tens of people to the team in much the same way: give them a real problem, get them to use Codex immediately, and let them learn by doing.

This was all week 1.

I remember laughing to myself at how much I had managed to get done. I don’t think I would’ve realized I could move that fast if I hadn’t been put in an environment where that was simply expected of me.

A month inside OpenAI can feel like a year. People move at an unbelievable pace, doing in weeks, and sometimes days, what would normally take months.

A week later, we launched Codex Web inside ChatGPT.

In two weeks, I’d gone from barely knowing how Codex worked to helping ship it, debugging reliability issues, working across the stack, and making changes to the product itself. That felt like a major victory.

Distribute Your Bets

Codex Web was quite different from the version of Codex most people know today. It was a web interface where you connected a repository, described a task, and handed it off to an agent running in the cloud. You could keep doing your regular work while it handled a small fix or worked through a larger independent change.

On paper, at least, we thought this should’ve gone viral with engineers. What wasn’t to like? You could hand work off to an agent, let it run in the cloud, and come back when it was done.

Except it didn’t quite happen that way.

The launch itself was a success, but fewer people kept coming back than we wanted. Part of it was the interface, and part of it was that cloud agents were still complicated at the time. We were figuring out the infrastructure while also asking engineers to change how they worked and get comfortable handing work off asynchronously.

When you’re building something this new, you rarely know which interface will catch on. We kept several in development and watched how people used them.

We put serious effort into several different interfaces:

  • Codex CLI and VS Code. A couple of months after launch, it was becoming clear that working interactively with an agent felt more natural to many engineers. They could see what it was doing, redirect it, and stay inside an environment they already knew. We took that as a signal and started putting more effort into Codex CLI, and later the VS Code extension.
  • Codex Web. It was still very much alive, especially for incident investigation and larger independent changes. But most active usage was increasingly coming from the CLI and VS Code. People were showing us where an agent naturally fit into their work.
  • Slack. We noticed internally that engineers would discuss a problem in a thread, move that context into Codex Web, get a pull request back, and then take it into GitHub for review. Oddly, the context for one problem was spread across multiple systems. So we built Codex in Slack. People could delegate work directly from the conversation where the problem was already being discussed. I had a lot of fun wrangling Slack APIs, figuring out which context mattered, and trying to keep the agent focused. It was an early glimpse of multiplayer development, with people and an agent working from the same thread. Claude Tag is a more recent take on that same idea. We also built a common runtime that made integrations like Codex in Linear easier to develop.
  • The Codex app. As the models improved, people internally started giving agents longer and more complicated tasks. More terminals were open, more agents were running in parallel, and more of the human work became deciding what to delegate, coordinating tasks, and evaluating what came back. That became one of the strongest signals behind the Codex app, which launched in February 2026 as a command center for agents.

Building Codex inside OpenAI gave us one unfair advantage: we could watch our own company figure out how to use agents before the behavior was obvious anywhere else. That made internal usage one of the strongest signals for where the product needed to go next.

Taking Codex Beyond Software Engineering

Pretty soon, people across OpenAI were giving Codex work well beyond writing code. We saw it help with design documents, Slack conversations, feedback intake, and all kinds of work that didn’t really look like software engineering anymore.

The same ability to gather context, use tools, and work through a task was useful well beyond code.

Skills gave Codex instructions for particular kinds of work. Plugins brought those instructions together with connections to the tools and information people already used. As we added more general tools and interfaces that didn’t assume you were a developer, Codex started becoming useful across design, research, operations, and other roles.

Our team was pretty obsessed with users. We watched how colleagues used Codex and what people said online, then used that feedback to make small changes that helped people who weren’t engineers use it.

One project that was especially close to me was ChatGPT Sites.

With Sites, someone who isn’t technical can create a website, refine it through conversation, and share it directly from ChatGPT. Generating code had become much easier; turning it into something another person could open and use still meant dealing with hosting, deployment, databases, and storage. Sites handles that work so the whole process can stay in the conversation.

I worked on Codex Sites end to end, from inception through launch. Four or five of us built much of the early version, and more people joined as we brought it to users and kept improving it.

I could probably talk for hours about the tiny design decisions and all the back and forth that went into making the experience feel seamless. When you’re making dozens of painfully small product decisions, their effect compounds. You need something to anchor them.

For Sites, that anchor was actually pretty simple: does this make it easier for someone, especially someone who isn’t deeply technical, to just build?

Funny enough, that anchor came straight from the product tagline. I’d never really understood why people cared so much about product taglines before, but when you’re staring at five reasonable options for a decision, having one clear product principle to come back to really helps.

You Can Just Build Things — Codex
You Can Just Build Things, Codex

That influenced everything from hosting to storage to how much of the underlying system we exposed. Even today, I think Sites is one of the cleanest examples of letting Codex handle the machinery while the person stays focused on what they’re trying to create. I’m glad we made a lot of those decisions the way we did.

This broader direction eventually brought ChatGPT and Codex together in the same app. ChatGPT Work gave people a place to use these capabilities across many kinds of work, while Codex remained focused on software development. By then, the audience for what we had been building had grown far beyond the engineers we started with.

For us, “just build things” meant that whenever we had a choice between exposing more complexity or taking work off the user’s plate, the answer was usually pretty clear.

Lots of Other Fun Side Quests

During my time on Codex, I picked up whatever needed to get done, which was pretty normal on the team. OpenAI felt remarkably flat to me, especially within Codex, and that made it easy to work across teams as the needs changed. I learned a lot of things I hadn’t expected to work on when I joined.

When I joined, there were probably fewer than 15 of us. Even as the team grew, people had a lot of independence to pick up a problem and build something. For a while, we had a Friday demo tradition that was one of my favorite parts of the week. Someone would show a prototype in a part of the product they didn’t normally work on, a fix for an annoying papercut, or something that simply made the team faster. A surprising number of those prototypes eventually shipped.

It would take forever to go through all the projects I worked on, but one I remember especially well was Code Review.

As agents started taking on more work, we also had to figure out how to review what they produced. Part of the motivation came from alignment and safety work around checking whether an agent had actually done what the user intended.

The research team absolutely cooked on this one. They made Codex really good at catching issues, and the product work had to evolve alongside it. Reviews needed to be useful enough that people would keep running them. Catching more bugs helped, but commenting on every possible issue would quickly make the review exhausting to use.

I worked on a lot of the context engineering around those reviews: avoiding repeated findings, collecting implicit feedback from users, and using that feedback to improve the product. Larger teams also needed controls over when reviews ran, what they looked at, and what they ignored. Working closely with research taught me how much the context, feedback loops, and product decisions could change the experience, even with the same model underneath.

I still think Code Review in Codex is underrated, and I’m leaving with a bit of an itch that we probably could’ve pushed harder in this direction to build out a comprehensive product. It also got a fair amount of attention for finding bugs in code written by our dear competitors’ tools, which was always fun.

From there, I moved into GPU capacity and fleet management for Codex. I learned a lot about inference management and what it took to keep the product running. Some of that learning involved being woken up at 1 a.m. Pacific to deal with GPU capacity issues.

We were already using Codex constantly ourselves, including for incident triage. That worked beautifully until the infrastructure incident also took Codex down. Those were the nights when we suddenly had to remember where all the dashboards were, which commands to run, and how to query the logs ourselves. I was terrible at some of those queries. Having to figure them out while dealing with a capacity problem was quite an education.

Using our own product and hearing directly from colleagues made the learning fast, even when it was uncomfortable. I want to carry that way of working into what I build next.

Sweet Pain

I said something on Lenny’s podcast that ended up going kind of viral: “Pain is the new moat.”

At the time, I was talking about building AI products. But I think it describes a lot of my own experience at OpenAI too.

A lot of my time at OpenAI was sweet pain. I was stretched, slightly uncomfortable, and learning faster than I’d like. That’s where most of my judgment came from.

Nobody could’ve handed me that judgment. I couldn’t have read my way into it or reasoned my way into it from first principles. I had to build, get things wrong, fix them, and do it again. That’s why I still believe pain is the new moat.

What I Want to Build Next

We now have an incredible amount of intelligence available, but the path from that intelligence to a company actually using AI well is still far from obvious.

There’s still a lot of groundwork: figuring out where an agent should help, giving it the right context and tools, and knowing when to trust it and when to check its work. Companies also need systems built around these capabilities and help changing how they work.

At OpenAI, I got to see what happens when people use these systems deeply. I also saw how much work sits between having access to a great model and actually changing how an organization operates.

That made me much more convinced that there’s a huge amount left to build here, especially outside the AI bubble. And I mean bubble in the nicest possible way :)

Codex also completely changed my sense of what a small number of people can take on.

I want to test that belief for myself now.

I want to work across a wider range of problems, build closely with companies, understand where these systems actually work and where they don’t, and share what we learn along the way.

That’s why I’m joining Aishwarya Naresh Reganti to build LevelUp Labs together. We help companies build and adopt AI through engineering, strategy, and education. I’ll be working closely with teams to understand their problems, build systems they can actually use, and help them learn to work with those systems themselves.

Leaving OpenAI is bittersweet.

I’m deeply grateful to the Codex team1 for trusting me with problems far beyond what I thought I knew how to solve, helping me through the unfamiliar parts, and making it so much fun to build together.

If you’d asked me a few years before I joined OpenAI whether I could imagine working on something like Codex, watching it become such a major part of OpenAI, and then choosing to leave to build something of my own, I probably would’ve laughed.

Who in their right mind would leave? But I think I understand now.

I’m leaving because of what this experience gave me, not despite it.

When you spend the longest year and a half of your life seeing how much a small team can build, how quickly your own limits can move, and how many things that once felt impossible can become normal, it changes what you’re willing to attempt next.

In a strange way, you start wanting the sweet pain again, because now you know what can be on the other side of it.

A year and a half ago, Codex was the deep end.

Now I’m choosing another one.

This time, I know a little better what can happen when you jump in.

Wham!

  1. A special thank-you to Alexander Embiricos, Thibault Sottiaux, Fouad Matin, Rasmus Rygaard, Michael Zeng, Will Wang, Xin Lin, Anton Panasenko, Joey Trasatti, Alec Barber, DDR (David de Regt), Sam Tang, Rohan Varma, and Jon Abrams. Thanks also to Maja Trębacz and Sam Arnesen from the research team for their work on Code Review, and to the many others I’m sure I’ve missed. ↩

Enjoyed this article? Leave a like.

Comments

Share a thought or a question. Just your name, no account needed.

Your name and comment will be public.