Facebook Pixel
DevOps Maturity Model: 5 Stages to AI-Native Delivery

DevOps Maturity Model: 5 Stages to AI-Native Delivery

A five-stage DevOps maturity model for engineering leaders, from manual releases to AI-native delivery, built on DORA's 2025 research.

Muhammad Arslan Aslam
••7 min read

Every DevOps maturity model I have read has the same shape. A team starts with manual deploys and ends somewhere called "optimized" or "elite." Each stage sits on top of the one before it, like floors in a building. The shape is useful. It is also a little too tidy, and the best-known research program in this field dropped it last year.

For most of a decade, DORA sorted teams into low, medium, high and elite performers. In its 2025 report, DORA replaced those tiers with seven team profiles, built from a cluster analysis of nearly 5,000 technology professionals. Some of those profiles do not fit a ladder at all. About 7% of teams land in one DORA calls "high impact, low cadence," which ships relatively rarely yet reports strong product performance, alongside high friction and burnout. A linear model has nowhere to put that team.

I still think a five-stage model is worth writing, because the order matters when you decide where the next quarter of engineering time goes, and that is the only job I would give this one. It is a tool for choosing the next investment. It works badly as a score, and one codebase will usually sit at several stages at once. The checkout service might be at stage four while the billing cron job nobody has touched since 2022 sits firmly at stage one.

AI makes the order matter more than it used to. DORA's 2024 research found that every 25% increase in AI adoption came with an estimated 7.2% drop in delivery stability, as TechTarget reported. By 2025, throughput had recovered while instability stayed elevated, and DORA described AI as an amplifier of whatever system it enters. My reading is that AI rewards the stages underneath it. A team that buys agents at stage two gets its stage-two problems sooner and in larger batches.

Stage 1: Releases depend on specific people

At stage one, a deploy is an event. Someone with the right access and the right memory runs a sequence of steps, some written down and some not, usually outside business hours. Configuration lives on the server itself. When that person is on leave, releases wait.

Most teams at this stage know it, and most got here for sensible reasons. The product grew faster than anyone had time to formalize. The way out is mostly unglamorous work: a deploy script, configuration in version control, and a second person who can release without help. That last condition is the honest test. If only one person can deploy, the team is at stage one regardless of the tools it pays for.

Stage 2: The pipeline exists, and trust in it is low

Stage two teams have CI/CD. Every merge triggers a build, and a deploy is one command or one button. A lot of teams stop here, because the obvious pain went away.

The remaining pain is about trust. Test coverage is thin or uneven, so a green pipeline proves the code compiles and that a few happy paths still work. Engineers compensate with manual checks and release freezes before holidays, and they batch changes into large weekly releases because each release feels risky. Large batches then make each release riskier. That loop keeps teams at stage two for years.

This is also the stage where AI tools tend to arrive, because they are the cheapest and most visible thing a leadership team can buy. Faros AI's 2025 telemetry on more than 10,000 developers found that on high-adoption teams, pull requests grew 154% larger and review time rose 91%. Put that on top of a pipeline the team already doubts, and the weekly release gets bigger and harder to trust.

Getting out of stage two is mostly a testing project. If end-to-end tooling is the open question, this comparison of Playwright, Cypress and Selenium covers the trade-offs. I would still put unit and integration coverage on the money paths first, since those are where a silent failure costs the most.

Stage 3: Delivery is measured and changes are small

Stage three is where a team starts managing delivery with numbers instead of feel. DORA's metrics are the standard set, and DORA now tracks five: change lead time, deployment frequency, failed deployment recovery time, change fail rate, and the newer deployment rework rate.

The numbers matter less than the habits that move them. Changes get small enough to review in one sitting, and short-lived branches or trunk-based development replace month-long feature branches. Feature flags let a team merge unfinished work without releasing it to users. Somebody owns on-call, and production has enough logging and tracing that an incident starts with evidence instead of guesswork.

DORA's work on small batches found that they amplify the positive effect of AI adoption on product performance. Put plainly, stage three is the first stage where AI-generated code has a reasonable path to production. Diffs stay reviewable, and when quality slips the metrics show it within days.

You can usually tell a team has reached stage three when a Friday afternoon deploy stops being a topic of conversation.

Stage 4: A shared platform carries the practices

At stage four, the habits from stage three stop living in each team's heads and move into a shared internal platform. New services start from a template that already includes CI, logging, secrets handling and a deploy path. Environments get created on demand, and security and compliance checks run as policy in the pipeline, so nobody has to remember them.

This is now the common case, at least on paper. DORA's 2025 report found that 90% of organizations have adopted at least one internal platform and 76% have dedicated platform teams. The same research found that platform quality changes what AI is worth. Where the platform was good, AI adoption showed a strong positive effect on organizational performance, and where it was poor the effect was close to nothing.

"Adopted a platform" covers a lot of ground, though. I would be skeptical of any team that counts itself at stage four because it has a Kubernetes cluster and a wiki page. DORA's guidance is to run the platform as an internal product, with developers as its customers. A platform that engineers route around is a stage-two pipeline with extra YAML.

A smaller company can reach stage four without a platform team. One well-kept service template, one deploy path that every repo uses, and infrastructure defined in code will do it. The ZapBook rebuild in the SociiLabs case studies gives a sense of the scale. A venue booking platform with five years of history had run on a single VPS before it moved to GCP infrastructure as part of a full rebuild.

Stage 5: AI works inside the delivery system

"AI-native" gets used loosely, so here is what I mean by it. At stage five, AI takes part in how code is written, reviewed, tested and operated, and the delivery system was adjusted to carry the extra volume. At SociiLabs, engineers write code with Cursor in the loop, but the editor is the least interesting part.

Review capacity gets sized for AI output. PR size limits run in CI, automated checks pass before a human looks, and an AI reviewer takes a first pass. Human attention is left for the code around authentication and payments. The earlier roadmap on AI transformation goes through that sequencing in detail.

The model also needs the context senior engineers carry. Conventions files and architecture decision records sit where the tools can read them, and each one has an owner so it stays current. Anas Shahid's piece on how senior engineers get better output from AI covers the human side of this.

Agents take bounded work, such as dependency upgrades, test backfill and mechanical refactors, under a daily ceiling on spend and actions. For anything that touches production data or infrastructure, an agent can draft the change while a person approves and runs it.

Most teams miss that AI features in the product need their own tests and their own telemetry. A unit test cannot tell you whether a model gave a good answer. Stage-five teams run evaluation suites in CI whenever a prompt or a model version changes, and they trace model calls in production with a tool such as Langfuse. If that sounds unfamiliar, Muhammad Ali Saeed's piece on why testing AI is nothing like testing normal software is a reasonable place to start.

None of this needs a frontier-model budget. It does need the four stages underneath.

What skipping stages looks like

The clearest public example I know is from July 2025. Jason Lemkin, the founder of SaaStr, was building an app on Replit when its AI agent deleted his live production database during a code freeze he had explicitly set. Replit's response included automatic separation between development and production databases. That fix is a stage-four control, environment separation, added after a stage-five tool had already been given production access. The full story is in Should Startups Vibe Code in Production?

Most skipped stages fail less dramatically. A team at stage two adds agents, the release gets larger, and someone adds a release freeze. Six months later the team ships less often than it did before the tools arrived. Nobody decides that outcome. It accumulates.

I also doubt that teams move one stage per quarter, whatever the slide deck says. My guess is that the jump from stage two to stage three takes longest, because it changes habits more than tools. A pipeline can be bought in a week. Getting forty engineers to merge small changes daily takes much longer, and some of them will not like it.

Where to start

Find the lowest stage among the two or three services that matter most to the business, and work on that one. A stage-four platform does little for a billing service that only one person can deploy. And if the team is already stretched, it is fine to fix one service properly before touching the rest.


Maybe the jump from stage two to stage four is the part your team cannot staff right now. That is the kind of migration and scaling work a 30-minute call can scope. Bring the problem. Leave with a plan.

Book a 30-min scoping call

Related reading: AI Transformation for Engineering Teams: A Practical Roadmap and Should Startups Vibe Code in Production?

Subscribe To Our Newsletter

Real talk on building software that ships.

MVP scoping, tech decisions, and the stuff agencies won't say out loud. Every two weeks.

We respect your inbox. Unsubscribe anytime.
By clicking 'Subscribe' you are confirming that you agree with our Terms and Conditions.