> ## Content Index
> Fetch the complete content index at: https://sociilabs.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# AI Transformation for Engineering Teams: A Practical Roadmap
- URL: https://sociilabs.com/blog/ai-transformation-engineering-teams
- Published: 2026-10-05T00:11:55.000Z
- Updated: 2026-10-05T00:11:55.000Z
- Description: Your team adopted AI and delivery barely moved. A five-step roadmap for engineering leaders, built on DORA and Faros data.
- Author: Muhammad Arslan Aslam
- Tags: Artificial Intelligence, Code Review, Engineering Best Practices, Developer Tooling, Quality Assurance

Most engineering teams finished adopting AI a while ago. DORA's [2025 State of AI-assisted Software Development report](https://dora.dev/research/2025/dora-report/?ref=blog.sociilabs.com) surveyed nearly 5,000 technology professionals and found that 90% use AI at work. The median user spends about two hours a day with it. If your team already pays for coding assistant seats, you are in the large majority.

So why do so many of those teams ship at roughly the speed they did two years ago?

Faros AI has one answer, drawn from telemetry on more than 10,000 developers across 1,255 teams in its [2025 AI Productivity Paradox report](https://www.faros.ai/blog/ai-software-engineering?ref=blog.sociilabs.com). Developers on heavy-AI teams completed 21% more tasks and merged 98% more pull requests. Review time on those pull requests rose 91%, and the researchers found no significant link between AI adoption and improvement at the company level. Most of the extra output, in other words, ended up waiting for the same senior reviewers who were already busy.

DORA's 2025 numbers come at this from another direction. AI adoption now correlates with higher delivery throughput (in 2024 it did the opposite). It also correlates with higher delivery instability, meaning more failed changes and more rework. DORA sums this up by calling AI an amplifier of whatever system it lands in. Inside one company that usually means both effects at once, with the well-tested corner of the codebase speeding up while the module everyone avoids gets worse.

At [SociiLabs](https://sociilabs.com/?ref=blog.sociilabs.com) I run an engineering team where [Cursor](https://cursor.com/?ref=blog.sociilabs.com) sits in the coding loop, and I have read most of this research. My view is that AI transformation for an engineering team is mostly a change to the delivery system around the code. The coding step got cheaper much faster than review, testing, model context, and safe release did. The rest of this piece goes through those, in about the order I would tackle them.

## Why faster coding has not shown up in delivery numbers

A reviewer's focused hours per week stayed the same when the team bought AI seats. Each engineer now opens more pull requests, and Faros found the average one on high-adoption teams is 154% larger. Review load grows faster than output under those conditions. Large diffs are also harder to read carefully, so a reviewer either slows down or starts skimming. Skimming is the worse of the two, because it stays invisible until something breaks in production.

Codebase age matters as well. Stanford research cited in [DORA's 2026 ROI report](https://dora.dev/ai/roi/report/?ref=blog.sociilabs.com) found AI gains of 35 to 40% on simple greenfield tasks and 10% or less on complex legacy code. A team with a live product spends most of its week in the legacy case, while most of the demos you see are greenfield.

I should be honest about the sources. Faros sells engineering measurement software, so a finding that teams need better measurement suits its business. DORA relies mostly on surveys, which means self-reported data. I still take both seriously. They agree with each other, and they agree with a [randomized trial METR ran in 2025](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/?ref=blog.sociilabs.com), in which experienced open-source developers expected a 24% speedup from AI. On their own repositories they were 19% slower. Any one of these studies could be wrong. All three pointing the same way is harder to dismiss. Output per developer looks like a poor number to manage by.

## A written stance and a baseline, before anything else

The first item in the [DORA AI Capabilities Model](https://dora.dev/ai/?ref=blog.sociilabs.com) is a clear and communicated AI stance. Engineers who are unsure what is allowed tend to avoid the tools or use them in ways nobody reviewed. One page fixes a surprising amount of that. Mine would cover the approved tools and the repositories they can see. It would also list what an agent may never do without a person (running a migration against production, for one). The most useful line is about ownership. Whoever merges a change owns it, regardless of what wrote it.

Measurement comes next, and 2026 already offers a cautionary example. Early in the year some companies started ranking engineers on internal leaderboards by AI token consumption, a practice that became known as tokenmaxxing. DORA [responded in June](https://dora.dev/insights/finding-balance-in-the-era-of-tokenmaxxing/?ref=blog.sociilabs.com) with a piece citing reports of developers running autonomous agents on pointless projects to stay above the company average. Its point was that token count measures activity and is easy to game. Shopify ran into this directly. As [The Pragmatic Engineer reported](https://newsletter.pragmaticengineer.com/p/the-pulse-tokenmaxxing-as-a-weird?ref=blog.sociilabs.com), the company replaced its token leaderboard with a usage dashboard and added circuit breakers that catch runaway agents and spend spikes. Starburst never set token targets in the first place. Its VP of Engineering, Jitender Aswani, [told Trending Topics](https://www.trendingtopics.eu/tokenmaxxing-is-ai-token-consumption-a-productivity-metric-or-vanity-trap/?ref=blog.sociilabs.com) that the team judges AI against DORA metrics, code quality, and incident resolution times. Starburst reportedly cut time to production by 60% since December. That figure is self-reported, but the measurement choice behind it is the useful part.

For a team of 15 or 50 engineers, a reasonable baseline is lead time from commit to production, change failure rate, and DORA's newer deployment rework rate. I would add median PR size and time to first review. Those two move weeks before the delivery numbers do. Most of it already sits in the Git host and the CI system. Things will probably get somewhat worse first, and only a baseline tells a temporary dip from a real failure.

## Fix code review before adding more AI

Most teams skip this part because it does not feel like AI work. I think it is where most of the return gets decided.

Batch size is the first lever. [Working in small batches](https://dora.dev/capabilities/working-in-small-batches/?ref=blog.sociilabs.com) is one of DORA's seven AI capabilities, and the mechanics are easy to see. AI made a 900-line diff cheap to produce, and that diff costs a reviewer as much attention as it always did. Before buying any AI review tool, I would add a CI check that asks for a written reason whenever a PR passes a few hundred changed lines. It is a crude rule. It changes behavior anyway, mostly because nobody enjoys writing the reason.

Automated gates come next, with tests, type checks, linting, and dependency and secret scanning all running before a person looks at the diff. An AI reviewer can sit on top of those as a first pass for style problems and obvious bugs. That frees human attention for the code where a quiet mistake is expensive, such as authentication, payments, tenant isolation, data access and migrations.

[Sonar's 2026 State of Code survey](https://www.sonarsource.com/blog/state-of-code-developer-survey-report-the-current-reality-of-ai-coding/?ref=blog.sociilabs.com) covered more than 1,100 developers. It found that 96% do not fully trust AI-generated code, yet only 48% say they always check it before committing. Sonar sells code review software, so the same caveat as Faros applies. Put those two numbers side by side and you get a team that pays for its doubt in slower reviews while still shipping whatever slips past.

And then there is test coverage, where I would be selective. A coverage percentage across the whole repo is mostly a vanity number. What matters is coverage on the paths where a silent failure costs money or customer data. Writing those tedious tests also happens to be one of the better first jobs for an AI tool.

## Context for the model, then agents

At the start of each session, an AI coding tool has no idea why the billing service talks to the queue the way it does. It also has no idea which pattern the team abandoned last spring. DORA treats AI-accessible internal data as its own capability for that reason. In practice this is unglamorous documentation work. Most current coding tools read a conventions file at the root of a repository. That makes the file a natural place to record the patterns the team wants and the ones it retired. Short architecture decision records and current runbooks help too if the tools can reach them, and new hires benefit from the same documents.

Upkeep is where this usually fails. A conventions file written in one burst by one motivated engineer tends to go stale within a quarter, faster if that engineer changes teams. Giving the file an owner, the same way a service has an owner, is the only fix I know of that holds.

With review and context in decent shape, agents become worth the trouble. Early candidates are jobs a test suite can check and a revert can undo, like dependency upgrades, test backfill, mechanical refactors and lint cleanups. Agents can draft changes that touch production data or infrastructure, but for now a person should approve and run those. Shopify's circuit breakers are a good operating model here. Put a hard daily ceiling on what an agent can spend or do, and alert someone as it gets close.

## What the first few months look like

DORA's 2026 ROI report describes adoption as a J-curve, where delivery performance dips before it improves. It gives three reasons. People are learning new workflows, and reviewing AI output adds what the report calls a verification tax. Downstream steps like testing and change approval also have to absorb more code. The report's sample calculator assumes a 15% dip lasting three months. Treat that as a default someone typed into a calculator. The authors themselves call their figures a high-uncertainty estimate, and your dip could run shorter or much longer.

The same sample model has change failure rate rising from 5% to 6% after adoption and counts the rise as a real cost. DORA's prescription is more automated testing and smaller batches, with no slowdown in adoption. It also argues against using AI as a reason to cut headcount, because retraining people costs less and keeps institutional knowledge in the building.

Here is a made-up example of why the dip is hard to read from the inside. A 20-person team turns on agents for everyone in January. By February the review queue has roughly doubled, and two senior engineers are approving large diffs faster than they would like to admit. The conventions file came from one engineer who then went on parental leave, so half of it describes patterns the team already dropped. Lead time worsens for six weeks. Leadership looks at a cost line that went up and a delivery chart that went down, and asks whether the tools work at all. Maybe the team adds PR limits and a file owner and recovers by April. Maybe the recovery is only partial, which is the more realistic ending. Either way, without a December baseline nobody in that meeting can tell a dip from a hole.

## When the team is already at capacity

Plenty of teams can run all of this with the people they have. For a team stretched across a feature roadmap with customers waiting, review fixes and context documents will keep losing the priority fight. That is a defensible call. Embedded engineers who have done this work before can help there, on one condition. The practices should stay after the outside engineers leave.

If I had to compress the roadmap into one instruction, it would be about order. Stance and baseline first, review second, context and agents last. Teams that start with agents usually get a better first month and a harder quarter after it.

---

If your team has the tools and the delivery numbers have not moved, a 30-minute call is enough to look at where the time is going. Bring the problem. Leave with a plan.

[**Book a 30-min scoping call**](https://cal.com/sociilabs/30min?ref=blog.sociilabs.com)

*Related reading:* [*Should Startups Vibe Code in Production?*](https://sociilabs.com/blog/vibe-coding-production?ref=blog.sociilabs.com)