Written by Gowtham Raj, Director at TartLabs, who helps clients size custom software and dedicated engineering teams for products that have to keep running after launch.
The Short Answer
In our view, possibly a few fewer for the same scope, but not dramatically fewer, and not zero. AI makes producing a first draft of code cheap. It does not make deciding what to build, checking that it is right, and keeping it running any cheaper. Those jobs still belong to people, and they scale with the amount of software you own, not with the amount of typing.
So the useful question is not "how many engineers does AI replace?" It is "where does my team's time go, and which of those hours did AI actually remove?" Our reading of the evidence below is that AI removed some of the writing, left most of the verifying in place, and did not touch accountability.
Key Takeaways
- 84% of developers use or plan to use AI tools (Stack Overflow 2025 survey, 49,000+ respondents), yet only 3.1% highly trust the accuracy of what they produce and 45.7% actively distrust it.
- METR's first controlled trial found a 19% slowdown for experienced developers in early 2025. Its February 2026 follow-up could not produce a reliable new number and says its own data is "only very weak evidence."
- DORA 2025 found AI adoption raises both delivery throughput and delivery instability. More code arrives faster, and delivery becomes less stable.
- BLS now projects 10% growth for software developers, QA analysts and testers from 2025 to 2035, against 3% for all occupations, with a May 2025 median wage of $135,980 for software developers specifically.
- Team size follows what you own and how risky it is, not how fast code is typed. Review capacity is the constraint to size first.
Each figure above is sourced in the sections below. Our own planning heuristics are labelled as such.
What the Evidence Actually Says About Speed
Most of the headcount debate rests on one assumption: that AI makes each engineer several times faster. The controlled evidence is thinner than the headlines.
METR ran a randomized controlled trial in early 2025 with 16 experienced open-source developers working on 246 real issues, mostly using Cursor Pro with Claude 3.5 and 3.7 Sonnet. According to METR, developers took 19% longer with AI than without it, even though they expected a speed-up. That page now carries a note that the results are out of date.
The February 2026 follow-up used late-2025 tools. It estimated an 18% slowdown for returning developers (confidence interval -38% to +9%) and a 4% slowdown for newly recruited ones (-15% to +9%). METR then explains why it distrusts those numbers: many developers would not work without AI even at $50 an hour, and 30 to 50% avoided submitting tasks they expected AI to speed up dramatically. Its own conclusion is that the data is "only very weak evidence" of the size of the increase, and it reads the result as a lower bound.
Read together, these two studies do not prove AI is slow. They show something more useful for planning: nobody has a trustworthy multiplier yet. Any headcount plan that assumes 2x or 3x per engineer is built on a number no controlled study supports.
Where the Time Goes Instead
Writing code was never the only job. Our view, from delivery work, is that for most product teams it is a minority of the week, with the rest going to requirements, review, testing, deployment, incidents and meetings. When one slice gets cheaper, the others dominate.
The DORA research programme reports that 90% of technology professionals use AI at work and more than 80% believe it has raised their productivity. It also reports that 30% of developers have little to no trust in AI-generated code. Its central finding is that higher AI adoption is associated with increases in both delivery throughput and delivery instability. DORA's analysis of engineers' comments adds that time saved generating code is often re-allocated to verification and prompting.
The Stack Overflow 2025 survey, with more than 49,000 respondents, shows the same pattern from the developer side. 84% use or plan to use AI tools and 51% of professional developers use them daily. But only 3.1% highly trust the accuracy of the output, 45.7% actively distrust it, and 66% name "almost right, but not quite" as their biggest frustration. That last number is the headcount story in one phrase. Output that is nearly correct is likely the costliest kind to review.
Why Review, Not Writing, Sets Your Headcount
If code generation gets cheaper and nothing else changes, three things grow.
- Volume. More changes per week means more diffs to read, more tests to run and more releases to watch.
- Review load. Someone with enough context to judge a change has to read it. In our experience that person is usually your most senior engineer, and there are few of them.
- Surface area. Software that is cheap to create is also cheap to over-create. Every extra service, script and integration is something a person is eventually on call for.
This is why we think the right first number to size is review and ownership capacity, not raw output. A team of four that ships twice the code it used to may need the same four people, plus clearer rules about what gets merged. Our guide to setting AI coding standards covers how to write those rules down.
The demand side has not collapsed either. The US Bureau of Labor Statistics projects employment of software developers, QA analysts and testers to grow 10% from 2025 to 2035, against 3% for all occupations, adding about 185,400 jobs. Median pay for software developers was $135,980 in May 2025. BLS lists the expansion of AI as one of the growth drivers. Labour-market projections are not a guarantee for your company, but they are not what a profession on its way out looks like. The 10% and 185,400 figures cover software developers, QA analysts and testers together.
A Practical Way to Size Your Team
These are our planning heuristics, not published benchmarks. Use them to structure a conversation, then replace them with your own measurements.
Step 1: Count what you own. List every product, service and integration that must stay running. Headcount follows this list far more closely than it follows feature velocity.
Step 2: Find your review ceiling. How many changes can your reviewers read carefully in a week? If AI doubles the change count and review capacity stays flat, you have not doubled output. You have built a queue.
Step 3: Rate the risk of each area. Payments, health data and anything regulated need a senior human on every merge. An internal dashboard does not. Spend senior time where mistakes are costly.
Step 4: Pilot before you cut. Pick one team, give them the tools and track cycle time, change failure rate and review wait time for a full quarter. DORA's instability finding is the reason to watch failure rate, not just throughput. If it holds up, then adjust.
Step 5: Re-balance the mix, not just the count. Teams that lean on AI tend to need relatively more people who can specify, review and test, and relatively fewer doing routine implementation. In our experience that points toward a smaller group of strong engineers and a flexible layer around them. A dedicated team can supply that layer without a permanent hiring commitment, and our comparison of dedicated developers versus an in-house team explains the trade-offs.
In short:
| Step | Question to answer | Signal to track |
|---|---|---|
| 1. Count what you own | How many products and services must stay running? | Systems per engineer |
| 2. Find your review ceiling | How many changes can reviewers read carefully per week? | Review wait time |
| 3. Rate the risk | Which areas need a senior human on every merge? | Change failure rate |
| 4. Pilot before you cut | Does the team ship faster without more incidents? | Cycle time and failure rate |
| 5. Re-balance the mix | Do you have enough people who can specify, review and test? | Senior-to-junior ratio |
For the broader question of where AI sits in your product and engineering strategy, our AI strategy guide for software companies lays out how these decisions connect.
What Not to Do
- Do not cut headcount on a vendor's speed-up claim. METR's own 2026 update shows how hard a reliable multiplier is to measure, and your codebase is not theirs.
- Do not measure lines of code or pull request counts. DORA's finding that instability rises with throughput means volume can improve while quality falls.
- Do not remove your senior reviewers first. They are the people turning "almost right" into "correct."
- Do not freeze hiring for juniors without a plan. Today's juniors are tomorrow's reviewers. A pipeline that stops producing them creates a seniority gap a few years out. This is our judgement rather than a measured result.
The Bottom Line
AI has moved the cost of writing code, but the cost of owning software has stayed where it was. The controlled evidence on speed is weak in both directions: METR measured a slowdown, then said selection effects left its follow-up as only very weak evidence. DORA found that more throughput came with more instability.
So plan around review capacity and risk, and treat any claim of a fixed multiplier with caution. Pilot on one team, measure failure rate as well as speed, and adjust the mix of people before you adjust the number. If you want a second opinion on how to size a team for your own product, or a dedicated team to run the pilot, get in touch. We will tell you plainly if the right answer is to change nothing yet.




