Replace code reviews with verified intentAviator Verify
Aviator
All episodes
Code reviews October 8, 2026

Why Teams Only Pretend to Review AI Code with Dr. Michaela Greiler

Dr. Michaela Greiler, software engineering researcher and consultant, talks about code review surrender, her SCOPE model for matching oversight to risk and alignment, and why an engineer admitting they ship code they don't understand is a green flag.

About the guest

Dr. Michaela Greiler

Dr. Michaela Greiler

Researcher and consultant, Codalytics Consultancy

Dr. Michaela Greiler is a software engineering researcher, consultant, and educator. Over 15 years, she has combined hands-on advisory with original academic research. Her background includes roles at Microsoft and Microsoft Research, consulting for 125+ organizations (including BMW, Wikimedia, and National Instruments), and a PhD in Software Engineering from Delft University of Technology.

Same Principles, More Pressure

There was definitely a point, more than a year ago, when I was wondering, is code review holding up to all of that? Is everything that I know about code reviews still valuable, or do I have to rethink everything completely? I really tried to reimagine code reviews from scratch, and the funny thing is that I'm reaching a place where I'm not far off from the values and the principles that I had before.


Things are changing, definitely, and I think code review also has to change. But I've been working with teams on code reviews, and the bottleneck that code reviews can become, for 15 years now. It's not a new problem.

It's not that code reviews were such a pleasant or easy practice before and now we are struggling. It's a very valuable practice, but it's a hard practice. You have social challenges, technical challenges, organizational challenges, and all of that existed before large language models.

But now things are getting more serious. There is definitely more pressure on code reviews. We see it in the data: larger PRs, more PRs. Teams that don't have their ducks in a row definitely have more problems than they had before.

Start With Why You Review

One of the principles that I taught in workshops five years ago was really thinking about why we are doing this, and what the outcome is. When a team was coming and struggling, one of the first things that I clarified with them was, "What is your goal, actually, in doing code reviews?

Code review is not one thing. There are different techniques. You can review in person, you can have a meeting, you can have asynchronous code reviews. It depends on what part of the codebase you're looking at. It's not one thing, and it has never been.

Review Practices Are More Fragmented Than Ever

When I went back to the drawing board, I also went back to research. I'm interviewing a lot of developers, working with teams, and really studying the problems they are facing and the techniques they are using, to make an informed, empirical decision on where this should go.


And I definitely see people have an even more fragmented review practice. With AI adoption, not everybody is on the same page, to put it lightly. Even within the same team, the same organization, AI adoption is very different. Not just in maturity or how much, but in the techniques people use. One person has a one-shot approach: I'm planning in planning mode, then I have one shot, and this is what I'm submitting. Another person is more iterative, going over it and changing here and there. So there are very different approaches to programming now, or to code generation, and the same for code reviews.

"Code Review Is a Business Decision"

It's coming back to team culture. If the team has good values on engineering practices, they somehow find their way together: what are we doing about AI slop, and how do we value code reviews? But there's definitely pressure from outside too.

I have a couple of engineers who tell me, "Don't ask me about code reviews, because it's a business decision," which I found a very interesting take. It says: as an engineer, I would like to review the code. But in the end, it's a business decision.

If you don't want to spend money or time on me doing code reviews, then you have to live with the trade-offs: quality that's going down, understanding that's going down.

When you put things in place that people agree on, everybody feels a little bit more secure. I think trust and psychological security are very important parts of this conversation.

If people trust each other in a team, they can say, "I don't understand this." This is super important. Then you can experiment, then you can also push large language models to the limit, if you're in a team where you can raise your hand and say, "We tried this, I don't understand, this is too much for me. Let's go one step back."

I would probably start with that: really trying to find out where people are daring to speak their mind and where they are not.

Culture Decides How Hard This Will Be

This is likely the oldest story of all. It has really nothing to do with large language models or AI. Even before AI, when I was going into a company, giving a workshop, or consulting, I could tell very much from the beginning how easy or how hard this would be, just based on the culture in the company, its structures, and even who hired me, and who wants me to be there.

The best experiences are when engineers really reach out and want me to be part of their team and help the team understand, consolidate, and share. When that happens, you can move a lot. People come together, you get them talking, and maybe they see each other's sides. Not saying that everything goes very smoothly all the time, but there's definitely common ground, and people start to see that it's often not a right-or-wrong thing. It's just different perspectives.

It's seldom that someone says, "I actually want this company to go down," or "I want this codebase to be full of technical debt." This is not how it normally works. It's more that maybe you have too much time pressure to actually read the code. Or I'm getting so frustrated by all this AI slop that's sent my way that I say, "Well, I'm also not reading that anymore."

When we have bad behaviors or strange behaviors, there might be reasons for it. If we understand and can talk to each other, then we can also help go in the right direction.


When the culture is different, with less ability to speak our minds, you can quickly tell these are the leaders, the technical lead, or the managers, and they very openly mention their opinions and how things have to go, and others are more quiet or not talking. It's very difficult to work with them. It's not impossible, but it takes much more time and different approaches. And I see the same for AI.

Same Problems, Just Bigger

Before AI, the problems we had in code reviews were slow turnaround times and people not giving good feedback. With AI, I would say it's the same problems, but they are much heavier. And it's not only that we have more problems; we actually have more tools as well. You can use AI to review your code. We have more PRs, but we also have good tools. Large language models are great for reviewing.

I wouldn't go and say, "They review, and I'm not reviewing," because this brings other problems: for understanding, for building a mental model of your codebase, and for alignment. You're missing out on all of that. But they can really get rid of a lot of the problems that maybe don't need human attention.

We had good static analysis: linters, and style checkers. And a lot of teams wouldn't have that and would have people look for those things. And I'm like, why would you do that?

Spend your human attention where it's really needed. Now we have different tools, but the same problems, just bigger.

Code Review Surrender

One phenomenon I noticed when working with teams and doing research is code review surrender, which is giving up on code reviews. But it's not a deliberate thing. It's not, "We have this great strategy in place, so we don't need code reviews anymore." It's more, "This is not working." So it's a window-dressing activity.

We just pretend to do code reviews. I'm wading through this large PR, and I try to understand what's going on and be smart, but the fact is that I don't really know what's going on.

It could be silent, so people are not talking about it. It's just one "looks good to me" after the other, and maybe I find one thing I can nitpick on so that I seem to be participating. Or it could be that teams really say, "We can't do it anymore," and they circumvent it and do something else, but without a strategy in place.

Code reviews are possibly a little bit overloaded as well. We have bug finding, accident prevention, mentoring and learning, tracing and tracking, knowledge sharing. We want all of that out of code reviews, which I thought was crazy even before. You want all of that, all the time, at the same time, which is almost impossible; you would have to spend so much time doing it. So let's be a little bit more choosy here. Which is hard. But life is hard.

Can We Stop Doing Code Reviews?

I'm investigating that heavily right now. My take is that if we let go of code reviews completely, we would miss out on a lot. And it's not possible. We would have to imagine a world where everything is done by agents. But that also means requirements engineering done by agents, the idea of the product, the creation done by agents. Because where do you stop?

What I see now is that we are shifting it left and right.

SCOPE: Staged Code Oversight with Proportional Escalation

SCOPE is the new code review model that I'm proposing. It stands for staged code oversight with proportional escalation. Sounds very fancy. What it means is that instead of code review, I'm thinking more that we have code oversight, or change oversight, and that we have it at different stages.

A lot of the stages will now be before implementing, which is weird. It's not code review anymore. We start looking at something before it's even created in the coding phase. Planning, for example, would be a good place to start thinking about the code and deciding about the oversight it should have. The proportional escalation part means I don't think each code change needs the same oversight. Some code changes are higher risk; they need more oversight. Some are lower risk; they need less.

Risk is not the only thing we should use to decide how much review we do. Alignment as well. Understanding. How much understanding does this need? How much alignment does this need?

It could be low risk because it's not impacting a lot of users, but for our architecture it has a lot of impact. So we want engineers, good people, to be there and align on this decision, to understand that this happens, and to be okay with it.

The code review itself, I think, will get less. And I think that's okay. If I looked at the plan and reviewed the plan, and I know that there is enough detail to it, and then I have the output of the agent and have AI agents review that, and maybe I look at the output of those, then not every code change will need me to go through each line of code. Some simple things I think we can maybe just hand over.

Steering an Agent Is Already a Review

Some people really have this factory model: there's a ticket, nobody looks at the ticket, an agent implements it, another agent reviews it, and out comes something that goes directly to production. Where are we there? We're very much at the end. That could be very expensive as well, when you decide you have to roll that back, even with AI. I wouldn't go for that one.

Adopting changes from an agent is still very different from steering an agent. When you're steering, there's more control. There's still a developer who maybe looked through the plan very carefully, then gave the plan to a large language model of their choice, and then looked a little bit through the code. Even there, you don't have to look at every line of code.

This back and forth between the large language model and the developer who steers it, that's part of a review. It's different from writing code for three weeks, because this is done in half an hour. So you're actually reviewing the code. And I think this should be honored.

If you're bringing in a peer, they probably don't have that context. Something that used to be created over a week is now created in half an hour, but it's still very difficult to comprehend. So they shouldn't replicate what the developer did. They should bring a different challenge, an independent challenge, their own background.

For some changes, we don't need another developer. It's enough if one developer intimately knows what's going on, depending on risk and alignment. For other changes, it would be very valuable, and it would be a bad choice not to have that. Not saying that you can't do it, but there are always costs to it. Do you want to swallow them or not?

Where Teams Should Start

Code review is quite different from other engineering practices like writing code, or now generating code, or even writing tests. You can do those on your own. And partly, and this is really new with large language models, you can now do code reviews on your own too. You just ask your favorite large language model, "What do you think about security in this area?" or "Did I make that modular or not?" and you actually get feedback.

But what differentiates code review from many other practices is that it's a team practice. It's a lot about sharing, being open, being transparent. Team leads and engineering managers have a crucial role: to be vulnerable themselves, to model good behavior, to model that they sometimes don't know.

I've interviewed so many senior engineers, principals, super smart people, and they don't know. They don't have a crystal ball. We don't know whether code will be important three months from now, three years from now, whether we will really read it. Whoever tells you they know, they're lying.

So be very open and transparent that we don't know right now what the right thing is, and invite others to share.

If you don't feel secure enough to say, "I have so much pressure, I'm shipping code that I don't understand," that's a problem. I think that's a red flag. It happens to a lot of people. The question is, are they saying it? And how do you react to it?

If somebody tells me that, it's not a red flag; it's actually a green flag. They dare to say, "Hold on, we are shipping at a pace that I don't understand, and I don't feel good about what we're doing here." I think this is a very, very healthy sign for a team.