Breaking News

AI Is Making Software Faster to Build – and Harder to Understand

https://ift.tt/Q4S3s2u

As a CEO at a startup, getting back to building has been fun. Last night, I wrote code for our upcoming release. I helped fill a gap so the team didn’t have to take the entire load, and we could hit our deadline. Of course, I used AI tools. They made it easy to execute the idea. But when I opened the file later and traced how it got called, I realized something else: AI had made it easier to write the code than to explain it.

The question came up in our stand-up. The team needed me to explain what I had done so they could interface with what was built. In my role, I explain things for a living, to investors and customers. This was harder. I could tell them what the code did. I traced that in a couple of minutes. What I couldn’t explain was why it looked like that and not some other way. There was a condition in the middle I couldn’t account for. I didn’t know whether it was core to what we were building or an artifact of something we had abandoned.

I wrote it, but nothing in me remembered writing it.

That experience exposed a problem I’ve been thinking about from the other end of the software lifecycle: incident response. We have become remarkably good at knowing when something is wrong. We are not getting better at understanding why.

Thinking differently about investigation

For years, engineering organizations have invested heavily in reducing detection time. Observability tools can surface anomalies quickly. Alerts tell us that latency jumped, an error rate changed or a service is behaving differently than expected. But an alert is only the beginning of an incident. Detection tells you that you have a problem. Investigation tells you what problem you actually have. And increasingly, investigation is where the time goes.

To understand an incident, an engineer may have to move between telemetry, code repositories, deployment histories, tickets and conversations. They need to know what changed, who changed it and, often, why a decision was made in the first place. That context may exist across multiple tools and multiple people, or it may not have been recorded at all. This is why I think we need to start thinking about investigation speed differently.

Investigation speed is a ratio between the amount and velocity of change happening in a system and the context an organization has retained about those changes. Historically, that ratio was kept somewhat in check by the limits of how quickly humans could produce software. The tedious parts of development also created understanding. You mapped out methods. You built the class structure. You ran into problems. You changed your approach. By the time the code went into production, somebody had accumulated a detailed mental model of why it worked the way it did.

AI changes both sides of that equation. It dramatically increases how quickly we can produce software, while reducing the amount of context engineers naturally accumulate while producing it. The same work that once required three people to understand and implement is now completed by one person working with AI. That’s a productivity gain. But when something breaks six months later, the investigation may depend on reconstructing context that nobody ever really absorbed.

Productivity gains vs. investigation cost

We are measuring productivity gain. We aren’t measuring the investigation cost. You can see this problem in a familiar incident response ritual. Something breaks and the first question is whether we have seen it before. One person remembers something similar from last week. Someone else says this looks different. Eventually, somebody calls in the engineer who has been at the company for 10 years.

That person has the context of everything. Your single point of failure that everyone is grateful for has arrived. They have other work to do, but now is their time to lend the company their brain. If they’re on vacation, the investigation takes longer. If they leave the company, much of that accumulated knowledge leaves with them.

AI makes this problem more acute because the people producing changes aren’t necessarily accumulating the same depth of context as they did when they had to work through every step themselves. More changes enter the system, while less context about each individual change may stick with the humans responsible for it. Yet our operational metrics don’t make this particularly visible.

Mean time to resolution (MTTR) tells us how long it takes to get the customer unblocked. Mean time to detect (MTTD) tells us how quickly we know something has gone wrong. Both matter.

But there is an increasingly important interval between them: the time between knowing something is wrong and understanding why.

Today, much of that interval disappears inside MTTR. Once we understand the problem, we fix it and move on. The explanation may live in a Slack thread, a ticket or, frequently, in the head of the person who solved it. Then the next incident starts and the reconstruction begins again.

Multiply that across thousands of changes and hundreds of engineers, and investigation becomes more than an incident response problem. It becomes a measure of how well an engineering organization can keep up with the software it is creating. We’ve spent years optimizing how quickly we know something is wrong. The next challenge is measuring how quickly we can understand why. 

AI: Gains made now, discuss what and why later

Engineering leaders should start pulling that interval apart. How long does it take after an alert fires before the team understands what actually happened? How many people need to become involved before they get there? How much of that time is spent finding evidence or reconstructing past decisions? How often does the investigation depend on locating the one person who remembers why something was built a particular way? Those questions tell us something that MTTR alone cannot: how effectively an organization understands the software it is operating.

This will matter more as AI increases the volume and velocity of software change. Faster development isn’t going away, nor should it. The gains are real. I experienced them myself when AI helped me get that release work done. But so is the tradeoff I encountered the next day. 

I didn’t get to the why, because the fix was what mattered and I needed to move on. What was being built was finally working. I opened Slack and let the team know it was working. The change went in, so changes went up, but the context went down. This gains me time right now but will cost me when I need to discuss what was built. The next item on the list is up; I don’t need to worry about this. Until I have some explaining to do.

The post AI Is Making Software Faster to Build – and Harder to Understand appeared first on SD Times.



Tech Developers

No comments