Walk around any Cloud event and you’ll hear no shortage of AI claims. Every platform is becoming “AI-powered,” every roadmap includes autonomous operations, and every vendor seems convinced the future involves fewer humans watching dashboards.
Through Large Language Models and Machine Learning, AIOps (AI for IT Operations) has already earned a place in Cloud operations, but not for the reasons many people expected. The biggest gains aren’t coming from replacing engineers; they’re coming from removing the low-value work that keeps engineers from doing what they’re actually paid for: making good decisions.
That distinction is important because Cloud operations have always depended on two very different skills: processing huge amounts of information, and accepting responsibility when something goes wrong. LLMs and agentic AI are becoming exceptionally good at the first, while the second still belongs to people. The organizations getting the best return from AI are the ones that understand exactly where that line sits.
Let’s look at three AI uses where you may want to say ‘Go, go go!’, and three where it’s more likely to be ‘No, no no!”
Three Places AI and LLMs Deliver Real Value
Making observability useful again
Modern infrastructure generates more telemetry than any human can realistically consume. Logs, traces, metrics, and alerts all compete for attention, while monitoring platforms insist that everything is urgent.
AI thrives in exactly this environment. Rather than simply surfacing anomalies, they can correlate events across systems, explain why today’s behavior differs from last week’s, and present engineers with a coherent narrative instead of forcing them to assemble one themselves.
The benefit isn’t that AI spots issues nobody else could find; it’s that engineers spend less time searching for context and more time solving problems. In Cloud operations, reducing investigation time is often more valuable than surfacing one extra alert.
Capturing knowledge before it disappears
Every incident teaches something, but the lesson often disappears into Slack threads, war-room chats, and half-finished postmortems before anyone turns it into something reusable. Every operations team promises to update the runbook afterward; and then almost every operations team gets pulled onto the next customer migration, maintenance window, or support escalation before they do.
LLMs are remarkably effective at closing that gap. They can generate incident timelines while events are still unfolding, draft postmortems immediately after recovery, and update operational documentation before hard-won knowledge fades. Nobody loves writing documentation, but everyone benefits when the next on-call engineer doesn’t have to rediscover the same solution at two o’clock in the morning.
Turning technical detail into customer confidence
Outages don’t just create technical problems; they create communication problems. Infrastructure engineers, support teams, executives, and customers all need different versions of the same story, yet producing those updates traditionally pulls senior technical people away from resolving the incident itself.
This is one of the least glamorous but most valuable uses of LLMs. They can translate complex infrastructure failures into clear, audience-specific communications without sacrificing technical accuracy. For service providers, that means engineers stay focused on restoring services while customers receive faster, more consistent updates. Their work will need checking, but that takes a fraction of the time it takes to create it from scratch.
These three examples demonstrate how AI excels when it helps people understand systems faster, not when it makes decisions on their behalf.
Three Places AI and LLMs Don’t Add Value
Accepting production risk
Modern Cloud platforms already automate deployments, scaling, and rollbacks, but they do so within carefully engineered guardrails. An LLM or AI agent can recommend delaying a deployment because customer traffic is unusually high, or suggest rolling back a problematic release. What it shouldn’t do is decide, independently, that production is the right place to experiment.
The issue here isn’t capability; it’s accountability. Production changes affect customers, revenue, and reputation, and those decisions deserve deterministic policies and human judgment, not a probability model that’s very good at sounding confident.
Chaining together autonomous remediation
Agentic AI has made autonomous operations one of this year’s hottest topics, and rightly so: the ability for AI to invoke tools and execute workflows on its own is genuinely exciting. It’s also where things get dangerous.
Restarting an unhealthy service or replacing a failed instance is one thing. Allowing an AI agent to string together multiple operational actions based on an incorrect assumption is something else entirely, since small mistakes can quickly become large outages when systems move faster than humans can intervene.
As AI agents become more capable, operational guardrails become more important, not less. The safest infrastructure won’t be the one with the most automation; it will be the one with the clearest limits on what that automation is allowed to do./art
Owning security and compliance
LLMs are excellent assistants during security reviews. They can identify suspicious configurations, compare policies against best practice, and help draft compliance documentation. What they cannot do is sign off on risk.
Security, governance, and compliance ultimately require someone who can stand behind a decision when an auditor, regulator, or customer asks difficult questions, and probability models don’t attend board meetings or regulatory hearings. LLMs have become very good at sounding convincing; infrastructure teams should remember that convincing and correct are not the same thing.
The Real Opportunity
Every technology generation promises to replace operators, and so far, each has instead changed what operators spend their time on. AI looks set to follow the same path.
The biggest productivity gains won’t come from removing humans from Cloud operations, but rather from removing repetitive work that keeps experienced engineers from applying their expertise where it matters most: every hour not spent searching through dashboards, reconstructing incident timelines, or rewriting customer updates is an hour spent improving resilience and preventing tomorrow’s outage.
The question for Cloud leaders is no longer whether AI belongs in IT operations, because it’s already here. Instead, it’s all about whether you’re using it to explain what your infrastructure is doing or letting it place bets that your business will have to pay for.
Sign up for the CloudFest newsletter (below) for industry insights and free event registration.
