Hello readers! What if you had a tool that could find out the root cause of a problem rather than just monitoring it?
Imagine the situation where there is a sudden drop in performance in the production application, and Prometheus detects an anomaly in the metrics. Before, it was the job of an engineer who needed to be alerted, look at dashboard metrics, analyze several services, compare metrics and logs, and identify the root cause of the issue.
But what if one could connect Prometheus with an autonomous DevOps agent to investigate the issue, retrieve relevant metrics, correlate them, and perform automatic root cause analysis?
The value lies in combining Prometheus with AWS DevOps Agent. Knowing about AI-driven root cause analysis is also essential.
Amazon Managed Service for Prometheus offers scalable monitoring compatible with Prometheus and PromQL for queries of operational metrics. AWS DevOps Agent is capable of analyzing incidents and determining the potential causes of the incident by using connected tools and telemetry. AWS just showcased such a pattern by connecting Amazon Managed Service for Prometheus and its MCP tools along with an alert pipeline.
Here comes the shift from problem detection to automatic investigations.
What does Prometheus Exactly Mean?
Prometheus is an open-source monitoring and alerting solution that is built for collecting and querying time series metrics. It has become extremely popular in container/Kubernetes and cloud-native application monitoring environments.
Whereas traditional monitoring tools scan the applications for any issues, Prometheus collects metrics such as CPU usage, memory usage, number of requests per second, latency, and count of errors on a continual basis.
It can be done with the help of PromQL – the Prometheus Query Language, which allows filtering, aggregating, and analyzing those metrics. The AWS-managed Prometheus service retains Prometheus-like features but without the need for any infrastructure management.
What Is AWS (Amazon Managed Service) for Prometheus?
Amazon Managed Service for Prometheus is a managed Prometheus-based monitoring service provided by AWS. It provides support for containerized workloads such as Amazon EKS and self-managed Kubernetes, and scales metrics collection, storage, and querying according to workloads.
Users can ingest metrics from existing Prometheus environments using the remote write functionality. AWS offers a Prometheus-compatible API for querying metrics and working with monitoring data.
The key difference is important because you don’t have to throw away your existing Prometheus knowledge.
The Definition of AWS DevOps Agent
AWS DevOps Agent is an autonomous agent that troubleshoots issues related to operations and DevOps. The telemetry data collected by the connected devices is analyzed. It also determines the root cause of the problems. The tool also concentrates on why it occurred in the first place.
Component | Primary Role |
AWS for Prometheus | Offers managed Prometheus-compatible monitoring |
Prometheus | Gathers and stores application metrics |
Prometheus MCP | Provides the agent with access to Prometheus signals |
PromQL | Analyzes metrics and queries |
Alert pipeline | Triggers investigation when an anomaly happens |
AWS DevOps Agent | Correlates signals and investigates incidents |
From Detection to Investigation
Now consider the situation where there is a spike in the response time of an application.
With Amazon Managed Service for Prometheus, such an anomaly can be detected, and an event workflow initiated that triggers AWS DevOps Agent. By connecting to Prometheus MCP through its integration, the agent can analyze metrics by running PromQL queries and investigate the anomaly.
The engineers do not need to go through different dashboards to investigate the problem; instead, they can leave the investigation to the agent. It is important for you to know about the model context protocol.
This was shown by AWS using the example of Amazon EKS, Amazon Managed Service for Prometheus, Amazon SNS, AWS Lambda, API Gateway, Amazon Cognito, and a Prometheus MCP server.
Incident Detection and Investigation Workflow of Prometheus and AWS DevOps Agent
It is easy to comprehend the workflow of incident detection and investigation with the help of an incident.
Step 1: Metrics Generated by the Application
In this stage, the application or a Kubernetes workload generates metrics for throughput, latency, errors, and resource consumption.
These metrics are collected and stored in Amazon Managed Service for Prometheus while supporting the Prometheus data model and PromQL.
Step 2: Metrics Reach the Monitoring Workspace
Here, the metrics reach an Amazon Managed Service for Prometheus workspace, where queries can be performed on these metrics with the help of PromQL and visualization through Amazon Managed Grafana.
With the help of visualization, it is possible to have visibility; however, visibility alone does not clarify the root cause of the incident.
Step 3: Anomaly Detection Detects an Anomaly
The use of Random Cut Forest (RCF) in AWS's demo for detecting anomalies instead of threshold-based approaches.
As an example, if a service is processing 1,000 requests per second under normal circumstances, a reduction to 700 will not meet the threshold of 500.
Step 4: The Alert Triggers an Investigation
The alert can travel through Alert Manager, Amazon SNS, and AWS Lambda before reaching an AWS DevOps Agent webhook.
This makes sure that the investigation is automatically triggered and does not need any action from the engineer to pass the alert.
Step 5: The AWS DevOps Agent Queries Prometheus
The AWS DevOps Agent can use a Prometheus MCP server to query monitoring data stored in the Amazon Managed Service for Prometheus workspace. This AWS MCP guide explores this.
It can ask queries using PromQL to determine when and how the signals have changed.
Step 6: Correlation of the Signals by the Agent
Usually, one metric cannot explain the incident.
For example, if the application throughput decreases, the agent can investigate other signals like latency, error rates, CPU, and memory usage. This correlation will allow it to identify the cause of the incident.
Monitoring is the evidence. Investigation is the connection between evidence.
Step 7: Root Cause Analysis Generated by the Agent
After analyzing the telemetry data, the AWS DevOps Agent can create a report with the root cause of the problem and the potential mitigations.
In the September 2026 demo by AWS, the DevOps agent identified the 5G issue within 25 seconds and generated the investigation report in less than 60 seconds, correlating low RSRP, SINR, and CQI levels to low throughput, hence poor coverage or link budget.
The aim is straightforward - to go from “something is wrong” to “we know why.”
Why This Combination is Important?
It Can Lessen Manual Investigations
Modern applications may include many services; thus, manual investigations may take too much time. An automated agent can start checking necessary metrics immediately after receiving an alert.
It Can Lessen Alert Fatigue
Static alerts can trigger lots of noise due to fluctuations in the workload. An anomaly detection system can recognize abnormal activity, while AWS DevOps Agent can add more information about the signals behind the alert.
It Combines Metrics and Logic
Prometheus collects and queries telemetry data, while AWS DevOps Agent analyzes signals relevant to the metric and forms an investigation out of them.
Prometheus provides evidence. The agent provides an explanation for this evidence.
When Should You Opt for This Approach?
The solution is highly appropriate in cases when complex distributed systems utilize metrics from Prometheus, such as those running in Kubernetes.
It can produce numerous time series data metrics, which makes manual analysis difficult and time-consuming. AWS DevOps Agent can be useful for analysis of these metrics in various AWS, multicloud, and on-premises environments.
Nevertheless, an artificial intelligence investigation agent cannot substitute proper observability. Instrumentation and metrics must remain top priorities. Cloud monitoring also helps the cause.
What Should Be Done Before Connecting Them?
Start With Useful Metrics
Collect metrics which will be able to give insight into the behavior of your applications like request rate, latency, errors, resource usage, and valuable information about your application.
Create Useful Labels
Create useful labels using Prometheus for your metrics depending on the service, instance, environment, and other aspects. Correct labels will guarantee that the agent collects only necessary components.
Ensure Appropriate Access
Provide appropriate access but keep the permission under control. AWS suggests setting up Agent Spaces that define appropriate access boundaries for resources.
Develop Investigation Manuals
Establish a unified approach to investigations. For instance, the agent will be able to define the affected service, look through the related metrics, compare the timelines, verify the possible cause, and summarize its analysis.
What are the Limitations?
This is great technology, but it is not a silver bullet!
First, the agent requires good telemetry. Poor telemetry may prevent an investigation.
Second, correlation is not causation. Simply because two metrics change at the same time doesn't necessarily mean one causes the other.
Third, security and permissions must be considered carefully. Allowing broad access may introduce unnecessary operational and security risk.
Fourth, the engineers should review significant investigations. Root cause analysis generated by the agent should have supporting evidence and reasoning.
Putting Things into Perspective
It all comes down to moving from observability to operational reasoning.
Current observability platforms capture metrics, logs, traces, and alerts in complex modern applications. What they don't do is analyze all that information fast enough.
A connection between the agent and the platform makes it possible for the agent to start organizing and investigating the information from the very moment an alert goes off.
Perhaps the future of observability is not about collecting ever more data, but about building systems that interpret data, investigate anomalies, and give engineers a starting point.
Conclusion
In other words, how does Prometheus get along with AWS DevOps Agent?
Prometheus offers telemetry capabilities and query capabilities, which the AWS DevOps Agent can leverage for root cause analysis of operational issues.
With Amazon Managed Service for Prometheus, users can leverage a managed Prometheus-based monitoring solution, supporting PromQL queries and scalable metrics collection. Prometheus MCP integration will enable AWS DevOps Agent to leverage the data collected by this monitoring solution for automated root cause analysis.
The transition itself is straightforward but very powerful.
Rather than just saying that "something is not right," the process could take a different direction – something is wrong, here is the proof, and here is the possible cause of the problem.
For modern cloud-based applications, this approach to handling incidents may be extremely helpful.
However, it is important to treat the agent not as a replacement for proper observability but as an investigation tool.