AWS Cloud Operations Blog
Category: Management & Governance
Analyze Application Load Balancer Logs with Amazon CloudWatch Logs
Summary Amazon CloudWatch Logs now supports Application Load Balancer (ALB) logs as vended logs, giving you out-of-the-box visibility into the health and performance of your ALB. All three ALB log types (access, connection, and health check) are delivered as structured JSON with named fields. This allows teams to attribute 5xx errors to the load balancer […]
Use AWS DevOps Agent to triage and route AWS Health event impact
Triaging the impact of AWS Health events is one of the most repetitive jobs in cloud operations, and it is exactly the kind of work AWS DevOps Agent can take on. Scheduled maintenance, operational issues, and Trust & Safety notifications (alerts about resources that may violate the AWS Acceptable Use Policy) land in your inbox […]
Multi-Cloud Observability with Amazon CloudWatch Using Bearer Token Auth and OpenTelemetry
Organizations running serverless workloads across multiple cloud providers face a specific observability challenge. There is no persistent compute to host a telemetry collector, no sidecar to attach, and no daemon running between invocations. The standard OpenTelemetry deployment model (application to local collector to a telemetry backend) does not apply in this environment. Authentication presents an […]
Use CloudWatch syslog and Log Alarms to give AWS DevOps Agent on-premises visibility
Your on-premises firewalls, routers, and switches emit syslog that record device events such as denied connections, tunnel state changes, and routing changes. Network devices send their logs over syslog rather than the Amazon CloudWatch Logs API, so bringing that data into AWS takes extra components. A common approach has been to run a collection tier […]
Getting per-resource alarm notifications with Amazon CloudWatch
Introduction As organizations scale their AWS footprint across multiple accounts and Regions, operations teams increasingly rely on Amazon CloudWatch Metrics Insights alarms with GROUP BY to monitor entire fleets from a single alarm. When this single alarm tracks hundreds of resources, a critical operational question arises: how do you ensure that every individual breach is […]
Autonomous Root Cause Analysis for AWS Systems Manager Patch Failures Using AWS DevOps Agent
Introduction When AWS Systems Manager Patch Manager reports failures across hundreds of managed nodes spanning multiple accounts and regions, operations teams face a time-consuming investigation: logging into each account, correlating events, analyzing logs, and determining whether failures share a common root cause. A single patch cycle failure can consume hours of engineering time. In this […]
Deploy OpenTelemetry Gateway on AWS: Monitoring Your Observability Pipeline
Deploy an OpenTelemetry gateway on Amazon EKS, export metrics to CloudWatch over native OTLP, and monitor the pipeline’s own health with PromQL dashboards and alarms.
Turn Your Amazon CloudWatch Alarms into Actionable Signals
Your alarm fires at 2 AM. You grab your phone, squint at the notification, and see: “ALARM: my-service-alarm has transitioned to ALARM state.” No context. No application. No hint about which of your 200 instances is the problem, or whether it even matters. I’ve been there. We’ve all been there. Alarm frustration often comes from […]
Using Amazon S3 Server Access Logs with Amazon CloudWatch Logs
TL;DR What if you could go from raw Amazon S3 server access logs to a complete security dashboard without building a custom pipeline? The dashboard below is deployed using the CloudFormation template provided in this post. Figure 1: Amazon S3 Server Access Logs Security, Compliance & Audit Dashboard Until now, getting security visibility from Amazon […]
Log analysis with facets, correlation, enrichment, and automation in Amazon CloudWatch Log Analytics
Teams working with distributed applications accumulate logs across multiple log groups, including application logs, access logs, and audit trails. When something needs investigating, an engineer opens the console and starts writing queries from scratch. The same query gets written differently by different people. The results lack context because the log event does not contain who […]









