Your Monitoring Setup is Waking You Up at 3 AM for Nothing
Why your pager goes off for a 2ms latency spike, how to fix it, and why you’ll still be afraid to sleep through an alert.
3:14 AM. Your phone lights up. The pager is screaming. You stumble out of bed, heart racing, already imagining the database corruption, the security breach, the angry customers.
You open the alert: “Average response time exceeded threshold (200ms) on web-server-3 for 2 minutes.” Current value: 202ms. Threshold: 200ms.
Two milliseconds. Someone coughed near the server and you’re standing in your kitchen at 3 AM in your underwear staring at a graph.
This is the story of every over-alerted DevOps team. And it’s killing you — figuratively, and eventually literally.
The Alert Noise Problem
Most monitoring setups alert on symptoms, not causes. Latency spikes 2ms? Alert. CPU hits 80% for 30 seconds? Alert. A single error in a log? Alert.
The problem is that these are normal operating variations. Servers have bad moments. Networks hiccup. Databases breathe. If you alert on every micro-blip, you train your team to ignore alerts. And when a real problem happens at 3 AM — the database is actually corrupt, the payment system is actually down — nobody responds because they assume it’s another false alarm.
This is called “alert fatigue,” and it’s more dangerous than having no monitoring at all.
How to Fix It
1. Alert on urgency, not severity. Severity is “how bad is this?” Urgency is “how fast do I need to wake up?” A non-critical service being down at 3 AM is not urgent. A database running low on disk space at 3 AM is urgent because by 7 AM it will be full. Alert on urgency.
2. Use proper thresholds with hysteresis. Alert when CPU is over 90% for 5 minutes, not 30 seconds. Use “alert at 90% for 5 min, resolve at 80% for 5 min” — the gap between alert and resolve prevents flapping.
3. Implement maintenance windows. If you deploy every Tuesday at 2 PM and the app is slow for 3 minutes afterward, that’s expected. Silence alerts during known disruptive events.
4. Create a tiered escalation policy. Tier 1: automated self-healing (restart the container, scale up, clear the cache). Alert only if self-healing fails. Tier 2: page one person during business hours only. Tier 3: wake someone up at 3 AM. Most alerts should be Tier 1 or 2.
5. Review and tune your alerts monthly. Go through every alert that fired. Ask: “Did this require human action?” If no, tune or delete it. Do this for three months and your alert volume drops by 80%.
The Real Test
You’ll know your monitoring is healthy when you can sleep through an alert because you trust that it’s not worth waking up for. And you’ll also know you’re a real DevOps engineer when you’re still anxious about sleeping through an alert even though you know it’s fine.
The anxiety never goes away. But at least the 3 AM false alarms can. Fix your thresholds. Your sleep schedule will thank you.
Stop waking up for nothing: Our Cloud Cost Optimization and monitoring courses help you build systems that alert on real problems, not 2ms latency spikes.