The instinct after an incident is to add an alert. The higher-value move is usually deleting three that have never led to action — and here’s how to make that argument without losing it.
The incident looks like an application bug for the first hour. Why data failures propagate silently, why they cost more than outages, and what treating data quality as an availability concern changes.
Every label you add to a metric multiplies its cost. The failure mode is gradual, invisible, and ends in a budget review — here’s how to decide what not to measure.