Notes
Short pieces on application analytics for data system reliability metrics — written for operators who prefer clarity over chart theatre.
When uptime percentages lie about reliability
A 99.9% figure can hide long tail pain if your application analytics never see the journeys that actually stalled.
Error budgets need clean events before they need dashboards
Burn rates collapse when duplicate fires and missing properties pollute the reliability stream.
Measure recovery, not only failure
Reliability metrics that ignore time-to-restore leave operators flying blind after the first alert.
Hong Kong ops windows and metric cadence
Regional traffic peaks change which reliability windows matter — and how often packs should refresh.