Observability that earns its keep
Every team I have joined had monitoring. Fewer had alerting anyone trusted. The gap between those two things is where incidents live.
Scope
What I take on
Instrumentation
Datadog, New Relic, Prometheus and Grafana, plus Splunk and ELK on the logging side. Getting the signals out of the system in the first place.
Custom checks
The checks that matter are usually specific to your application, not the ones that ship in the box. Writing those is most of the value.
Alert hygiene
Cutting the alerts nobody acts on. An on-call rotation that has learned to ignore its pager is worse than no pager.
Dashboards people actually open
Fewer panels, chosen to answer the questions asked at three in the morning.
Method
How it runs
-
01
Start from the incidents you have had
Your last six incidents tell you what to instrument far better than a vendor checklist does.
-
02
Alert on symptoms, not causes
Page when users are affected or about to be. Everything else is a dashboard or a ticket, not a phone call at 03:00.
-
03
Every alert needs an action
If there is nothing to do about it, it is not an alert. That one rule removes most of the noise in a typical estate.
Evidence
Where this has been done before
- LeasePlan Designed a Datadog alerting layer with custom checks that surfaced performance bottlenecks proactively, on infrastructure where availability was contractual.
- 3D Hubs Prometheus and the ELK stack across rebuilt AWS environments.
- Knab New Relic across a regulated banking platform, and Splunk in the earlier engagement.
Fit
Probably worth a conversation if
- You find out about outages from customers.
- Your on-call rotation has quietly stopped reading the alerts.
- You are paying for an observability platform and using ten percent of it.
- Something is slow, intermittently, and nobody can prove where.
Contact
Tell me what is actually broken
The contact details live on the front page. If I am not the right person for this, I will say so.