· Protocolzone
The worst way to learn about an outage is from the customer it hit. By the time their message arrives they have already lost money, already lost patience, and already started wondering whether anyone is watching the platform at all.
On a betting, casino, trading or payments platform the gap between the fault and the complaint is where trust goes. A wallet that stops settling, a game provider that stops answering, an order gateway returning errors: each of these is visible in the platform’s own logs and metrics well before a customer picks up the phone. The information exists. It just reaches the wrong person last.
Why the usual setup does not close the gap
Most platforms have monitoring. The problem is what the monitoring is wired to.
Alerts go to a channel, not to an owner. A busy alert channel scrolls. An error that matters arrives between forty that do not, and nobody is sure whether someone else already picked it up.
One fault becomes a thousand alerts. A broken dependency throws the same exception on every request. If each occurrence pages someone, the team learns to mute the alerts, which is exactly the habit that hides the next real incident.
A flapping connection looks like an outage every few seconds. External links to game providers, price feeds, payment gateways and partner APIs drop and recover constantly. Treat every drop as an incident and the desk spends its night on noise.
Nobody knows which customer is affected. On a platform serving several brands or operators, “the wallet service is failing” is not actionable for support. “The wallet service is failing for this operator” is. Without that link, the desk cannot tell the right customer anything, so it tells nobody.
What we build
On a multi-tenant betting platform we build and operate, we built a monitoring service for our own support desk that addresses each of those points. It reads the platform; it never changes it.
- Watch what the platform is already saying. The service reads application logs, JVM and Spring Boot metrics, Kubernetes pod health, Kafka broker readiness and the health of the datastores, every minute.
- Fingerprint each error. An error is identified by where it came from and what caused it, not by its exact text. The first occurrence opens a ticket. Every repeat adds to a counter on that same ticket instead of raising a new alert.
- Give connections a state, not a reflex. Each external connection moves through up, suspect, down and recovering. It has to stay down for a set time before it counts as down, and stay up for longer before it counts as recovered, so a link that flickers does not page anyone.
- Open the ticket with an owner and a customer. Each ticket carries the affected tenant, or is marked platform wide. Errors that an outage causes further down are folded into that outage’s ticket. It is posted to the support channel with its severity and the evidence, and the channel is reminded on a fixed interval for as long as the error is still live. When the error stops, the ticket says so, and if it comes back the same ticket reopens rather than starting from zero.
- Let the customer see their own tickets. Users from each tenant can see that tenant’s tickets in the support console, and only that tenant’s.
- Use a model for the first read, not for the fix. A language model summarises the evidence on a ticket for the engineer on shift. It has no ability to change the platform. The engineer decides.
The step that makes it customer facing
Detection and ticketing are software. Telling the customer first is a desk discipline, and it is the part we agree with each client before go live:
- Who is told. A named contact per customer, per severity.
- When. The first notice goes out once the engineer has confirmed real customer impact, not on every automated ticket. Being early matters; being wrong in front of a customer costs more than being a few minutes later.
- What it says. What is affected, what is not, and when the next update comes. No root cause guesses in the first message.
- How often. Updates on an agreed interval while the ticket is live, even when the update is “still working on it”.
- How it ends. A resolution note with what happened and what changes, taken from the ticket record rather than reconstructed afterwards.
Automating the first notice is possible once the ticket knows the tenant, and for some fault types it is the right call. We would not start there. A human confirming impact before the customer hears about it keeps the notices trustworthy, and trust is the whole point of sending them.
What it takes on your platform
- Read access to your logs and metrics. The monitoring service needs no write access to production, and should not have it.
- A tenant or customer identifier on your log lines, or a reliable way to derive one. This is the step most platforms are missing, and usually the first piece of engineering work.
- A severity map and a contact matrix, agreed with you before the first incident rather than during it.
- A desk on shift. The 24×7 Operations Desk is the team that reads the tickets, confirms impact and sends the notices. The engineers on it can read the code behind the alert.
Where this has been done
The detection, fingerprinting, automatic ticketing, connection state handling and per tenant ticket visibility described above are built on a multi-tenant betting platform we develop and operate. The customer notice process is how we run the 24×7 Operations Desk for clients; the monitoring service does not send customer messages on its own. We do not publish response time or uptime figures for this work.
Related reading: how a deploy failure is detected from its logs and fixed on one confirmation, and why AI output needs an approval path.
Running a platform where customers find the faults first? Talk to us about the 24×7 Operations Desk, on a platform we build for you or one you already run.
- incident-response
- monitoring
- operations
- multi-tenant