A regional leader at a multi-location retail brand recently found out, almost by accident, that one of their stores had been effectively closed for the better part of a week.
Not from a monitoring alert. Not from a support ticket summary. From a conversation about lost sales.
By the time it reached someone with the authority to act, the location had been down long enough that the technical fix, which turned out to be fairly routine, was almost beside the point. The real damage had already been done, and it had nothing to do with the network itself.
Stories like this play out more often than most multi-site operators would like to admit. And the details vary, a different vendor, a different region, a different root cause, but the pattern is remarkably consistent. Which is exactly why it’s worth examining.
The technology usually isn’t the failure
In most versions of this story, the underlying technical issue is small. An ISP-side connectivity problem. A misconfigured failover. A device that needs a reboot no one is physically present to give it. These things happen constantly across a distributed footprint, and in isolation, none of them are alarming. Multi-site infrastructure is built with the expectation that individual components will occasionally fail.
What’s alarming is how long it can take for a small, fixable problem to reach someone who can actually fix it, and how little anyone above the initial point of contact knows about it while that clock is running.
That’s not a technology failure. It’s a process one.
Anatomy of a missed escalation
Here’s roughly how it tends to go. A report comes in, a store manager, a regional lead, sometimes a customer-facing employee noticing the registers are down. It gets logged. It gets assigned to someone. That person starts working it, maybe loops in a vendor, maybe waits on a callback.
And then, without anyone deciding this should happen, it just… sits. Not because anyone was careless. Usually because there was no clear rule for when a stalled issue needs to be pushed up a level rather than continue riding with whoever picked it up first.
Most multi-site IT organizations have monitoring. Fewer have an actual escalation SLA which is a defined answer to “if this isn’t resolved or even meaningfully updated within X hours, who gets notified automatically, and how loudly?” Without that second layer, monitoring only tells you something is wrong. It doesn’t guarantee the right person finds out in time to matter.
The real blind spot: “online” isn’t the same as “operational”
There’s a deeper issue underneath the escalation gap, and it’s arguably the more important one.
Most network monitoring is built to answer a narrow question: is the device responding? Is there a heartbeat? That’s useful, but it’s not the same question the business actually cares about, which is: is this location able to transact right now?
A site can look perfectly “up” on a dashboard — the router’s pinging, the switch shows green — while the point-of-sale system, the payment processor, or some other dependency downstream is completely non-functional. From an infrastructure standpoint, nothing is on fire. From a business standpoint, the doors might as well be locked.
Closing that gap means pairing device-level monitoring with business-impact signals: transaction volume, POS heartbeat, payment processing status, something tied to whether the location is actually doing what it exists to do, not just whether its hardware is reachable.

A single-location business rarely has this problem, because someone is standing in the building. If the registers go down, a person notices within minutes, because it’s happening in front of them.
At 20 locations, that direct feedback loop starts to strain. At 200 or 2,000, it’s gone entirely. Awareness has to be manufactured through process, because it’s no longer going to happen organically. This is one of the underappreciated costs of scale: the things a single site handles through simple human proximity have to be deliberately engineered everywhere else.
That’s not a criticism of any particular team. It’s just what distributed operations require, and it’s easy to under-invest in until a gap like this actually costs something.
What a better version of this looks like
A few things consistently make the difference for multi-site operators who’ve closed this gap:
-
-
- Defined escalation SLAs. If an issue hasn’t been meaningfully updated within a set window, it escalates automatically, not because someone remembered to flag it, but because the system requires it.
- Business-impact alerting, not just uptime alerting. Tie monitoring to transaction data or POS status, not only device connectivity.
- A named backup path. Escalation shouldn’t depend on one inbox or one person being available. If the primary contact doesn’t respond, there needs to be a predetermined next step that fires without anyone having to think of it in the moment.
- A regular tabletop exercise. Periodically ask, deliberately, “if a site went completely dark right now, how would we find out, and how fast?” It’s a cheap exercise that tends to surface exactly these kinds of gaps before they cost anything.
-
The real lesson
Go back to that regional leader, finding out almost by accident that a location had been down for days. The lesson isn’t that outages happen, because they will, at any meaningful scale, regularly and unavoidably.
The lesson is that the cost of an outage is determined far less by the outage itself than by how quickly the right people know about it. Technology will fail sometimes. Whether that failure costs you an afternoon or a week is almost entirely a question of process, and it’s one worth answering before the next outage, not during it.
