
At 08:17 on a Monday morning, a 28-person professional services firm could not access its line-of-business application, shared files or internet telephony. Staff could still switch on their computers, but they could not do the work those computers were intended to support. This small business downtime case study examines how a single infrastructure failure affected a working day, why the impact spread so quickly, and what changed afterwards.
The organisation in this representative case study has been anonymised, but the circumstances are typical of many small and mid-sized businesses. It had a busy office, a mixture of cloud and on-premises systems, remote workers, and no dedicated internal IT department. Its technology had evolved gradually as the company grew. It worked well until it did not.
Small business downtime case study: the incident
The business environment
The firm handled time-sensitive client work and relied on a central file server, a hosted customer relationship management platform, Microsoft 365, cloud backups and internet-based phone services. Most day-to-day processes depended on a stable network connection and access to shared folders.
Its main network switch was more than six years old. It had not failed previously, so replacement had been deferred in favour of more visible business expenditure. There was a firewall in place and endpoint protection on staff devices, but the network was not being actively monitored. The business also had a backup solution, though the team had not recently tested how quickly files or systems could be restored.
This was not reckless decision-making. It was a familiar small-business trade-off: equipment that appears to be working is difficult to prioritise against recruitment, premises and client delivery. The problem is that ageing infrastructure rarely gives a convenient warning before it becomes an operational issue.
What happened that morning
A power fluctuation overnight did not cause a full building outage, but it exposed a fault in the core network switch. By the time the first employees arrived, the switch was repeatedly restarting. Wireless access points, desk phones, printers and several wired workstations were disconnected or unstable.
The immediate assumption was that the broadband service had failed. That was understandable, as internet access was the first visible symptom. However, checks showed that the external connection was live. The fault was inside the office network, where a failing device was interrupting traffic between users, servers and cloud services.
The lack of central monitoring meant no alert had been raised overnight. The business first became aware of the issue when staff began reporting that systems were unavailable.
A simplified timeline shows how quickly the disruption expanded:
- 08:17: First staff member reports no network access and phones unavailable.
- 08:35: Office manager confirms the issue affects most users and escalates to IT support.
- 09:20: Remote diagnosis identifies instability at the core switch rather than a broadband fault.
- 11:10: A replacement device is sourced and configuration information is recovered.
- 14:45: Network services are restored, with remaining checks completed later that afternoon.
The physical replacement took less time than the overall incident. Diagnosis, sourcing compatible hardware, rebuilding configuration and validating each dependent service accounted for the delay. This distinction matters. Downtime is not simply the time needed to replace a failed component. It is the time required to understand the fault, make a safe change and restore confidence that core systems are working correctly.
The cost was wider than lost working hours
It would be easy to calculate the cost by multiplying eight hours by 28 employees. That gives a useful starting point, but it does not show the full commercial effect.
Some staff were able to work from mobile devices or personal hotspots. Others could prepare work offline. Client meetings still took place, although employees could not retrieve supporting documents or update records during those meetings. The finance team delayed invoicing, the reception desk had to use mobile phones, and several client calls were missed before call diversion was arranged.
The firm estimated that around 110 productive hours were lost or materially reduced. At an internal cost of £32 per hour, that represented approximately £3,520 before considering replacement hardware, emergency support and management time. The direct figure was uncomfortable, but the secondary effects were more significant.
The secondary impact
A delayed client deliverable did not result in a contractual penalty, but it required senior staff to spend the following day managing expectations and reprioritising work. One prospective client called during the telephone outage and was unable to reach the office. It is impossible to prove whether that contact would have converted, but it exposed a weakness in the firm’s ability to remain reachable.
There was also a confidence cost within the business. Staff were frustrated because they did not know whether to wait, work elsewhere or contact clients. Leaders had limited information during the first hour, making it harder to give clear instructions. In operational incidents, uncertainty can be as disruptive as the technical fault itself.
A case such as this should not be used to claim that every outage has a precise price. It depends on the business model, staffing pattern, contractual obligations and the systems affected. A manufacturing firm may face immediate production losses. A legal practice may have court deadlines. A retailer may be unable to take payments. Yet even an office-based company can experience meaningful exposure when communications, documents and customer records become inaccessible at once.
Why the outage lasted most of the day
The root cause was a failed switch, but the underlying causes were broader. There was no current lifecycle plan for network equipment, no alerting for device health, and no documented recovery priority for essential services. Configuration records existed, but they were not maintained in a single, readily accessible location.
The business had also treated backup and continuity as the same thing. Backups protect data. They do not automatically keep staff productive when the network, identity services or communications platform is unavailable. Recovery arrangements need to consider the order in which systems return, who makes decisions, how staff communicate and what temporary methods are available.
The company did have a support contact, but its arrangement was mainly reactive. That was sufficient for routine password resets and occasional desktop issues. It was less suitable for managing the health of infrastructure that underpinned every part of the operation.
The changes made after the incident
The response was not to buy every available technology product. The firm focused on reducing the chance of a repeat and shortening recovery if another failure occurred. The first priority was replacing the failed switch and reviewing other network equipment approaching end of life.
It then introduced a managed approach to the systems that mattered most. This included monitored network devices, documented configurations, patch and firmware management, and a clear escalation route for incidents. Alerts were configured so that unusual device behaviour could be investigated before staff reported a complete loss of service.
The continuity plan was also rewritten in practical terms. It identified critical systems, recovery targets, key supplier contacts and named decision-makers. The firm agreed how calls would be diverted, how staff would be updated and which employees could work securely from an alternative location if the office network was unavailable.
Four controls made the greatest difference:
- lifecycle planning for network, server and firewall equipment;
- active monitoring with alerts reviewed by technical specialists;
- tested backups and documented recovery procedures; and
- an incident communication plan for staff, clients and suppliers.
There are trade-offs. Redundancy costs money, and not every small business needs duplicate hardware for every device. A business with a short tolerance for disruption may justify a secondary internet connection, spare network equipment or a more advanced failover design. Another may decide that reliable monitoring, rapid support and documented recovery are a proportionate first step. The correct level of investment should follow the cost of interruption, not a generic checklist.
What decision-makers should take from this case
The central lesson is not that all hardware will fail. It is that a single point of failure remains invisible until it stops working. If one device, one connection or one undocumented configuration can prevent staff from serving customers, it deserves attention before an incident forces the issue.
Business leaders should be able to answer a few straightforward questions: Which systems must be available first? How long can the business operate without them? Who is responsible for responding? Is equipment monitored and within its supported lifespan? Have recovery procedures been tested rather than assumed?
A dependable IT partner can make these questions easier to manage by turning technical detail into a planned service: visible risks, clear priorities, maintained infrastructure and support when an incident occurs. The objective is not perfect immunity from failure. It is ensuring that a failure is contained, understood and recoverable before it becomes a lost day of business.