How Well Is IT Really Running?

A green dashboard can hide a surprising amount of work. 

Service-level agreements may be within target while employees rely on workarounds. Operations teams may restore the same service every week without removing the cause. A key application may technically be available even though the transactions that matter most are slow or unreliable. 

For CIOs and IT operations leaders, availability is only the starting point. The harder questions are whether services behave predictably, whether teams can recover from disruption, whether technology spending is reducing friction, and whether the environment is becoming easier to manage. 

Answering those questions requires a direct look at how services are owned, operated, measured, and improved. 


A Quiet Environment May Still Be Fragile 

Some IT organizations appear stable because serious outages are uncommon, while others rely on experienced employees who quietly prevent problems from becoming visible. The difference can be difficult to detect until those employees are unavailable or the underlying issue becomes more serious. 

A service may depend on one administrator who knows which job to restart every Monday. Support teams may close tickets promptly, only to see them reopened days later. Engineers may postpone maintenance because they do not trust the change process. Monitoring may generate hundreds of alerts while offering little warning about events that interrupt business activity. 

The operation may look calm even when the service itself is fragile. 

A useful assessment looks behind incident totals and service-level results. It examines whether teams understand recurring failures, whether changes create avoidable disruption, whether operational knowledge is documented, and whether staff have time for preventive work. 

The cost of overlooking these conditions can be substantial. In its 2024 outage analysis, Uptime Institute reported that sixteen percent reported a cost above $1 million. Four in five said their most recent serious outage could have been prevented through better management, processes, or configuration. 

When preventable failures remain embedded in daily work, operational discipline becomes increasingly important. 


Connect Metrics to the Business Experience 

Service-level agreements still have a role. Problems arise when they become the primary definition of good performance. 

A service desk can meet its response target while a department loses hours of productive time. Infrastructure can meet its availability target while a business-critical interface fails intermittently. An application can look healthy over a month even if it becomes unreliable during the organization’s busiest processing period. 

This can create two versions of IT performance: the results shown in formal reports and the experience of employees, customers, and business leaders. When the two diverge, stakeholders may begin to lose confidence in what the metrics are telling them.  

Operations teams may be working hard and meeting every target while stakeholders still believe technology is slowing them down. 

Reporting becomes more useful when it connects operational activity to service outcomes. For each critical service, leaders should be able to answer: 

  • Did the service support the business activity it was intended to support? 
  • How much productive time did an interruption take from employees or customers? 
  • Was the underlying condition addressed, or was service merely restored? 
  • Did a change create additional work, risk, or disruption? 
  • Is support demand rising, falling, or moving to lower-cost channels? 
  • What management decision should follow?

ServiceNow’s IT service management guidance takes a similar approach, recommending that organizations define success through business outcomes and select key performance indicators that demonstrate progress. 


Give Every Executive Metric a Job 

Executive IT dashboards tend to grow as each supplier, platform, and operating group adds measures, while few measures ever leave. 

A more useful dashboard starts with one test: What decision does this measure support? 

If a metric cannot be tied to a service, an owner, a business consequence, and a management decision, it probably does not belong in the executive view. 

A focused set of measures may include: 

  • Reliability: Is the complete service available and responsive when the business needs it? 
  • Recovery: How long does restoration take, and were recovery objectives met? 
  • Recurring demand: How much effort is consumed by repeat incidents, reopened tickets, manual corrections, and avoidable requests? 
  • Change performance: How often do changes cause incidents, rollback, or unplanned remediation? 
  • Business experience: What do employees and leaders encounter when they use the service? 
  • Risk and resilience: Which services rely on unsupported components, untested recovery procedures, concentrated knowledge, or a single supplier? 
  • Improvement capacity: How much effort goes toward automation, prevention, modernization, and service improvement? 

The exact measures will differ because a hospital, manufacturer, financial institution, and retailer do not share the same definition of a critical service. 

Software delivery teams may also use DORA metrics, including change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate. These measures should support learning and improvement, rather than becoming isolated performance targets. 


Follow the Service From Need to Decision

One way to test an IT operating model is to trace a business-critical service through a Service-to-Decision Chain: 

  1. Business need: What activity, customer interaction, or obligation does the service support? 
  2. Service experience: What do employees or customers encounter? 
  3. Operational evidence: What do incident, change, monitoring, capacity, and support data show? 
  4. Accountability: Who owns the end-to-end result and has authority to act? 
  5. Management decision: What should the organization maintain, correct, fund, retire, or redesign?

When the chain breaks, operating weaknesses often become easier to see. 

Monitoring teams may see a technical event but not know which business process it affects. The service desk may recognize a recurring issue but lack a route to problem management. A supplier may meet its contract while the full service remains unreliable. Leadership may receive reports without having an agreed threshold for funding corrective work. 

Replacing a tool rarely fixes an operating problem by itself. New platforms can organize work, improve visibility, and remove repetitive steps, but service ownership, competing priorities, unreliable data, and time for preventive work still require management decisions. 


Design Resilience Around Critical Services 

Operational resilience should begin with the business service, not the recovery technology. 

The National Institute of Standards and Technology describes cyber resiliency as the ability to anticipate, withstand, recover from, and adapt to adverse conditions involving systems that use or depend on cyber resources. 

That definition raises practical questions. Which services must continue during a disruption? What is the minimum acceptable level of service? Are dependencies documented? When was the complete recovery path last exercised? What did the exercise prove? 

A recovery document proves that a plan exists, not that recovery can be completed. 

Not every application needs the same recovery time, redundancy, or control structure. Leadership should direct resources toward services where disruption could create the greatest customer, operational, financial, regulatory, or safety consequence. 


Turn the Assessment Into a 90-Day Roadmap

An assessment has limited value if it ends with a lengthy deficiency register or an isolated maturity score. 

The first 90 days should focus on decisions that can help reduce exposure and change how the organization operates: 

  • Stabilize: Address recurring failures in critical services, unsupported components, control gaps, untested recovery procedures, and unclear escalation routes. 
  • Strengthen: Assign service ownership, improve operational data and reporting, clarify supplier accountability, and protect capacity for preventive work. 
  • Advance: Expand automation, observability, self-service, predictive analysis, and platform modernization after the operating foundation is sound. 

BDO can assist organizations with an independent review of service management, infrastructure operations, support, resilience, tooling, data, and governance. The review can help leadership identify where reported performance differs from business experience and establish a limited, owned set of actions for the next 90 days. 

A green dashboard is useful, but it means more when leaders and users can trust what it represents. The real measure of IT operations is whether people can rely on critical services, teams can respond without depending on heroics, and management can see evidence that the environment is becoming more predictable over time.

Get a Clearer View of How IT Is Really Running

A BDO IT Operations Excellence assessment or workshop can help you look beyond reported performance to evaluate service management, resilience, operational risk, and improvement priorities.