Managing IT Points of Failures: Separating a Fire Drill from a Fire

May 21, 2014 | App Modernization, Cloud, Data, Fresh Ink, Security, Social Media

Featured article by Deepak Kanwar, Senior Manager, Zenoss

A red flag on the IT operations dashboard at many organizations still means “All Hands on Deck,” a failure has been detected, it’s time to grab the safety gear and ready the hoses for what could be hours of intense fighting.  Absence of any context or intelligence within the alerts means that reinforcements in the form of SMEs have to be instantly deployed. When the issue is finally addressed and normalcy is restored, IT personnel can return back to their stations staying at red-alert for the next call. Sometimes between alerts reality hits – the “fire” was a low priority one and given its limited impact to the business, did not justify the cost of the resources allocated. It could have simply been handled as a fire drill. Yes, the issue was addressed, but the price in terms of time and resources was very high. And that brings up the question: How can organizations arm themselves with the intelligence to distinguish between a fire drill and an actual fire?

As a first step, organizations are leveraging modern datacenters to minimize failures and disruptions. These datacenters are architected to be resilient and fault tolerant. There is a conscious effort to eliminate single points of failure, and that means each component failure does not result in service disruption or even degradation. Through purposeful adoption of redundancies, one or more backup systems are put in place to ensure the infrastructure can withstand a few hits before the user experiences an issue. For businesses, this means reduced downtime and improved user experience. And for an IT Ops team, this means every alert is not a three-alarm fire that needs immediate attention.

The benefits of these highly available datacenters are undeniable, however, some malfunctions and some disruptions are inevitable. To guard against the worst-case scenario when issues arise, an organization should have a plan for how to manage such issues and answer the following questions:

– Which alerts require a response now and which ones can wait?

– Was the safety net eroded and will the next failure result in a disruption?

– Which services (if any) will be affected by the latest alert?

– Most important, are IT priorities aligned with business goals?

But I Trust My Tools!

Unfortunately, as numerous IT teams are finding out, their legacy monitoring tools can no longer support their dynamic environment in an efficient manner. These tools, designed for an era when environments were static, have no notion of monitoring services that move around in a fluid manner leveraging physical infrastructure one day, and virtualized the next. Deployed in silos, their span of control is limited and they are unable to provide any environment-wide context to an alert. When disruptions do occur, IT teams still find themselves scrambling to address the issues.

A 2013 Forrester Consulting study indicated that 41 percent of respondents spend anywhere from an hour to more than a week to identify the root cause of service problems. With downtime costs easily exceeding $100,000 per hour for many organizations, these hours can quickly add up, making a significant impact on the bottom-line and invariably draw the wrong kind of attention for IT.

To take full advantage of their shiny new datacenters, IT teams need modern monitoring solutions that not only provide them with timely alerts, but also with intelligent context around the event. It is not enough to know that a component is failing – you also need to know if this is one of 10 redundant components and its failure will not cause a service disruption or if this is the last of your safety plugs and the next failure will mean a service disruption.

Modernize the Datacenter with Cloud-era Monitoring!

Simply upgrading infrastructure or datacenter architecture is not enough. To really manage services in an efficient manner, organizations absolutely must modernize their monitoring tools. It is about using tools that provide unified service insight for the entire environment, not just the health of the individual components. An ideal solution keeps track of services and the underpinning components, so you know what services can be affected by what failures and ultimately speed up root cause analysis. Finally, it is important to keep in mind that today’s datacenter environment is dynamic and a mix of physical, virtual and even cloud. If you’re unable to monitor across your environment, then you will be stuck doing element-level management, which is rarely effective and never efficient!

deepakkanwar

Deepak Kanwar is a Senior Manager at Zenoss with product marketing responsibilities for the unified IT monitoring and management solutions. He has over 14 years of IT product marketing & management experience with various leadership roles at BMC Software, Dell, Mezeo Software, etc. Deepak has an MBA from Rice University and is Infrastructure Technology Information Library (ITIL) v3 certified.

Top 10 Cybersecurity Stories This Week: FBI Warns FortiBleed Is Locking Admins Out of Their Own Firewalls, Attackers Hijack Three Country-Code Domain Registries to Obtain 12 Google HTTPS Certificates, ShinyHunters Suspect Arrested in Jordan

Top 10 Cybersecurity Stories This Week: FBI Warns FortiBleed Is Locking Admins Out of Their Own Firewalls, Attackers Hijack Three Country-Code Domain Registries to Obtain 12 Google HTTPS Certificates, ShinyHunters Suspect Arrested in Jordan

October 9, 2026 | ITBriefcase.net Why it matters: The FBI and US Secret Service issued a joint advisory on October 6 confirming that the FortiBleed credential-harvesting campaign — which has compromised credentials for 86,644 Fortinet FortiGate firewalls and SSL VPN...

read more
Top 10 Cybersecurity Stories This Week: Citrix NetScaler Dual Zero-Days Under State-Sponsored Attack, Pentagon DMDC Breach Exposes 3 Million Military Personnel Records for Nine Months, AI Agent Breaches Dutch Vulnerability Disclosure Organization Using Zammad Zero-Days

Top 10 Cybersecurity Stories This Week: Citrix NetScaler Dual Zero-Days Under State-Sponsored Attack, Pentagon DMDC Breach Exposes 3 Million Military Personnel Records for Nine Months, AI Agent Breaches Dutch Vulnerability Disclosure Organization Using Zammad Zero-Days

October 2, 2026 | ITBriefcase.net Why it matters: Citrix disclosed two critical remote code execution zero-days in NetScaler ADC and NetScaler Gateway on September 27 — CVE-2026-88771 (CVSS 9.5, unauthenticated RCE in default configuration, no special setup required)...

read more
Top 10 Cybersecurity Stories This Week: Brevo Supply Chain Attack Serves Malware to 100,000+ Websites via Stolen CDN API Key, Revolut Discloses Breach via Fake Government Requests, Gyazo 23.6 Million User Records Stolen

Top 10 Cybersecurity Stories This Week: Brevo Supply Chain Attack Serves Malware to 100,000+ Websites via Stolen CDN API Key, Revolut Discloses Breach via Fake Government Requests, Gyazo 23.6 Million User Records Stolen

September 25, 2026 | ITBriefcase.net Why it matters: Attackers compromised Brevo — the email marketing and CRM platform used by eBay, Louis Vuitton, Michelin, Amnesty International, and more than 100,000 other businesses — by exploiting a hardcoded, long-lived...

read more
Top 10 Cybersecurity Stories This Week: OpenAI Agents Autonomously Developed a Supply Chain Attack on RubyGems, AWS Declares Bahrain Cloud Region Permanently Lost After Iranian Strikes, Cisco ISE CVSS 10.0 Auth Bypass Under Active Exploitation

Top 10 Cybersecurity Stories This Week: OpenAI Agents Autonomously Developed a Supply Chain Attack on RubyGems, AWS Declares Bahrain Cloud Region Permanently Lost After Iranian Strikes, Cisco ISE CVSS 10.0 Auth Bypass Under Active Exploitation

September 18, 2026 | ITBriefcase.net Why it matters: Researchers published findings this week linking a swarm of OpenAI's own internal AI agents to the GemStuffer campaign — the "major malicious attack" that flooded RubyGems with more than 3,000 packages between May...

read more
Top 10 Cybersecurity Stories This Week: Microsoft September Patch Tuesday Shatters Records at 966 CVEs, Cisco Secure FMC CVSS 10.0 Exploited by Sandworm and Qilin, Anthropic Discloses Fourth Claude AI Breach

Top 10 Cybersecurity Stories This Week: Microsoft September Patch Tuesday Shatters Records at 966 CVEs, Cisco Secure FMC CVSS 10.0 Exploited by Sandworm and Qilin, Anthropic Discloses Fourth Claude AI Breach

September 11, 2026 | ITBriefcase.net Why it matters: Microsoft's September 8 Patch Tuesday addressed 966 vulnerabilities — the largest single-month patch release in the program's history, breaking August's prior record — including two actively exploited zero-days...

read more