Back to Intelligence

Your On-Call Tech Got 300 Alerts Last Night — and the Real Problem Was One Disk

SA
AlertMonitor Team
September 15, 2026
11 min read

Alert fatigue isn't a volume problem. It's a signal quality problem. Here is what that means for your on-call rotation, and how to fix it at the root instead of swiping notifications forever.

The Industry Conversation IT Ops Should Be Paying Attention To

The Register recently published a piece titled Open weights are not open source: Why AI's favorite label is under dispute. The core argument is simple and uncomfortable: downloading a model is increasingly easy. Understanding how it was made, or changing the system at its root, is another matter entirely. A label on the package does not guarantee transparency under the hood.

Swap the words around and you have described the state of alerting in most IT departments and MSPs today.

Getting alerts is easy. Every RMM, every monitoring platform, every helpdesk on the market will happily email, SMS, and push-notify your on-call technician at 2 AM. The label on the box says intelligent alerting. But understanding why an alert fired, whether it actually matters, and changing how your alerting behaves at its root — that is another matter entirely. Most tools hand you the output of an opaque engine and call it a day. You receive the noise. You cannot inspect the reasoning. You cannot fix the behavior without clicking into hundreds of individual sensors, one at a time.

If you have ever opened your RMM on a Monday morning to find 1,400 alerts from the weekend and one actual incident, you already know the label and the reality do not match.

The Problem in Depth: Black-Box Alerting Is a Design Flaw, Not Bad Luck

What your current stack is actually doing

Threshold alerts with zero context. The classic page: CPU above 90% for 5 minutes on SRV-APP-01. What the alert does not tell you: that server sits at 85-90% every night at 23:00 because that is when the backup job and the nightly AV scan overlap. The threshold does not know what healthy looks like for that machine. So it fires, your tech acknowledges it, and by the twelfth time, your tech starts acknowledging without reading. That is not a discipline problem. That is a system teaching a human to ignore it.

Cascades instead of root causes. A Hyper-V host brownouts at 02:14. One incident. Your queue shows: host unreachable, agent offline, 14 services down, 3 network checks failed, 2 printers unreachable, an ODBC timeout, a failed backup job. Forty-plus critical alerts for one root cause, all firing within 90 seconds, all paging the same exhausted human. The actual incident count: 1. The page count: 40.

Deduplication that does not exist. The same printer-offline check fires every 5 minutes until someone manually clears it. Over a weekend, that is 200 alerts about one print server in a client's spare room.

Email-to-ticket gravity. Monitoring lives in PRTG or Nagios. Ticketing lives in ConnectWise Manage, Autotask, Zendesk — or, let's be honest, a shared Outlook mailbox. The alert becomes a ticket only when a human forwards it. Which means your mean time to acknowledge at 3 AM is not 5 minutes. It is whenever the tech wakes up, glances at a wall of noise, and manages to find the one real alert.

Suppression that is not suppression. You scheduled a patch reboot for Saturday 02:00. The monitoring tool — which was never told, because it lives in a different product than your patching tool — pages the on-call tech for exactly the reboot you planned. Every single patch cycle.

Tuning that requires archaeology. Want to fix the noise? Open the sensor list, find the checks, adjust thresholds one by one across dozens of devices and clients. There is no view that tells you which 5 rules generate 80% of your pages. The engine is a black box: you can receive its output, but changing the system at its root is effectively impossible. So nobody does, and the noise compounds quarter after quarter.

Why the gaps exist

  • Siloed architecture. Monitoring, RMM, helpdesk, and patching were bought separately, from different vendors, in different years, and glued together with email forwarding and CSV exports. The alert engine cannot know what the patch engine is doing, so it pages you for planned maintenance.
  • Legacy alert engines. Static thresholds and per-sensor configuration were designed when a network was 20 servers in one server room. They do not scale to 2,000 endpoints across 30 client sites, and they were never designed to correlate, suppress, or explain.
  • Missing context by construction. The monitoring tool does not know which client owns the device, who is on call, what changed 10 minutes before the alert, or what normal looks like for that endpoint. Without that context, no amount of cleverness at the notification layer can save you. Garbage signal in, garbage pages out.

What it costs — in numbers a practitioner will recognize

  • Response time. A file server's C: drive starts filling at 2:05 AM from a runaway IIS log directory. The disk threshold fires into a queue already holding 200 overnight alerts. A human triages it at 6:40 AM. Users start calling the helpdesk at 7:15. Resolved at 9:00. Seven hours for a problem a contextual alert — C: grew 3 GB in the last hour, 4% free, predicted full by 02:40 — would have surfaced and routed at 2:06.
  • Desensitization. When 90% of pages are noise, technicians swipe notifications reflexively. One night, the swiped notification is the domain controller. Everyone in this industry knows a story like this. Nobody believes it will be their team. Until it is.
  • Ticket flooding and SLA distortion. One incident generating 40 tickets wrecks your open-ticket counts and your SLA reporting. And because monitoring data and ticket data live in separate systems, your IT manager cannot produce one honest number for how long the team actually took to respond to client X's outage last month. The data exists. It just lives in four places that do not talk.
  • Burnout and turnover. Nobody stays in an on-call role through years of nightly false alarms. The rotation becomes a punishment, the good techs leave, and replacing a sysadmin costs far more than fixing your alerting ever would.

This is the through-line from that Register article, applied to your own tooling: it is not enough to receive the output of a system. You need to understand it, and you need to be able to change it at its root.

How AlertMonitor Solves This

AlertMonitor was designed around a specific insight: alert fatigue is not a volume problem — it is a signal quality problem. Volume is what you get when the tool lacks context. Here is what changes when context is built in.

Every alert carries the full story. Device identity, client and site, what changed immediately before the alert — a config change, a patch install, a service stop, a topology shift — and the baseline: what healthy looks like for that specific machine. An AlertMonitor alert reads like a sentence a senior engineer would say, not a threshold tripwire. Disk alerts show the growth trend and a predicted-full time. Service alerts show the last change and whether it correlates with Tuesday's patch run.

Smart deduplication and correlation. When a host drops, AlertMonitor raises one root-cause alert and groups the downstream symptoms under it as child events. One page. The full blast radius is visible on the network topology map, so the tech sees the impact without opening five tools. The printer that fires every 5 minutes becomes a single open alert that updates — not 200 entries burying everything else.

Maintenance window suppression that actually works. Because patch management and monitoring are the same system, AlertMonitor knows when a reboot is planned. Schedule the window once per client or site; every affected alert during that window is suppressed and logged — visible in history, never a page. The loop closes itself: patch installs, endpoint reboots, services verified back up, alert history annotated. No Saturday morning apologies to whoever was on call.

Multi-level on-call routing and escalation. Policies, not per-sensor archaeology: L1 tech paged immediately; no acknowledgement in 10 minutes escalates to the team lead; 15 more minutes goes to the escalation manager. Configurable per client, per severity, per time of day. When the noise pattern changes, you change the policy — the behavior of the system at its root — instead of editing 800 individual checkboxes.

One console from signal to resolution. The alert becomes a ticket in the integrated helpdesk with one click, carrying all of its context. The tech remediates from the built-in RMM remote session, runs the cleanup script, verifies the disk, and the ticket, the alert, and the response timeline stay linked automatically. SLA reporting finally reflects reality, because monitoring, ticketing, and response data were never separate systems to begin with.

The workflow, side by side

The old fragmented way: pager buzzes → tech opens email → 40 alerts, manual triage → creates a ticket in ConnectWise or Autotask by hand → opens VPN → opens a second remote tool → fixes the issue → documents in a third tool → closes everything. A routine 2 AM disk cleanup: 45-60 minutes, four tools, three opportunities to forget the documentation.

The AlertMonitor way: one correlated alert with trend context routes to on-call → runbook link attached → ticket auto-created → remote session in one click → cleanup script runs → verified resolved, alert and ticket auto-linked. Eight to twelve minutes, one console, documentation that essentially wrote itself.

Fewer overnight pages. Faster response. And a team that is not being burned out by its own monitoring tools.

Practical Steps You Can Take This Week

1. Measure the noise before you touch anything. Export the last 30 days of alerts and count them by rule and source. In nearly every environment, five rules generate the majority of pages. You cannot fix a black box until you can see its output distribution.

2. Replace dumb thresholds with trend-aware checks where you can. Even a scheduled sweep with context beats a static 90% page. Here is a PowerShell disk sweep across your Windows Servers that reports free percentage and absolute headroom:

PowerShell
$servers = Get-Content .\servers.txt
$report = foreach ($srv in $servers) {
    Get-CimInstance -ComputerName $srv -ClassName Win32_LogicalDisk -Filter 'DriveType=3' |
        Select-Object @{n='Server';e={$srv}},
                      @{n='Drive';e={$_.DeviceID}},
                      @{n='FreeGB';e={[math]::Round($_.FreeSpace/1GB,1)}},
                      @{n='SizeGB';e={[math]::Round($_.Size/1GB,1)}},
                      @{n='FreePct';e={[math]::Round(($_.FreeSpace/$_.Size)*100,1)}}
}
$report | Where-Object { $_.FreePct -lt 15 -or $_.FreeGB -lt 10 } |
    Sort-Object FreePct | Format-Table -AutoSize
$report | Export-Csv .\disk-report.csv -NoTypeInformation

3. Self-heal before you page anyone. A surprising share of overnight pages are service blips a script can fix. Check critical services, restart them, and escalate only if the restart fails:

PowerShell
$critical = @('DNS','DHCP','Spooler','wuauserv')
foreach ($name in $critical) {
    $svc = Get-Service -Name $name -ErrorAction SilentlyContinue
    if ($svc -and $svc.Status -ne 'Running') {
        Write-Output "$($env:COMPUTERNAME): $name is $($svc.Status) - attempting restart"
        try {
            Start-Service -Name $name -ErrorAction Stop
            Write-Output "$name restarted successfully"
        } catch {
            Write-Warning "$name failed to restart - escalate to on-call"
        }
    }
}

4. Do the same for your Linux fleet. Disk usage plus failed systemd units, in one pass:

Bash / Shell
#!/bin/bash
THRESHOLD=85
for srv in $(cat servers.txt); do
    pct=$(ssh -o ConnectTimeout=5 "$srv" "df -P / | awk 'NR==2 {print \$5}'" | tr -d '%')
    if [ -n "$pct" ] && [ "$pct" -gt "$THRESHOLD" ]; then
        echo "$srv: root filesystem at ${pct}%"
    fi
    failed=$(ssh -o ConnectTimeout=5 "$srv" "systemctl --failed --no-legend | wc -l")
    if [ "$failed" -gt 0 ]; then
        echo "$srv: $failed failed systemd unit(s)"
    fi
done

5. Write the escalation ladder down before you need it. Who is paged first, the acknowledgement timeout, who is next, who is last. If the escalation path lives in someone's head, it does not exist at 3 AM. In AlertMonitor, this is the escalation policy editor: define the tiers, set the timeouts, scope it per client and severity, done.

6. Make suppression automatic, not manual. Calendar the patch windows per client in AlertMonitor's patch management module and let maintenance window suppression handle the rest. If your current stack requires someone to manually clear alerts before every patch night, that step will get skipped on the busiest weeks — which are exactly the weeks with reboots.

7. Audit quarterly. Alert noise creeps back. Re-run the 30-day export, prune the top offenders, and retire checks for decommissioned printers nobody ever told monitoring about.

The Register's point about AI labels applies to your tooling too: the question is not whether the box says intelligent alerting. The question is whether you can see how it thinks — and change it when it is wrong. Your on-call rotation lives or dies on that difference.

Related Resources

AlertMonitor Alert Management & On-Call Operations

AlertMonitor Platform Overview

Book a Demo

Alert Management & On-Call Operations Resources

alert-fatiguealert-managementon-callescalation-policyalertmonitoron-call-operationsmsp-operationsmonitoring-noise

Is your security operations ready?

Get a free SOC assessment or see how AlertMonitor cuts through alert noise with automated triage.