Closing a false positive today is easy. Keeping it closed, correctly, for the next eighteen months is the real job.

That is the lens for any effort to reduce false positives in Microsoft Sentinel. SIEM alert fatigue is not a volume problem you can hire your way out of. Microsoft and Omdia’s State of the SOC report, published in February 2026, found that 46% of alerts are false positives and 42% go uninvestigated [10]. Meanwhile, the ISC2 workforce gap reached nearly 4.8 million in 2024, up 19% in a year, while the workforce itself grew just 0.1% [11]. The gap is growing faster than anyone can hire, so the noise must be removed at the source.

The human cost is well documented. In interviews with SOC analysts, AlAhmadi et al. found that constant false alarms lead to burnout and, eventually, desensitisation [1]. Sundaramurthy et al. link analyst burnout to turnover and poor judgement [2].

There are two ways to remove that noise: manual rules and automated tuning. This guide explains what each one is, where each breaks down over time, what Sentinel already gives you, and how to combine them. The short version: use rules where the answer is certain and the risk is high, use automated tuning for the volume, and re-check both regularly. We use eighteen months as the yardstick throughout, because it is long enough for IP ranges, roles, vendors and rules to change at least once.

Why SIEM alert fatigue happens in Sentinel

The causes are structural. Too many low-value analytics rules fire on broad thresholds. Repetitive, benign patterns (service principals, vulnerability scanners, admin tooling). Alert detections are written for worst-case behaviour. AlAhmadi et al. found most false positives are “benign triggers”: true detections of legitimate activity [1].

Layer on the Sentinel/Defender overlap, where the same activity surfaces as multiple correlated alerts, and analysts spend their day validating noise. A 2021 IDC survey of 300+ US IT executives found staff spend around 30 minutes on each actionable alert and 32 minutes on each false lead [12]. Analysts learn to expect noise, and that expectation is what lets a real incident slip through [6].

Manual vs. automated tuning: what’s the difference?

Manual tuning means a person writes deterministic rules (KQL exclusions, automation rules, watchlists) that close the same alert the same way every time. Automated tuning derives the decision from data: baselines, clustering or a classifier trained on your analysts’ past classifications. Rules are exact but decay as the environment changes; models scale but drift unless re-validated.

Both terms get thrown around in vendor pitches and forum threads. The difference between them is specific, and it decides what you end up maintaining.

Manual tuning means rule-based tuning: a person identifies a pattern, writes a rule once, and maintains it from then on. The person may judge differently from one day to the next; the rule does not. Once written, it behaves deterministically, giving the same outcome for the same input every time. Microsoft documents three ways to do it: automation rules or logic apps, KQL exclusions in analytics rules, and watchlist allow-lists [15]. Logic Apps playbooks with hand-written conditions, threshold edits and suppression belong to the same family. It is “if Y then X”, with every branch enumerated by hand.

Automated tuning derives the tuning from data. Instead of enumerating scenarios, the system learns tendencies: statistical baselining per rule or entity, anomaly detection, clustering and similarity of alerts, supervised classification trained on analyst classifications (Sentinel’s incident Classification and ClassificationReason fields are natural labels), and entity risk scoring. It looks at patterns, not lists.

Two things often get called automated tuning but are not what we mean here. LLM-based triage is non-deterministic: strong for investigation, weak for consistent suppression. And Microsoft’s built-in ML (Fusion, UEBA) is automated, but it detects and correlates rather than tunes, and you mostly switch it on or off.

 Manual (rule-based) tuningAutomated tuning
LogicDeterministic, “if Y then X”Derived from data patterns
AuthoringA human writes each ruleThe system learns from history
ExamplesKQL exclusions, automation rules, watchlists, suppressionBaselining, clustering, supervised classification on dispositions
StrengthExact, auditable, safe for high-risk detectionsScales across rules and entities
WeaknessBlind to context; upkeep grows with every exceptionDrift, explainability, inaccurate labels

Reducing false positives with manual tuning: precise today, obsolete tomorrow

A manual exclusion is correct the day you write it. Take a user who works from an approved VPN: “this IP range with this user is fine”. That exception now has to exist in every rule the user can trip: new-country sign-ins, impossible travel, unfamiliar sign-in properties, MFA anomalies. Then the VPN provider changes its ranges, the user changes role, or a contractor with the same pattern joins. Each change touches every one of those rules, and the logic quickly becomes a nest of if/then/else.

The blunt version is just as common: auto-close every informational-severity incident. It removes a large share of the queue overnight, and it is a rule nobody revisits. The day attacker activity surfaces as an informational alert, it is closed before anyone sees it. The rule did exactly what it was told; it had no way of knowing the context had changed.

The problem scales with detection surface, not headcount. A five-person team and a fifty-person team both drown, because the combinatorics live in the rules, not the organisational chart. Research on intrusion detection rules points the same way. Vermeer et al. found that most rules in a SOC’s ruleset never trigger [3], and a follow-up study found rule upkeep is manual and time-consuming, with better documentation the improvement practitioners asked for most [4].

Doesn’t Sentinel already do this? Fusion and UEBA

Sentinel ships with machine learning, so it is a fair question. Fusion correlates alerts from different sources into multistage incidents, using ML trained on 30 days of historical data [13]. UEBA baselines each entity against its own history, its peers and the organisation [14]. Both are automated, but neither is tuning: they create or group alerts, they do not learn which of your existing alerts are benign and close them. Correlation can lower the incident count by merging related alerts, but the noise underneath is still there. And you have little control over either. The built-in ML feature that exposes thresholds you can adjust is customizable anomalies [19].

Reducing false positives at scale: the options, from rules to ML

  • Manual rules: Complex Logic Apps and automation rule workflows that weigh many signals. This is rule-based tuning taken as far as it goes: powerful, but still a set of conditions someone has to maintain.
  • LLMs in the triage path: fast to prototype and shown to work on narrow tasks such as phishing triage [9], but expensive to operationalise and non-deterministic.
  • Custom ML built for alert classification: trained on your own analysts’ classifications and controllable, as research systems like DeepCASE demonstrate [5], but you own the pipeline.
  • The Defender XDR correlation engine: capable, but Microsoft’s model, not yours [16].
  • Security Copilot’s alert triage agent: in preview, covering a subset of alert types, and priced per Security Compute Unit [17]. In our experience it is only as good as the process around it, and it does not remove false positives on its own.
  • Statistical baselining per rule or entity: thresholds derived from your own environment. Sentinel’s own tuning recommendations already mine your false-positive classifications this way [18].

The research backs custom ML and baselining. DeepCASE, a model that learns from analysts’ past decisions, filtered 86.72% of events and cut operator workload by 90.53%, underestimating risk in under 0.001% of cases [5]. NoDoze scored alerts by how unusual their surrounding activity was and shrank investigation graphs by two orders of magnitude [6].

Why not just put an LLM on it?

It is the obvious question, so it deserves a proper answer. LLMs are good at investigation and weak at consistent suppression.

Benefits: low barrier to start, tolerance for unnormalized data, and criteria you can describe in plain language instead of code.

Drawbacks: they need heavy context (expensive), infrastructure around them, and deep knowledge of your data and of what a good outcome looks like.

Take the multi-country login rule. You give an agent the incident and one instruction: compare this sign-in with the user’s last 30 days in SigninLogs and AADNonInteractiveUserSignInLogs; if the country, device and IP range all appear in that history, close it as a false positive, otherwise escalate. The instruction is one sentence. Making it reliable is not. The agent has to run the right query over the right window, decide what “appears in the history” means (once? regularly?), and leave a record of why it decided what it did. Ask it the same question twice and you can get two different answers.

Microsoft’s own trial of its Copilot phishing triage agent shows where LLMs do work. 167 professional analysts were randomly assigned to work with or without the agent, so the difference in results can be put down to the agent and not to who happened to use it. Agent-assisted analysts found up to 6.5 times as many true positives per minute and improved their F1 score, a measure that balances catching threats against raising false alarms, by 77%. Most of that gain came from the agent closing benign reports on its own [9].

The lesson is in the conditions, not the headline. Phishing reports are one narrow task with a clear verdict. Most Sentinel analytics rules are not: they span identities, devices and networks, and whether an alert is benign depends on context that keeps changing. There, an LLM earns its place as an investigation assistant that gathers context for an analyst, not as the thing that closes incidents.

For a wider look at which AI belongs where in the SOC, and what the research says about LLMs versus classical ML for triage, see Agentic AI for the SOC: which AI, for which problem.

When to tune by hand, and when to let ML do it

Combine both. When a case is genuinely deterministic and you know it is normal (if Y then X, always), close it with a rule on precise terms. Rules are also the right tool for high-risk detections, where every change needs careful validation and an audit trail. Volume is not their problem: one rule can close thousands of alerts. Context is. A rule closes everything that matches, whether or not the circumstances still make it benign, and keeping hundreds of rules aligned with a changing environment is where the effort goes.

For recognising patterns at volume, the research above points to purpose-built models over manual effort [5][6]. That is also why custom machine learning for alert classification is chosen rather than a language model: it gives the same answer to the same alert, it learns from your analysts’ classifications, and you can see and control what it does.

Be honest about ML’s limits, too. Models decay. Pendlebury et al. showed that malware classifiers which score well in testing lose much of their accuracy on data from later months, because threats and environments change over time [7]. For tuning, that means a model trained on last quarter’s alerts slowly drifts away from this quarter’s environment unless it is re-checked against fresh data. Arp et al. list ten common pitfalls in security ML, including training on inaccurate labels and on data that would not be available in real use, which make models look better in the lab than they perform in production [8]. Automated tuning needs the same discipline: models that are not re-validated on fresh data lose accuracy over time [7][8].

For the step-by-step version (baselining, multi-factor exclusions, watchlists and re-validation), see our Microsoft Sentinel alert tuning checklist.

Back to the eighteen-month question

Manual tuning fails slowly, by accumulation: every exception is right when it is written and becomes a liability as the environment moves. Automated tuning fails quietly, by drift, when nobody re-checks it. What holds up is the combination: rules for the deterministic and the high-risk, automated tuning for the volume, and a regular check on both.

Start by measuring your per-rule false positive rate over 30 days. You cannot tune what you have not baselined, and you cannot know whether today’s decision still works in eighteen months if you never wrote down where you began.

Start with your own numbers.


See your per-rule false-positive rate for the last 30 days and which rules to tune first, by hand or by model.

 

Seculyze is a Danish cybersecurity SaaS company offering AI-powered Security Operations on top of Microsoft Sentinel. It reduces alert noise, improves detection quality and cuts Sentinel cost, while your team keeps control of the SOC. seculyze.com

REFERENCES

RESEARCH PAPERS

[1] B. AlAhmadi, L. Axon, I. Martinovic. “99% False Positives: A Qualitative Study of SOC Analysts’ Perspectives on Security Alarms.” USENIX Security 2022. https://www.usenix.org/conference/usenixsecurity22/presentation/alahmadi

[2] S. C. Sundaramurthy et al. “A Human Capital Model for Mitigating Security Analyst Burnout.” SOUPS 2015. https://www.usenix.org/conference/soups2015/proceedings/presentation/sundaramurthy

[3] M. Vermeer, M. van Eeten, C. Gañán. “Ruling the Rules: Quantifying the Evolution of Rulesets, Alerts and Incidents in Network Intrusion Detection.” ACM AsiaCCS 2022. https://doi.org/10.1145/3488932.3517412

[4] M. Vermeer, N. Kadenko, M. van Eeten, C. Gañán, S. Parkin. “Alert Alchemy: SOC Workflows and Decisions in the Management of NIDS Rules.” ACM CCS 2023. https://dl.acm.org/doi/10.1145/3576915.3616581

[5] T. van Ede et al. “DEEPCASE: Semi-Supervised Contextual Analysis of Security Events.” IEEE S&P 2022. https://doi.org/10.1109/SP46214.2022.9833671

[6] W. U. Hassan et al. “NoDoze: Combatting Threat Alert Fatigue with Automated Provenance Triage.” NDSS 2019. https://doi.org/10.14722/ndss.2019.23349

[7] F. Pendlebury et al. “TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time.” USENIX Security 2019. https://www.usenix.org/conference/usenixsecurity19/presentation/pendlebury

[8] D. Arp et al. “Dos and Don’ts of Machine Learning in Computer Security.” USENIX Security 2022. https://www.usenix.org/conference/usenixsecurity22/presentation/arp

PREPRINTS

[9] J. Bono. “Randomized Controlled Trials for Phishing Triage Agent.” arXiv:2511.13860, 2025. Preprint, not peer-reviewed. https://arxiv.org/abs/2511.13860

INDUSTRY & DOCUMENTATION

[10] Microsoft & Omdia. “State of the SOC: Unify Now or Pay Later” (survey conducted 2025). Microsoft Security Blog, 17 Feb 2026. https://www.microsoft.com/en-us/security/blog/2026/02/17/unify-now-or-pay-later-new-research-exposes-the-operational-cost-of-a-fragmented-soc/

[11] ISC2. 2024 Cybersecurity Workforce Study. https://www.isc2.org/Insights/2024/10/ISC2-2024-Cybersecurity-Workforce-Study

[12] IDC for Critical Start (2021), reported by Forbes, 8 Nov 2021. https://www.forbes.com/sites/edwardsegal/2021/11/08/alert-fatigue-can-lead-to-missed-cyber-threats-and-staff-retentionrecruitment-issues-study/

[13] Microsoft Learn. Advanced multistage attack detection (Fusion) in Microsoft Sentinel. https://learn.microsoft.com/en-us/azure/sentinel/fusion

[14] Microsoft Learn. User and Entity Behavior Analytics (UEBA) in Microsoft Sentinel. https://learn.microsoft.com/en-us/azure/sentinel/identify-threats-with-entity-behavior-analytics

[15] Microsoft Learn. Handle false positives in Microsoft Sentinel. https://learn.microsoft.com/en-us/azure/sentinel/false-positives

[16] Microsoft Learn. Exclude Microsoft Sentinel analytics rules from correlation (Defender XDR). https://learn.microsoft.com/en-us/defender-xdr/exclude-analytics-rules-correlation

[17] Microsoft Learn. Security Copilot Security Alert Triage Agent in Microsoft Defender (Preview). https://learn.microsoft.com/en-us/defender-xdr/security-alert-triage-agent

[18] Microsoft Learn. Get fine-tuning recommendations for your analytics rules in Microsoft Sentinel. https://learn.microsoft.com/en-us/azure/sentinel/detection-tuning

[19] Microsoft Learn. Use customizable anomalies to detect threats in Microsoft Sentinel. https://learn.microsoft.com/en-us/azure/sentinel/soc-ml-anomalies