Building a Sentinel detection lab · Part 3
Building a Sentinel Detection Lab, Part 3: My Detection Flagged Legitimate Windows Behavior. Here's How I Tuned It Without Going Blind
Registry persistence, a textbook false positive, and the single most important lesson in detection engineering.
Part 3 of a series on building a detection engineering lab in Microsoft Sentinel. This post: registry persistence, a textbook false positive, and the single most important lesson in detection engineering.
The threat
To survive a reboot, malware commonly writes itself into a registry autostart location—or a Winlogon shell key—so it launches automatically at every login. Sysmon logs registry modifications as Event ID 13.
The naive detection
The logic seems obvious: flag any write to an autostart location.
Event
| where Source == "Microsoft-Windows-Sysmon"
| where EventID == 13
| extend TargetObject = extract(@"TargetObject:\s+(\S+)", 1, RenderedDescription)
| extend Details = extract(@"Details:\s+(.+?)\s+(?:User:|Image:)", 1, RenderedDescription)
| extend Image = extract(@"Image:\s+(\S+)", 1, RenderedDescription)
| where TargetObject has_any (@"CurrentVersion\Run", @"CurrentVersion\RunOnce", @"Winlogon\Shell", @"Winlogon\Userinit")
| project TimeGenerated, Computer, TargetObject, Details, Image
I ran it. It fired straight away—three hits. All of them were userinit.exe writing ctfmon.exe to a run key.
The false positive
ctfmon.exe is the legitimate Windows text/input service (keyboard layouts, handwriting, and speech). Windows itself registers it to autostart via the registry. This happens on every Windows machine, constantly. It is 100% normal.
My detection wasn't wrong—it correctly matched "a write to an autostart location.""It was too broad. And this is the lesson:
A detection that fires on benign activity is worse than no detection at all.
Here's why. If a rule cries wolf on normal behavior, analysts learn to dismiss its alerts. Then, when a real attack trips the same rule, nobody looks. That's alert fatigue, and it's how real intrusions get missed in environments that are technically "monitored." Tuning out benign noise isn't cleanup work you do at the end—it is the detection engineering.
The fix: a narrow allowlist
The tempting fix is to exclude userinit.exe. That's the wrong fix—an attacker who abuses userinit.exe (a real technique) would then sail straight through. Good exclusions are narrow: they drop the exact benign pattern, not a whole process.
| where not(Image endswith "userinit.exe" and Details has "ctfmon.exe")
This excludes only the specific combination of userinit.exe writing ctfmon.exe. Anything else writing to those keys—including malware abusing userinit.exe — still fires.
Validating both sides
This is the part people skip. After tuning, I confirmed the detection went quiet on the ctfmon false positive. But "it stopped firing" proves only half a detection. The other half: Does it still catch real attacks? Over-tuning is the silent failure mode—you suppress the noise and accidentally blind the rule.
So I planted a fake persistence entry (harmless — the referenced file doesn't exist):
New-ItemProperty -Path "HKLM:\Software\Microsoft\Windows\CurrentVersion\Run" -Name "FakeTestMalware" -Value "C:\temp\evil.exe" -PropertyType String -Force
The detection caught it—while still ignoring the ctfmon write. True positive fires, false positive suppressed. That's a complete detection.

The scheduled rule then raised a high-severity incident, MITRE-mapped to T1547.001, with an attack-story graph linking the host. (I removed the fake entries afterward.)

The takeaway that matters most
The first two detections taught me how to write detections. This one taught me how to operate them:
-
Firing on benign activity is a failure, not a success. A noisy rule trains people to ignore it.
-
Allowlist narrowly. Exclude the exact benign pattern, never a whole process.
-
Validate both sides. A detection that stopped firing on noise but no longer catches attacks is worse than the noisy version—at least the noisy one worked.
That third point is the difference between someone who writes queries and someone who runs detections. You can't claim a detection works until you've watched it catch the real thing and ignore the fake one.
Full lab and all three detections on GitHub.
Catching Reconnaissance, and the Case-Sensitivity Bug That Made It Look Broken
Spotted an error or have a suggestion?
Email me