Building a Sentinel detection lab · Part 4
Building a Sentinel Detection Lab, Part 4: Catching Reconnaissance, and the Case-Sensitivity Bug That Made It Look Broken
The first detection in the lab that looks for a pattern of behavior, and the case-sensitivity bug that made it look broken.
Repo: github.com/Kajal-Dhanjal/sentinel-detection-lab
This is Part 4 of the series where I build a Microsoft Sentinel detection engineering lab from scratch and write down what actually happened — including the parts where I got it wrong. Parts 1–3 covered a brute-force detection (volume threshold), a suspicious-PowerShell detection (signature), and a registry persistence detection (signature plus a false-positive I had to tune out).
This one is different. It's the first detection in the lab that doesn't look for a thing—it looks for a pattern of behavior. And it's the one that taught me the most, because for a while it just sat there returning nothing, and I was convinced I'd broken it.
The detection that returned nothing
Here's where I'll start, because it's the most useful thing in this post.
I wrote the rule. I ran the recon commands on my lab VM to generate test data. I ran the query. Empty. No results.
My first instinct — the wrong one — was to start rewriting the logic. Maybe my threshold was off. Maybe the time window was wrong. I poked at it for longer than I'd like to admit.
The actual problem: Windows doesn't log process names consistently. Some commands showed up in Sysmon as arp.exe. Others showed up as ARP.EXE. My filter was matching against a lowercase list, so every uppercase entry silently fell through. The data was right there—my detection just couldn't see half of them because of letter casing.
The fix was one function:
| extend ProcessName = tolower(tostring(split(Image, "\\")[-1]))
Normalize everything to lowercase before you compare. Obvious in hindsight. But the lesson stuck harder than the fix: when a detection returns empty, don't assume your logic is wrong—check whether your assumptions about the shape of the data are wrong. A detection that silently under-counts is more dangerous than one that errors out, because it looks like it's working. It'll sit quiet during a real attack, and you'll never know it missed.
That's the whole job, really. The query is easy. Knowing why it lied to you is the skill.
What I was actually trying to catch
Reconnaissance. The phase right after an attacker lands on a machine, before they move anywhere—when they stop and ask the box where am I? who am I? and what's around me.
They run things like whoami, net.exe, nltest, systeminfo, ipconfig, tasklist, hostname, nslookup, arp, route. In MITRE ATT&CK terms this spans a few discovery techniques—T1057 (Process Discovery), T1082 (System Information Discovery), and T1016 (System Network Configuration Discovery).
Here's the problem that makes this detection interesting: not one of those commands is malicious. A sysadmin runs ipconfig 50 times a day. whoami is harmless. If I alerted on any single recon command, I'd drown the SOC in noise, and the alert would get muted within a week.
So the signal isn't the command. The signal is varied in a short window. A real human admin doesn't fire off ten different discovery tools inside ten minutes. An attacker — or an automated recon script — does. They want the full picture, fast.
A third kind of detection: behavioural clustering
My first three detections were a volume threshold and two signatures. This one needed a different shape entirely. I wasn't counting how many times one bad thing happened, and I wasn't matching one known-bad string. I was measuring how many distinct recon tools showed up together.
That's dcount() — count distinct—not a plain count(). Here's the tuned query:
Event
| where Source == "Microsoft-Windows-Sysmon"
| where EventID == 1
| extend Image = extract(@"Image:\s+(\S+)", 1, RenderedDescription)
| extend ProcessName = tolower(tostring(split(Image, "\\")[-1]))
| where ProcessName in ("whoami.exe", "net.exe", "net1.exe", "nltest.exe", "systeminfo.exe", "ipconfig.exe", "tasklist.exe", "hostname.exe", "nslookup.exe", "arp.exe", "route.exe")
| summarize ReconCommands = make_set(ProcessName), DistinctReconCount = dcount(ProcessName) by Computer, bin(TimeGenerated, 10m)
| where DistinctReconCount >= 5
Walking through it: pull Sysmon process-creation events (Event ID 1), extract the image path, and normalize the process name to lowercase (the fix from earlier); keep only known recon tools, then bucket everything into 10-minute windows per machine. make_set() collects which tools appeared and dcount() counts how many distinct ones. If five or more different recon commands cluster inside one 10-minute window on one host, it fires.
The make_set is there on purpose—when this alerts, the analyst immediately sees which tools ran, so they're not starting their investigation from zero.
Tuning in the opposite direction
In Part 3, my registry detection fired on a legitimate Windows process, and I fixed it by adding a narrow exclusion — I told the rule to ignore one specific benign pattern.
This detection's false-positive risk is completely different, so the fix is too. Here, the thing that might trip it innocently is a sysadmin genuinely multitasking—patching, troubleshooting, and running a handful of diagnostic commands in a burst. There's no single benign pattern to exclude; the behavior itself overlaps with legit admin work.
So instead of an exclusion, I tuned the threshold. I started at 4 distinct commands, and it felt slightly twitchy, so I raised it to 5. Fewer false alarms, and a real recon sweep still trips it easily because attackers don't stop at four.

The meta-lesson I took from running both detections back to back: the direction you tune depends on where the false positives come from. A specific benign actor → exclude it. A fuzzy overlap with normal behavior → move the threshold. Same goal, opposite move.
Validating it

I ran a burst of distinct recon commands on the lab VM, waited for ingestion (Sentinel log latency is real — if you run the query the instant you fire the commands, you'll scare yourself with an empty result that's just lag, not failure), then confirmed the rule fired once the distinct count crossed 5. Detection and the resulting incident are both screenshotted in the repo under evidences/.
What I took from this one
Three things:
-
An empty result is not proof the logic is correct. Check the data's actual shape before you touch the query. Case sensitivity, field formatting, and silent undercounting will fool you every time if you let them.
-
Some attacks have no bad ingredient — only a bad recipe. No single recon command is suspicious; the cluster is. Detection engineering is often about catching combinations, not items.
-
Tuning is a decision, not a default. Exclude a known-benign actor; raise a threshold against fuzzy overlap. Knowing which is which is the part that takes judgement.
Next up, I'm moving off the endpoint entirely and into identity—Entra ID detections, starting with attackers registering their own MFA methods on compromised accounts to keep access after a password reset. Different log source, different threat model. That's Part 5.
Full lab on GitHub.
MFA Registration, a Missing License, and the First Detection That Talks Back
Spotted an error or have a suggestion?
Email me