Detection engineering · Blog

What your logs can't see

Nine common sources and the gaps they carry by design.

The premise

"A detection rule can only fire on evidence that reached the platform."

That sounds obvious written down, and it's routinely skipped in practice. When coverage gets audited, it gets audited at the rule layer, which usually consists of questions like how many techniques we've mapped, how many analytics are enabled, and which ATT&CK cells are green. That's one layer too high. Underneath every rule is a log source, and every log source is a lens: built to answer a particular question well, which means built not to answer others.

Those aren't bugs, and they aren't vendor failures. They're design boundaries, chosen deliberately for cost, performance, privacy or licensing. But they set the ceiling on what your detections can ever see, and most of them stay invisible until the day something doesn't fire.

Below is what nine commonly trusted sources cover and what each one structurally cannot.

The table

Log sourceWhat it answers reliablyWhat it structurally cannot seeWhy it's built that way
Entra ID sign-in logsWho authenticated, from where, on what client, and how conditional access and MFA resolvedWhat the token holder does once the token is issued. An attacker who steals a session or refresh token can impersonate the user until that token expires or is revoked, carrying a session that already satisfied MFA [1]Sign-in logs are an authentication record, not an activity record. Activity lives in a different log. Retention is also licence-bound: seven days on Entra ID Free and thirty on P1/P2, and the change isn't retroactive [2], [3]
Windows Security log event 4688That a process was created, its image path, its parent, and the account that spawned itWhat the process was actually told to do. The Process Command Line field is empty unless you explicitly turn it on, so you get powershell.exe ran without the argument that mattered [4]Command lines are logged in plain text and often contain secrets, so Microsoft ships the field disabled and leaves the trade-off to you [5]
SysmonProcess lineage, network connections, image loads, file and registry activity: the richest native Windows telemetry availableAnything its configuration excludes. Sysmon is disabled by default and produces almost nothing until you install it and give it an XML config [6]Sysmon is a filter engine, not a firehose. Community baselines are deliberate about what they drop to keep volume survivable [7] and a filtered-out event is silently absent, not empty
EDRBehavioural detection, process trees and response actions everywhere the agent is installedEverywhere the agent isn't. Network appliances, hypervisors, unmanaged VMs, contractor laptops and BYOD aren't lightly covered; they're absent. Microsoft found that over 90% of ransomware attacks reaching the ransom stage used unmanaged devices for initial access or remote encryption [8]Agent-based tooling requires something to install the agent on. Devices you don't manage are, by definition, devices you can't instrument
Firewall / NGFWConnections: source, destination, port, volume, allow or deny, and session durationWhatever crossed the connection. TLS hides the payload, so without inspection your firewall sees an encrypted blob to a legitimate-looking destination [9]. Behind NAT, egress logs give you the gateway address, not the hostDecryption is compute-intensive, breaks certificate-pinned applications, and carries privacy and legal weight, so most environments inspect selectively or not at all [9]
Email gatewayInbound mail at the delivery boundary: sender, attachments, links, verdictsEverything after delivery, and everything that never crosses the boundary. Mailbox rules, auto-forwards, reads, and internal-to-internal mail don't transit the gatewayThe gateway is a perimeter control. Post-delivery actions are recorded by mailbox auditing in a different system entirely, and only if that auditing is configured
Microsoft 365 unified audit logTenant-wide user and admin activity across Exchange, SharePoint, Teams and EntraAnything past your retention window, and historically events your licence tier didn't emit. Audit (Standard) retains 180 days; Audit (Premium) extends to one year, but only for specific workloads and only for users holding an E5 or equivalent licence [10], [11]This gap isn't technical; it's commercial. Which is what makes it easy to miss during design and expensive to discover during an incident
AWS CloudTrailControl-plane operations: who created, modified or deleted a resource, and who signed in to the consoleWhat happened inside the resource. Data events, meaning S3 object reads and writes, Lambda invocations and DynamoDB item operations, are not logged by default and must be explicitly enabled [12]Data events are high-volume and separately billed. The default is the cheap one, and the default is where most accounts stay
DNS resolver logsName resolution that passes through your resolver; a near-complete map of what your endpoints try to reachResolution that goes around it. With DNS-over-HTTPS, the query is encrypted inside ordinary HTTPS on port 443, bypasses your resolver entirely, and gateway logging is usually not possible [13]DoH was designed for user privacy against network observers. Your enterprise resolver is a network observer.

How to read this

Two failure modes fall out of the table, and they look nothing alike from the console.

  • The first is a rule written against a source that can't carry it. The logic is sound, the ATT&CK mapping is correct, and the syntax validates, but the evidence never arrives. The rule stays silent, and silence gets read as safety.

  • The second is confidence drawn from a source covering less than assumed, almost always because a default was never changed. The rule fires, the tuning looks healthy, and the blind spot sits inside the alerts you're already receiving. This one is worse, because the alerts create the impression of coverage.

Neither failure produces an error message. That's the point of writing them down.

The gap that costs people the breach

The clearest documented case is Storm-0558. In June 2023, a US federal agency detected suspicious activity in its Microsoft 365 tenant by spotting MailItemsAccessed events carrying an unexpected application ID [14]. That single event type was how the intrusion surfaced at all.

Other affected organisations couldn't run the same check. At the time, MailItemsAccessed was gated behind premium licensing, so tenants without it had no way to see the mailbox access that had already happened [14]. The detection logic wasn't the variable. The log was.

Microsoft subsequently expanded the event to Standard-tier customers and raised default audit retention from 90 to 180 days [10], [14]. The fix is real, and it arrived after the incident that proved it was needed, which is the whole argument in one example: nobody noticed the boundary until an intrusion sat on the wrong side of it.

What actually changes

For anyone running detections across client environments, four things follow directly:

  • Name the source under every rule. For each detection you own, write down which log it depends on and which gap that log carries. If you can't name the source, you can't state the coverage.

  • Audit the defaults you inherited, not the ones you chose. CLI auditing, Sysmon config, mailbox auditing, and CloudTrail data events are all off or minimal out of the box, all cheap to enable, and all invisible until they matter.

  • Treat licence tiers as an architecture decision. Retention windows and event availability are coverage decisions made by procurement, usually without anyone in detection being asked.

  • Add sources to close named gaps, not to raise volume. More telemetry isn't more coverage. Telemetry that answers a question you couldn't previously answer is.

There's a reason this keeps biting. In the 2026 SANS detection engineering survey of 307 practitioners, cloud-native environments were named the number one coverage gap – more than two and a half times any other environment [15]. That isn't a rule-writing problem. It's a visibility problem wearing a rule-writing costume.

Your coverage stops where your telemetry does. It's worth knowing exactly where that is before an incident tells you.


References

[1] Microsoft, "Token Protection by using Microsoft Entra ID," Microsoft Community Hub, Jun. 2025. [Online]. Available: https://techcommunity.microsoft.com/blog/coreinfrastructureandsecurityblog/token-protection-by-using-microsoft-entra-id-/4302207

[2] Microsoft, "Microsoft Entra data retention", Microsoft Learn. [Online]. Available: https://learn.microsoft.com/en-us/entra/identity/monitoring-health/reference-reports-data-retention

[3] Microsoft, "Microsoft Entra monitoring and health FAQ", Microsoft Learn. [Online]. Available: https://learn.microsoft.com/en-us/entra/identity/monitoring-health/reports-faq

[4] Microsoft, "4688(S) A new process has been created," Microsoft Learn. [Online]. Available: https://learn.microsoft.com/en-us/previous-versions/windows/it-pro/windows-10/security/threat-protection/auditing/event-4688

[5] Microsoft, "Command line process auditing", Microsoft Learn. [Online]. Available: https://learn.microsoft.com/en-us/windows-server/identity/ad-ds/manage/component-updates/command-line-process-auditing

[6] Microsoft, "Enable and configure Sysmon in Windows", Microsoft Learn. [Online]. Available: https://learn.microsoft.com/en-us/windows/security/operating-system-security/sysmon/how-to-enable-sysmon

[7] SwiftOnSecurity, "sysmon-config: A Sysmon configuration file template with default high-quality event tracing," GitHub. [Online]. Available: https://github.com/SwiftOnSecurity/sysmon-config

[8] Microsoft, "10 essential insights from the Microsoft Digital Defense Report 2024", Microsoft Security Insider. [Online]. Available: https://www.microsoft.com/en-us/security/security-insider/threat-landscape/10-essential-insights-from-the-microsoft-digital-defense-report-2024

[9] Amazon Web Services, "TLS inspection configuration for encrypted traffic and AWS Network Firewall", AWS Security Blog. [Online]. Available: https://aws.amazon.com/blogs/security/tls-inspection-configuration-for-encrypted-traffic-and-aws-network-firewall/

[10] Microsoft, "Learn about auditing solutions in Microsoft Purview", Microsoft Learn. [Online]. Available: https://learn.microsoft.com/en-us/purview/audit-solutions-overview

[11] Microsoft, "Manage audit log retention policies", Microsoft Learn. [Online]. Available: https://learn.microsoft.com/en-us/purview/audit-log-retention-policies

[12] Amazon Web Services, "Logging data events", AWS CloudTrail User Guide. [Online]. Available: https://docs.aws.amazon.com/awscloudtrail/latest/userguide/logging-data-events-with-cloudtrail.html

[13] National Security Agency, "Adopting Encrypted DNS in Enterprise Environments," Cybersecurity Information Sheet, Jan. 2021. [Online]. Available: https://media.defense.gov/2021/Jan/14/2002564889/-1/-1/0/CSIADOPTINGENCRYPTEDDNSUOO102904_21.PDF

[14] Cybersecurity and Infrastructure Security Agency, "Microsoft Expanded Cloud Logs Implementation Playbook", CISA, Jan. 2025. [Online]. Available: https://www.cisa.gov/sites/default/files/2025-01/microsoft-expanded-cloud-logs-implementation-playbook-508c.pdf

[15] SANS Institute and Anvilogic, "The State of Detection Engineering 2026: What the Data Reveals About Accuracy, Automation, and AI Adoption," Jun. 2026. [Online]. Available: https://www.sans.org/white-papers/state-detection-engineering-2026

Kajal Dhanjal

Kajal Dhanjal

I build detections in my own Microsoft Sentinel lab and write up what holds and what doesn’t. About