Federated search sounds great on paper. Detection is where it breaks.

Security teams going up-market ask for federated search by name. It’s a reasonable request. All of the logs already live somewhere, whether it’s an S3 bucket, a Snowflake warehouse, or a dozen SaaS tools with their own APIs. Federated search promises to leave logs where they are and to push the question to the data instead of moving the data to the question. In practice, this means there is nothing to copy or migrate, and there is no second storage bill. For a team whose SIEM invoice is priced per gigabyte and already covers a tenth of what they would like to watch, it’s hard to argue with that pitch.
The goal behind federated search is right, but it falls short in some ways, especially now that AI swarms are compressing the time between intrusion and damage.
What follows is where the approach stops working, as we measured, and what has to change for the goal to hold. We argue that the solution is to keep the goal and change one thing about how it is met: read the logs once and build a federated index, then run every question and every detection rule against the index instead of the source.
What teams ask for
Federated search is a shorthand for three requirements:
- The source of truth stays where it is.
- No need to pay for a SIEM to store logs that are almost never queried.
- Every log is reachable during an investigation.
Every log being reachable includes the high-volume sources like VPC flow logs and DNS that get dropped first when the ingest bill climbs.
If an architecture fails any one of these then it’s not the answer to your problems. And an architecture that meets all three of these requirements under a name other than federated search is worth considering.
Federated indexing, not just search, is that architecture. It solves the fundamental problems with the bottlenecks that occur with federated search, and the rest of this post is the case for it.
Your investigation is as fast as your slowest source
Federated search breaks down in three places. The first is latency. This breaks differently based on where the logs live. Every query you make on your data has to wait for whichever source is the slowest to answer; results can’t be assembled until the last one is returned.
Logs in object storage queried through an engine like Trino are fast only if someone has imposed a tabular structure ahead of time so the query engine has tables to work with. Because logs have flexible schemas, if fields were not defined when the tables were being structured, those fields end up being invisible to the investigation. Logs in SaaS tools are reachable only through the tool’s API, which was built for the tool’s own use and comes with rate limits, throttling, and retention windows that were not designed for a security investigation. Logs in warehouses like Snowflake or BigQuery are reachable through a system built for large analytical questions that are occasionally asked. A security investigation is the opposite kind of workload consisting of many small queries in quick succession and each query building on the last.
A security engineer who is trying to follow a lead through five or ten pivots has to pay for latency at every step. When they are only halfway there they end up losing the context for what makes the next question worth asking.
Latency is the failure a security team notices first. The next failure broadens who feels the pain, as platform engineering’s invoice balloons.
The bill lands on the wrong team
The second break is organizational. When the security team’s queries run against the platform team’s warehouse, they compete with the platform team’s own workloads for the same compute, and the cost shows up on a budget the security team does not own.
Because no one planned for the spend, the platform team’s first move is to ask what the security team is doing, and its second is to ask them to do less of it. Instead of easing the friction between two teams, going with federated search exacerbates it.
Both of these are survivable. A team can pad a query with patience and a budget with a conversation. The third failure is the one there is no working around.
Detection is where it falls over
A security investigation involves asking questions once, but detection is asking the same question over and over for years. A detection library is a loop: in our benchmark, 1,175 rules re-evaluated every minute, which comes to more than 1.6 million evaluations a day. Federated search cannot carry this workload because it re-does the expensive part of the work, which is finding and reading the relevant logs for every continuous rule.
Our team at Scanner measured this. We took the free, public SigmaHQ Windows process-creation rule pack, translated 1,175 of its 1,182 rules into Snowflake SQL, verified the translations against a known capture, and ran the pack once a minute against one terabyte a day of endpoint telemetry. The smallest warehouse that kept pace with a 60-second cycle was a Large, at $11,680 a month on list pricing. This is because a rule set that fires every minute never leaves the warehouse idle long enough to pause, causing a 24/7 workload in practice. To detect on 1 TB of new logs a day, that warehouse scanned between 85 and 416 TB a day, because each recent minute on every cycle needed to be re-read for every rule.
Then we tried the workarounds a capable team would try. Folding all 1,175 rules into one query cut the data scanned by 1,200 times and was slower regardless of warehouse size. Rebuilding on Snowflake’s own incremental tooling, streams and tasks, produced a task that reported success on every run while processing rows at 73% of the rate they arrived, so detection drifted 20 to 40 minutes behind live with no error reported anywhere. Turning on the Enterprise-tier performance features meant repricing every credit the company buys by 50%. The rules that were slow stayed slow, because those features speed up selective lookups, and the slow rules are broad pattern matches where there is nothing to skip.
A security leader should be most concerned about the fact that no one gets paged when there is such a large drift. Across roughly 50,000 rule queries in every configuration that fell behind, none returned an error, and every dashboard showed success while detection continued to get staler.
The three failures have one cause
Slow investigations, surprise bills, and ineffective detection share a single cause. A detection library asks thousands of questions a minute for years, and federated search must pay the full cost of locating and reading the logs on every single question. Because the cost lives in the read pattern, there is nothing you can do at query time or with warehouse sizing.
Federated indexing changes the read pattern. It reads each log once, wherever it lives, and builds a compact index in the customer’s own storage bucket. Every later question, and every detection rule, runs against the index instead of the source, which never moves. The expensive part of the work happens once, at ingest, and after that cost scales with how much data arrives, not with how many rules you run or how far back they look. A 30-day behavioral baseline costs about the same to evaluate as a 30-minute one, because you don’t need to access and read the logs again.
But… you moved my data
This is the objection federated indexing has to answer, and it is fair. The index is a copy. It is also a small fraction of the raw size, it lives in the customer’s own bucket, and it can be deleted or rebuilt from the source at any time.
In our benchmark, there was no measurable cost to export logs from Snowflake to S3 on the warehouse that already handled ingest. A federated search that copies nothing detects nothing. Between a copy that costs nothing to make and a no-copy rule that costs five figures a month to honor, the copy is the easier position to defend.
Federated search for the question you ask once. Federated indexing for the questions you never stop asking.
