What is federated indexing?

Two construction workers walk into a museum gallery. From inside the gallery, nothing about this looks strange. Construction crews come through museums all the time. What is unusual here is how they enter: cutting open a window with a disc cutter. But there are no cameras on the outside of the building. The security team learned a heist was underway when the thieves tripped the gallery alarms and threatened the guards. By the time backup arrived, the thieves were gone with the French Crown Jewels.
The Louvre took a lot of heat after the October 2025 heist. There were no security cameras on the outside, and only 39% of its rooms were covered by CCTV, with some of those cameras badly positioned. The behavior that would have identified the heist was visible the whole time, just from a vantage point that wasn’t being monitored.
For all the criticism the Louvre took over its inadequate monitoring, this is exactly what cybersecurity looks like today: most organizations can watch only 10% of their security telemetry, and the other 90% remain a blind spot where attackers can slip by unnoticed. This month, OpenAI disclosed six recent incidents that indicate that the Hugging Face swarm attack was not a standalone event. These rogue agents are getting through in a similar way to how the thieves did: enter where no one is looking. The logs for the Artifactory package manager were not monitored, so the AI swarm built a communication network there and coordinated a significant attack. We found out about the attacks weeks or months later, long after the crown jewels were gone.
The nature of offensive agents is changing, and the defensive posture has to change with it. No attack should succeed because logs and telemetry were being thrown away. Federated indexing is the structural answer to this industry challenge.
The real cost of the status quo
The monitoring blind spot starts to appear with how the SIEM is priced. Ingest has a significant price per gigabyte, often in the range of a few dollars USD, so security teams can only feasibly send a small share of their total log volume into the SIEM. EDR alerts make it into the SIEM, but the raw telemetry behind those alerts is rarely included. VPC flow logs and DNS logs are usually the first sources cut, and because ingest has been expensive for years, most teams have never had the opportunity to monitor anomalies in that telemetry in the first place. When an incident finally forces the question of how far an attack reached, the data that would answer it was discarded months earlier as a straightforward budgeting decision.
Security teams that have experienced this blind spot and the pain of not knowing have tried to work around it in two ways.
The first is to move the logs into a data lake. Teams build a pipeline that extracts logs, transforms them into SQL tables, and loads them into a traditional data lake like Amazon Athena or Google BigQuery. The ingestion economics of this genuinely work. Object storage sits underneath, so cost per gigabyte falls well below SIEM ingest pricing, and there is no longer a budget reason to leave a source out. The cost moves to data engineering instead. Logs fit awkwardly into SQL tables, so getting them there is substantial work, and the work does not end at the initial load, because vendors add and remove fields and the schemas have to be maintained as they drift. ETL pipelines like this can have ingestion delays of tens of minutes to hours. What the team has at the end is slow. Queries can take minutes or hours, query costs ratchet up as more people use the system, and continuous detection querying is not something the architecture supports.
The second is federated search, which removes the infrastructure entirely by connecting to data wherever it already lives and querying it live whenever a question comes up. For example, a federated search tool might query a data warehouse like Snowflake over here, a SaaS tool like Okta over there, and so on, giving you the ability to query many tools from one place. For the occasional question asked during an investigation, this works well enough, although your search query’s speed will be limited by the slowest tool you are pulling from. Another difficulty is that modern detection consists of hundreds of standing rules evaluating continuously. The underlying tools that federated search interacts with were never designed for that kind of sustained load. Continuous detection queries can cause Snowflake compute costs to skyrocket, and SaaS tool rate limits to slow your detections to a crawl. This causes your detection rules to run late or not at all. Federated search was not designed for continuous detection at scale, so it is inadequate to keep you covered in threat detection.
The two approaches differ in whether they keep a second copy of the data, but they fail in the same place. Both make full coverage affordable. Both pay for it by sacrificing the two things the SIEM was bought for: queries fast and powerful enough to be usable for serious security work, and detections that run continuously. Federated indexing gives you the best of both worlds: the economics of a data lake, and the query and detection power of a SIEM.
What changes when you index instead of search
Federated indexing reads each log once, wherever it lives (S3, GCS, Azure Blob Storage, SaaS APIs, Kafka or Kinesis streams), and writes it into a single compact index stored in object storage. The index is highly optimized for minimizing storage size, search speed, and high-scale threat detection. The raw sources stay where they are and remain the system of record. The index contains a compressed copy of the logs with high-speed index structures built in, which comes in around 15% of the raw volume and sits in storage billed at commodity rates, rather than a full duplicate priced per gigabyte by a vendor.
Anyone can read an S3 pricing page - so it is clear to see that the storage costs will be quite low compared to a traditional SIEM. The reason the cost is low is the same reason why S3 has high per-request latency: it runs on magnetic spinning hard-drives with mechanical moving parts, so there can be a big delay between when you request data and when it starts streaming back. Time to first byte on S3 runs to hundreds of milliseconds regardless of how much data the request returns, and a conventional inverted index, which depends on a large number of small reads, performs poorly against that profile.
An index purpose-built for object storage has to reverse the pattern: a small number of large reads, parallelized across many workers, with the indexing and querying paths kept from interfering with one another as load rises. Each of those requirements is well understood on its own. Meeting all three simultaneously, at the scale of hundreds of TB of logs ingested per day, with continuous detection load, is an engineering path that is not for the faint of heart. But if you can achieve it, the results are profound: fast petabyte-scale search at low cost per GB, meaning that it becomes feasible to search and monitor the entire estate of logs you have, not just the lucky 10% that you can afford to ingest into the SIEM.
Standing detections run inside the indexing pipeline as data arrives, with a streaming engine that computes a partial result from each batch of events and caches it in S3 so that subsequent evaluations build on prior work rather than starting over. The source is read once, at ingest, and never re-scanned for the sake of a rule. To compare recent activity against a 30-day baseline, you do not need to re-run a full 30-day query: you can rapidly query the interval tree cache. Ad-hoc questions are served by ephemeral compute that spins up, reads the relevant log segments in parallel, and returns results in seconds.
The result is an architecture with two distinct parts: infrastructure that runs continuously to write new data into index files and serve standing detections, and on-demand compute for everything else, both operating over a copy of the data small enough to be indefinitely retained. The SIEM indexes only the fraction of logs the budget allowed in. The conventional data lake holds everything in object storage but is hard to load, answers slowly, and cannot run standing detections. Federated search falls apart under continuous load. Federated indexing covers the full log volume and holds up under the pressure of agentic workloads.
Why we built Scanner
Before this problem had a name, we were living it. The architecture in this post is what we built when we could find no other solution. Scanner indexes logs directly from S3, runs full-text search that feels instant, and evaluates detections continuously as new data lands.
Our customers inspire us, and they have put the case to us more plainly than we had put it to ourselves: no cyberattack should succeed because of unmonitored logs. For most of SIEM’s history there was no way to honor that principle, because monitoring everything cost more than any security budget could carry, and teams were forced to accept a coverage gap that was based on cost rather than security. Just as an art heist succeeds where security cameras aren’t pointed, AI swarm attacks will operate in more obscure parts of the infrastructure where their activity is only captured in logs that nobody watches. Our obsession is to make that scenario a thing of the past. In the future, it will be easy and comfortable to use 100% of your “security cameras,” monitor all of your telemetry and logs for threats, and not the mere 10% that is too common today. Our mission is not over until we and the industry make that a reality.
