Wazuh + AWS Bedrock: S3 Vectors for SOC Knowledge (Part 4)

Moving the knowledge corpus out of the Indexer

Part 3 ended with a condition, not a recommendation: the vector index belongs inside the Wazuh Indexer while the corpus is small and the clients already speak OpenSearch, and belongs somewhere else when the vector engine matters, when the corpus outgrows a single node, or when alert ingestion already claims that node’s CPU and heap. This part tests the other side of that condition. The same four sources went into Amazon S3 Vectors through an Amazon Bedrock Knowledge Base, the in-Indexer branch was rebuilt from the same snapshot so both could be measured on the same day, and the same frozen questions ran against both.

Three things changed, and only one of them is the improvement you would expect. Retrieval coverage went up. The exact indicator path went away. And an ingestion job reported success while three documents in the index were not what I had uploaded.

Architecture diagram: a frozen snapshot of 7115 chunks is uploaded to a general purpose S3 bucket, a Bedrock Knowledge Base syncs that prefix through Titan Text Embeddings V2 into an S3 vector index, a client holding the reader role retrieves passages and reads whole playbooks, and a separate generation identity calls Claude with the stored context

The measured setup: two buckets, the Knowledge Base between them, and two client identities, because retrieval and generation do not need the same permissions

This series has four parts.

  • Part 1 - ML Commons, the Bedrock Claude connector, and the Dashboard chat agent
  • Part 2 - opensearch-mcp-server-py as a sidecar, and Claude Desktop over MCP
  • Part 3 - Titan Embeddings V2, a lucene k-NN index over four sources, and hybrid retrieval
  • Part 4 (this article) - the same corpus in Amazon S3 Vectors through Bedrock Knowledge Bases, measured against the rebuilt in-Indexer index

Every number below comes from runs on 2026-09-11 and 2026-09-12 in us-east-1, from one client, against one corpus of 7115 chunks and 36 frozen questions. The lab was deleted on 2026-09-12 and an independent listing of the account came back empty, so nothing here can be re-measured without rebuilding it.

Rebuilding a comparable baseline

The stand from Part 3 no longer existed. It was removed on 2026-09-06 together with its volumes, which means the 7118-document index Part 3 measured is historical evidence and not a baseline a new measurement can be compared against. The comparison had to be rebuilt, not resumed, and that is the first thing anyone repeating this should budget for.

The corpus was captured once and replayed offline. Each source body was fetched a single time and stored with its digest, and the chunks were rebuilt from those stored bytes through the loaders of the original lab; two replays produced byte-identical chunk files. Chunk size stayed at 3000 characters with 200 characters of overlap, unchanged from Part 3, so both branches split the text the same way and a difference in results cannot be blamed on a difference in boundaries.

The final corpus holds 7115 chunks: MITRE ATT&CK 1031, Wazuh documentation 3851, playbooks 24, MISP events 2209. Part 3’s index held 7118. The three missing chunks are not an accounting error, and the next section is about them.

The S3 and Bedrock path

The store is two buckets and one knowledge base. A general purpose bucket holds chunks/ with one .txt and one .metadata.json per chunk, and sources/ with the whole playbooks. A vector bucket holds the index soc-knowledge, float32, 1024 dimensions, cosine. The knowledge base reads the chunks/ prefix with chunkingStrategy NONE, because the corpus was already split by the Part 3 loader and letting the service split it again would have destroyed the parity the comparison depends on.

The service role is created before the knowledge base that it trusts, so its trust policy starts with a wildcard on knowledge-base/* and is narrowed to the exact ARN once the base exists. Between those two steps there is a race worth a bounded wait: CreateKnowledgeBase called about one second after CreateRole failed with “unable to assume the given role”, and the same call at 17:35:51Z went through on the first attempt. Retry it; it is not a permissions bug.

Each chunk carries five custom metadata keys. Through a knowledge base the budget is 1 KB and 35 keys per vector, and the largest custom metadata in the pilot batch was 345 bytes, so the headroom is real but not unlimited. The control that matters here went the other way on purpose: two documents with 1161 bytes and 36 keys of metadata. The ingestion job reported COMPLETE, 25 scanned, 0 new, 0 failed, 0 skipped, and ignored both documents. The only trace was the text in failureReasons. A job status of COMPLETE with zero failures does not mean every document was indexed.

The full load scanned 7118 documents and 7118 metadata files, indexed 7106 new ones, failed none, and took 1929 seconds.

Then I read the stored text back, and three documents did not match. They had been uploaded as .txt with Content-Type: text/plain; charset=utf-8, and they were stored as text converted from HTML: tags removed, paragraph breaks collapsed into spaces, a curly quote straightened, backslashes added. An XML example <reports> and placeholders in angle brackets such as <WAZUH_INDEXER_USERNAME> were gone from the stored text entirely. 2771 bytes were stored as 2173, 2967 as 1762, and 455 as 194.

Grouped bar chart of three chunks, uploaded bytes against stored bytes: 2771 to 2173, 2967 to 1762, and 455 to 194

What the readback found after a job that reported COMPLETE with zero failures: three documents stored shorter than they were uploaded.

The three are exactly the chunks whose text contains the string <title. Of 1519 chunks carrying tag-like tokens, only those three changed. To rule out the S3 connector I ingested the same three plus the 12 pilot chunks into a separate knowledge base through a CUSTOM data source, sending the text inline. The stored text came back byte-identical to the main index, and all 15 vectors matched it at cosine similarity 1.0. These runs do not establish whether the embedding is computed before or after that conversion; separating the two would take embedding the original and the converted text directly and comparing both against the stored vector, and after the lab was deleted there are no vectors left to try it on.

The documentation says the default parser handles text in .txt, .md, .html and other text files, but it does not describe how the type is detected, so <title is an observed correlation and not a documented rule. I excluded those three chunks from both branches. The corpus went from 7118 to 7115, no question in the frozen set depends on them, and a second sync scanned 7115 documents, deleted 3 from the index, reported 0 new, 0 modified and 0 failed, and finished the readback in 275 seconds.

Retrieval and complete playbook reads

The client retrieves as a reader role holding bedrock:Retrieve on the knowledge base and s3:GetObject on sources/. It holds nothing on the embedding model, and it does not need to: the knowledge base embeds the question itself. Generation is a different identity, because Converse requires bedrock:InvokeModel, which the reader role does not have, so the 33 generation calls ran under the operator identity. One client, two identities, and the diagram above splits them for that reason.

The operational question was the one a tier 1 analyst asks mid-incident: what does our runbook prescribe for an SSH brute force against web-server-01. The knowledge base returned two of that playbook’s six chunks in its top five, at positions 1 and 3; the other three were a chunk of a different playbook and two chunks of the documentation page on active response. The in-Indexer neural search returned no playbook chunk at all in its top five - four chunks of that same documentation page and one MITRE technique.

Neither top five contains the whole Containment section, which is what the analyst actually asked for. A better score would not fix that, because the section is longer than the passages a retrieval layer returns. What fixes it is that each chunk carries its doc_id, so the client reads the whole playbook out of sources/: 1906 characters, SHA-256 equal to the captured file.

With that full document in the context the model answered the question operationally - block the source address on the firewall of web-server-01 through the firewall-drop active response, for 12 hours - and cited the passages it used. The five knowledge base passages produced the same answer. The in-Indexer neural passages produced a refusal: the model said there is no runbook for web-server-01 in the context, which was an accurate statement about what it had been handed.

What changes for MCP and the Dashboard agent

The Part 2 sidecar does not carry over, and this is worth checking before you plan around it. The installed MCP server (package 0.11.0, serverInfo.version 1.29.1, protocol 2025-06-18) exposes nine tools: ListIndexTool, IndexMappingTool, SearchIndexTool, GetShardsTool, GenericOpenSearchApiTool, ClusterHealthTool, CountTool, MsearchTool and ExplainTool. All nine talk to OpenSearch. GenericOpenSearchApiTool will call an arbitrary OpenSearch API path, but not an AWS service, so there is no route to a Bedrock knowledge base through any of them.

The path therefore needs an adapter you write. The lab’s own stdio server on the official MCP SDK exposes a single tool, soc_knowledge_retrieve, taking question, source, version and k. Asked the same playbook question through an SDK client it returned the same five passages as the direct client; with source=playbook all five came from playbooks; and an invalid filter value arrived at the client as a tool error, not as an empty result, which is the behaviour you want when a filter name drifts.

The Dashboard path runs through ML Commons, with a limit that decides whether it is useful to you. An aws_sigv4 connector with service_name bedrock against POST /knowledgebases/{id}/retrieve, signed with the reader role’s temporary keys including session_token, plus a remote model and a flow agent with MLModelTool, executed on the first attempt: inference_results present, 1063 ms, the same five passages as the direct client. But MLModelTool passes the Retrieve response through unprocessed, so the agent hands back JSON containing retrievalResults, location and metadata instead of an answer.

Two boundaries I did not cross: the Dashboard chat binding and a check in the browser were not done, and the reader role’s temporary keys live one hour, so anything built this way needs a credential refresh story before it is more than a demonstration.

Semantic search and exact indicators

Coverage is the measurement the move is supposed to improve, and it did. Over 24 English questions at k=5 the knowledge base put a supporting passage in the top five for 21 of them, MRR 0.778, against 18 of 24 and 0.681 for the in-Indexer neural search. At k=10 it was 23 of 24 (0.791) against 21 of 24 (0.707). Over 8 Russian questions the gap was wider and identical at both depths: 7 of 8 (0.667) against 5 of 8 (0.417).

Two panels showing questions with a supporting passage in the top k, and mean reciprocal rank, for the Knowledge Base against the in-Indexer neural search in four conditions

Coverage and MRR over the frozen question set. The Knowledge Base leads in every condition, and the gap is widest on Russian.

The indicator control is where the managed store stops being equivalent. Part 3’s hybrid query found the carrier of a SHA-256 at position 1, and only for the exact value: with the last digit changed, and with a value that is not in the corpus at all, the carrier was not in its top five. That is what you want from an indicator lookup, because a near miss is a miss.

The knowledge base returned the same carrier at position 4 for all three values - exact, near and absent. It answered by topic, not by value, which is the same thing as not answering the question. The in-Indexer neural search returned the carrier for none of the three.

Matrix of three retrieval paths against three indicator values: the hybrid query returns the carrier at position 1 for the exact value only, the Knowledge Base at position 4 for all three, and the neural search never in the top five

The same SHA-256 asked three ways. Only the hybrid query separates an exact value from a near miss.

Generation does not rescue it. Across all six indicator contexts the model said the hash could not be determined from the context, and for the exact value that was correct about what it had been given: the indicator is recorded only in the indicators metadata field of the carrier chunk - a MISP event with 180 characters of text and 6 indicators - and the retrieved passages carry text only. The hash never reached the prompt. A semantic store has no exact lookup, so if your workflow asks whether you have seen a hash before, keep a term clause somewhere that can answer it.

Freshness, missing sources, and lifecycle

A retrieval layer that always returns five passages will return five passages when the right answer is not in the corpus at all, so the controls are paired on purpose.

For staleness, a control base held Wazuh documentation 4.7 and 4.14 together, 23 chunks. Without a filter the top five were three 4.14 chunks and two 4.7 chunks, with the renamed capability on both sides: vulnerability-detection second at 0.7932 and vulnerability-detector third at 0.7852, all five inside the band 0.7591 to 0.7946. No score separates the current document from the superseded one. With version=docs-4.14 all five came from 4.14, and with version=docs-4.7 all five from 4.7. A metadata filter is the thing that separates them, not ranking.

For a missing source, the corpus was replaced with 1098 chunks that do not contain the answer and asked the same question. None of the top five contains either term, and the scores sit between 0.7546 and 0.7945 - the same band as the run that did contain the answer. Absence does not show up as a low score, which is the single most useful thing to know before trusting a threshold.

Access control behaved: the reader role, holding Retrieve on the main base only, got AccessDeniedException against the control base, and generation was never called.

Updates behaved too. Editing the text of one chunk and deleting another, then syncing, reported 1 modified and 1 deleted with no failures, and the readback over the remaining 1097 chunks confirmed that the new text is what is stored. A second sync with nothing changed reported 0 new, 0 modified and 0 deleted.

Deletion was the part that cost me three attempts. A data source whose dataDeletionPolicy is DELETE removes its vectors gradually while it reports DELETING, at roughly 46 vectors a minute on this corpus: 352 of 1121 were left after three minutes in the control index, and 4283 of 7115 after about an hour in the main one. Two bounded waits, 900 and 3600 seconds, expired on that. Deleting the vector index instead cleared the remaining 4283 at once, but it also pulled the store out from under the running deletion, and the data source and then the knowledge base went to DELETE_UNSUCCESSFUL. The service names its own remedy in failureReasons: set dataDeletionPolicy to RETAIN and retry. After that the data source went in about five minutes and the knowledge base in five more. The lab finished with three indexes, a vector bucket, a documents bucket of 16428 objects and four roles deleted, and an independent listing of the account found no lab name left.

Measured cost and latency

Latency, English, k=5, 24 questions, 10 repeats after 3 warmups, 0 failures in 720 calls: the knowledge base at p50 481.4 ms and p95 658.8 ms, the in-Indexer neural search at p50 279.7 ms and p95 696.4 ms. Russian, 8 questions and 160 calls: 477.8 and 517.1 against 276.0 and 304.9. The knowledge base is a network call to us-east-1 and the Indexer was local to the client, so this compares two deployments and not two algorithms.

Dot plot of p50 to p95 latency for five measurement series, ranging from 276 to 1359 milliseconds

Latency from one client. The Knowledge Base is a call to us-east-1 and the Indexer was local, so this compares deployments.

One number I cannot explain: a direct Titan call from the same client, embedding the same question, took p50 706.7 ms, which is slower than the knowledge base’s Retrieve that embeds the question itself and then searches. I am reporting it and leaving the theory alone.

The whole experiment cost between 0.46 and 0.75 USD by receipt: Titan embeddings 0.04 to 0.09 over 2.11 to 4.44 million tokens for both branches, the control base and the probe; S3 PUT requests 0.08; S3 reads 0.02; S3 Vectors writes 0.01 to 0.22 for 8255 vectors; generation 0.31 to 0.34 depending on which Claude tariff applies. Retrieval requests came to less than 0.01, and deleting resources costs nothing.

Horizontal bars of cost by component: generation 0.31 to 0.34 dollars, S3 Vectors writes 0.01 to 0.22, S3 PUT requests 0.08, Titan embeddings 0.04 to 0.09, S3 reads 0.02

Where the bill went. Storage and search are the cheap part; the model is not.

Generation dominates that total, and it is not close: 33 calls at temperature 0 with maxTokens 800 produced 71595 input and 6505 output tokens. Storage and search are the cheap part of a knowledge base; the model is the bill.

Temperature 0 did not make the answers reproducible. Three repeats matched word for word in only one of the eleven contexts, and the output token count moved with the text - 194 and 310 tokens for the same context in two repeats. If you plan to diff model answers between runs, that is the noise floor you start from.

Which corpus belongs outside Wazuh

Move the corpus out when semantic coverage is what you are short of, when you would rather not have embedding and k-NN competing with alert ingestion for one node’s heap, and when a managed store with per-vector metadata filtering is worth a network hop of roughly 200 ms over a local index. Keep it in when the questions are about values, not topics, because the exact indicator path exists only where a term clause does, and a knowledge base on S3 Vectors has none.

The honest answer for a SOC is that these are two stores with different jobs, and this experiment measured the cost of running the semantic one at 0.46 to 0.75 USD for a corpus of 7115 chunks, which is small enough that the choice is rarely about money.

One rule applies to either store, and it is the one I did not expect to be writing: an ingestion job that reports COMPLETE with zero failures is not evidence that the index holds what you uploaded. Read the stored text back, and read failureReasons even when nothing failed.


See also