Run RAG Without Sending Data to a Public AI

Run RAG Without Sending Data to a Public AI

Summary

  • Regulated enterprises can run private RAG entirely inside their AWS VPC by pairing a customer-managed Bedrock Knowledge Base with an OpenSearch Serverless collection that has a VPC-only network policy.
  • Bedrock Knowledge Bases do not require AWS Agents; the AgentsforBedrockRuntime.retrieve and retrieve_and_generate APIs allow direct, deterministic, auditable retrieval and generation.
  • Use a fixed chunking strategy and enable APPLICATION_LOGS so ingestion and retrieval results are reproducible and audit-ready.
  • Bedrock ingestion limits (1,000 files per sync, 25 files manual, 5 concurrent requests) must be managed with batching and retries.
  • Jinba Flow orchestrates the full private pipeline — ingestion, retrieval, generation, and logging — inside your VPC.

Regulated enterprises in finance, healthcare, and government face a concrete blocker when evaluating Retrieval-Augmented Generation: the documents that would make RAG valuable are exactly the documents that cannot leave the environment. Sending contract terms, patient records, or internal policy to a public AI endpoint is not a trade-off to weigh. It is a compliance violation.

The result is a familiar stall. Teams either skip RAG entirely or build workarounds that fragment the workflow and introduce new risk. Neither outcome is acceptable when the underlying technology is otherwise ready.

The architecture that resolves this combines an AWS Bedrock Knowledge Base with a private vector store and an orchestration layer that enforces deterministic execution and records every step. Jinba provides that orchestration layer, running the full pipeline — document ingestion, embedding, retrieval, and generation — inside the organization's own AWS VPC.

The Compliance Barrier

The constraint is not just legal. It is architectural. Public AI services process queries on shared infrastructure. Even when a provider offers a data processing agreement, the data transits a public endpoint and is handled by systems outside the enterprise's control boundary. For workloads governed by GDPR, HIPAA, or sector-specific mandates, that boundary matters independently of what the provider does with the data.

Enterprises evaluating cloud AI services for internal document search consistently run into a second problem: the AWS service model itself is not immediately transparent. A recurring pattern in enterprise evaluations is the assumption that querying an AWS Bedrock Knowledge Base requires AWS Agents. It does not. The AgentsforBedrockRuntime.retrieve API and the AgentsforBedrockRuntime.retrieve_and_generate API are available for direct programmatic use. The Agent layer is optional orchestration, not a required dependency.

This distinction matters for compliance architecture. When the retrieval call is made directly and deterministically, rather than delegated to an agent that decides its own path, the query, the retrieved chunks, and the generated response can all be logged explicitly. That is the audit trail regulators expect.

AWS Bedrock Knowledge Base: The Foundation

Amazon Bedrock Knowledge Bases is Amazon's managed RAG service. It connects an LLM to an internal data source by handling chunking, embedding, indexing, and retrieval. Enterprises do not build or maintain the vector pipeline. Bedrock manages it.

There are two deployment models, and the choice between them determines how much control the enterprise retains over where data rests.

Managed Knowledge Base: Amazon Bedrock AgentCore manages the underlying vector store and infrastructure. The model reached general availability in June 2026. It supports KMS encryption and removes infrastructure overhead, but it offers less configurability over vector search and network placement.

Customer-managed Knowledge Base: The enterprise supplies and owns the vector store. This model is the correct choice for regulated workloads. The enterprise provisions the vector index, configures its network policy, and retains full control over where vectors are stored and how they are accessed.

For compliance-sensitive deployments, AWS guidance points to a specific configuration: a customer-managed knowledge base paired with an Amazon OpenSearch Serverless collection, with a private network policy applied to the collection. This configuration keeps the vector index inside the customer's VPC. No vector data is accessible over a public endpoint.

A technical clarification worth noting: traffic between Bedrock and S3 travels on the AWS private backbone network even when public endpoints are used. That said, applying a private network policy to the OpenSearch Serverless collection is the configuration that satisfies strict compliance mandates, because it enforces isolation at the vector store level rather than relying on transit behavior.

The Jinba Workflow: Deterministic by Design

Jinba is an orchestration layer, not an LLM agent. The distinction is consequential for regulated enterprises. An agent decides its own execution path at runtime. Jinba executes a defined workflow with a fixed sequence of steps. Every ingestion job, retrieval call, and generation request follows the same path and produces a log that reflects exactly what happened.

The pipeline breaks into four steps.

Step 1: Configure the private foundation

Create a customer-managed AWS Bedrock Knowledge Base backed by an Amazon OpenSearch Serverless collection. During collection setup, apply a network policy that restricts access to the VPC. No traffic to the vector index should route over a public endpoint.

The Bedrock API primitives for this stage are create-knowledge-base and create-data-source. The data source points to an S3 bucket that also resides within the private network boundary.

Step 2: Automate data ingestion

Once the knowledge base and data source are configured, Jinba orchestrates ingestion by calling the StartIngestionJob API on a defined schedule or in response to document events.

A known operational constraint here: Bedrock Knowledge Bases quotas cap a single sync job at 1,000 files, limit manual ingestion to 25 files per request, and allow no more than 5 concurrent requests per knowledge base. These limits do not prevent large-scale ingestion. They require it to be managed. Jinba handles batching, sequencing, and retry logic so that document sets larger than a single job's capacity are processed reliably without manual intervention.

Chunking strategy is a configuration decision that directly affects retrieval determinism. Bedrock Knowledge Bases support Default, Fixed-size, and No-chunking strategies. Selecting a specific strategy means the same source document always produces the same chunks. That consistency is what makes retrieval reproducible and audit results meaningful. A fixed chunking configuration applied at ingestion time removes ambiguity about what the model saw when it generated a given answer.

Step 3: Private retrieval and generation without agents

Querying the knowledge base does not require AWS Agents. Jinba calls the retrieval APIs directly.

AgentsforBedrockRuntime.retrieve returns the document chunks most semantically relevant to the query. Those chunks are then passed as context to an LLM hosted on Bedrock — Claude, Llama 3, or another model available in the region. The generation step runs on Bedrock infrastructure within the same AWS environment.

For lower latency, AgentsforBedrockRuntime.retrieve_and_generate combines both operations into a single API call. Retrieval and generation complete in one request rather than two. Jinba selects the appropriate call based on the workflow configuration.

The critical privacy property: at no point in this sequence does any data transit a public endpoint. The document chunks live in the private OpenSearch Serverless collection. The LLM is Bedrock-hosted. The API calls route within AWS infrastructure.

Step 4: End-to-end audit logging

Bedrock Knowledge Bases support ingestion logging via APPLICATION_LOGS. These logs record the file-level status of every document processed in an ingestion job. Log delivery is configured using PutDeliverySource, PutDeliveryDestination, and CreateDelivery, with delivery targets including CloudWatch Logs, Amazon S3, or Firehose. Note that log delivery is not supported for knowledge bases created with a structured data store.

Because Jinba orchestrates the retrieval and generation calls explicitly, the query, the retrieved chunks, and the generated response are all available for logging at the workflow level. The result is a complete record: what documents were ingested, when, at what status; what query was submitted; what context was retrieved; and what the model returned. That record is what a compliance audit requires.

Comparing the Alternatives

Public AI services are the baseline for comparison. Sending internal documents to ChatGPT or a similar public endpoint provides no control over where the data is processed, no visibility into what the model retains, and no audit trail. For regulated workloads, this is not an architecture choice. It is a disqualification.

DIY AWS-native orchestration is a legitimate alternative. AWS has published a reference pattern that combines EventBridge Scheduler and Step Functions to automate knowledge base sync. The pattern works. It also requires building and maintaining the scheduling logic, error handling, batching coordination, and any logging or governance layer on top. None of that is provided by the pattern itself.

Jinba delivers the same orchestration capabilities with built-in audit logging, role-based access control, and deterministic workflow execution. The development and maintenance overhead of the DIY approach is the cost Jinba replaces.

What to Build Next

A private, auditable RAG system on AWS is achievable with three components in place: a customer-managed AWS Bedrock Knowledge Base, an Amazon OpenSearch Serverless collection with a private network policy, and an orchestration layer that enforces determinism and logs every operation.

For enterprises ready to proceed:

  1. Provision an OpenSearch Serverless collection and apply a VPC-only network policy before creating the knowledge base.
  2. Create the customer-managed knowledge base using create-knowledge-base and create-data-source, pointing the data source at a private S3 bucket.
  3. Configure a fixed chunking strategy on the data source to ensure retrieval reproducibility.
  4. Enable APPLICATION_LOGS delivery to CloudWatch Logs or S3 before running the first ingestion job.
  5. Use Jinba to schedule ingestion jobs, manage batching against the 1,000-file sync limit, and orchestrate retrieval and generation via the direct Bedrock APIs.

The architecture keeps every component inside the AWS environment the enterprise already controls. Regulated enterprises do not have to choose between compliance and capable AI over internal data. The choice is in how the pipeline is built and governed.

Frequently Asked Questions

Does AWS Bedrock Knowledge Base require AWS Agents for private RAG?

No. AWS Bedrock Knowledge Bases can be queried directly with the AgentsforBedrockRuntime.retrieve and AgentsforBedrockRuntime.retrieve_and_generate APIs. AWS Agents are an optional orchestration layer, not a dependency. Direct API calls make the retrieval path deterministic and easier to log for compliance.

Can a Bedrock Knowledge Base be fully private?

Yes, but only with a customer-managed knowledge base paired with a vector store such as Amazon OpenSearch Serverless configured with a private network policy. This setup keeps vectors inside your VPC and prevents access over a public endpoint. AWS guidance points to this configuration for private, compliance-sensitive workloads.

What is the difference between managed and customer-managed Bedrock Knowledge Bases?

A managed knowledge base uses Amazon Bedrock AgentCore to manage the vector store and infrastructure, reducing operational overhead but limiting control over network placement and vector search. A customer-managed knowledge base requires you to provision and own the vector store, which gives you the control needed to enforce VPC-only access and meet strict compliance requirements.

Which vector store should I use for a private Bedrock Knowledge Base?

For regulated workloads, the recommended approach is Amazon OpenSearch Serverless with a private network policy. This keeps the vector index inside your VPC and isolated from public endpoints. It is the configuration AWS guidance points to for fully private Bedrock Knowledge Bases.

How does Jinba make RAG audit-ready?

Jinba runs the full pipeline — ingestion, retrieval, and generation — inside your AWS VPC and logs every step explicitly. It calls the Bedrock APIs directly, so the query, retrieved chunks, and generated response are all recorded at the workflow level. Combined with Bedrock ingestion logging, this creates an end-to-end audit trail.

What are the ingestion limits for AWS Bedrock Knowledge Bases, and how does Jinba handle them?

Bedrock caps a single sync job at 1,000 files, manual ingestion at 25 files per request, and concurrent requests at 5 per knowledge base. Jinba manages batching, sequencing, and retry logic so larger document sets are processed reliably without manual intervention.

Why is a fixed chunking strategy important for compliance?

A fixed chunking strategy ensures the same source document always produces the same chunks. That consistency makes retrieval reproducible and audit results meaningful. You can prove exactly what context the model received for each generated answer.

Is traffic to Bedrock and S3 private by default?

Traffic between Bedrock and S3 travels on the AWS private backbone network even when public endpoints are used. However, for strict compliance you should still apply a private network policy to the OpenSearch Serverless collection, because that enforces isolation at the vector store level rather than relying only on transit behavior.

Is a private Bedrock Knowledge Base architecture compliant with HIPAA or GDPR?

The architecture supports compliance with HIPAA and GDPR because it keeps data inside your AWS environment, uses private endpoints, and provides audit logs. However, compliance also depends on your data classification, region settings, model selection, and formal agreements with AWS. This architecture removes the technical blockers, but it should still be reviewed by your compliance team.

Does Jinba replace AWS Step Functions or EventBridge Scheduler for RAG orchestration?

Yes, Jinba provides the same orchestration capabilities as a DIY AWS-native approach — scheduling, batching, error handling — with built-in audit logging, role-based access control, and deterministic execution. That removes the development and maintenance overhead of building and governing your own orchestration layer.

Build your way.

The AI layer for your entire organization.

Get Started