- First-party - Your application owns the documents and manages permissions directly. You write tuples to OpenFGA as part of your normal application flow (e.g., when a user creates a folder or shares a document).
- Third-party - Documents and permissions live in an external system (Google Drive, Confluence, SharePoint, etc.). You synchronize content into your vector database and permissions into OpenFGA, keeping both in sync with the source system. The authorization model and filtering approaches are the same - the difference is that tuples come from a sync pipeline rather than your application.
Authorization model
A typical RAG knowledge base contains documents organized in folders, with access controlled at both levels. The following model represents this structure: Afolder has owner and viewer relations. A document belongs to a folder and inherits its viewers - anyone who can view the folder can view all documents inside it. You can also grant direct access to individual documents.
Writing tuples
Set up the folder structure, document ownership, and user access:annecan view all documents in the engineering folder (as owner).bethcan view all documents in the engineering folder (as viewer).carlcan only view the roadmap document.
Filtering approaches
There are two main approaches to integrate OpenFGA into a RAG pipeline. Both ensure that the LLM only sees documents the user is authorized to access.Post-filtering
Query the vector database first, then filter results by checking permissions with OpenFGA. This is the most common approach and works well when the vector search returns a manageable number of candidates. The flow is:- The user sends a query to the RAG pipeline.
- The pipeline retrieves candidate documents from the vector database.
- For each candidate, call OpenFGA to check whether the user can view it.
- Filter out unauthorized documents.
- Pass only the authorized documents to the LLM as context.
BatchCheck API to check multiple documents in a single request. For example, if a vector search returns three documents for user:carl:
Only document:roadmap is returned as allowed. The pipeline filters out the other two documents before passing context to the LLM.
Pre-filtering
Retrieve the list of documents the user can access first, then pass those IDs as a filter to the vector search. This approach works well when the user has access to a relatively small number of documents. The flow is:- Call the
ListObjectsAPI to get all document IDs the user can access. - Pass those IDs as a metadata filter to the vector database query.
- The vector search only returns results from authorized documents.
- Pass the results to the LLM as context.
user:carl can view:
Pass the resulting document IDs as a filter to your vector database. Most vector databases support metadata filtering - use the document ID stored in each vector’s metadata to restrict the search.
Choosing an approach
For detailed guidance on choosing between these approaches and handling more complex scenarios, see Search With Permissions.
When using post-filtering, request more candidates than you need from the vector database (e.g., 2-3x your target count) to account for documents that will be filtered out.
Framework integration
The filtering patterns above are framework-agnostic. Here is how to apply them in popular RAG frameworks:- LangChain (Python/JS): Implement a custom retriever that wraps your vector store retriever. After retrieving candidates, call OpenFGA
BatchCheckand filter the results before returning them to the chain. - LlamaIndex: Use a post-processing step or a custom node postprocessor that checks permissions against OpenFGA before passing nodes to the response synthesizer.
- Custom pipelines: Insert the authorization check between the retrieval and generation steps of your pipeline.
Further reading
These resources explore RAG authorization patterns with OpenFGA in more detail:- RAG and Access Control: Where Do You Start?
- Building a Secure RAG with Python, LangChain, and OpenFGA
- Build a Secure LangChain RAG Agent Using Auth0 FGA and LangGraph on Node.js
- Securing AI Document Agents with LlamaIndex and Auth0
- Securing Agentic RAG Pipelines
- Building a Permissions System For Your RAG Application
Related Sections
Take a look at the following sections for more information.Search With Permissions
Detailed guidance on integrating authorization into search, with trade-off analysis for different approaches
Task-Based Authorization
Grant agents scoped permissions to perform specific actions without permanent access
Conditions
Add time-based expiration or other conditions to document access grants