Multiple sources, one index
Document loaders ingest local PDFs, website URLs, and Google Drive files. Recursive splitting uses 1,000-character chunks with 200-character overlap.
Answers grounded in your documents.
A document-aware assistant connecting local PDFs, websites, and Google Drive to a retrieval pipeline with conversation history.
A language model does not know your documents by default. This project brings heterogeneous document sources into a searchable index and adds retrieved context to a conversation.
Select a stage to see its responsibility.Conceptual flow · based on the documented implementation
Load local PDFs, website URLs, and Google Drive documents into a common retrieval pipeline.
The details that make this an engineering system.
Document loaders ingest local PDFs, website URLs, and Google Drive files. Recursive splitting uses 1,000-character chunks with 200-character overlap.
A document hash detects changes before rebuilding the FAISS vector store. An unchanged corpus loads the saved index.
The backend retrieves relevant chunks and combines them with conversation memory. Per-chat JSON files persist history across interactions.
FastAPI handles chat creation, history, and queries. Streamlit supplies the user interface. This is a public prototype, without published production benchmarks.
A working retrieval-oriented implementation is available for code review. No accuracy, latency, or scale claims are made.
Review the implementation on GitHub