Selected workCase study 02
Public source · RAG prototype

Document retrieval assistant.

Answers grounded in your documents.

A document-aware assistant connecting local PDFs, websites, and Google Drive to a retrieval pipeline with conversation history.

TECHNOLOGY
PythonFastAPILangChainFAISSStreamlit
EVIDENCE

Public repository · backend_api.py, frontend.py

ENGINEERING FOCUS

A document-change hash reuses the FAISS index instead of rebuilding an unchanged corpus.

01 / THE PROBLEM

The system challenge.

A language model does not know your documents by default. This project brings heterogeneous document sources into a searchable index and adds retrieved context to a conversation.

02 / SYSTEM FLOW

Trace the architecture.

Select a stage to see its responsibility.Conceptual flow · based on the documented implementation

STAGE 01

Load local PDFs, website URLs, and Google Drive documents into a common retrieval pipeline.

01

Multiple sources, one index

Document loaders ingest local PDFs, website URLs, and Google Drive files. Recursive splitting uses 1,000-character chunks with 200-character overlap.

02

Avoid rebuilding unchanged content

A document hash detects changes before rebuilding the FAISS vector store. An unchanged corpus loads the saved index.

03

Retrieval with conversation context

The backend retrieves relevant chunks and combines them with conversation memory. Per-chat JSON files persist history across interactions.

04

A separated interface and API

FastAPI handles chat creation, history, and queries. Streamlit supplies the user interface. This is a public prototype, without published production benchmarks.

04 / OUTCOME & EVIDENCE

What the work demonstrates.

A working retrieval-oriented implementation is available for code review. No accuracy, latency, or scale claims are made.

Review the implementation on GitHub