Automated Evaluation Metrics
This system was mathematically evaluated using the ragas framework and LangSmith against a ground-truth dataset, proving its enterprise readiness and guardrail efficacy.
- Faithfulness (Hallucination Guardrail): 97.06%
(System successfully refuses to invent code when documentation is missing for a specific OS/Framework). - Context Precision: 100% (On Valid Contexts)
(Cross-Encoder reranking ensures the top returned node is the exact match for the user's environment). - Observability: 100% of pipeline executions, token usage, and cross-encoder latencies are traced and monitored via LangSmith.
System Architecture
Core Features
- Intelligent Ingestion (document_parser.py): Deep content analysis that scans markdown headers and code blocks to dynamically assign os, language, and framework metadata tags to vector chunks before embedding.
- Dynamic Metadata Filtering: Prevents context contamination. A macOS/React developer will never be served Windows/Python installation instructions.
- Cross-Encoder Reranking: Utilizes HuggingFace ms-marco-MiniLM-L-6-v2 to mathematically score and rerank retrieved nodes, boosting precision by filtering out semantically similar but contextually irrelevant chunks.
- Asynchronous Backend: Built on FastAPI with asyncio.to_thread execution, ensuring the API remains highly responsive during computationally heavy ML embedding and reranking tasks.
Tech Stack
- LLM: Google Gemini 2.5 Flash / (You can also switch to ollama if Gemini API is not feasible)
- Orchestration: LlamaIndex
- Vector Database: Qdrant (Local Backend)
- Embeddings: HuggingFace (BAAI/bge-large-en-v1.5)
- Reranker: Sentence Transformers (cross-encoder/ms-marco-MiniLM-L-6-v2)
- Backend: FastAPI, Uvicorn, Pydantic
- MLOps / Evaluation: LangSmith, Ragas, Datasets
- Frontend: Vanilla HTML/CSS/JS (Glassmorphism UI)
Quick Start
1. Prerequisites
Ensure you have Python 3.10+ installed. Clone the repository and install the dependencies.
git clone https://github.com/AbdullahWali007/DocuSync.git
cd DocuSync
pip install -r requirements.txt
2. Environment Variables
Create a .env file in the root directory:
GEMINI_API_KEY="your_gemini_key"
LANGCHAIN_API_KEY="your_langsmith_key" # Only required for Phase 4 tracing
LANGCHAIN_PROJECT="DocuSync_AI_Production"
3. Initialize the Vector Database (Phase 1)
Run the ingestion script to parse your documents, generate embeddings, and populate Qdrant.
python ingest.py
Note: This single command internally imports and uses the supporting modules config.py, document_parser.py, and qdrant_setup.py – you do not need to run them manually.
4. Start the Application (Phase 5)
Spin up the FastAPI backend server:
python phase5_api.py
Once the server is running on http://localhost:8000, open index.html in your web browser to interact with the Multimodal UI.
Pipeline Execution Order (Detailed)
For those who wish to run the optional evaluation or tracing phases, below is the full recommended sequence:
| Step | File | Purpose |
|---|---|---|
| 1 | ingest.py | (Mandatory) Loads, chunks, enriches metadata, embeds, and stores documents in Qdrant. |
| 2 (Optional) | phase2_baseline.py | Runs a set of benchmark queries against a simple Python-only retriever; exports results to phase2_control_state.json. Useful for before/after comparisons. |
| 3 (Optional) | phase3_orchestration.py | Can be run directly (has __main__) to test a single query against the 4‑Agent pipeline without the API. Mainly used for development. |
| 4 (Optional) | phase4_tracing.py | Wraps the orchestrator with LangSmith @traceable decorators and executes a sample query. Requires LANGCHAIN_API_KEY. |
| 5 (Optional) | phase4_ragas_eval.py | Runs the Ragas evaluation suite against a ground‑truth dataset; outputs a CSV with per‑sample metrics (faithfulness, context precision, answer relevancy). |
| 6 | phase5_api.py | (Production) Starts the FastAPI server. All other phases are integrated; this is the final service entrypoint. |
Important: The supporting library files (config.py, document_parser.py, qdrant_setup.py) are not executed directly; they are imported by the scripts above.
Evaluation & Observability (Phases 4)
To run the automated evaluation and tracing:
# Trace a query with LangSmith (ensure LANGCHAIN_API_KEY is set)
python phase4_tracing.py
# Evaluate the pipeline with Ragas on a ground-truth dataset
python phase4_ragas_eval.py
Results are saved locally (phase4_ragas_metrics.csv) and visible in your LangSmith dashboard.