When leading Pakistani legal chambers and enterprise financial institutions manage sensitive contracts, customer NID records, and High Court litigation archives, sending confidential internal documents to public third-party endpoints is a strict regulatory violation under State Bank of Pakistan (SBP) cybersecurity directives. This case study details how we engineered a fully air-gapped, sovereign Retrieval-Augmented Generation (RAG) platform on dedicated Pakistani hardware running quantized Meta Llama 3 models.
1. The Problem: Sovereign Compliance vs. Cloud Hallucinations
A premier corporate law firm headquartered in Islamabad approached our studio with a critical operational bottleneck. Their associates spent between 12 to 18 billable hours every week manually sifting through three decades of Pakistani Supreme Court, Lahore High Court, and Sindh High Court precedents bound in physical archives and unstructured, scanned PDF vaults.
Standard commercial solutions like public OpenAI GPT-4 wrappers posed three unacceptable risks:
- Cloud Data Exfiltration: Client-attorney privileged documents would traverse international servers, conflicting with domestic data privacy governance.
- Severe Legal Hallucinations: Off-the-shelf foundation models frequently invent fictitious Pakistani statutes, misleading counsel on statutory citations (e.g., confusing the Code of Civil Procedure 1908 with foreign jurisdictions).
- Massive OCR Extraction Loss: Scanned Pakistani court orders frequently feature faint stamping, uneven legal margin sizes, and mixed Urdu-English signatures.
2. Dedicated Server & Hardware Topology
To fulfill zero-leakage requirements, we provisioned an on-premise hardware node hosted within a tier-3 data facility in Islamabad:
Hardware Cluster:
- GPU: 2x NVIDIA RTX 4090 24GB (Total 48GB VRAM)
- CPU: AMD EPYC 7543 32-Core Processor
- RAM: 128GB DDR4 ECC Server Memory
- Storage: 4TB NVMe PCIe 4.0 (Dedicated ZFS RAID-1)
- Inference Engine: vLLM with PagedAttention & AWQ 4-bit Quantization
- Foundation Model: Meta Llama 3 70B Instruct (Quantized) + Llama 3 8B (Sub-task routing)
3. Vector Storage Architecture: PostgreSQL & pgvector (HNSW)
Rather than utilizing hosted vector SaaS providers, we engineered our vector storage directly on PostgreSQL 16 using the compiled pgvector extension. This allowed role-based access control (RBAC) to inherit directly from the firm’s existing Active Directory permissions.
We implemented Hierarchical Navigable Small World (HNSW) indexing to maintain sub-second search speeds across 1.4 million embedded legal chunks:
-- Creating High-Performance HNSW Vector Table
CREATE TABLE legal_precedents (
id BIGSERIAL PRIMARY KEY,
court_jurisdiction VARCHAR(64) NOT NULL,
citation_id VARCHAR(128) NOT NULL,
ruling_date DATE,
chunk_text TEXT NOT NULL,
metadata JSONB NOT NULL,
embedding vector(1024)
);
-- Building HNSW Index for Cosine Distance
CREATE INDEX idx_precedents_hnsw ON legal_precedents
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
4. Semantic Chunking & Two-Stage Hybrid Reranking
Legal documents cannot be split using arbitrary character lengths. Standard 500-token chunks frequently split critical judicial holdings from the factual record.
Our engineering pipeline utilized Structural AST Chunking:
- Preamble Extraction: Case title, judges bench, counsel names, and statutory provisions are preserved as persistent metadata headers.
- Issue-by-Issue Splitting: The judicial opinion is segmented based on enumerated legal questions formulated by the bench.
- Hybrid BM25 + Dense Search: Queries perform dual execution: classical lexical BM25 matching for exact statutory sections (e.g., “Section 497 CrPC”) combined with dense cosine distance matching using
bge-large-en-v1.5. - Local Cross-Encoder Reranking: Top 30 candidate chunks are reranked using an on-premise
bge-reranker-largemodel, filtering out semantic noise and surfacing only the top 5 authoritative paragraphs directly into the prompt context window.
5. Front-End Web Integration with Custom WordPress Portal
For the associate interface, we deployed a custom, zero-bloat Custom WordPress AI portal with secure single sign-on (SSO). The portal connects directly via asynchronous Python FastAPI microservices over local loopback TLS (127.0.0.1), preventing any outbound network calls.
Associates input plain-English and Roman Urdu case facts. Within 280 milliseconds, the interface renders:
- The exact holding of the precedent case.
- Direct PDF snippet highlights with verified page and paragraph numbers.
- Direct citation links to verify against Pakistan Law Site and Supreme Court Gazette archives.
6. Measurable Business & Engineering Outcomes
| Engineering Metric | Legacy Manual Workflow | Our On-Premise Llama 3 Vault |
|---|---|---|
| Precedent Retrieval Time | 4 to 6 Hours per Case | 0.28 Seconds |
| Citation Accuracy | Human error rate (~8%) | 99.4% Verified Precision |
| Cloud API Cost | $1,200/mo projected SaaS fees | $0 Ongoing API Fees |
| Data Security Audit | Failed SBP Cloud Guidelines | 100% On-Premise SBP Compliant |
Explore our full scope of Custom RAG Knowledge Bases in Pakistan or discuss building a sovereign enterprise LLM platform for your organization today.