Multilingual Retrieval-Augmented Generation Systems

Large language models can generate useful answers, but they do not inherently have access to a user's private document collection or specialized knowledge.

Problem Statement

Large language models can generate useful answers, but they do not inherently have access to a user’s private document collection or specialized knowledge. This becomes particularly challenging when the source material is written in Urdu or Arabic, where users may need to retrieve specific information from documents without manually reading through them.

A further challenge is the trade-off between cloud-based AI capabilities and data privacy. A cloud architecture can provide access to powerful embedding and inference models, while sensitive documents may require a solution where the entire retrieval and generation process remains locally controlled.

The problem was therefore to build a system that could allow users to ask natural-language questions about Urdu and Arabic documents, while exploring both cloud-based and fully local approaches to document retrieval and question answering.

Proposed Solution

Two independent Retrieval-Augmented Generation (RAG) pipelines were developed to investigate these approaches.

01 — Cloud RAG Pipeline

The first system uses a cloud-based architecture combining:

Google Gemini Embeddings → ChromaDB → Groq Inference

Documents are converted into vector representations using Google Gemini embeddings and stored in ChromaDB for semantic retrieval. When a user asks a question, relevant document content is retrieved and passed to a language model through Groq inference to generate a contextual answer.

02 — Local RAG Pipeline

The second system was designed around a fully local architecture:

bge-m3 Embeddings → Local Vector Retrieval → Ollama / Qwen2.5:7B

Instead of relying on cloud inference, the system uses Ollama running Qwen2.5:7B alongside bge-m3 embeddings, allowing the document-questioning workflow to operate locally.

A Streamlit interface provides the user-facing layer for interacting with both systems.

Why two architectures?

The project explores two different approaches to the same fundamental problem:

Cloud-based RAG
Powerful hosted models and inference services.

Local RAG
Greater control over the processing environment and a privacy-oriented workflow.

The comparison makes the project particularly interesting because it isn’t simply a RAG chatbot—it demonstrates how the same multilingual document-retrieval problem can be approached through both cloud and locally operated AI architectures.

Core Idea

Ask questions in natural language. Retrieve the relevant knowledge from Urdu and Arabic documents. Generate answers grounded in the source material.

Technology Stack

Python · Google Gemini · Groq · ChromaDB · Ollama · Qwen2.5:7B · BGE-M3 · Streamlit · RAG