Building the RAGfoundation behindmunicipal assistants.
I worked with Neuraflow's CTO on the ingestion, retrieval, deployment and evaluation infrastructure behind public-facing assistants used across approximately 20 German municipal customers.
- ~20
- Weekly
- Isolated
- Public

Municipal information is difficult to navigate.
Residents should not need to understand a municipality’s information architecture before they can ask, “How can I register my dog?” or “My passport is expiring. What do I need to do?”
The assistants let people ask those questions in normal conversational language, across multiple languages. Answers were grounded in each municipality’s own service information and delivered inside its public website.
A shared platform, isolated municipal data.
We did not build an independent RAG stack for every customer. One reusable platform supported separate municipal deployments, while each municipality’s retrieval data remained isolated.
Source formats varied by customer. The ingestion layer turned structured APIs, scraped pages and documents into a common path toward retrieval.
Multiple municipalities feed a shared ingestion and retrieval platform, while their datasets and resident-facing assistants remain separate.
Retrieval quality started at the source.
The LLM was one stage in a longer production path. Most of my work lived in the layers that determined what evidence reached it.
- Often XML-based municipal service data
- Primarily collected through Apify scraping
- PDFs and other files entering the same retrieval path
Municipal APIs, websites and documents pass through ingestion, normalization, chunking, OpenAI embeddings, Pinecone indexing, retrieval, prompt construction and an OpenAI language model before a resident receives an answer.
The model wasn't the hardest part.
The most common input was structured municipal data exposed through APIs, often as XML. Other customers depended on website content, PDFs and general documents. Apify handled most website collection.
These sources did not arrive in one clean, consistent format. The engineering work was turning different inputs into dependable retrieval material without assuming every municipality published information the same way.
A correct page can still be bad context.
Large municipal service pages often mixed unrelated topics. Vector search could retrieve the broadly correct page for “How do I register my dog?” while burying the useful passage inside a large amount of irrelevant context.
We reduced oversized page-level documents into smaller retrieval units. A large page could become roughly four or five pieces, making returned context more focused without relying on the model to find one useful detail in a wall of text.
Correct page, noisy context.
Focused evidence for the answer.
The model could not compensate for consistently poor retrieval.
Public information does not stay static.
Opening hours, administrative processes and municipal pages change. Scheduled jobs refreshed source data approximately once per week; this was a recurring ingestion process, not a one-time knowledge import or real-time synchronization.
Debugging RAG means finding the failing layer.
Evaluation was primarily manual: run a repeatable set of questions, inspect retrieved content, check the answer and use production feedback. A bad answer was not automatically an LLM failure.
Public-facing, not a prototype.
Residents used these assistants inside real municipal websites. Delivery therefore included more than model and retrieval code: repeatable Azure infrastructure, Docker, Terraform, scheduled jobs and monitoring.
A separate development environment let the team test prompt and behaviour changes before collaboratively deploying updates to public systems.

What I owned
I worked closely with Neuraflow’s CTO throughout the architecture and delivery of the platform. My responsibility was substantial and collaborative — not sole ownership of the company’s entire system.
- Ingestion architecture
- Retrieval
- Prompt construction
- Deployment
- Monitoring
- Scheduled refresh jobs
- Evaluation
- Vector and data schema
- Backend and API development
- Internal admin tooling
The pipeline was the product.
Production RAG changed how I thought about applied AI. Answer quality depended on source ingestion, document structure, chunking, tenant isolation, retrieval, prompt behaviour, freshness, evaluation and deployment infrastructure.
The biggest practical lesson was that improving a RAG system often meant fixing the pipeline before changing the model.