Practical example · AI & Automation
An AI knowledge base that keeps your data in-house
Ask questions in plain language and get an answer with a source citation drawn from your own documents, with no text ever sent to a cloud service. I built and measured this in a test environment on my own hardware.
- Hardware
- 6 CPU cores, 10 GB RAM, no graphics card
- Document set
- 131 documents, 353 sections
- Response time
- 9 to 40 seconds
- Cloud AI
- none, everything runs locally
Principle
How it works
The documents are split into sections and stored as numeric vectors in a database. For a question, the system searches for the most relevant sections, re-ranks them with a second model, and passes the best four to a language model. It formulates the answer only from these sections and cites them as its source.
This method is called Retrieval-Augmented Generation (RAG). The advantage over a chatbot without its own data: the answer can be checked against the source, and if the system finds nothing relevant, it says so instead of making something up.
The test data set consisted of the pages of this website as well as the sections of the German Federal Data Protection Act (BDSG) and Telecommunications-Digital-Services-Data-Protection Act (TDDDG) in the official version from gesetze-im-internet.de.
Technology
Architecture
- Database
- PostgreSQL 17 with pgvector and an HNSW index for similarity search.
- Embedding
- bge-m3 (multilingual, MIT license) converts text into vectors.
- Re-ranking
- bge-reranker-v2-m3 scores the 20 search hits; both rankings are merged.
- Language model
- Qwen3-4B-Instruct (Apache-2.0) via llama.cpp, CPU only.
- Containers
- Rootless Podman with Quadlet: every service is a systemd unit with a health check and starts itself again after a restart.
- Access
- Caddy with HTTPS in front of Authelia and LLDAP: no login, no access, with permissions per user group.
Measurement
Measured values
- Converting a question into a vector
- 30 to 70 milliseconds
- Searching 353 sections
- around 10 milliseconds
- Re-ranking 20 hits
- 5 to 6 seconds
- Full answer with sources
- 9 to 40 seconds, depending on length
- Ingesting all 131 documents
- 190 seconds
- Ready again after a restart
- 85 seconds, all 8 services
- Backup with restore test
- scheduled daily; 353 of 353 sections restored during setup
- Access without login
- redirected to the login page
Measured on 29.09.2026 in the test environment. Response times would be substantially shorter with a graphics card; that was not measured here.
Test
Four test questions, four outcomes
- Pricing question about this website
- Answered correctly, with the pricing page as its source.
- Question on employee data protection (§ 26 BDSG)
- Answered correctly, with the matching section as its source.
- Question on cookie consent
- Found the correct source (§ 25 TDDDG), but the answer included a sentence the source does not support. This exact error is listed below under limits.
- Question outside the data set
- On Austrian income tax, the system found nothing relevant and said so, instead of making up an answer.
Limits
Limits, stated honestly
- A small language model can still summarize incorrectly despite having sources. In one test it added a claim the source does not contain. That is why every answer shows its sources, and a human reviews important answers.
- On CPU alone, an answer can take up to 40 seconds. A team asking many questions per day would need a graphics card or a larger model.
- This is a test environment, not a production deployment. Real-world use would add monitoring, document updates and an access-rights concept for confidential content.
- The test answers about legislation are not legal advice. Data-protection questions about your specific use case should be clarified with your law firm or data protection officer.
For you
What this means for your business
Manuals, price lists, work instructions or contract templates can be made searchable the same way. I start with a small data set and one measurable question, such as how long it currently takes to find a piece of information, and only expand once it proves worthwhile.
More on AI & AutomationWant your own documents made searchable?
Tell me briefly what documents your team searches for daily. I will tell you whether a pilot is worth it.
Quickest via WhatsApp or email, or give me a call:Phone: +49 5677 9218771Open WhatsApp chat with +49 5677 9218771 (Phone) (new window; data is only sent to WhatsApp once you click)Mon–Fri 9am–6pm, Sat 9am–2pm, closed Sundays
