One search across 90,000 documents, and a plan to make the company findable
My thesis project at a Swedish fintech: an on-premises search system over Confluence, Jira and GitHub that cut average task time from 294 to 96 seconds in a user study, plus an audit of how the company shows up in search engines and AI assistants.
The challenge
Vermiculus builds clearing and trading technology for exchanges. Its knowledge was spread across Confluence, Jira and GitHub, more than 90,000 documents in total. Finding an answer meant jumping between systems, an average of 3.3 switches per task, and the answer was often in a system the person didn't think to check. The company also had a second problem: its expertise was hard to find from the outside. Reasonable searches for what it does returned little, and AI assistants had almost nothing to say about it.
Approach
For internal search, I built a retrieval-augmented generation system from scratch. It combines keyword search (BM25) with vector search, merges the results with reciprocal rank fusion and re-ranks them with a cross-encoder. A product-and-client filter narrows the search before retrieval, so people only see material relevant to the product and client they're working on. For external visibility, I audited how Vermiculus appears in search engines and AI assistants and wrote an improvement plan. I did the whole audit myself and presented the plan to the Head of Marketing, who agreed with it and started working on it.
Technical details
The system indexes more than 90,000 documents, including over 83,000 Confluence pages. Embeddings come from Instructor-XL, and answers are generated with Qwen 32B, both running on-premises on an NVIDIA GPU, so no company data leaves the building. Access is role-based, storage is encrypted, and the system has a REST API and a web interface. Alongside the build, I trained around 150 colleagues, mostly technical, in how to use it and how to think about AI answers.
Results
Search system
- 85% smaller search space with product-and-client filtering
- 65% faster search and 15% higher precision compared with unfiltered search
- 90,000+ documents indexed, including 83,000+ Confluence pages
User study
- Average task time down 67% (294 → 96 seconds)
- System switches per task: 3.3 → 0
- Usefulness rated 4.6 out of 5 and reliability 4.2 out of 5
- 5 participants, same people with and without the system, run during a busy release week
Visibility
- Full audit of search-engine and AI-assistant visibility, with an improvement plan
- Plan presented to the Head of Marketing, who agreed and began implementing it
What this means
Building the retrieval system taught me exactly how retrieval systems choose their sources, and that directly shaped the visibility audit. The same principles that make a document findable inside a company knowledge base make a business findable to Google and ChatGPT.
At a glance
- Client: Swedish fintech, clearing and trading technology
- Type: Thesis project, October 2025 – January 2026
- Delivered: on-premises RAG system and a visibility audit with improvement plan
- Recognition: Best Thesis 2026, IHM Business School