Enterprise knowledge search: from hours to a minute
Large manufacturing group
- Over 10,000
- items available through one search system
- 1,500 outdated
- documents identified in the knowledge base
- Hours to a minute
- to find the information employees need
A large manufacturing group had accumulated more than ten thousand items in its corporate knowledge base, yet employees struggled to find answers. We deployed Clairon on the company’s own infrastructure. Getting information that once took hours now takes less than a minute. Along the way, we identified around 1,500 documents containing duplicates, outdated information or contradictions, and helped the team responsible put the knowledge base in order.
The knowledge exists. Finding it is another matter
Each division had its own part of the shared knowledge base: linked articles, illustrations and attachments ranging from Word and PDF files to spreadsheets and presentations. Staff maintained the content and tried to keep it organised. The information was there, and employees had access.
In practice, a search returned a long list of documents. The employee then had to work out which one answered the question and which version to trust. Only a few people knew the knowledge base well, so colleagues kept asking them for help.
Many others did not ask at all. They spent hours finding things out another way, even though the answer was already stored in the company’s systems. Years of work collecting knowledge were not translating into easy access to it.
An answer you can trace to its source
We connected Clairon to the knowledge base and tuned its search to the company’s terminology and documents. An employee asks a question in ordinary language; the system finds relevant material, compares the information and produces an answer with source extracts and links. The user can open the original document or ask a follow-up question.
The model’s role is to help people navigate company knowledge. Its answers stay connected to the source material. If two documents disagree, Clairon shows both versions and explains the difference.
Several retrieval steps make this possible. Clairon first interprets the question and identifies ways the information might be expressed in the documents. It combines vector search, which finds related meaning, with optimised text search for exact terms, model names and product identifiers. A reranking model tuned to the client’s data then selects the most relevant passages for the answer.
This works across Word documents, Excel spreadsheets, PowerPoint slides, PDFs including scanned documents, images and screenshots. Files can be attached to articles or stored separately: their content becomes available for search and answers.
But reading every word is not enough to understand a document. On a slide, the position of blocks and the arrows between them may explain a sequence or a dependency. A table’s rows and columns give meaning to its numbers. A chart shows relationships and changes that its labels alone do not describe. Extracting only the text would lose some of the knowledge the document was created to convey.
During indexing, specialised models interpret this visual information and describe the connections, processes and relationships it contains. Those descriptions become part of the search index. An answer can therefore use information from a diagram in an attached presentation, even when the article itself never mentions it.
Instead of repeatedly opening search results that turn out to be irrelevant, an employee gets a meaningful answer in 20–30 seconds, with the passages and links needed to check it. They can then continue the conversation to clarify the details.
Company data stays inside the company
The entire system, including the language models, runs on-premises. Documents and employee questions never leave the client’s infrastructure. We selected and optimised the models for the task, so deployment did not require a large computing setup.
The index receives incremental updates every night: only new and changed material is processed. This keeps search up to date without rereading the entire knowledge base or placing unnecessary demands on the hardware.
All major divisions use the system. Access follows user groups in the corporate Active Directory, so employees receive answers based on documents they are allowed to read and can check those sources themselves.
Knowledge collected over many years became much easier to use in everyday work.
Exposing years of accumulated confusion
Bringing the material into a shared semantic index revealed how many copies, historical versions and contradictions had accumulated. Some documents were duplicated in full; others repeated individual passages or gave conflicting answers to the same question. Each article could look reasonable on its own. The problem became visible when it was compared with the rest of the knowledge base.
We identified around 1,500 such documents and mapped the problems for the team responsible: duplicate content, old versions and information that needed to be reconciled. Staff could go straight to the relevant material rather than manually read and compare thousands of pages.
They used this map to remove duplicates, reduce outdated content and resolve contradictions. Deploying search also became an AI-assisted review of the company’s accumulated knowledge, exposing the confusion that had made it hard for employees to know which document to rely on.