Work / Context
Context
Ask your company's documents a question and get a cited answer, from a system that never leaves your network.
- Type
- AI model
- Status
- In Production, 2026
- Our role
- Retrieval architecture, model selection and tuning, indexing pipeline, interface
- Timeframe
- 2026 — ongoing
Context is an internal answer engine. It indexes the places a company's knowledge actually lives: file shares, wikis, ticket systems, contract folders, engineering drawings and the PDF archive nobody has opened since the last audit. It lets anyone ask a question in plain language and get back a short answer with a link to every passage it relied on.
It runs entirely on infrastructure the company controls, a single server or a private cloud tenant, with open-weight models and no calls to any outside AI provider. The firms that most need this kind of tool are often the ones whose contracts, clients or regulators forbid sending documents to one.
- Calls to outside AI providers
- 0
- Calls to outside AI providers
- Source passage linked per claim
- 1
- Source passage linked per claim
- Of answers filtered by the asker's access
- 100%
- Of answers filtered by the asker's access
The problem with asking the cloud
The value of an AI assistant at work grows with how much of the company it can read. So does the risk. Pasting a client contract, a pricing model or a design file into a hosted chatbot sends it to someone else's servers, under someone else's terms. Many firms have, sensibly, banned it outright, and their staff go back to searching folder names and asking whoever has been there longest.
Context is designed around that constraint rather than around it being waived. The models, the index and every query stay inside the company's network, and the system has no outbound route to anything.
Finding the right passage
Answer quality depends almost entirely on retrieval. Context indexes documents in place and splits them along their real structure, such as sections, clauses, table rows and slide notes, rather than fixed-size chunks. That way a retrieved passage is a coherent unit and not half of two paragraphs.
Each query runs two searches at once. Keyword search catches exact part numbers, clause references and names. Vector search catches the same idea phrased differently. The combined candidates are rescored by a reranking model, and only the strongest passages reach the language model, which is instructed to answer from them alone and to say so when they don't contain an answer.
Permissions are part of the search
An internal search tool that ignores access control is a data leak with a search box. Context reads the permissions on each source, including file-share ACLs, wiki spaces and group memberships, and filters at retrieval time. A passage the user couldn't open directly can never be used to answer their question, and can't be summarised around either.
Every answer is shown with numbered citations that open the source at the exact passage. When the documents disagree, Context says so and shows both, rather than choosing one silently.
Running it on your own hardware
Context uses current open-weight models, quantised to run well on a single workstation or server GPU, so a mid-sized firm can run it without a data centre. The index updates incrementally as files change. Models can be swapped as better open-weight releases appear, with no change to the rest of the system and no new vendor to approve.
Where it stands
Ingestion, hybrid retrieval with reranking, permission filtering and cited answers are working end to end in the lab against large mixed document sets. Current work is on connectors for common file stores and wikis, and on evaluation: a fixed set of real questions with known answers, scored on every change, so retrieval quality is measured rather than assumed.
Next is handling drawings and scanned documents as first-class sources, so a question about a detail on a drawing can be answered from the drawing itself.
Technology
Open-weight LLMs (Qwen, Llama families), vLLM / llama.cpp, Hybrid BM25 + vector search, Cross-encoder reranking, PostgreSQL + pgvector, Python, Next.js
Join the Context beta waitlist
We're taking a small number of early deployments. Tell us roughly what your documents look like and where they live, and we'll be in touch.