AI automation
How much does a RAG chatbot for company documents cost, and should you build or buy?
A RAG chatbot on company documents costs roughly $4,000–$12,000 for a one-source MVP and $15,000–$80,000 for production scope, according to 2026 price guides, plus about $22 a month in model fees for 20 staff. Ready tools charge about $18–$25 per seat a month. VITON13 builds a bounded version from $113.
What is a RAG chatbot, in plain words?
RAG stands for retrieval-augmented generation. When someone asks a question, the system first searches your documents, then hands the best passages to a language model and asks it to answer from them, with links to the sources.
Two things follow. Nothing is retrained: your files sit in a search index outside the model. And the answer is only as good as the search, so duplicate, outdated or badly scanned files give fluent wrong answers.
How much does it cost to build a RAG chatbot?
Published 2026 price guides disagree by a factor of ten or more, because each prices a different project. None is a quote.
| Source and type | Scope as stated | Build price | Running cost as stated |
|---|---|---|---|
| Hamza Shabbir, freelance engineer, 9 Jun 2026 | MVP: one data source. Production: several sources, evaluation, monitoring | $4,000–$12,000 MVP; $15,000–$40,000+ production | $5–$30 per 1,000 questions plus $25–$150 a month |
| SoluLab, development agency, 31 Jul 2026 | Basic: one source. Mid-level: 2–4 sources, custom interface | $8,000–$25,000 basic; $25,000–$50,000 mid-level | Not priced |
| gmware, software firm, 22 May 2026 | Production chatbot on one knowledge base; proof of concept with evaluation | $30K–$80K production; $50K–$150K proof of concept | $400–$6,000 a month |
| VITON13 | Up to 200 approved documents, 3 sources, 3 user groups, 50 test questions, 1 interface | From $113 AI-assisted or $173 human-led | Model, vector database and hosting fees go to the providers |
Read the VITON13 row as a narrower scope, not a discount: the scope caps documents, sources, user groups and test questions, and leaves scanning, integrations and provider fees outside the price.
Why do two quotes for the same chatbot differ so much?
- Data preparation: Hamza puts it at 30–40% of a build budget and gmware at 30–50%. In Hamza’s experience a folder of clean Markdown takes a day and years of scanned PDFs take weeks.
- Permissions: if staff may read different documents, the rules must travel from the source system into the search layer. gmware asks: does the sales team see HR’s documents?
- Evaluation: Hamza prices a 50–100 question test set at $1,000–$3,000 and calls it the best money in a build.
- Data location: OpenAI and Anthropic both add about 10% for regional data-residency endpoints; private hosting adds operations.
- Who builds it: Hamza says agencies quote two to three times a solo engineer for the same scope, and his own list is the low end of the table.
Can a ready-made tool replace a custom RAG build?
For a team that already lives in one suite, often yes. The table shows what each vendor prints on its own pricing page today.
| Tool | Price on the vendor page | What it reads | Check before buying |
|---|---|---|---|
| ChatGPT Business | $20 per seat a month billed annually, $25 monthly; 2–200 employees; Enterprise by quote | Files in shared Projects; Google Workspace, Slack, GitHub, Microsoft 365 | Role-based access controls are listed for Enterprise only; the reasoning window is about 320 pages |
| Microsoft 365 Copilot Business | $18 per user a month paid yearly (promotion for purchases 1 Jul–31 Dec 2026, first year only; list $21); $25.20 billed monthly | Microsoft 365 content | Add-on for an eligible Microsoft 365 plan, up to 300 users |
| Notion Business | €19.50 per member a month (page priced in euros; yearly billing saves up to 20%) | Notion pages and connected apps such as Slack | Agents beyond the free trial: $10 per 1,000 credits |
| Chatbase | $40, $150 or $500 a month for 700, 4,000 or 15,000 message credits | Content you upload, for a bot on your site | Extra credits cost $40 per 1,000 |
| Intercom Fin | $0.99 per outcome; helpdesk seats from $19 a month | Your help centre and support content | Built for support teams; a resolved conversation is the billing unit |
Seats add up. Twenty people on ChatGPT Business cost $4,800 a year billed annually or $6,000 billed monthly; on the Copilot Business offer, $4,320 for the first year (our arithmetic). That is the range of a small custom build, so ask what a seat cannot do: run inside your website or Telegram bot, answer the public, or apply document-level rules.
What does a RAG chatbot cost to run each month?
Three meters run after launch: indexing your files, storing the index and paying tokens per question. Indexing is nearly free: 200 documents of 20 pages at roughly 600 tokens a page is 2.4 million tokens, and OpenAI’s text-embedding-3-small costs $0.02 per million, about five cents. Storage is small too: OpenAI’s file search includes 1 GB free.
Questions are the main meter. Assume 3,000 input tokens (instructions, five retrieved passages, the question) and 400 output tokens, the sizes Hamza uses, and 20 people asking five questions a day for 22 working days: 2,200 questions.
| Model (price per million input / output tokens) | One question | 2,200 questions |
|---|---|---|
| GPT-6 Luna, OpenAI ($0.10 / $0.50) | $0.0005 | $1.10 |
| Claude Haiku 4.5, Anthropic ($1 / $5) | $0.005 | $11 |
| Claude Sonnet 5.5 and GPT-6.1 Sol ($2 / $10) | $0.01 | $22 |
| Claude Opus 5.5 ($4 / $20) | $0.02 | $44 |
| Claude Fable 5.1 and GPT-6 Astra ($10 / $50) | $0.05 | $110 |
So the model bill for a small team is tens of dollars. Add $25–$150 a month of hosting, Hamza’s range. The cost that matters is a person who keeps documents current and reads the questions the assistant failed.
Where do RAG chatbots still give wrong answers?
- The wrong passage is found: a table or clause split across chunks is answered from half of it, and a scan without a text layer gives the search nothing.
- Old and duplicate files: gmware notes that retrieval returns the 2022 price sheet as readily as the current one. Give every document an owner and a date.
- Hallucination survives: Stanford researchers found RAG-based legal research tools from LexisNexis and Thomson Reuters hallucinated 17–33% of the time, less than a general chatbot but far from the elimination vendors had claimed.
- Access leaks: OWASP lists weak access control on embeddings, and shared vector stores that leak context between users, as an LLM top-10 weakness.
- Poisoned content: a document with hidden instructions can steer answers; OWASP advises validating sources and logging retrieval.
A company answers for what its assistant says. In Moffatt v. Air Canada (2024 BCCRT 149) a British Columbia tribunal rejected the airline’s argument that its chatbot was a separate entity and ordered it to pay CAD 812.02 in damages, interest and fees after wrong fare advice. That bot faced the public, but staff act on an internal assistant’s answers too.
What do GDPR and the CNIL say about a documents assistant?
If your documents hold personal data, a model provider that processes it is your processor. GDPR Article 28 then asks for a written contract binding it to your instructions, and Article 32 for security measures that fit the risk. France’s regulator, the CNIL, adds that a company connecting an assistant to its own knowledge base answers for that processing, and that for personal data or sensitive documents an on-premise deployment is generally the safer choice.
This is not legal advice. Ask your data protection officer or lawyer where the provider processes your data, what it retains and whether files are used for training.
Build, buy or skip it: which checklist fits your case?
- Skip RAG if your whole library is under about 100–200 pages and rarely changes. Hamza suggests a cached prompt, and the ChatGPT Business page lists a reasoning window of about 320 pages.
- Buy a seat plan if everyone works in Microsoft 365, Google Workspace or Notion, users number a few dozen and the suite’s own permissions suffice.
- Use a bot builder such as Chatbase if the audience is website visitors and the content is public.
- Build custom if different groups may read different documents, answers must link to the passage, sources sit in several systems, you need your own interface or a re-runnable test set, or data must stay in a chosen region.
How VITON13 builds a documents assistant
The VITON13 package is the narrow version of the build above. We agree the sources, who may see what and the questions people ask most often; index the approved documents and wire in the access rules; then run 50 test questions with expected answers, fix wrong or unsupported answers and set the threshold for a clear ‘not enough information’ reply. After your team’s test and the revision rounds, the assistant goes live with the evaluation report. In the AI-assisted mode AI does the first pass and a person checks it; in the human-led mode a specialist leads each step and uses AI as support.
RAG knowledge assistant
- Price
- $113 AI-assisted or $173 human-led
- Timeline
- 10–18 working days
AI chatbot with human handoff
- Price
- $73 AI-assisted or $93 human-led
- Timeline
- 5–8 working days
The final price is fixed in writing after a short scope review.
AI agents for business
- Price
- $73 AI-assisted or $93 human-led
- Timeline
- 5–8 working days
AI-assisted: AI drafts inside defined steps, a person directs and checks every result. Human-led: a specialist does the work and AI assists.
What is included
- Up to 200 approved documents indexed from up to 3 sources
- Access rules for up to 3 user groups
- 50 test questions with expected answers and a report you can re-run after changes
- One interface, a web chat or a messenger bot, with a source link under every answer
- One update schedule that re-reads changed documents and drops removed ones
- Delivery: 10–18 working days after the brief is approved
Not included
- Model, vector database and hosting fees, which you pay to the providers
- Scanning, OCR clean-up or rewriting of documents that are hard to read
- Actions in other systems, such as creating tickets: that is an AI agent, with its own scope
- Any promise of accuracy beyond the tested questions
Revisions: 2 revision rounds on the deliverables
Get an estimate in 1 working dayWritten estimate within 1 working day
Example
PILOTPilot in progress
No published documents-assistant case yet
VITON13 has not yet published a case for a RAG assistant on a client’s documents with measured answers, so this guide invents no accuracy figures. Every number comes from a source below or from arithmetic shown on the page.
What this does not prove
- The model-fee table is arithmetic on list prices for one assumed question size. It does not show how accurate any assistant is or what your usage will cost.
Questions and answers
Is a RAG chatbot the same as training ChatGPT on my documents?
No. Your files are indexed outside the model and passages are fetched when a question arrives, so nothing is retrained and a deleted file stops being used once the index refreshes. Fine-tuning mainly changes style and format, as Hamza notes.
How many pages can I give ChatGPT before I need RAG?
The ChatGPT Business page lists a 256K-token reasoning window, about 320 pages of input, and the answer must fit in it too. Hamza suggests loading material under 100–200 pages into a cached prompt. Past that, or when files change weekly, searching first pays off.
Can the chatbot show different answers to different employees?
Yes, if access rules sit inside retrieval, so a passage a person may not read is never fetched. Filtering after the model has seen it is too late. OWASP warns about shared vector stores that leak context between users. The VITON13 package applies rules for up to 3 user groups.
How do I test a RAG assistant before launch?
Write real questions with the expected answer and source, run them after every change, and check retrieval and answer separately: was the right passage found, and is the answer faithful to it? Hamza prices a 50–100 question set at $1,000–$3,000. The VITON13 package includes 50.
Do I need a vector database such as Pinecone?
Not always. Pinecone’s Standard plan has a $50 monthly minimum and its Builder plan is $20. Hamza says Postgres with the pgvector extension handles most knowledge bases up to a few million vectors for $0–$50 a month extra. Vector database fees are outside the VITON13 package.
What do you need from me to start?
Send the approved documents or their locations, who may see which of them, and the questions people ask most often. Scans and hard-to-read files need cleaning before indexing, and that clean-up is outside the package.
Sources
- OpenAI, API pricing —
- Anthropic, Claude pricing —
- OpenAI, ChatGPT Business pricing —
- Microsoft, Copilot business pricing —
- Notion, Pricing —
- Chatbase, Pricing —
- Intercom, Pricing —
- Pinecone, Pricing —
- Hamza Shabbir, Cost to Build a RAG Chatbot in 2026 —
- SoluLab, RAG Development Cost —
- gmware, RAG Implementation Cost —
- Magesh et al., Hallucination-Free? (Stanford) —
- OWASP, LLM08:2025 Vector and Embedding Weaknesses —
- Moffatt v. Air Canada, 2024 BCCRT 149 —
- EUR-Lex, GDPR, Articles 28 and 32 —
- CNIL, Questions and answers on generative AI —