AI automation

How much does a RAG chatbot for company documents cost, and should you build or buy?

A RAG chatbot on company documents costs roughly $4,000–$12,000 for a one-source MVP and $15,000–$80,000 for production scope, according to 2026 price guides, plus about $22 a month in model fees for 20 staff. Ready tools charge about $18–$25 per seat a month. VITON13 builds a bounded version from $113.

What is a RAG chatbot, in plain words?

RAG stands for retrieval-augmented generation. When someone asks a question, the system first searches your documents, then hands the best passages to a language model and asks it to answer from them, with links to the sources.

Two things follow. Nothing is retrained: your files sit in a search index outside the model. And the answer is only as good as the search, so duplicate, outdated or badly scanned files give fluent wrong answers.

How much does it cost to build a RAG chatbot?

Published 2026 price guides disagree by a factor of ten or more, because each prices a different project. None is a quote.

Published build prices for a RAG chatbot on company documents (read 2026-10-04)
Source and typeScope as statedBuild priceRunning cost as stated
Hamza Shabbir, freelance engineer, 9 Jun 2026MVP: one data source. Production: several sources, evaluation, monitoring$4,000–$12,000 MVP; $15,000–$40,000+ production$5–$30 per 1,000 questions plus $25–$150 a month
SoluLab, development agency, 31 Jul 2026Basic: one source. Mid-level: 2–4 sources, custom interface$8,000–$25,000 basic; $25,000–$50,000 mid-levelNot priced
gmware, software firm, 22 May 2026Production chatbot on one knowledge base; proof of concept with evaluation$30K–$80K production; $50K–$150K proof of concept$400–$6,000 a month
VITON13Up to 200 approved documents, 3 sources, 3 user groups, 50 test questions, 1 interfaceFrom $113 AI-assisted or $173 human-ledModel, vector database and hosting fees go to the providers
Quoted as each source prints them, before tax. The firms selling the work wrote these pages, so they show the spread, not a market average.

Read the VITON13 row as a narrower scope, not a discount: the scope caps documents, sources, user groups and test questions, and leaves scanning, integrations and provider fees outside the price.

Why do two quotes for the same chatbot differ so much?

  • Data preparation: Hamza puts it at 30–40% of a build budget and gmware at 30–50%. In Hamza’s experience a folder of clean Markdown takes a day and years of scanned PDFs take weeks.
  • Permissions: if staff may read different documents, the rules must travel from the source system into the search layer. gmware asks: does the sales team see HR’s documents?
  • Evaluation: Hamza prices a 50–100 question test set at $1,000–$3,000 and calls it the best money in a build.
  • Data location: OpenAI and Anthropic both add about 10% for regional data-residency endpoints; private hosting adds operations.
  • Who builds it: Hamza says agencies quote two to three times a solo engineer for the same scope, and his own list is the low end of the table.

Can a ready-made tool replace a custom RAG build?

For a team that already lives in one suite, often yes. The table shows what each vendor prints on its own pricing page today.

Ready-made options that answer from your documents (vendor pages read 2026-10-04)
ToolPrice on the vendor pageWhat it readsCheck before buying
ChatGPT Business$20 per seat a month billed annually, $25 monthly; 2–200 employees; Enterprise by quoteFiles in shared Projects; Google Workspace, Slack, GitHub, Microsoft 365Role-based access controls are listed for Enterprise only; the reasoning window is about 320 pages
Microsoft 365 Copilot Business$18 per user a month paid yearly (promotion for purchases 1 Jul–31 Dec 2026, first year only; list $21); $25.20 billed monthlyMicrosoft 365 contentAdd-on for an eligible Microsoft 365 plan, up to 300 users
Notion Business€19.50 per member a month (page priced in euros; yearly billing saves up to 20%)Notion pages and connected apps such as SlackAgents beyond the free trial: $10 per 1,000 credits
Chatbase$40, $150 or $500 a month for 700, 4,000 or 15,000 message creditsContent you upload, for a bot on your siteExtra credits cost $40 per 1,000
Intercom Fin$0.99 per outcome; helpdesk seats from $19 a monthYour help centre and support contentBuilt for support teams; a resolved conversation is the billing unit
Prices exclude tax.

Seats add up. Twenty people on ChatGPT Business cost $4,800 a year billed annually or $6,000 billed monthly; on the Copilot Business offer, $4,320 for the first year (our arithmetic). That is the range of a small custom build, so ask what a seat cannot do: run inside your website or Telegram bot, answer the public, or apply document-level rules.

What does a RAG chatbot cost to run each month?

Three meters run after launch: indexing your files, storing the index and paying tokens per question. Indexing is nearly free: 200 documents of 20 pages at roughly 600 tokens a page is 2.4 million tokens, and OpenAI’s text-embedding-3-small costs $0.02 per million, about five cents. Storage is small too: OpenAI’s file search includes 1 GB free.

Questions are the main meter. Assume 3,000 input tokens (instructions, five retrieved passages, the question) and 400 output tokens, the sizes Hamza uses, and 20 people asking five questions a day for 22 working days: 2,200 questions.

Model fees for 2,200 questions a month at list prices (read 2026-10-04, our arithmetic)
Model (price per million input / output tokens)One question2,200 questions
GPT-6 Luna, OpenAI ($0.10 / $0.50)$0.0005$1.10
Claude Haiku 4.5, Anthropic ($1 / $5)$0.005$11
Claude Sonnet 5.5 and GPT-6.1 Sol ($2 / $10)$0.01$22
Claude Opus 5.5 ($4 / $20)$0.02$44
Claude Fable 5.1 and GPT-6 Astra ($10 / $50)$0.05$110
Cached input costs a tenth or less: Sonnet 5.5 cache hits are $0.20 per million, which helps with the fixed instructions.

So the model bill for a small team is tens of dollars. Add $25–$150 a month of hosting, Hamza’s range. The cost that matters is a person who keeps documents current and reads the questions the assistant failed.

Where do RAG chatbots still give wrong answers?

  • The wrong passage is found: a table or clause split across chunks is answered from half of it, and a scan without a text layer gives the search nothing.
  • Old and duplicate files: gmware notes that retrieval returns the 2022 price sheet as readily as the current one. Give every document an owner and a date.
  • Hallucination survives: Stanford researchers found RAG-based legal research tools from LexisNexis and Thomson Reuters hallucinated 17–33% of the time, less than a general chatbot but far from the elimination vendors had claimed.
  • Access leaks: OWASP lists weak access control on embeddings, and shared vector stores that leak context between users, as an LLM top-10 weakness.
  • Poisoned content: a document with hidden instructions can steer answers; OWASP advises validating sources and logging retrieval.

A company answers for what its assistant says. In Moffatt v. Air Canada (2024 BCCRT 149) a British Columbia tribunal rejected the airline’s argument that its chatbot was a separate entity and ordered it to pay CAD 812.02 in damages, interest and fees after wrong fare advice. That bot faced the public, but staff act on an internal assistant’s answers too.

What do GDPR and the CNIL say about a documents assistant?

If your documents hold personal data, a model provider that processes it is your processor. GDPR Article 28 then asks for a written contract binding it to your instructions, and Article 32 for security measures that fit the risk. France’s regulator, the CNIL, adds that a company connecting an assistant to its own knowledge base answers for that processing, and that for personal data or sensitive documents an on-premise deployment is generally the safer choice.

This is not legal advice. Ask your data protection officer or lawyer where the provider processes your data, what it retains and whether files are used for training.

Build, buy or skip it: which checklist fits your case?

  • Skip RAG if your whole library is under about 100–200 pages and rarely changes. Hamza suggests a cached prompt, and the ChatGPT Business page lists a reasoning window of about 320 pages.
  • Buy a seat plan if everyone works in Microsoft 365, Google Workspace or Notion, users number a few dozen and the suite’s own permissions suffice.
  • Use a bot builder such as Chatbase if the audience is website visitors and the content is public.
  • Build custom if different groups may read different documents, answers must link to the passage, sources sit in several systems, you need your own interface or a re-runnable test set, or data must stay in a chosen region.

How VITON13 builds a documents assistant

The VITON13 package is the narrow version of the build above. We agree the sources, who may see what and the questions people ask most often; index the approved documents and wire in the access rules; then run 50 test questions with expected answers, fix wrong or unsupported answers and set the threshold for a clear ‘not enough information’ reply. After your team’s test and the revision rounds, the assistant goes live with the evaluation report. In the AI-assisted mode AI does the first pass and a person checks it; in the human-led mode a specialist leads each step and uses AI as support.

RAG knowledge assistant

Price
$113 AI-assisted or $173 human-led
Timeline
10–18 working days

AI chatbot with human handoff

Price
$73 AI-assisted or $93 human-led
Timeline
5–8 working days

The final price is fixed in writing after a short scope review.

AI agents for business

Price
$73 AI-assisted or $93 human-led
Timeline
5–8 working days

AI-assisted: AI drafts inside defined steps, a person directs and checks every result. Human-led: a specialist does the work and AI assists.

What is included

  • Up to 200 approved documents indexed from up to 3 sources
  • Access rules for up to 3 user groups
  • 50 test questions with expected answers and a report you can re-run after changes
  • One interface, a web chat or a messenger bot, with a source link under every answer
  • One update schedule that re-reads changed documents and drops removed ones
  • Delivery: 10–18 working days after the brief is approved

Not included

  • Model, vector database and hosting fees, which you pay to the providers
  • Scanning, OCR clean-up or rewriting of documents that are hard to read
  • Actions in other systems, such as creating tickets: that is an AI agent, with its own scope
  • Any promise of accuracy beyond the tested questions

Revisions: 2 revision rounds on the deliverables

Get an estimate in 1 working day

Written estimate within 1 working day

Example

PILOTPilot in progress

No published documents-assistant case yet

VITON13 has not yet published a case for a RAG assistant on a client’s documents with measured answers, so this guide invents no accuracy figures. Every number comes from a source below or from arithmetic shown on the page.

What this does not prove

  • The model-fee table is arithmetic on list prices for one assumed question size. It does not show how accurate any assistant is or what your usage will cost.

Questions and answers

Is a RAG chatbot the same as training ChatGPT on my documents?

No. Your files are indexed outside the model and passages are fetched when a question arrives, so nothing is retrained and a deleted file stops being used once the index refreshes. Fine-tuning mainly changes style and format, as Hamza notes.

How many pages can I give ChatGPT before I need RAG?

The ChatGPT Business page lists a 256K-token reasoning window, about 320 pages of input, and the answer must fit in it too. Hamza suggests loading material under 100–200 pages into a cached prompt. Past that, or when files change weekly, searching first pays off.

Can the chatbot show different answers to different employees?

Yes, if access rules sit inside retrieval, so a passage a person may not read is never fetched. Filtering after the model has seen it is too late. OWASP warns about shared vector stores that leak context between users. The VITON13 package applies rules for up to 3 user groups.

How do I test a RAG assistant before launch?

Write real questions with the expected answer and source, run them after every change, and check retrieval and answer separately: was the right passage found, and is the answer faithful to it? Hamza prices a 50–100 question set at $1,000–$3,000. The VITON13 package includes 50.

Do I need a vector database such as Pinecone?

Not always. Pinecone’s Standard plan has a $50 monthly minimum and its Builder plan is $20. Hamza says Postgres with the pgvector extension handles most knowledge bases up to a few million vectors for $0–$50 a month extra. Vector database fees are outside the VITON13 package.

What do you need from me to start?

Send the approved documents or their locations, who may see which of them, and the questions people ask most often. Scans and hard-to-read files need cleaning before indexing, and that clean-up is outside the package.

Sources

  1. OpenAI, API pricing —
  2. Anthropic, Claude pricing —
  3. OpenAI, ChatGPT Business pricing —
  4. Microsoft, Copilot business pricing —
  5. Notion, Pricing —
  6. Chatbase, Pricing —
  7. Intercom, Pricing —
  8. Pinecone, Pricing —
  9. Hamza Shabbir, Cost to Build a RAG Chatbot in 2026 —
  10. SoluLab, RAG Development Cost —
  11. gmware, RAG Implementation Cost —
  12. Magesh et al., Hallucination-Free? (Stanford) —
  13. OWASP, LLM08:2025 Vector and Embedding Weaknesses —
  14. Moffatt v. Air Canada, 2024 BCCRT 149 —
  15. EUR-Lex, GDPR, Articles 28 and 32 —
  16. CNIL, Questions and answers on generative AI —

About the author

Vitalii Tarasov

Founder of VITON13 Studio · Dubai, UAE

Vitalii Tarasov founded VITON13 in 2026 and runs VITON13 Studio from its Dubai office. He is responsible for every studio project: the written scope and price, the direction of the work and the check before delivery. Studio prices in this guide are read from the service pages, so they match what you would be quoted there.

About VITON13

Services this guide leads to