200-400 ms

response time, answering only from approved documents

The problem

A voice assistant can help customers explore company information in a natural conversation. Retrieval from approved documentation gives it relevant context, while output evaluation and human escalation remain important.

The project delivered a real-time voice agent with:

Near-instant responses (200-400 milliseconds)
Natural conversations where users can interrupt anytime
Information pulled exclusively from your company content
Seamless experience across desktop and mobile

Grounding responses in approved documentation

The heart of this assistant is RAG (Retrieval-Augmented Generation). It supplies relevant source material to help the assistant answer; it does not eliminate the possibility of errors.

What is RAG?

RAG is like giving your assistant a locked filing cabinet containing only your approved company documents. Before answering any question, the assistant:

Searches your filing cabinet: It scans your documents, FAQs, product guides, internal policies
Retrieves relevant information: It selects the passages that answer the question
Responds only with what it found: The prompt directs it to use retrieved sources and acknowledge missing information

Why this matters for your business

Without RAG, an AI assistant can:

Hallucinate answers that sound confident but are completely wrong
Share outdated information
Misrepresent your products or policies
Create legal and compliance risks

With RAG, you get:

Source control: Only your approved documents are used as sources
Maintainable knowledge: New documents become searchable after ingestion and indexing
Reduced risk of unsupported answers: Grounding and evaluation help reduce errors; they do not guarantee accuracy
Source traceability: Retrieved passages can support answer review

The tech stack

Here's the n8n workflow that powers the RAG system:

Original n8n workflow connecting incoming questions, an AI agent, embeddings, and a Supabase vector store
Original project workflow · n8n

The architecture includes:

Anthropic Claude: Powers the Master Knowledge agent for intelligent responses
Supabase Vector Store: Secure document storage with semantic search via pgvector
OpenAI Embeddings: Converts text into vector representations for fast retrieval
WebSocket: Persistent connection for lightning-fast back-and-forth
Next.js + React: Modern, responsive interface

How it works in practice

  1. You feed the knowledge base

    Upload your documents (FAQs, product guides, HR policies, etc.). The system chunks them and indexes them for search.

  2. A user asks a question

    “What's your return policy?”

  3. The assistant searches your documents

    It retrieves relevant passages in your documentation about returns.

  4. It formulates a natural response

    Using the retrieved documents, the assistant generates a clear, conversational answer.

  5. It responds with voice

    The response arrives in 200–400 milliseconds.


Use cases that drive ROI

Intelligent customer support: Support routine questions using approved source material. Your human agents focus on complex cases that need a personal touch.
Automated onboarding: New hires can ask any question about processes, policies, and tools—and retrieve relevant information, with escalation when needed.
Product knowledge assistant: Sales teams get instant voice access to all specs, pricing, and talking points. No more scrambling through documents during calls.
Training content assistant: A voice coach that knows all your training content and can answer learner questions on the fly.

Why this matters for SMBs and startups

Controlled costs

The assistant supports routine questions and helps the team reserve time for cases that need personal judgment.

Capacity planning

Capacity, latency, and response quality need testing against the expected volume of conversations.

A shared source of information

A maintained knowledge base provides a consistent reference; answers still need appropriate evaluation.

Modern experience

Your customers and employees get a smooth voice experience that matches the personal assistants they use every day.

The numbers that matter

200-400ms

End-to-end response time

<100ms

Reaction time when users interrupt

24/7

Designed for ongoing availability, subject to service operation

Source grounding

Responses are instructed to use official documentation


Ready to transform how people access your knowledge?

This technology isn't just for enterprise anymore. With the right tools and a well-designed architecture, any SMB or startup can deploy an intelligent, secure voice assistant.