Product Build Note
Building Ask Davida RAG chatbot with memory, guardrails and observability
Written by
David Davila
August 20, 2026
Ask David started as a local Python experiment with ChromaDB. The production version needed a different shape: something that could live inside the PGP website, run safely on Vercel, store usage data, and evolve into a monitoring portal for AI product experiments.

Product Principle: Start small and get good data
Ask David is not meant to behave like a generic assistant with a polished voice. The product decision was to make it narrow, transparent and useful: it answers questions about my experience, product philosophy and working style, and it admits when the recorded knowledge base does not contain enough context.
That boundary matters. The chatbot can use conversation memory to understand the visitor, but it cannot use that memory to invent new facts about my experience. Claims about my experience must come from the curated corpus only.
Architecture: Next.js, OpenAI and Supabase
The production architecture runs inside the existing Next.js site. The modal calls a server-side endpoint which handles OpenAI calls, vector search and logging. Secrets stay server-side in Vercel environment variables.
Supabase was chosen as the persistence layer because the same Postgres project can support vector search with pgvector, structured run logs, session memory and future dashboards for usage, cost and quality monitoring.
Knowledge base and ingestion
After preparing several questions that I asnwered in recorded meetings, I used the transcripts to create the curated corpus A TypeScript ingestion script reads those files, chunks them by Q&A blocks, generates embeddings with text-embedding-3-small and upserts the chunks into ask_david_documents.
The local Python RAG remains a lab for experimentation. Validated ideas are reimplemented in TypeScript before reaching production.
Retrieval, typo correction and answering
Before embedding a visitor question, the pipeline applies a lightweight typo normalizer as if a question doesn't match the question in the corpus then it was ignoring useful answers. That vocabulary is generated from the corpus and protected with rules for common product, pricing and strategy terms so correction helps retrieval without rewriting user intent.
The query embedding is searched against Supabase using cosine similarity. Only chunks above the similarity threshold are used. If nothing reliable is found, Ask David returns a fallback instead of forcing an answer.
Logs as product observability
Every run is stored in ask_david with the question, normalized question, answer, retrieval success, chunks used, models, token counts, latency and estimated cost. Input and output tokens are separated because generation pricing depends on direction.
This makes the chatbot measurable from day one. The same schema can later power a dashboard by project, model, day, session, cost, retrieval quality and fallback rate.
Session memory without modifying the corpus
Conversation memory is stored separately in ask_david sessions and messages. After each turn, the system updates a compact memory summary of the visitor's context: company, goals, product challenge, constraints and previous questions.
On later turns, that memory helps tailor the answer to the visitor. It does not change the Ask David corpus, and the prompt explicitly says that memory is only for user context, not for facts about myself.
Guest flow, bot prevention and data collection
The UI lets a visitor ask three questions as Guest. After that, the chat input pauses and asks for name, email and company. Each field must be confirmed with an inline check before the conversation continues. Also, it applies a Tursntile check to prevent bots from impersonating visitors.
The product intent is to give value before asking for data. Once the visitor has context, the system asks for enough information to support follow-up and better session memory without making the first interaction feel like a lead form.
How to build it at a glance
Runtime
Next.js API route, OpenAI embeddings and generation, Supabase vector search, server-side secrets in Vercel.
Data
Markdown corpus, allowed words, vectorized chunks, run logs, messages and compact session memory.
Controls
Similarity threshold, fallback response, no source notes in UI, no corpus mutation from user conversations.
Measurement
Model names, input/output tokens, embedding tokens, latency, retrieval success and estimated cost per run.
“From local experiment to implemented product. A small assistant can still have serious product architecture.”

Written by David Davila
About the author
David Dávila Arimuya is a product leader and founder of P&G Partners, focused on product strategy for Cloud, AI and B2B2C companies.