Skip to main content
Mukerem Shifa
Back to projects

AI/ML · 2026 — Present

Conversekit: Multi-Tenant AI Chat Platform

Drop-in conversational AI chat widget for any website with multi-tenant support, eleven AI vendor integration, RAG over custom documentation, deployed on Cloudflare Workers with Supabase backend.

Role
Sole engineer, from architecture through deployment
Timeline
2026 — Present
Team
Solo project
Status
In progress
Stack
  • TypeScript
  • React
  • Hono
  • Cloudflare Workers
  • Supabase
  • Tailwind CSS
  • OpenAI
  • Anthropic
  • Google Gemini
  • Groq
  • Workers AI
  • RAG
The ConverseKit widget open on a demo site, answering a visitor’s question about which AI vendors it supports, beside the install snippet — one script tag — and the headline “Drop-in AI chat for any website.”

Overview

ConverseKit is a production multi-tenant platform deployed as four separate Cloudflare Workers—the API, dashboard, widget CDN, and landing page—each with its own hostname, cache policy and blast radius. Onboarding a new client is a single SQL row insertion and a script tag; no redeploy, no build step on the customer side.

The platform abstracts eleven different AI vendors behind a single interface, from OpenAI and Anthropic to free options like Gemini Flash Lite and Workers AI. Each bot maintains its own knowledge base built from documents, markdown, and URLs through hybrid lexical and vector search with reranking.

Tenancy is enforced at the database layer through row-level security keyed off organization membership, not application code. Each bot is origin-locked to an explicit allowlist, and every request goes through multiple layers of isolation before the model is ever called.

Key features

  • Eleven AI vendors, one interface

    OpenAI, Anthropic, Gemini, Groq, OpenRouter, Mistral, Workers AI, DeepSeek, Together, Ollama and LM Studio plus any OpenAI-compatible endpoint. Switching is a dropdown per bot; the platform defaults to Gemini Flash Lite for chat and Workers AI for embeddings, verified to run at no cost.

  • Answers from your documents

    Text, markdown and URLs become a searchable knowledge base. At question time the query is embedded and matched against that bot's chunks by cosine similarity, and the best passages go into the prompt.

  • Captures leads mid-conversation

    The model collects a name, email or phone number during conversation and files it for CSV export, converting chat interactions into business outcomes.

  • Streaming with fallback

    Replies arrive over Server-Sent Events; a transport failure falls back to a buffered endpoint so the visitor still gets an answer even if streaming disconnects.

  • Isolated by the database

    Row-level security keyed off organization membership, not application code that a refactor can quietly drop. A permission change cannot be forgotten at any of a dozen call sites.

  • Origin-locked deployments

    Each bot allows an explicit list of origins; a request from anywhere else is refused before the model is ever called. No CORS surprises, no open vector indexes.

  • One environment, deployed from a laptop

    There is no staging. One Supabase project, one set of four Workers, one branch. What `npm run dev` runs is what is deployed. The database is real—nothing local is a rehearsal.

  • Hybrid retrieval with rank fusion

    Lexical index catches exact terms and identifiers that embeddings blur. Vector index catches paraphrase. Two ranked lists are fused before a cross-encoder reranks whatever survives.

Lessons learned

  • Four Workers with separate hostnames contain blast radius better than one. Each has its own cache policy and can be deployed independently.
  • Tenancy is an infrastructure concern, not a business logic one. Push it to the database layer so it cannot be forgotten.
  • Vendor abstraction at the interface level is cheap; running eleven vendors in production is not. Test the free tier path before making it the default.
  • The cost of a second environment is the cost of a second project, second set of secrets, and second thing to be wrong. One real environment early saves investigation time later.
  • Onboarding a client as a single row insertion with zero deployment friction scales the business model as much as the code does.

Case study

Challenge
A naive RAG implementation returns the right answer but tells the visitor what documents they are not allowed to read. Filtering results after retrieval still shifts the ranking of everything else, leaking information through the shape of the response.
Decision
Move the boundary into the database. Tenant and role scoping became arguments to the index query and a row-level security policy, so the application cannot forget them and the mistake is one policy in one place rather than missing conditions at a dozen call sites.
Outcome
Results are now deterministic per role. The same question from two users with different permissions is two separate queries, not one result and a filter. The coupling has held up—changes to permissions now require re-running the query rather than re-filtering a cached list.

Tell me what you are building.

I read everything that arrives. If you are hiring, scoping a project, or stuck on a multi-tenancy or LLM integration problem, a few sentences about the constraint is enough to start a useful conversation.