AI platformsthat don't break.
I design the system, build it with the team, and own the part where it has to survive production — 50+ enterprise tenants on one platform, fully isolated, 99.9% uptime.
312
req / sec now
118
p95 ms
50+
tenants
01 — selected work
Platforms in production
01 / 05 · names withheld
Also delivered
Digital twins of industrial plants · multi-plant SSO with CyberArk · IPSec tunnels into client environments · an AIOps tool extending HolmesGPT across GCP VMs and Kubernetes · MLOps and CI/CD pipelines · a spec-driven development workflow adopted team-wide.
Architecture · Platform engineering · Team lead
Multi-tenant chatbot training SaaS
The problem
Enterprise customers needed to train and deploy chatbots on their own private data, with strict isolation between organisations and no shared state between tenants.
What I built
A multi-tenant platform on Google Kubernetes Engine, with a dedicated namespace per tenant and isolated data stores — MongoDB StatefulSets, Redis, and Qdrant vector collections per organisation. Horizontal pod autoscaling and an Nginx ingress controller handle concurrent peak load; Chargebee runs subscription billing and plan enforcement.
Edge
Nginx ingress · TLS · rate limits
Namespace
API pods · training workers · HPA
Stores
MongoDB · Redis · Qdrant — per tenant
Platform
Chargebee · plan enforcement
Outcome
50+ enterprise clients, each with 1,000+ users — sustained low-hundreds requests per second at peak, with node and pod autoscaling absorbing burst traffic.
02 — architecture
One platform, fifty-two tenants
1 connection · click any node
Client and infrastructure
Edge and network
Build and release
Production GKE cluster · one namespace per tenant
Master namespace · shared control plane
Tenant boundary · fully isolated stack per client
Shared platform · managed and third party
Client surface
Embedded chat widget
Served on client websites
- tenancy
- One script tag per tenant, keyed to their namespace.
- data
- Carries no secrets — the tenant key resolves server-side.
- failure mode
- Fails closed and hides itself, so a broken widget never breaks the host page.
- scaling
- Cached at the edge. The heaviest traffic on the platform and the cheapest to serve.
03 — live console
What operating it looks like
representative data · built for this page · no client data used
Enterprise · asia-south1 · last 24 hours
requests
22,232
+12.4% vs previous
p95 latency
161 ms
SLO 250 ms
error rate
0.11 %
budget 0.5%
active users
1,534
unique in window
Requests
requestsp95
Top intents
04 — engagement
How the work runs
01 / 06
01 — Assessment
Designs and requirements read in full, so scope, effort and risk are fixed before anything is quoted.
05 — about
Nine years, one direction

What I work with
Architecture & platform
01AI & data
02Backend
03Frontend
04Cloud & infrastructure
05Security & networking
0606 — contact
Let's talk about what you're building.
If you're building a platform, I can take it from architecture to production — or step in wherever it currently stands. Either way it starts with understanding the system properly before anything gets committed to.
A first call is usually thirty minutes: what you're building, where the hard parts are, and whether I'm the right person for it.