
Forward Deployed AI Infrastructure Specialist
Taking AI solutions from business problems to production-grade systems inside client VPCs, air-gapped networks, and Kubernetes GPU clusters.
11 Real Systems
Local Codebase Repos
100% Private
Air-Gapped & On-Prem
vLLM & Rust
GPU Inference Engine
Full Lifecycle
Problem → Cost Opt
Data Ingestion
OCR, Parsing, Chunking, Embeddings
Model Engine
Open-Weight LLMs, vLLM, Fine-Tuned LoRA
AI Application
Agents, RAG, Typed Tool Calling
APIs & Gateways
FastAPI, Semantic Router, Rate Limiter
Enterprise Systems
Core Banking, ERP, CRM, Identity
Kubernetes Cluster
GPU Autoscaling, vLLM Node Pools
Cloud & On-Prem
Private VPC, Air-Gapped Datacenters
Monitoring & Tracing
OpenTelemetry, Self-Hosted Evals
Alerts & Governance
Hallucination Checks, Drift Detection
Cost Optimization
Semantic Caching, Quantization

Senior AI Engineer | Forward Deployed Engineer
Singapore · Founder @ KaviAI
As a Senior AI Engineer & Forward Deployed Engineer (FDE), I sit inside customer engineering repos, infrastructure clusters, and compliance reviews to build, deploy, and scale AI systems.
My specialism is Sovereign AI & Private On-Premises Infrastructure — tuning and serving open-weight models (Llama 3, Qwen2, DeepSeek, Mistral) in private VPCs, air-gapped data centers, or local GPU nodes for banks, healthcare systems, and regulated enterprises.
From fine-tuning open weights to containerized vLLM serving, hybrid vector retrieval, agent tool calling, and Kubernetes GPU scaling — I deliver fully operational AI systems with complete cost optimization and self-hosted observability.
Click any project card to open an interactive deep dive into its architecture flow, API endpoints, performance metrics, and tech stack.
Turn your PC, Mac, or Linux box into a self-hosted private AI server with a single command
A self-hosted AI deployment platform built around 24 bundled Docker service manifests, hardware-accelerated overlays (NVIDIA, AMD, Apple Silicon, Intel Arc), a control dashboard, LiteLLM gateway, RAG pipeline, local voice STT/TTS, and privacy tools.
Local-first personal assistant with SQLite memory state, local web cockpit, and 95-line plain Python loop
A local-first personal AI assistant framework demonstrating the four pillars of serious agents: Harness, Reasoning Loop, Stateful Memory (SQLite), and LLM-as-Judge Evals. Features a local web dashboard cockpit at localhost:7777.
Desktop application with Tauri v2, xterm.js terminals, and KaviSwarm multi-agent pipeline
A Tauri v2 + Rust desktop application and multi-agent development environment that autonomously writes, builds, deploys, and live-demos full-stack applications through a three-phase AI swarm pipeline (KaviSwarm).
AI agent platform planning, generating, and publishing content across 10+ platforms automatically
An enterprise AI agent platform for growth, marketing, and distribution. Uses Google Gemini with LiteLLM gateway and OpenAI fallback to automate SEO, Reddit, LinkedIn, X, Instagram, and YouTube publishing.
Self-hosted AI development infrastructure for secure, governed agentic coding
A self-hosted developer platform providing containerized workspace environments, AI coding agents, and governance controls. Allows developers and AI agents to code side-by-side inside controlled sandbox environments.
Production-grade Retrieval-Augmented Generation for enterprise knowledge management
A complete corporate organization RAG system that ingests internal documents, enables hybrid search across organizational knowledge, and provides intelligent Q&A through agentic retrieval with LangGraph.
Handwriting-aware cheque processing with automatic fraud detection for regulated banking
A production-ready Bank Cheque OCR API automating the extraction and verification of critical information from scanned bank cheques. Designed for high-volume banking back-offices processing 5,000+ cheques daily.
Production-ready realtime voice agent for customer service call centers
A complete customer service voice agent system capable of handling inbound and outbound telephone calls with sub-second speech recognition, intelligent tool-calling responses, knowledge lookup, and conversation tracing.
Production-ready n8n automation workflows powered by AI agents for enterprise operations
A curated collection of 17 enterprise-grade n8n automation workflows integrating LLM agents, automated triage, email notification generators, IT ticket processors, and daily reporting systems.
High-throughput document parsing with vLLM, Rust API gateway, and async GPU queues
A multi-stage asynchronous invoice processing system engineered with vLLM vision model inference, a high-concurrency Rust API gateway, and async task queues for enterprise accounting teams.
Model Context Protocol (MCP) server giving AI agents web capture, screenshots, and page inspection
An open-source Model Context Protocol (MCP) server connecting AI coding assistants (Cursor, Windsurf, Claude Desktop, Cline) to PageBolt capture APIs for screenshotting, PDF generation, and page inspection.
Not just an AI developer. An FDE who owns the full lifecycle from initial business discovery to GPU cost optimization.
Identify workflow bottlenecks, regulatory constraints, ROI targets, and latency requirements.
Design model strategy, open vs closed weights, vector storage, context windows, and safety barriers.
Fine-tune open-weight models (LoRA/QLoRA), build RAG pipelines, typed agent tools, and evaluation harnesses.
Connect models to ERPs, core banking APIs, CRM database queues, and OAuth/RBAC identity systems.
Containerize with vLLM/Triton, deploy on Kubernetes, GPU autoscaling, and zero-downtime rollouts.
Self-hosted tracing, evals, hallucination monitoring, semantic logging, and token usage analytics.
Automated failover, load balancing, continuous eval gates, and human-in-the-loop fallback queues.
Semantic caching, model distillation, dynamic batching, and GPU node pool scale-to-zero.
Models add value when integrated cleanly into business workflows, databases, and core APIs.
Proven technology choices enabling enterprises to run AI models on-premises or in private clouds without third-party API dependencies.
From regulated banking change control to SaaS GPU platforms and forward-deployed client engagements.
KaviAI · kaviagentic.com
Enterprise SaaS
Payments
Banking
Available for Senior AI Engineer, Forward Deployed Engineer (FDE), or Sovereign AI Infrastructure contracts and permanent roles across Singapore, APAC, and EMEA.