Available for AI/ML collaboration

Building intelligent systems
that reason, see & scale.

Applied AI engineer & team lead who takes frontier research to production — RL post-training (GRPO) of SOTA LLMs & VLMs, multimodal pipelines fusing 5+ models, and architecture-level tuning for fast inference. Full stack: from model weights to Kubernetes.

▋
View Work → ↓ Download Résumé Get in touch
✉thegigasurgeon@gmail.com 📞+91-9915503368
Delivered Value

Outcomes, not just models.

0
PDFs processed daily by AI platform I lead
0
Document extraction accuracy — up from 35%
0
Developer-hours saved weekly via LLM orchestration
0
Inference pipeline speedup from optimization
0
FPS gain — real-time video analytics on the same GPUs
0
Post-processing time eliminated in geospatial CV
0
Engineers led across validator jurisdictions
0
SOTA models fused in one production pipeline
About

Research-grade AI, engineered for production.

I'm an AI engineer working across NLP, vision, and everything multimodal — at the intersection of deep-learning research and large-scale systems. I pretrain and fine-tune state-of-the-art LLMs and VLMs (LLaMA 3.3, Qwen 3, Gemma, Gemini 3.5 Pro, GPT-4o) using LoRA/QLoRA and RL post-training with GRPO to build reasoning and thinking capabilities into domain models.

Today I lead the development of an enterprise-scale, agentic AI document intelligence platform processing 80K+ PDFs a day — multi-agent text + visual reasoning that fuses 5+ specialized models (detection, OCR, VLMs, reasoning LLMs, embeddings) in a single microservices-based application, with tool-calling, subagents, and skills. Along the way I've built RAG systems, multi-LLM orchestration infrastructure, and extraction pipelines that took accuracy from 35% to 99.85%.

I also go below the framework line — tinkering with model architectures, quantization, and graph-level optimization to hit aggressive inference targets. Those instincts come from years of computer vision: real-time sports analytics, satellite segmentation, pose estimation, and edge deployment. I care about the whole stack, from a novel architecture to a quantized engine shipping 3× the FPS.

At a glance

RoleAI/ML Engineer
LocationDelhi, India
FocusNLP · Multimodal · CV
TeamLed 8 engineers
EducationB.Tech CSE · 8.28 GPA
Featured Work

Systems I've shipped in production.

// Work is proprietary — the visuals below are architecture illustrations, not source or screenshots.
Experience

Where I've built.

AI/ML Engineer · Real Brokerage
Oct 2024 — Present
  • Pretrained & fine-tuned SOTA LLMs and VLMs with reasoning capabilities (LLaMA 3.3, Qwen 2.5 72B, Qwen 3, Gemma) using LoRA/QLoRA and GRPO-based RL post-training, tailored for real-estate auto-review.
  • Lead development of an enterprise agentic AI document intelligence platform processing 80K+ PDFs daily — 5+ models fused per pipeline via multi-agent text + visual reasoning with tool-calling, subagents & skills, deployed as microservices.
  • Built scalable Multi-LLM Orchestration Infrastructure, saving thousands of developer hours weekly.
  • Fine-tuned Gemini 3.5 Pro/Flash and GPT-4o for robust multi-LLM agentic workflows.
  • Engineered an end-to-end RAG pipeline: intent routing, semantic chunking, hybrid vector retrieval, multi-source context aggregation & structured output.
  • Integrated the Claude Agent SDK for automated PR reviews, incident investigation & webhook-driven workflows.
  • Boosted extraction accuracy from 35% → 99.85% with YOLOv9, Doctr OCR & custom VLMs; achieved 200% pipeline speedup via graph optimization, batched inference & quantization.
Deep Learning Engineer · Streamingo AI
Nov 2022 — Oct 2024
  • Developed temporal event-understanding systems with Swin, VideoMAE, MVD, ViT, BEiT.
  • Built a real-time sports analytics engine — object detection, multi-object tracking, pose estimation & court localization.
  • Designed zero-shot & few-shot classification with CLIP, Detic & Prototypical Networks for rapid scaling to unseen classes.
  • Deployed optimized inference with TensorRT & ONNX, delivering 3× FPS.
Computer Vision Research Engineer · Attentive AI
Apr 2021 — Nov 2022
  • Built & trained SOTA networks for satellite image segmentation (HRNet, Vision Transformer).
  • End-to-end ownership of CV products (Falcon & Respod) across the company.
  • Improved model accuracy by 12% and cut post-processing time by 90%.
Software Engineer · Nokia
Jul 2020 — Mar 2021
  • Created a meta-learning platform for dynamic model selection & hyperparameter optimization.
  • Integrated Explainable AI (SHAP, LIME) for model interpretability & compliance.
Under the Hood

How I take models to production.

🚀 Deployment

Multi-cloud and self-hosted — each model lands where its cost, latency, and data-privacy profile fits best.

AWS · EC2 / ECS / SageMakerGCP · Vertex AI / GKE Kubernetes + DockerMicroservice architecture AutoscalingSelf-hosted GPU clusters

⚡ Inference Optimization

The 200% pipeline speedup and 3× FPS came from stacking these — measured at every step, not guessed.

INT8 / FP16 quantizationAWQ / GPTQ TensorRT enginesONNX + graph fusion Continuous batchingKV-cache reuse

🎯 Model Adaptation

Right-sizing intelligence: RL post-training for reasoning, low-rank adaptation for the domain, and architecture surgery when off-the-shelf isn't fast enough.

GRPO (RL post-training)LoRA / QLoRA DistillationPruningArchitecture modification LLaMA / Qwen / GemmaGemini 3.5 Pro / GPT-4o tuning
Toolkit

Technologies I work in depth.

🧠 LLMs, NLP & Post-Training

LLaMA 3.3Qwen 3Gemma Gemini 3.5 ProGPT-4o GRPO (RL)LoRA / QLoRA Reasoning modelsMultimodal fine-tuning

🤖 Agentic AI

Multi-agent systemsTool-callingSubagents Claude Agent SDKAgent skillsLangChain Webhook-driven agents

🔎 RAG & Retrieval

Intent routingSemantic chunkingHybrid retrieval Vector storesEmbeddingsContext optimization Structured output

👁️ Vision & VLMs

YOLOv9Doctr OCRSwinVideoMAE ViT / BEiTCLIPDeticHRNet Detectron2Pose estimation

⚡ Inference & MLOps

TensorRTONNXQuantization Batched inferenceGraph optimizationMicroservices Kubernetes / DockerEdge / Raspberry Pi

💻 Languages & Frameworks

PythonC++JavaScript PyTorchTensorFlowOpenCV SLAM / ORBv3
Selected Projects

Things I've built for fun & research.

01

Basketball Video → 3D Mapping

Reconstructed player trajectories in 3D from monocular sports footage using deep video transformers and classical SLAM. Added temporal event summarization via transformer-based description generation.

PyTorchVisual SLAMORBv3BlenderVideoMAE
02

Tekken with Pose

A camera-driven interface that maps real-world body actions into virtual gameplay using pose estimation, latency-aware gesture recognition, and real-time control mapping.

YOLO-NASPose EstimationRaspberry Pi
03

Charlie — AI Virtual Assistant

An edge-deployable assistant with multi-modal recognition: speaker ID, facial login, offline speech-to-text, and contextual response generation — all running on-device.

Raspberry PiSTTFaceNetPython
Contact

Let's build something intelligent.

Open to conversations about AI/ML engineering, applied research, and hard problems in LLMs, VLMs, and computer vision.

Email thegigasurgeon@gmail.com
Phone +91-9915503368