0 %

漏 2026 | Loading Samarth's world. 3D, systems, and AI.

Attention mobile users: "This website has heavy 3D components which may not be fully supported on your mobile, for the best interactive experience visit this site on desktop / laptop."

Software & AI Engineer

Samarth

Carnegie Mellon University CMU SCS 路 Scalable Systems 路 Edge AI

I build software systems that take AI from a notebook to production: inference infrastructure, agentic interfaces, and evaluation pipelines that other engineers can actually run.

img

About Me

I'm Samarth Shukla, a Software and AI Engineer and a Master's student in Software Engineering (Scalable Systems) at Carnegie Mellon University's School of Computer Science.

My work lives where research becomes product: WebGPU inference in the browser, agentic CLI workflows, evaluation harnesses on FastAPI and MongoDB, and export tooling that packages fine-tuned models for CUDA, Apple silicon, and edge backends.

That work has been recognized by the Meta Llama Impact Grant and NVIDIA Inception, published at IEEE and TechRxiv, and proven across nine national and one international hackathon win. I care about systems that are reliable, private by default, and ready for production.

Based in Pittsburgh. Built in Indore. Shipping software that holds up under load.

Major Projects

WebGPU Studio

WebGPU Studio:
Browser-side LLM inference on WebGPU and the Vercel AI SDK. Offline, privacy preserving execution that cut pipeline latency by 55% for 1,500+ users.

Robo Claw

Robo Claw:
An OpenClaw agent that turns plain English into executable CLI actions. Built in two weeks with dynamic skill loading, 15+ commands, and safe setup checkpoints.

HealthX

HealthX(AI):
On-device voice agent for early psychiatric screening and structured report generation. Privacy-first inference for clinics the cloud does not serve. IEEE IC-EETA 2025.

DisasterX

DisasterX:
Wildfire detection and emergency coordination with YOLO, OpenCV, React, Node, and MongoDB. One-click ambulance dispatch, live routes, and Twilio SMS alerts.

SOLO Export

SOLO Export:
Open-source LLM export CLI plus a Next.js UI. ONNX, TensorRT, TorchScript, CoreML, TFLite, Hugging Face. 3.2脳 faster TensorRT. IEEE TechRxiv preprint.

10

Hackathon wins. 9 national, 1 international, including Bit-n-Build on both stages.

55%

Lower inference latency on WebGPU Studio, serving 1,500+ users offline.

3.2脳

Faster TensorRT inference with 60% smaller models on Penn Treebank benchmarks.

1,500+

Users on privacy-preserving WebGPU execution, plus the Meta Llama Impact Grant.

Experience

Solo Tech

1 yr 8 mo

Mountain View, California 路 AI infrastructure, developer platforms

Founding Engineer
Dec 2025 - Present 路 9 mo

Promoted from intern to founding engineer to own Solo's AI platform end to end: the software that lets a team calibrate hardware, teach a policy, run inference, and put it in a customer's hands without a research bottleneck.

Python WebGPU Vercel AI SDK LLMs Agentic AI Edge AI

Architected a unified Python platform after promotion from intern: calibration, teleoperation, dataset recording, and policy training in one workflow. Standardized how engineering ships, and underpinned Solo Tech's CES 2026 enterprise showcase.

Built a no-code training studio: record demonstrations, select ACT, SmolVLA, Groot, Pi0 and other VLA policies, and train on managed Google Cloud GPUs.

Engineered WebGPU Studio with the Vercel AI SDK and MuJoCo for fully offline LLM/VLM inference. Privacy-first edge AI, then live talks and demos for 600+ attendees on a Pan-India tour.

Shipped Robo Claw: natural-language control that handles install, environment setup, calibration, and execution. Flagship capability on the CES 2026 floor.

Established AITR's first hands-on lab with Solo Tech, inaugurated by founder Dhruv Diddi, to expand practical AI learning.

Led AI infrastructure across training systems and agentic workflows, contributing to a Top 10 finish at SuperAI Genesis and $250K in AI cloud credits.

Software Engineer Intern
Jun 2025 - Dec 2025 路 7 mo

Turned fragmented LLM deployment into a single, measurable pipeline, then used that rigor to strengthen Solo's NVIDIA Inception story.

Python Next.js ONNX Runtime TensorRT CUDA Hugging Face

Pioneered SOLO Export, a Python framework that exports LLMs across ONNX, TensorRT, TorchScript, CoreML, TensorFlow Lite, and Hugging Face with FP16/INT8 quantization and benchmarking. Later published on IEEE TechRxiv.

Benchmarked deployments on Penn Treebank: 3.2脳 faster TensorRT inference and 60% smaller models, strengthening Solo Tech's NVIDIA Inception portfolio.

Led and mentored engineering interns while coordinating delivery for Bots & BEVS in San Francisco under conference-hard timelines.

Shaped product strategy across synthetic video, 3D reconstruction, and vision-guided manipulation, turning research ideas into implementation plans.

Standardized multi-platform LLM deployment workflows so teams could reproduce and ship AI products faster.

Software Engineer Intern
Jan 2025 - May 2025 路 5 mo

First tour at Solo: a continuous-learning pipeline for LLMs, a multimodal assistant, and customer-facing features that helped an emerging startup show up at CES 2025.

React FastAPI MongoDB Fine-Tuning RAG

Built the Auto Fine-Tune System, a continuous learning pipeline that later integrated synthetic data with Starfish Technologies, improving training efficiency by 1000脳 and contributing to the Meta Llama Impact Grant.

Integrated hybrid search, image generation, multimodal support, and speech using Serper, Stable Diffusion, LitServe, TTS/STT, and on-device LLMs.

Designed a Fine-Tuning Playground and feedback pipeline in React, FastAPI, and MongoDB so users could customize models while improving training data.

Shipped across frontend, backend, and AI infrastructure for Solo Tech's CES 2025 showcase.

Took customer-facing AI from concept to production: product engineering, full-stack development, and applied AI in one loop.

GMADP

North Brunswick, NJ 路 Remote

Full-stack donation platform for administrators

Software Engineering Intern
Jun 2024 - Jul 2024

Designed a scalable donations product with auth, caching, notifications, and receipts treated as one system, not four tickets.

Next.js Express.js MongoDB OAuth2 REST APIs TypeScript Redis

Designed a scalable full-stack donation management platform using Next.js, Express.js, and MongoDB with OAuth2 authentication and Redis caching, exposing REST APIs for CRUD operations.

Implemented Twilio-based notifications, donation tracking, and receipt generation, supporting 2,500+ active donors and 250+ monthly donations while automating administrative work and saving 12.5 hours per month.

Youngovator

3 mo

Indore 路 Web development, community, and a world-record workshop

Web Development Intern
Mar 2023 - May 2023

Engineering plus stage presence: helping run a record-setting workshop and a flagship innovation carnival while owning the digital craft around it.

Web Development Design Community

Contributed to setting a world record for the longest-running app development workshop during the event.

Served as an anchor and creative team member for Innovation Carnival 2.0, a milestone gathering for the community.

Played a key role in social post creation and design analysis, growing the event's online presence and engagement.

Stack I work with

Languages, systems, and AI I use to ship

Languages
Python JavaScript TypeScript SQL C++
Frameworks
FastAPI Node.js Express.js React Next.js REST APIs Typer OAuth2
Data
MongoDB PostgreSQL MySQL
Tooling
Docker Linux Git GitHub CI/CD Postman
Practices
Distributed Systems System Design Agile Unit Testing Integration Testing
AI and machine learning
LLMs VLMs VLA Models Agentic AI Edge AI PyTorch TensorFlow Hugging Face Transformers Fine-Tuning RAG WebGPU Vercel AI SDK ONNX Runtime TensorRT CUDA Ollama MuJoCo Imitation Learning

Publications

SOLO-EXPORT

Unified CLI for multi-format LLM export with post-export benchmarking. ONNX, TensorRT, TorchScript, TensorFlow Lite, and Hugging Face, with FP16 and INT8, device-aware artifacts, and a Penn Treebank harness. Up to 3.2脳 lower TensorRT latency and 60% smaller models.

IEEE TechRxiv preprint. Paper DOI GitHub

HEALTHX(AI)

Privacy-preserving on-device voice agent for early psychiatric screening and report generation. Edge LLMs, structured clinical summaries, and encrypted clinician handoff, built for places where the cloud is a luxury.

IEEE IC-EETA 2025, IEEE Xplore. Paper DOI GitHub

SMOTE 脳 XGBOOST

Money-laundering detection under extreme class imbalance. A SMOTE plus XGBoost pipeline that lifts recall and F1 without drowning compliance teams in false positives. Software systems research applied to financial ML.

IEEE ACROSET 2025, IEEE Xplore. Paper DOI GitHub

Honors

Hackathons 路 10 wins

Kriyeta 3.0, Skitech Innothon, Bit-n-Build (national and international), and Prayatna 2.0, plus 5 further national titles. 9 national and 1 international, against 2,000+ participants.

Industry recognition

Meta Llama Impact Grant, NVIDIA Inception, CES 2025 and CES 2026, SuperAI Genesis Top 10 with $250K in AI cloud credits, IEEE Xplore and TechRxiv papers, and 1,500+ users on privacy-preserving WebGPU inference. Recommendations on LinkedIn

Early academic

Best Project Award at Jr. Bal Vigyan, a national science presentation, and Head Boy, leading the student body before the internships began.

A few things people ask

What are you working on right now?

I am a Master's student in Software Engineering (Scalable Systems) at Carnegie Mellon University SCS, and I build production AI software: WebGPU inference, model export, evaluation pipelines, and agentic CLI systems.

What is the thread through your work?

Taking research-grade AI and making it operable: export it, quantize it, evaluate it, run it offline, and let a customer use it. Software systems, AI/ML pipelines, and research that ships. Not slides.

Where can I read the papers?

SOLO-Export is on IEEE TechRxiv. HealthX(AI) is in IEEE IC-EETA / Xplore. The SMOTE-XGBoost AML work is in IEEE ACROSET / Xplore. Links live in the Publications section next to the 3D canvas.

Are you open to conversations?

Yes, especially around AI infrastructure, on-device inference, evaluation pipelines, and ambitious product engineering. LinkedIn, email, and the resume linked in the hero are the fastest paths. Pittsburgh-based. CMU SCS.

How do I try the 3D experience?

Desktop or laptop is best. Click the equalizer to start audio. Hold Shift to fire the bot's weapons. Scroll to move the camera through the scene. The models, the particles, and the rocket path are all still here.