Software engineering · distributed data · retrieval · ML systems

Atharva Gupta

Software engineer working on distributed data, retrieval, agent evaluation, and ML systems.

Currently an AI Engineering Intern at Cisco. Previously built medical-imaging software at PulseImaging.AI and geospatial data pipelines at CASFER. Studying Computer Science and Applied Mathematics at Case Western Reserve University.

28K+ records / cycle Cisco external-data ETL; employer-scoped measurement
2,229 watersheds national CASFER data and publication scope
<100 ms PulseImaging measured batch inference latency
3 walkthroughs product artifacts, code, results, and explicit limitations

01 / Selected technical work

Three implementation walkthroughs

Each page starts from an artifact: a running interface, a publication figure, or source code. Metrics are labeled by their actual scope.

Robodex desktop workspace showing generated robot candidates, a 3D robot model, and a degraded-mode warning

Robotics · agent evaluation · local software

Robodex

A robot-design workspace and a protected evaluator for checking agent-produced robot files independently of the agent that wrote them.

Implemented
Frozen artifacts, structural and MuJoCo graders, replay digests
Verified
6 references pass; 6 seeded failures caught; 12 outcomes replay 3×
Unresolved
Live two-profile comparison remains unrun
Read the Robodex walkthrough →
Directed network of nitrogen watershed classifications with source-to-sink transitions highlighted in red

Geospatial data · distributed ingestion · research

CASFER national nutrient data

Go and Spark data work supporting a national analysis of nutrient sources, sinks, and adjacency across HUC-8 watersheds.

Implemented
Concurrent crawlers, raster normalization, identifier recovery
Scale
2,229 watersheds; estimated 20M+ nitrate/phosphate cells
Output
Second-author paper in Resources, Conservation and Recycling
Read the CASFER walkthrough →
PulseImaging desktop viewer rendering a generated, non-clinical CT volume in axial, sagittal, and coronal planes

Electron · DICOM · model serving

PulseImaging desktop viewer

A TypeScript/Electron viewer that preserves DICOM hierarchy, builds a cached 3D volume, and renders linked axial, sagittal, and coronal views.

Implemented
Three-plane MPR, 8-slice loading batches, segmentation UI
Serving
PyTorch U-Net through NVIDIA Triton and FastAPI on AWS
Artifact
Real product code rendered with a generated 96-slice fixture
Read the PulseImaging walkthrough →

02 / Experience

Engineering experience

This section stays at resume parity. Proprietary work is described only at the public scope of the resume.

Jun 2026—Present
San Jose, CA

AI Engineering Intern

Cisco

  • Built a Kubernetes deep-research system coordinating planner, retriever, and synthesis agents with MCP tools for source-cited answers.
  • Built PySpark/PyArrow ETL for 28K+ external records per quarterly cycle, live feeds, and Common Crawl into a Blob-backed evidence lake and semantic index; orchestrated updates with Azure Functions.
  • Checkpointed agent state by session, sandboxed runtimes in Podman, and streamed LangSmith traces for grounding, tool use, latency, failure, and cost evaluation.
  • Designed Cisco/Splunk’s external-signal churn platform, selected among 7+ providers, and worked with vendor engineers to join external evidence with internal telemetry.
Sep 2025—May 2026
Cleveland, OH

Full-Stack Software Engineer Intern

PulseImaging.AI

  • Built a TypeScript/Electron DICOM viewer with axial, sagittal, and coronal MPR plus 8-slice concurrent loading.
  • Engineered DICOM ETL for university-hospital studies, validating study/series/instance hierarchy and normalizing metadata and pixels into reproducible datasets.
  • Served PyTorch U-Net segmentation through NVIDIA Triton/FastAPI on AWS, with sub-100-ms batch latency and full scans under two minutes.
Jun 2025—Aug 2025
Cleveland, OH

Software Engineering Intern

CASFER

  • Wrote concurrent, distributed Go crawlers with a shared manifest for ingestion across 2,229 HUC-8 watersheds and multiple APIs.
  • Built PySpark/PyArrow ETL on CWRU’s Slurm cluster for an estimated 20M+ nitrate/phosphate raster cells, producing Geohash-partitioned data and recovering 22 HUC-8s lost to identifier mismatches.
  • Prepared 60K+ geolocated AFO/CAFO image records with PySpark on AWS Glue and transfer-learned YOLOv8 for detection and classification.

03 / Research

Publication

Nitrogen watershed transition probability network from the published study

Journal article · second author

“Towards Circular Nutrient Economies”

O. D. Akanbi, A. Gupta, et al. Resources, Conservation and Recycling 226 (2026), 108697.

I curated and validated data for a national study spanning 2,229 watersheds. The figure at left is the paper’s agricultural nitrogen transition network.

Paper and DOI ↗

04 / Education

Case Western Reserve University

B.S. in Computer Science and Applied Mathematics; 3.6 GPA; expected December 2027.

Languages

Python, Go, TypeScript / JavaScript, SQL

Systems and data

Distributed systems, concurrency, Unix/Linux, information retrieval, PySpark, PyArrow, ETL, data lakes, Slurm

Agents and ML

Deep-research agents, HGM, LangGraph, MCP, LangSmith, LLM-as-a-judge evaluation, PyTorch, NVIDIA Triton

Cloud and applications

Azure, AWS, Kubernetes, Docker/Podman, Next.js, Electron, FastAPI, PostgreSQL

Contact

Software engineering internships for 2027