Data / Models / Pipelines

I make messy data confess

SiddhiRohan

Data Scientist / Data Engineer. MS Data Science, University of Maryland. Three-plus years shipping GenAI, pipelines, and the guardrails around them.

↓ keep digging

Open to work · San Jose, CA · --:--

First, the mess.

Then the pipeline…

Hebrew soil records across 900 tables. Ten enterprise systems that never agreed on anything. A chatbot that made things up. I turn that into pipelines, models, and AI systems that can show their work.

and hold up under scrutiny →

40%fewer LLM
hallucinations
20+Airflow pipelines
orchestrated
5K+monthly queries
in production
15K+soil records
translated & standardized
01

Generative AI

  • RAG
  • LLM Fine-Tuning
  • LoRA
  • FAISS
  • Vector Search
  • Evaluation
02

Data Engineering

  • Python
  • SQL
  • Apache Airflow
  • PostgreSQL
  • MongoDB
  • Supabase
  • ETL
03

AI Governance

  • RBAC
  • Output Scanning
  • FERPA / HIPAA
  • FastAPI
  • Audit Trails

the unglamorous part. my favourite.

04

ML & Analytics

  • PyTorch
  • TensorFlow
  • Scikit-Learn
  • Pandas
  • Tableau
  • Power BI
  • ArcGIS

Currently / in the workshop

Next up.

Still half-built…

Private beta · final fixes

BuildTurn

A coding orchestrator for AI agents. More when it's ready for strangers.

coming soon

In build

Recall for Agent Work

When you correct something your AI agents relied on, we find the affected customer interactions and help you put them right.

coming soon

Want early access? ↗

Built some.

One teaches itself…

01 AI Governance / Open Source

Arbiter

a bouncer for your LLM

Bidirectional AI governance middleware for LLM applications. Enforces role-based access control before the model ever sees a query, detects inference-channel and cross-query accumulation attacks, and scans outputs for policy violations in FERPA/HIPAA contexts.

  • FastAPI
  • RBAC
  • LLM Security
View on GitHub ↗
A brass balance scale weighing glowing fibre-optic threads against a steel padlock
arbiterfig. 01

02 Self-Improving Local LLM

Morpheus

it fine-tunes itself. mid-conversation.

A local LLM (Qwen 2.5-3B) that fine-tunes its own LoRA adapters in real time during conversation, using Claude Sonnet as an automated critic to generate training signal. Concurrent inference and training on a single consumer GPU with under 2GB VRAM overhead.

  • Qwen 2.5-3B
  • LoRA
  • Claude Critic
View on GitHub ↗
A cracked plaster head with glowing circuitry inside the cracks, beside a soldering iron
morpheusfig. 02

03 AI + Healthcare

MedPal

a study buddy that notices burnout

AI-powered wellness companion with adaptive study planning, PDF summarization, flashcard generation, and real-time burnout detection using sentiment analysis.

  • Gemini API
  • Streamlit
  • Plotly
View on GitHub ↗
A glass anatomical heart on a desk beside textbooks and a stethoscope
medpalfig. 03

04 Social Impact

MealBridge

leftovers → logistics

Food redistribution platform connecting donors with NGOs through AI-powered matching, Google Maps route optimization, and automated volunteer coordination.

  • Flask
  • Google ADK
  • Maps API
View on GitHub ↗
Bowls of steaming food on a wooden plank bridging two concrete blocks
mealbridgefig. 04

05 NLP + Finance

BTC Sentiment

vibes, quantified

Real-time sentiment pipeline combining Selenium scraping, spaCy NLP, and VADER scoring with price correlation analysis in a Dockerized environment.

  • spaCy
  • Selenium
  • Docker
View on GitHub ↗
A gold coin balanced on crumpled newspaper lit by a glowing price chart
btc sentimentfig. 05

Experience / 2022 → now

The part with paychecks.

2025 → Now

Data Engineer

UMD / Environmental Science & Technology

Built a Python/NLP ETL pipeline translating and standardizing 15,000+ Hebrew soil records across 900+ relational tables. Loaded into PostgreSQL on Supabase powering a trilingual web app that cut retrieval time ~60%, with ArcGIS integration for spatial soil-profile queries.

2024 → 2025

Data Engineer / Research Data Scientist

UMD / Agricultural & Resource Economics

Redesigned a 50+ table schema and built the team's first repeatable data-quality framework; validation pipelines improved accuracy 20% and cut preprocessing 35%. Analyzed 200+ automotive production records across 5 regions, quantifying a 15% output decline via trend decomposition and spatial clustering.

2023 → 2024

Data Scientist

StackNexus / Hyderabad, India

Architected a RAG pipeline (vector search, chunking, LLM generation) reducing hallucinations 40% and improving accuracy 25% across 5,000+ monthly queries. EDA and segmentation across 12+ cohorts drove A/B tests with a 17% engagement lift; feature-extraction modules cut manual data entry 50%.

2024

AI Engineer

Kridha AI / Stanford, CA

Prototyped and validated classification and decision-automation pipelines for an AI product, testing 200+ edge cases to lift output accuracy 18%; logic shipped to early adopters.

2023

Data Scientist

Adventaus Technologies / Bengaluru, India

Engineered ML anomaly detection over infrastructure telemetry, cutting alert-review effort ~20%. Trained classification models on ~50K incident records; NLP ticket classification hit ~85% accuracy across 30K+ records; time-series forecasts supported capacity planning.

2022 → 2023

Data Engineer

Adventaus Technologies / Bengaluru, India

Built Python/SQL ETL consolidating 10+ enterprise systems into centralized analytics repositories. Orchestrated 20+ Apache Airflow pipelines, trimming 8-10 hours of weekly manual reporting; data-quality validation across millions of records; query optimization cut dashboard load times 25-35%.

Also, on paper.

Literally.

MS

University of Maryland

Master of Science, Data Science

BE

Osmania University

Bachelor of Engineering, Mechanical

IEEE · ICCD 2023

Neural Network-Based Synthesis for Digital Sound Processing

Intl. Conference on Cognitive Computing and Complex Data

read it ↗

so… got messy data?

Bug me with it.

Open to work · Data Science / Data Engineering / AI

© 2026 Siddhi Rohan