profile
Machine learning engineer — Atlanta, GA
Prudhvi Vuppalapati
I build ML systems that stay correct in production.
Four years building, shipping and monitoring production ML on AWS. At CGI I own marketing ML end to end: an XGBoost propensity model that lifted lead conversion 18% in A/B tests, churn scoring rebuilt as a Kafka and Spark Streaming pipeline handling 500K+ records a day at sub-second latency, and the MLflow, feature store and drift-monitoring standards the team now works to.
Before that, at Lumen, a credit-risk model at 0.91 AUC-ROC served through FastAPI, and a fine-tuned transformer routing support tickets at 87% accuracy.
The repositories below are the same discipline applied in the open: features that respect time, models that are calibrated, explained and documented, LLM systems evaluated before release, and services that report when they start to drift. Each runs from a clean clone against synthetic data, with no cloud account required.
pipeline
Features, train, serve, monitor
Ten repositories, arranged the way a model actually moves: features built so they respect time, training that is calibrated, explained and documented, services that sit behind a real endpoint, and monitoring that catches the decay before a stakeholder does.
- Features 1 repo Built as of the decision, never after it.
- Train 3 repos Calibrated, explained and documented from the artifact.
- Serve 4 repos Behind a real endpoint, with retrieval that is measured.
- Monitor 2 repos Drift caught, prompts regression-tested, decisions recorded.
-
features
most complete
PointInTimefeature store
Training data that can only be built as of the moment the decision was made.
A credit model that scores 0.91 offline and 0.68 in production almost always learned from the future: a balance as of today joined onto a decision from eight months ago. This store makes point-in-time-correct joins the only way to retrieve training data. Facts are held bitemporally, as-of reads run through SQL window functions, and a streaming materialiser replays the fact log under a watermark so the online values match the offline ones.
A planted future-leaking feature is part of the test suite: the store refuses to serve it, and an automated train/serve skew check compares offline and online values for the same entity and timestamp.
-
train
CalibrateKitprobability calibration
A model that says 0.8 should be right about 80% of the time.
Accuracy says nothing about whether a probability can be trusted, and risk decisions are made on the probability itself. CalibrateKit measures miscalibration with expected calibration error and Brier score, draws reliability diagrams against the diagonal, and refits probabilities with isotonic regression or Platt scaling for any scikit-learn classifier.
Built as a library and a CLI, with the reliability mathematics covered by its own tests rather than demonstrated in a notebook.
-
train
InterpretabilitySHAP attribution
Why the model said no, in terms a reviewer can act on.
Global and local explanation for tabular models: SHAP values for individual predictions, feature-importance ordering across the dataset, dependence plots showing how a feature moves the output, and the failure cases where a single number hides an interaction.
The same attribution path that turns a churn probability into a reason code an operations team can put in front of a customer.
model-interpretability-for-machine-learning-models — open Interpretability on GitHub
-
train
ModelCardgenerated documentation
Model documentation derived from the artifact, not typed by hand.
Model cards are recommended everywhere and maintained almost nowhere, because writing them by hand guarantees they drift from the model. This reads the trained estimator and its evaluation metrics and renders the card: intended use, training data, metrics table, feature importances, limitations and ethical considerations.
Output is byte-identical for identical inputs, so the card can be committed and diffed in CI, and the demo runs fully offline.
-
serve
most complete
MLOps Append-to-end lifecycle
The whole path from training run to served prediction, wired up.
A full-stack MLOps reference: data in BigQuery, experiments and models tracked in MLflow with a registry gating promotion, a FastAPI inference service containerised with Docker, infrastructure declared in Terraform, and GitHub Actions running the tests and the deployment.
The point is the seams between the parts, registry to service and CI to infrastructure, which is where most ML projects stop and most production incidents start.
-
serve
Sentiment on SageMakerhosted inference
A trained model behind a real endpoint, reachable from a web page.
A sentiment model trained on the IMDB review set and deployed to a SageMaker endpoint, then made reachable from a browser: a Lambda function holds permission to call the endpoint, API Gateway exposes that Lambda as a URL, and a small web page posts a review to it and renders the positive or negative verdict.
The interesting part is the wiring rather than the model - endpoint permissions, the Lambda contract, and the gateway in front - which is the same path any hosted model has to travel to reach a user.
sentiment-analysis-webapp-sagemaker — open Sentiment on SageMaker on GitHub
-
serve
Advanced RAGretrieval that is measured
A grounded assistant, with the retrieval trade-offs made visible.
Document ingestion, vector search, prompt orchestration and LLM APIs assembled into an assistant that answers from a controlled knowledge base. Providers are swappable behind one interface: an offline default that needs no key, OpenAI, or a local Ollama, with embeddings configured separately as deterministic hashing or sentence-transformers.
Built so the engineering trade-offs are the subject - retrieval quality, latency, grounding and evaluation - rather than a demo that answers one question well. It runs end to end from a clean clone with no API key.
-
serve
Multi-Agent RAGrouting, self-correction, fact-checking
Eight specialised agents, and a router that decides which one answers.
A modular agentic RAG system built with LangChain and Groq behind a Streamlit chat interface. A router sends each query to the right path - internal retrieval, query reformulation, web search or a request for clarification - over a knowledge base assembled from uploaded PDF, DOCX and TXT files or from URLs.
It is self-correcting: a weak retrieval is reformulated and retried before it falls back to web search, generated answers have their factual claims extracted and verified against live search, and a safety pass revises or blocks unsafe output. The workflow, logs and parameters are all visible in the UI.
-
monitor
DriftGuarddrift and retraining
Retrains when the evidence says to, and records why.
Models decay quietly after release. This measures drift between a reference and a monitored period with PSI and KS tests, compares predictions against observed outcomes, and retrains only when the drift is significant, registering the new model in MLflow, promoting it by alias, then re-measuring.
On the demand-forecasting case the inputs barely moved while demand rose sharply, the concept drift a feature-only monitor misses; retraining cut held-out error by about 45%, and every decision is written out as an auditable before-and-after record.
-
monitor
PromptProofpytest for prompts
Declare what an LLM's output must look like, then run it like a test suite.
Edit a system prompt, swap a model or nudge the temperature, and an extraction pipeline three services downstream breaks quietly. PromptProof turns prompt behaviour into a tested contract: YAML suites of cases with assertions, run against Ollama, Anthropic or a fake echo provider for dry runs.
It is sampling-aware, because LLM output is nondeterministic and pretending otherwise is how regressions ship: each case runs N times and has to clear a required pass rate. Exit codes and JUnit XML mean CI treats a prompt regression like any other failing test.
record
4 yrs 3 mos of engineering
Backend and platform work on the operational systems that the pipelines above are built to read.
-
Feb 2025 — Present
1 yr 8 mo
Machine Learning Engineer
CGI · USA
- Own marketing ML models end to end on AWS — feature design, training, A/B validation and staged release — and set the team standards for MLflow versioning, feature store and drift monitoring, cutting iteration cycles from 2 weeks to 4 days.
- Shipped an XGBoost propensity model on real-time behavioural data in 6 weeks, lifting lead conversion 18% in A/B tests.
- Re-architected batch churn scoring into a real-time Kafka and Spark Streaming pipeline feeding 3 downstream CRM and marketing systems, at 500K+ records/day and sub-second latency.
- Built a production RAG assistant over internal campaign documentation with LangChain, LangGraph and FAISS, gated on Ragas answer-faithfulness scores before release.
- Segmented 2M+ subscribers into behavioural cohorts with K-Means and hierarchical clustering, feeding lifecycle retention campaigns.
-
Jan 2021 — Jul 2023
2 yr 7 mo
Machine Learning Engineer
Lumen Technologies · India
- Built a customer credit-risk scoring model (logistic regression and gradient-boosting ensemble) at 0.91 AUC-ROC, served through real-time FastAPI and Docker scoring APIs, saving roughly $120K a year in manual review.
- Fine-tuned a transformer NLP model (PyTorch, Hugging Face) to 87% accuracy on real-time support-ticket classification and routing, reducing misrouted tickets and resolution time.
- Cut ETL runtime from 8 hours to under 3 with Spark and Databricks pipelines across 8+ sources, including schema-drift detection, feeding model-ready features to the risk and NLP models.
- Built modular feature pipelines with automated validation gates and CI checks, replacing a manual multi-week retraining cycle.
Education
-
Aug 2023 — May 2025
Master of Science, Information Technology
Kennesaw State University · Georgia, USA
stack
Tools
Marked skills are used by a repository listed above.
Languages & data
Machine learning
Deep learning & NLP
Generative AI & LLMs
MLOps & deployment
Cloud & big data
Statistics & experimentation
contact
Get in touch
Open to ML engineer and AI engineer roles. The fastest way to reach me is email.
Vuppalapati.prudhvi7@gmail.com
Or find the code on GitHub, or take the one-page résumé (PDF).