profile

Machine learning engineer — Atlanta, GA

Prudhvi Vuppalapati

I build ML systems that stay correct in production.

Four years building, shipping and monitoring production ML on AWS. At CGI I own marketing ML end to end: an XGBoost propensity model that lifted lead conversion 18% in A/B tests, churn scoring rebuilt as a Kafka and Spark Streaming pipeline handling 500K+ records a day at sub-second latency, and the MLflow, feature store and drift-monitoring standards the team now works to.

Before that, at Lumen, a credit-risk model at 0.91 AUC-ROC served through FastAPI, and a fine-tuned transformer routing support tickets at 87% accuracy.

The repositories below are the same discipline applied in the open: features that respect time, models that are calibrated, explained and documented, LLM systems evaluated before release, and services that report when they start to drift. Each runs from a clean clone against synthetic data, with no cloud account required.

Read the lifecycle Email me

pipeline

Features, train, serve, monitor

Ten repositories, arranged the way a model actually moves: features built so they respect time, training that is calibrated, explained and documented, services that sit behind a real endpoint, and monitoring that catches the decay before a stakeholder does.

  • Features 1 repo Built as of the decision, never after it.
  • Train 3 repos Calibrated, explained and documented from the artifact.
  • Serve 4 repos Behind a real endpoint, with retrieval that is measured.
  • Monitor 2 repos Drift caught, prompts regression-tested, decisions recorded.
  1. features

    most complete

    PointInTimefeature store

    Training data that can only be built as of the moment the decision was made.

    A credit model that scores 0.91 offline and 0.68 in production almost always learned from the future: a balance as of today joined onto a decision from eight months ago. This store makes point-in-time-correct joins the only way to retrieve training data. Facts are held bitemporally, as-of reads run through SQL window functions, and a streaming materialiser replays the fact log under a watermark so the online values match the offline ones.

    A planted future-leaking feature is part of the test suite: the store refuses to serve it, and an automated train/serve skew check compares offline and online values for the same entity and timestamp.

    Architecture schematic for PointInTime. Training data that can only be built as of the moment the decision was made.
    • Python
    • SQL
    • SQLite
    • Feature store
    • Streaming
    • unittest
  2. train

    CalibrateKitprobability calibration

    A model that says 0.8 should be right about 80% of the time.

    Accuracy says nothing about whether a probability can be trusted, and risk decisions are made on the probability itself. CalibrateKit measures miscalibration with expected calibration error and Brier score, draws reliability diagrams against the diagonal, and refits probabilities with isotonic regression or Platt scaling for any scikit-learn classifier.

    Built as a library and a CLI, with the reliability mathematics covered by its own tests rather than demonstrated in a notebook.

    Architecture schematic for CalibrateKit. A model that says 0.8 should be right about 80% of the time.
    • Python
    • scikit-learn
    • NumPy
    • CLI
    • pytest
  3. train

    InterpretabilitySHAP attribution

    Why the model said no, in terms a reviewer can act on.

    Global and local explanation for tabular models: SHAP values for individual predictions, feature-importance ordering across the dataset, dependence plots showing how a feature moves the output, and the failure cases where a single number hides an interaction.

    The same attribution path that turns a churn probability into a reason code an operations team can put in front of a customer.

    Architecture schematic for Interpretability. Why the model said no, in terms a reviewer can act on.
    • Python
    • SHAP
    • scikit-learn
    • pandas
    • Matplotlib
  4. train

    ModelCardgenerated documentation

    Model documentation derived from the artifact, not typed by hand.

    Model cards are recommended everywhere and maintained almost nowhere, because writing them by hand guarantees they drift from the model. This reads the trained estimator and its evaluation metrics and renders the card: intended use, training data, metrics table, feature importances, limitations and ethical considerations.

    Output is byte-identical for identical inputs, so the card can be committed and diffed in CI, and the demo runs fully offline.

    Architecture schematic for ModelCard. Model documentation derived from the artifact, not typed by hand.
    • Python
    • scikit-learn
    • CLI
    • Model governance
    • pytest
  5. serve

    most complete

    MLOps Append-to-end lifecycle

    The whole path from training run to served prediction, wired up.

    A full-stack MLOps reference: data in BigQuery, experiments and models tracked in MLflow with a registry gating promotion, a FastAPI inference service containerised with Docker, infrastructure declared in Terraform, and GitHub Actions running the tests and the deployment.

    The point is the seams between the parts, registry to service and CI to infrastructure, which is where most ML projects stop and most production incidents start.

    Architecture schematic for MLOps App. The whole path from training run to served prediction, wired up.
    • Python
    • MLflow
    • FastAPI
    • Docker
    • Terraform
    • BigQuery
    • GitHub Actions
  6. serve

    Sentiment on SageMakerhosted inference

    A trained model behind a real endpoint, reachable from a web page.

    A sentiment model trained on the IMDB review set and deployed to a SageMaker endpoint, then made reachable from a browser: a Lambda function holds permission to call the endpoint, API Gateway exposes that Lambda as a URL, and a small web page posts a review to it and renders the positive or negative verdict.

    The interesting part is the wiring rather than the model - endpoint permissions, the Lambda contract, and the gateway in front - which is the same path any hosted model has to travel to reach a user.

    Architecture schematic for Sentiment on SageMaker. A trained model behind a real endpoint, reachable from a web page.
    • Python
    • PyTorch
    • SageMaker
    • AWS Lambda
    • API Gateway
    • S3
  7. serve

    Advanced RAGretrieval that is measured

    A grounded assistant, with the retrieval trade-offs made visible.

    Document ingestion, vector search, prompt orchestration and LLM APIs assembled into an assistant that answers from a controlled knowledge base. Providers are swappable behind one interface: an offline default that needs no key, OpenAI, or a local Ollama, with embeddings configured separately as deterministic hashing or sentence-transformers.

    Built so the engineering trade-offs are the subject - retrieval quality, latency, grounding and evaluation - rather than a demo that answers one question well. It runs end to end from a clean clone with no API key.

    Architecture schematic for Advanced RAG. A grounded assistant, with the retrieval trade-offs made visible.
    • Python
    • RAG
    • Vector search
    • OpenAI API
    • Ollama
    • sentence-transformers
    • Embeddings
    • Prompt engineering
  8. serve

    Multi-Agent RAGrouting, self-correction, fact-checking

    Eight specialised agents, and a router that decides which one answers.

    A modular agentic RAG system built with LangChain and Groq behind a Streamlit chat interface. A router sends each query to the right path - internal retrieval, query reformulation, web search or a request for clarification - over a knowledge base assembled from uploaded PDF, DOCX and TXT files or from URLs.

    It is self-correcting: a weak retrieval is reformulated and retried before it falls back to web search, generated answers have their factual claims extracted and verified against live search, and a safety pass revises or blocks unsafe output. The workflow, logs and parameters are all visible in the UI.

    Architecture schematic for Multi-Agent RAG. Eight specialised agents, and a router that decides which one answers.
    • Python
    • LangChain
    • Groq
    • Streamlit
    • Web search
    • Vector retrieval
    • LLMs
    • Tool calling
  9. monitor

    DriftGuarddrift and retraining

    Retrains when the evidence says to, and records why.

    Models decay quietly after release. This measures drift between a reference and a monitored period with PSI and KS tests, compares predictions against observed outcomes, and retrains only when the drift is significant, registering the new model in MLflow, promoting it by alias, then re-measuring.

    On the demand-forecasting case the inputs barely moved while demand rose sharply, the concept drift a feature-only monitor misses; retraining cut held-out error by about 45%, and every decision is written out as an auditable before-and-after record.

    Architecture schematic for DriftGuard. Retrains when the evidence says to, and records why.
    • Python
    • MLflow
    • FastAPI
    • scikit-learn
    • SciPy
    • Docker
  10. monitor

    PromptProofpytest for prompts

    Declare what an LLM's output must look like, then run it like a test suite.

    Edit a system prompt, swap a model or nudge the temperature, and an extraction pipeline three services downstream breaks quietly. PromptProof turns prompt behaviour into a tested contract: YAML suites of cases with assertions, run against Ollama, Anthropic or a fake echo provider for dry runs.

    It is sampling-aware, because LLM output is nondeterministic and pretending otherwise is how regressions ship: each case runs N times and has to clear a required pass rate. Exit codes and JUnit XML mean CI treats a prompt regression like any other failing test.

    Architecture schematic for PromptProof. Declare what an LLM's output must look like, then run it like a test suite.
    • Python
    • YAML
    • Ollama
    • Anthropic API
    • JUnit XML
    • CI
    • LLMs
    • LLM evaluation

record

4 yrs 3 mos of engineering

Backend and platform work on the operational systems that the pipelines above are built to read.

  1. Feb 2025 — Present

    1 yr 8 mo

    Machine Learning Engineer

    CGI · USA

    • Own marketing ML models end to end on AWS — feature design, training, A/B validation and staged release — and set the team standards for MLflow versioning, feature store and drift monitoring, cutting iteration cycles from 2 weeks to 4 days.
    • Shipped an XGBoost propensity model on real-time behavioural data in 6 weeks, lifting lead conversion 18% in A/B tests.
    • Re-architected batch churn scoring into a real-time Kafka and Spark Streaming pipeline feeding 3 downstream CRM and marketing systems, at 500K+ records/day and sub-second latency.
    • Built a production RAG assistant over internal campaign documentation with LangChain, LangGraph and FAISS, gated on Ragas answer-faithfulness scores before release.
    • Segmented 2M+ subscribers into behavioural cohorts with K-Means and hierarchical clustering, feeding lifecycle retention campaigns.
    • Python
    • XGBoost
    • MLflow
    • Kafka
    • Spark Streaming
    • LangChain
    • LangGraph
    • FAISS
    • AWS
  2. Jan 2021 — Jul 2023

    2 yr 7 mo

    Machine Learning Engineer

    Lumen Technologies · India

    • Built a customer credit-risk scoring model (logistic regression and gradient-boosting ensemble) at 0.91 AUC-ROC, served through real-time FastAPI and Docker scoring APIs, saving roughly $120K a year in manual review.
    • Fine-tuned a transformer NLP model (PyTorch, Hugging Face) to 87% accuracy on real-time support-ticket classification and routing, reducing misrouted tickets and resolution time.
    • Cut ETL runtime from 8 hours to under 3 with Spark and Databricks pipelines across 8+ sources, including schema-drift detection, feeding model-ready features to the risk and NLP models.
    • Built modular feature pipelines with automated validation gates and CI checks, replacing a manual multi-week retraining cycle.
    • Python
    • scikit-learn
    • PyTorch
    • Hugging Face
    • FastAPI
    • Docker
    • Spark
    • Databricks

Education

  • Aug 2023 — May 2025

    Master of Science, Information Technology

    Kennesaw State University · Georgia, USA

stack

Tools

Marked skills are used by a repository listed above.

Languages & data

  • Python (used by a repository above)
  • SQL (used by a repository above)
  • PySpark
  • pandas (used by a repository above)
  • NumPy (used by a repository above)
  • Bash
  • PostgreSQL
  • MongoDB
  • Snowflake
  • BigQuery (used by a repository above)

Machine learning

  • scikit-learn (used by a repository above)
  • XGBoost
  • LightGBM
  • Random Forest
  • Logistic Regression
  • Clustering (K-Means, hierarchical)
  • Time-series forecasting
  • Feature engineering
  • Hyperparameter tuning
  • Model evaluation
  • SHAP (used by a repository above)
  • SciPy (used by a repository above)

Deep learning & NLP

  • PyTorch (used by a repository above)
  • TensorFlow
  • Keras
  • Hugging Face
  • LSTMs
  • Text classification

Generative AI & LLMs

  • LLMs (used by a repository above)
  • RAG (used by a repository above)
  • LangChain (used by a repository above)
  • LangGraph
  • Tool calling (used by a repository above)
  • Prompt engineering (used by a repository above)
  • LLM evaluation (used by a repository above)
  • Fine-tuning (LoRA)
  • OpenAI API (used by a repository above)
  • Embeddings (used by a repository above)
  • Vector database
  • Ragas

MLOps & deployment

  • MLflow (used by a repository above)
  • Model registry
  • Model governance (used by a repository above)
  • Feature store (used by a repository above)
  • Docker (used by a repository above)
  • GitHub Actions (used by a repository above)
  • Terraform (used by a repository above)
  • FastAPI (used by a repository above)
  • Airflow
  • Model monitoring
  • Data drift detection
  • pytest (used by a repository above)

Cloud & big data

  • SageMaker (used by a repository above)
  • AWS Lambda (used by a repository above)
  • S3 (used by a repository above)
  • EC2
  • Azure ML
  • Databricks
  • Apache Spark
  • Spark Streaming
  • Kafka
  • Streaming (used by a repository above)

Statistics & experimentation

  • A/B testing
  • Experiment design
  • Hypothesis testing
  • Exploratory data analysis

contact

Get in touch

Open to ML engineer and AI engineer roles. The fastest way to reach me is email.

Vuppalapati.prudhvi7@gmail.com

Or find the code on GitHub, or take the one-page résumé (PDF).