Ask the portfolio

Applied AI / ML Engineer

I build retrieval systems, ML pipelines and models that hold up in production. The work is separating signal from everything else.

4+
Years shipping ML
30+
Models in production

How I build retrieval

01Retrieve

A corpus is just noise

Eleven thousand documents. Every one of them equally likely to be the answer, which is another way of saying none of them are.

11,431 chunks

02Score

Score every candidate

Embeddings put a number on relevance. The field stops being flat. Most of it is still wrong, and that is the point.

cosine + BM25

03Re-rank

Rank, then re-rank

A cross-encoder reorders the shortlist. This is where the retrieval metric actually moves, and where most systems stop too early.

NDCG@5 +18%

04Ground

Keep five. Cite them.

The generator only sees what survived. Unsupported claims fall by forty percent because there is nothing left to hallucinate from.

unsupported -40%

A diagram of a retrieval pipeline: a corpus of document chunks is scored for relevance, ranked, re-ranked by a cross-encoder, and cut to the top five results that are passed to the generator.

Three systems,and the decision behind each.

Production ML

TelcoChurn

A retention model tuned on money, not on accuracy.

0.860
ROC AUC
0.853
Recall
1.28x
Return
Problem

A telecom operator was losing subscribers with no way to see it coming. Retention budget went out uniformly, which means most of it reached people who were never going to leave.

System

A gradient-boosted classifier over behavioural and billing features, served from a BigQuery feature store on Cloud Run. SHAP values ship with every prediction, so the marketing team sees why an account was flagged and not just that it was.

Decision

Accuracy was the wrong objective. A missed churner costs far more than a wasted retention offer, so the decision threshold is tuned on expected value rather than F1. Drag the cutoff in the panel to see the trade: push it left and you catch more churners while burning budget on people who would have stayed.

Result

0.860 ROC AUC at 0.853 recall, and 1.28x return on retention spend against the uniform baseline. A live dashboard ranks at-risk accounts daily.

XGBoostSHAPCloud RunBigQueryVertex AI
0%
Churners caught
0
Accounts flagged
0%
Spend wasted

Move the cursor across the field to set the cutoff

Panel illustrates the threshold mechanism on a synthetic cohort. The metrics beside it are measured on the real model.

Agentic systems

PokerAgents

Five local models at one table, with an incentive to lie.

5
Agents
0
Cloud calls
SSE
Transport
Problem

Most multi-agent work is evaluated on cooperative toy tasks. I wanted a setting with hidden information, real stakes between agents, and a reason to misrepresent, to see how small local models actually behave.

System

Five Ollama-backed agents share a table through a FastAPI game engine. Every action is emitted as structured JSON against a schema, hands are scored with Treys, and the match streams to the browser over server-sent events.

Decision

Agent reasoning is constrained to a typed schema instead of free text. It costs some expressiveness and buys a parser that never fails mid-hand. That is the same trade I would make for any tool-calling agent going to production.

Result

Runs entirely on local models with no API cost. The turn-taking, message passing and failure modes map directly onto enterprise multi-agent workflows.

PythonFastAPIOllamaSSETreys

Real capture, not a mockup

Recorded locally, no cloud inference

Published research

GridDemand

An ensemble that wins by being deliberately uneven.

-16%
RMSE
-20%
MAE
IEEE
Published
Problem

Grid operators schedule generation against a demand forecast. Single-model forecasts plateau, and the leftover error turns into either idle capacity or a shortfall.

System

An ensemble over gradient-boosted learners, with feature engineering for daily and weekly structure plus weather coupling. Each member is biased in a different direction so that their errors partly cancel.

Decision

The members are not individually optimal, on purpose. Diversity across the ensemble beats per-model accuracy, which is why the combined residual collapses further than either member reaches alone.

Result

16% lower RMSE and 20% lower MAE against the baseline. Co-authored and accepted at IEEE CICT 2025.

XGBoostLightGBMForecastingResearch

RMSE 0.0610/errors cancel, so the pair beats either alone

Panel illustrates error cancellation on a synthetic series. The RMSE and MAE figures are the published results.

Before the models, the analytics

Twelve published Tableau workbooks. Most of what I know about framing a question for a model started here.

Experience

Research to production,four years of it.

2026Present

AI Engineer Intern

Changing The Present

LangGraph retrieval workflows with evaluated retrieval, MCP tool integrations and production guardrails for nonprofit donor matching.

LangGraphRAGFastAPICloud RunMCP
+18%
NDCG@5
-40%
Unsupported claims
20252026

ML Engineer Intern

heal.ID

Anomaly detection over wearable health streams. MLflow experiment tracking, Vertex AI deployment and BigQuery feature stores.

PyTorchMLflowVertex AIBigQuery
0.81
ROC AUC
500
Users
20222024

Data Scientist

EPAM Systems

Payment fraud detection at scale. PySpark feature pipelines, Airflow orchestration and A/B testing on risk thresholds.

PySparkXGBoostAirflowRedshift
0.86
ROC AUC
-12%
False declines
Vinay Chanamallu

New York/open to relocation

Evaluation is thewhole job.

Anyone can get a model to produce an answer. The engineering is in knowing whether that answer is right, what it costs when it is wrong, and how you would notice if it quietly stopped working.

I work across retrieval, fraud, health signals and forecasting, and the pattern repeats every time. Build the measurement first. The model is the easy part.

Based in New York. Currently on agentic patterns with LangGraph and MCP. Open to relocation or remote.

GenAI and agentsPlatform and dataModelling
LangGraphRAGRetrieval evaluationGuardrailsTool callingMCPOllama
Vertex AICloud RunBigQueryAWS S3RedshiftDockerAirflowFastAPIPySpark
PythonSQLXGBoostLightGBMPyTorchTensorFlowSHAPMLflowA/B testing