Stanislav Tsepa

Stanislav Tsepa

Senior ML Engineer & Researcher
London

← Back

Curriculum vitae

ML engineer / researcher, 8 years. Production LLM and agent systems end to end, with a focus on post-training improvement loops (SFT / RL), tool use and reasoning quality, retrieval, and scalable evaluation.

Experience

ML/AI Engineer

Meta

2025 – 2026 · London, UK · End-to-end AI support

  • Worked on enabling end-to-end AI support: LLM agents that resolve a user's issue from first message to resolution, rather than triaging and routing it to a human.
  • Built and scaled tool-using LLM agents on an SFT → RL optimisation loop, combining automated prompt/policy search, reward modelling and curated knowledge updates.
  • Shipped a reasoning-based decision policy for task decomposition and escalation, optimised with GEPA — 25% fewer regrettable transfers than the prior SFT baseline.
  • Built a multi-agent deep-triage system on Google ADK (knowledge-base retrieval + past-case retrieval + plan synthesis); +50% self-serve resolution on hard cases, validated offline and online.
  • Improved retrieval with hybrid RRF (BM25 + embeddings), taking Recall@20 from 75% to 90% over a 5k-document knowledge base.
  • Led migration from text-only to multimodal SFT on user attachments, improving issue-diagnosis accuracy by 2%.
  • Cut latency via prompt and inference optimisation with bounded reasoning budgets, at 10k requests/day.
  • Built the evaluation and reward layer for agent iteration: LLM-as-judge, regression monitoring, and expert-reviewed update workflows feeding repeated post-training cycles.

ML Engineer

Booking.com

2023 – 2025 · Amsterdam, NL · User intent modelling

  • Built and productionised user-intent ML systems end to end — training, retraining, serving, monitoring — delivering millions of predictions per day.
  • Led adoption of transformer-based intent models, improving next-action prediction to >75% on offline metrics.
  • Deployed user-embedding models for personalisation and ranking; +3% bookings CTR on targeted intents in online experiments.
  • Migrated ML services from on-prem to AWS, Kubernetes and SageMaker, improving reliability and deployment velocity.

ML Researcher

Yandex Self-Driving Cars

2019 – 2022 · Moscow, RU · L4 prediction & planning

  • Built production ML for the prediction and planning layers — forecasting how other road users would behave, and the decision logic consuming those forecasts — under real-world latency and safety constraints; contributed to the L4 autonomy technology presented at CES 2020.
  • Implemented behaviour classification for double-parked and stationary vehicles over upstream perception tracks — deciding whether to wait or route around — halving related disengagements in dense urban scenarios.
  • Designed Bayesian filtering and temporal smoothing over prediction outputs, driving a 90% accuracy improvement on key edge cases.
  • Optimised a real-time C++ inference pipeline to <0.1s latency.

Junior ML Engineer

AeroNet Lab, Skoltech

2018 – 2019 · Moscow, RU · CV inference

  • Built and maintained an inference service to store geo-data and run object-detection models.

Skills

LLMs & agents — post-training (SFT, RL-style optimisation, reward modelling), agentic workflows, tool use, retrieval (BM25 / embeddings / RRF), automated prompt & policy search (GEPA), LLM-as-judge, golden sets and regression gates, guardrails.

Software & infra — Python, C++, Java/Kotlin; CI/CD, Docker, Kubernetes, AWS, SageMaker, Google ADK, Modal; observability.

Data & stats — experiment design, A/B testing, data-quality checks, SQL, Spark/Hadoop.

Research & talks

First-author paper at IEEE ICDM with Maxim Panov — graph node embeddings for link prediction, node classification and clustering

Talk at Data Sanity on LLM-as-judge evaluation and automated prompt optimisation loops

Education

MSc, Data Science — Skoltech + MIPT · GPA 4.2

BSc, Applied Mathematics & Physics — MIPT