Experience
ML/AI Engineer
Meta
2025 – 2026 · London, UK · End-to-end AI support
- Worked on enabling end-to-end AI support: LLM agents that resolve a user's issue from first message to resolution, rather than triaging and routing it to a human.
- Built and scaled tool-using LLM agents on an SFT → RL optimisation loop, combining automated prompt/policy search, reward modelling and curated knowledge updates.
- Shipped a reasoning-based decision policy for task decomposition and escalation, optimised with GEPA — 25% fewer regrettable transfers than the prior SFT baseline.
- Built a multi-agent deep-triage system on Google ADK (knowledge-base retrieval + past-case retrieval + plan synthesis); +50% self-serve resolution on hard cases, validated offline and online.
- Improved retrieval with hybrid RRF (BM25 + embeddings), taking Recall@20 from 75% to 90% over a 5k-document knowledge base.
- Led migration from text-only to multimodal SFT on user attachments, improving issue-diagnosis accuracy by 2%.
- Cut latency 6× via prompt and inference optimisation with bounded reasoning budgets, at 10k requests/day.
- Built the evaluation and reward layer for agent iteration: LLM-as-judge, regression monitoring, and expert-reviewed update workflows feeding repeated post-training cycles.
ML Engineer
Booking.com
2023 – 2025 · Amsterdam, NL · User intent modelling
- Built and productionised user-intent ML systems end to end — training, retraining, serving, monitoring — delivering millions of predictions per day.
- Led adoption of transformer-based intent models, improving next-action prediction to >75% on offline metrics.
- Deployed user-embedding models for personalisation and ranking; +3% bookings CTR on targeted intents in online experiments.
- Migrated ML services from on-prem to AWS, Kubernetes and SageMaker, improving reliability and deployment velocity.
ML Researcher
Yandex Self-Driving Cars
2019 – 2022 · Moscow, RU · L4 prediction & planning
- Built production ML for the prediction and planning layers — forecasting how other road users would behave, and the decision logic consuming those forecasts — under real-world latency and safety constraints; contributed to the L4 autonomy technology presented at CES 2020.
- Implemented behaviour classification for double-parked and stationary vehicles over upstream perception tracks — deciding whether to wait or route around — halving related disengagements in dense urban scenarios.
- Designed Bayesian filtering and temporal smoothing over prediction outputs, driving a 90% accuracy improvement on key edge cases.
- Optimised a real-time C++ inference pipeline to <0.1s latency.
Junior ML Engineer
AeroNet Lab, Skoltech
2018 – 2019 · Moscow, RU · CV inference
- Built and maintained an inference service to store geo-data and run object-detection models.
Skills
LLMs & agents — post-training (SFT, RL-style optimisation, reward modelling), agentic workflows, tool use, retrieval (BM25 / embeddings / RRF), automated prompt & policy search (GEPA), LLM-as-judge, golden sets and regression gates, guardrails.
Software & infra — Python, C++, Java/Kotlin; CI/CD, Docker, Kubernetes, AWS, SageMaker, Google ADK, Modal; observability.
Data & stats — experiment design, A/B testing, data-quality checks, SQL, Spark/Hadoop.
Research & talks
First-author paper at IEEE ICDM with Maxim Panov — graph node embeddings for link prediction, node classification and clustering
Talk at Data Sanity on LLM-as-judge evaluation and automated prompt optimisation loops
Education
MSc, Data Science — Skoltech + MIPT · GPA 4.2
BSc, Applied Mathematics & Physics — MIPT