Meta AI
2019–present
I work on large-scale machine learning algorithms and infrastructure that power the majority of Meta’s revenue growth. The systems operate at trillion-parameter scale, with training spanning thousands of GPUs. My work includes generative models, sparse-model pretraining, reinforcement-learning-based post-training, and continual learning.
I am especially drawn to problems that are interesting, sometimes even mysterious, important, and hard. Examples include:
- How do we solve the severe overfitting problem that appears when a large-scale sparse model enters its second epoch of training?
- Why does a large production model regularly become over-calibrated for user traffic from a particular country at around 10:00 a.m. every day?
- What is the right reinforcement-learning recipe, both algorithmically and in terms of infrastructure, when ads recommendation presents decoding patterns and challenges distinct from those of LLMs?
- At tens of billions of examples per day, how should continual learning balance “fitting to the future” with “remembering the past” under minute-to-minute distribution shifts?
Collectively, these solutions and ML innovations have consistently delivered billions of dollars in revenue growth for Meta, making this one of the highest-ROI areas in monetization ML.