An older role: first posted 39d ago, and AI expert network still listed it when we checked today. Newer roles tend to fill faster. See the jobs hiring now.
What the work is
Evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate a frontier AI lab's models. You'll assess experiment design, model-selection reasoning, and evaluation methodology — and provide clear, rubric-based written feedback.
Basic Qualifications
- 3+ years hands-on applied/experimental ML (experiment design, model selection, hyperparameter tuning, evaluation methodology)
- Strong grasp of data-quality rigor: leakage detection, metric gaming, and train/test/CV hygiene
- Proficiency with standard ML frameworks (PyTorch, TensorFlow, scikit-learn, XGBoost)
- Ability to critique ML claims against evidence and reproduce results
Preferred Qualifications
- Competition / benchmark experience (e.g., Kaggle)
- Graduate research or publication record in applied ML
- Prior task-grading or peer-review experience
Note: this role evaluates applied/experimental ML rigor — it is not an LLM-application-building or MLOps role.
Pay
$70–90/hr, fully remote.