An older role: first posted 39d ago, and AI expert network still listed it when we checked today. Newer roles tend to fill faster. See the jobs hiring now.
What the work is
Evaluate the quality, correctness, and production-readiness of Kubernetes tasks used to train and evaluate a frontier AI lab's models. You'll assess cluster-operations scenarios, manifest correctness, and failure-mode troubleshooting — and provide clear, rubric-based written feedback.
Basic Qualifications
- 3+ years hands-on production Kubernetes experience (EKS/GKE/AKS or self-managed)
- Deep understanding of cluster internals: CNI/DNS/ingress, PV/PVC storage, RBAC, and failure modes (CrashLoopBackOff, OOMKilled, scheduling/eviction)
- Experience authoring and reviewing manifests / Helm charts and debugging live cluster incidents
- Proficiency in Go, Python, or TypeScript
Preferred Qualifications
- CKA / CKAD certification
- Service-mesh, autoscaling, and observability experience (Istio, HPA, Prometheus/Grafana)
- Prior SRE / platform-engineering or task-grading experience
Pay
$70–90/hr, fully remote.