Resources
Guides and playbooks for AI data teams
Practical guidance from our operations and quality teams on running high-quality human-in-the-loop pipelines.
Guide
Designing a rubric for RLHF preference data
How to structure scoring categories so reviewers and models agree.
Playbook
Running a red-team project end to end
From scoping adversarial goals to adjudicating edge cases.
Report
The state of human evaluation in 2026
Trends in reviewer agreement, task design, and dataset quality.