Resources

Guides and playbooks for AI data teams

Practical guidance from our operations and quality teams on running high-quality human-in-the-loop pipelines.

Guide

Designing a rubric for RLHF preference data

How to structure scoring categories so reviewers and models agree.

Playbook

Running a red-team project end to end

From scoping adversarial goals to adjudicating edge cases.

Report

The state of human evaluation in 2026

Trends in reviewer agreement, task design, and dataset quality.