Open Model Evaluations
Translation status
This English page provides a localized entry and navigation shell. The full article body is currently available in Chinese.
A reproducible reading series for comparing open models across task setup, version, inference conditions, cost, and deployment constraints.
Series directory
1. High-Quality Datasets in the AI Era
2. Introducing Harbor: Infrastructure for Agent Evaluation
3. Analyzing the Agent Evaluation Workflow
4. Why Evaluation Matters More Than Models in the Age of AI Agents
5. Model Capability Evaluation: From LLM Eval to Agent Eval
6. Building an Agent Eval Harness from Scratch: Task, Environment, Trial, Trajectory, and Grader
7. The Core Objects of Agent Eval: Task, Environment, Trajectory, and Grader