Skip to content

Open Model Evaluations ​

Translation status

This English page provides a localized entry and navigation shell. The full article body is currently available in Chinese.

A reproducible reading series for comparing open models across task setup, version, inference conditions, cost, and deployment constraints.

Series directory ​

1. High-Quality Datasets in the AI Era
2. Introducing Harbor: Infrastructure for Agent Evaluation
3. Analyzing the Agent Evaluation Workflow
4. Why Evaluation Matters More Than Models in the Age of AI Agents
5. Model Capability Evaluation: From LLM Eval to Agent Eval
6. Building an Agent Eval Harness from Scratch: Task, Environment, Trial, Trajectory, and Grader
7. The Core Objects of Agent Eval: Task, Environment, Trajectory, and Grader

Building a long-term knowledge base for enterprise AI systems.