Model Selection Should Follow the Evaluation Boundary
Choose a model against documented task criteria, then re-evaluate when the task, context, tools, or model version changes.
Start with the decision and constraints
Model selection is a product and engineering decision, not a leaderboard exercise. A team should identify the actual task, inputs, output format, acceptable latency, cost boundary, privacy constraints, and failure behavior. A model that performs well on a broad benchmark may still be a poor fit for a narrow workflow with strict formatting or source requirements.
The comparison should include the full system. Retrieval quality, tool schemas, prompt structure, human review, and retry behavior can change the result as much as a model choice. Treating the model as one component makes it easier to improve the workflow without replacing everything at once.
- Define task-specific success and failure cases before comparing models.
- Measure quality, latency, cost, and recoverability together.
- Record the model version and configuration used for an evaluation.
Use representative evaluation cases
A useful evaluation set contains examples drawn from the workflow, not only ideal prompts. Include incomplete records, adversarial formatting, ambiguous requests, expected abstentions, and cases that require a citation or structured output. A small, well-maintained set can reveal regressions that a larger but generic benchmark misses.
OpenAI's model documentation presents different model choices and capabilities, while its evaluation tooling frames evaluation as repeatable criteria and data. Potential integration targets, including Anthropic services, require separate compatibility and security review; this wiki does not claim a current integration. Those official materials are useful starting points, but they do not replace a team's own acceptance criteria or approval process.
- Version prompts, schemas, and evaluation cases together.
- Compare changes against a fixed baseline before broad rollout.
- Review examples that fail gracefully as well as examples that succeed.
Re-evaluation is part of maintenance
Models, contexts, dependencies, and source data change. A previous result is evidence about a specific configuration, not a permanent compatibility claim. Teams should define the events that trigger re-evaluation, such as a model update, a new tool contract, a changed safety boundary, or an important failure report.
This keeps selection honest. It allows a system to improve while preserving a clear record of what was tested, what was accepted, and what remains a hypothesis.