Model Selection When Everything Changes Weekly

Internal leaderboards beat model fandom

Choosing an AI model used to feel like choosing infrastructure.

Now it feels more like weather.

A model looks unbeatable for two weeks. Another one gets cheaper. A third one suddenly becomes better at code. A new release changes the safety behavior. Context windows grow. Latency shifts. Pricing changes. Someone finds a strange failure mode on the same day the release note says the model is smarter.

It is tempting to respond by chasing whatever is newest.

That is usually a bad selection strategy.

Benchmarks Are Not Your Work

Public benchmarks are useful. I read them. They help me decide what to test first.

But a benchmark is not a product decision.

Your product has its own shape: real prompts, real documents, real latency requirements, real failure costs, real users who do not write benchmark-style questions. A model can score well in public and still be the wrong fit for your workflow.

The more often models change, the more important it becomes to build a small internal leaderboard.

Not a giant research benchmark. Just a disciplined set of examples that represent the work you actually care about.

A Good Internal Leaderboard Is Boring

The best internal leaderboard starts with boring questions:

  • What are the top five workflows this model will handle?
  • What does a good answer look like?
  • What mistakes are acceptable?
  • What mistakes are expensive?
  • How much latency can users tolerate?
  • How much can each successful task cost?
  • Which languages, formats, and edge cases appear in real usage?

Then you turn those answers into test cases.

For a coding assistant, that might include a flaky test fix, a small refactor, a code review, and a feature request with ambiguous requirements.

For a document assistant, it might include long-context retrieval, table extraction, conflicting instructions, and messy PDFs.

For a moderation system, it might include multilingual harassment, sarcasm, borderline sexual content, self-harm support language, and harmless content that looks suspicious at first glance.

The point is not to cover everything. The point is to stop evaluating models with examples that have nothing to do with your product.

Cost and Latency Are Model Quality

People often talk about cost and latency as business constraints around the model. I think they are part of model quality.

A model that gives a great answer in 45 seconds may be excellent for offline analysis and unusable inside an interactive UI. A cheap model with a good-enough answer may beat a stronger model if it lets you run more retries, add verification, or keep the feature available to more users.

Model selection should include the whole system:

  • first-token latency
  • total response time
  • tool call behavior
  • context window reliability
  • output stability
  • price per successful task
  • failure recovery cost

The last one is easy to miss. A model with a low sticker price can become expensive if the team spends hours cleaning up bad outputs.

Treat Long Context Like a Black Box You Can Study

Long context windows are especially easy to overtrust.

A model may accept a huge input, but that does not mean it uses every part equally well. It may miss details in the middle. It may overweight recent instructions. It may summarize instead of retrieve. It may behave differently when the important fact appears once versus ten times.

The right mindset is to treat the model as a black box and design experiments around the boundary.

Put facts at different positions. Add distractors. Ask for exact citations. Change document order. Repeat the same test after a model update. If the product depends on long-context reliability, do not settle for "the context window is large enough."

Capacity is not the same as recall.

Version Everything

The least glamorous part of model selection is versioning.

Version the prompt. Version the model. Version the test set. Version the tool schema. Version the expected answer. Save the actual output.

This is annoying until the first regression.

Without versioning, a model update creates vague conversations:

"It feels worse."

"It seems better."

"Maybe the prompt changed?"

"Was this always failing?"

With versioning, the conversation becomes concrete:

"The new model improved summarization, but failed three of our long-context tests and became slower on tool use. We can use it for offline analysis, but not for the live workflow yet."

That is a much better decision.

Model Preference Is Allowed

One thing I have changed my mind about: preference is not irrational.

Some models fit certain teams better. One model's writing style may match the product. Another model may be easier to steer. Another may produce code that fits the local codebase with fewer corrections. Another may be less flashy but much more predictable.

Preference becomes a problem only when it is not tested.

It is fine to like a model. Just make it compete on your work.

The Takeaway

When models change weekly, the winning move is not to memorize the leaderboard of the internet.

The winning move is to build your own small, honest one.

Start with the workflows that matter. Include cost and latency. Keep failure cases. Re-run the tests when models change. Let public benchmarks tell you what to try, not what to ship.

In a fast-moving model market, the team with the better evaluation habit has the calmer roadmap.