Evaluation, fallbacks, and ownership matter more than the model name on a slide.
Production AI fails quietly when nobody owns evaluation. We treat prompts, datasets, and acceptance examples as product artifacts: versioned, reviewed, and re-run when the model or the data changes.
Every AI feature needs a fallback. That can be a human queue, a rules path, or a conservative default. Users forgive a slower answer more readily than a confident wrong one.
Before we scale, we agree who watches quality after launch. If that owner is missing, the feature is a demo, not a product.