Skip to content
Devinity Solutions

Questions to ask before hiring an AI agency

Nine questions that separate firms who ship measured AI systems from firms who ship demos, with the answers you should expect.
Haider Ali

Chief Technology Officer

3 min read

The single question that separates a firm who ships production AI from one who ships demos is: how will you measure whether this is right often enough to depend on? A supplier without a clear answer involving a test set and an accuracy baseline is selling you a prototype, whatever the contract says.

Here are the nine we would want asked of us.

1. How will you measure accuracy?

Expect: an evaluation set built from your real data, with agreed correct answers, and accuracy scored against it before and after every change.

Worry if: the answer is about the model rather than the measurement. Which model they use matters far less than whether they can tell you how often it is right.

2. What happens when it is wrong?

Every AI system is wrong sometimes. The design question is what happens then.

Expect: a confidence threshold set with you, low-confidence cases routed to a person, and a review interface that is actually pleasant to use.

Worry if: the system always produces an answer. Confidently wrong is the expensive failure.

3. Whose data trains what?

Expect: provider APIs configured to exclude your data from training, or self-hosted models where data cannot leave your infrastructure.

Worry if: this has not come up. It should be their question, not yours.

4. What does it cost to run at our volume?

Inference is a per-use cost, so it scales with usage rather than sitting as fixed overhead. At pilot volume it is negligible; on a customer-facing feature it can consume the product's margin.

Expect: cost per request instrumented from the start, and specific levers — caching, routing simpler requests to smaller models, per-account limits.

Worry if: cost is discussed only as project price.

5. Is AI the right tool here?

Expect: at least one thing on your list talked out of. Plenty of problems described as AI problems are better solved with rules, a database query, or fixing the process upstream.

Worry if: everything you propose is enthusiastically possible.

6. What happens when the model provider changes something?

Providers change model behaviour without notice. A prompt that worked in March behaves differently in September.

Expect: the evaluation set re-run on a schedule, accuracy tracked in production, and alerting on drift.

Worry if: the system is treated as finished at launch.

7. Who owns the code and the prompts?

Expect: your repositories, your cloud accounts, from the first commit — prompts and evaluation sets included, since those are the real intellectual property.

Worry if: anything is described as their platform unless you are knowingly buying a product.

8. Can you show something you built that is running now?

Expect: named systems and what they do specifically. Ours: FleetChart parses freight rate confirmations into structured fields; Nigraan predicts property values from transaction patterns; McGrocer turns an incoming order into an automated vendor stock order.

Worry if: the examples are all demos, pilots or unnamed.

9. What will you not do?

Expect: a clear exclusions list, and a boundary they will not cross on accuracy or on data.

Worry if: there is nothing they will not take on.

The short version

Most AI projects fail in the same place: the demo works, the pilot looks promising, and nobody can say whether the system is right often enough to depend on. Every question above is a way of asking whether the supplier has solved that, or is planning to find out on your budget.

  • procurement
  • evaluation
  • due-diligence
  • llm

Tell us what you are trying to build.

A thirty-minute call is usually enough to tell you whether we are the right firm for it. If we are not, we will say so and point you somewhere better.

Or reach us directly: [email protected] · +1 321 335 0265