Buyer’s guide
Most disappointing AI engagements are a mismatch of provider type, not a shortage of talent. Here is how to tell the difference before you sign.
To choose an AI development company, first match the provider type to your need — consultancy, specialist studio, general agency, offshore team, or product vendor solve genuinely different problems. Then test capability with specifics: how quality will be measured, what they would refuse to build, what broke in a system they shipped, and who owns the evaluation suite afterwards.
Start here
Most disappointing engagements are a mismatch of type, not a lack of skill. Work out which column you actually need before you shortlist anyone.
| Provider type | Best for | Watch out for |
|---|---|---|
| Global consultancy | Multi-year programmes, organisational change, procurement that requires a household name. | The people who sold it are rarely the people who build it. Ask who is actually on the keyboard. |
| Specialist AI studio | A working system in months, with senior engineers close to the problem. | Limited capacity. If you need fifty people next quarter, this is the wrong shape. |
| General dev agency | Conventional software with an AI feature attached. | Retrieval and evaluation are specialist skills. Ask what they measure, not what they have built. |
| Offshore body shop | Cost-sensitive scale-out of well-specified work. | Needs a strong technical owner on your side. Ambiguity gets built exactly as written. |
| Product vendor | A solved, common problem — transcription, standard support deflection. | If it nearly fits, the customisation gap can cost more than building. Check this honestly. |
Evaluation
The answer should involve a test set and a number. If a prospective partner cannot describe how quality will be measured before the build starts, quality will be a matter of opinion for the whole engagement — and opinions escalate badly at month four.
Anyone who says AI is the right answer to everything you described is selling. A partner worth hiring will name at least one thing on your list that is better solved with ordinary software, a purchased product, or a process change.
Demos are cheap. Ask what failed after launch and how they found out. A team with genuine production experience answers this immediately and specifically; a team without one changes the subject to architecture.
The test set, prompts, and evaluation harness should be yours, in your repository. If those live with the vendor, switching later means rebuilding your quality bar from scratch — which is a lock-in mechanism whether or not it is intended as one.
A partner who has run something in production will reach for numbers: tokens per request, model mix, caching strategy, what changes at ten times the volume. Vagueness here reliably predicts a nasty surprise in month three.
Model pricing and capability shift constantly. The architecture should let you move providers without rewriting the application. If the answer is that you would not want to, treat it as a lock-in warning.
Ask for names, seniority, time zone, and how much of their week you get. Then ask whether those same people wrote the proposal. The gap between the pitch team and the delivery team is the single most common cause of disappointment.
Red flags
A number like "95% accurate" means nothing without knowing what was measured, on which data, against what baseline.
If the conversation is entirely about models and prompts, they have not run a real system. Most quality problems live in retrieval.
"We tell the model not to do that" is not a security control. Permissions belong in code, enforced server-side.
A partner insisting on a twelve-month programme before proving anything is transferring risk to you.
Ask what was built and what it measures. A logo with no describable outcome behind it is decoration.
A shop tied to a single provider will recommend that provider regardless of what your workload actually needs.
Being fair about it
We are a small senior studio, so we are a good fit if you want a working system in months with the people who scoped it doing the building, and a poor fit if you need a hundred consultants and a formal change-management programme. We hold no reseller relationships with any model provider, which is why we can pick on measured cost and accuracy — but it also means we will not be the cheapest option against an offshore rate card.
If, after the questions above, the honest answer is that an off-the-shelf product solves your problem, that is worth knowing before you spend anything. We have given that answer before and expect to again.
Common questions
Match the provider type to your actual need first — a global consultancy, a specialist studio, and an offshore team solve different problems. Then test capability with specific questions: how quality will be measured, what they would refuse to build, what broke in a system they shipped, and who owns the evaluation suite. Vague answers on measurement and cost are the most reliable warning signs.
It varies far more with risk and integration count than with model choice. A feasibility prototype against real data is typically a few weeks of work; a production slice covering one workflow usually runs six to ten weeks. Be suspicious of quotes given before anyone has looked at your data, and of any proposal where the cost per request at production volume has not been estimated.
If AI is core to your product and you can hire senior people, in-house wins long-term. Many teams use a partner for the first system specifically to transfer the practice — the evaluation discipline and retrieval patterns are the durable part, and those can be handed over deliberately.
Short enough to cancel. A first engagement that cannot show something real inside about ten weeks is structured wrong. Sequence it so there is a decision point after feasibility, before the large commitment.
Less than seniority and overlap hours. What matters is whether the people building it can talk to the people who understand the business problem often enough. Distributed teams work well when working hours overlap; they fail when the only communication is a written specification.
Related
Assessment and feasibility, if you want help scoping first.
What we build when the answer is an agent.
Why question four on the list matters most.
Who we have done this for.
Let's build
Whether you're testing a hypothesis or scaling an established product, we'd be glad to spend a half-hour helping you think through the next step — no pitch deck required.