Method
The whole rubric, published. If you think a weight is wrong, you now have enough to say so precisely — which is more useful to me than agreement.
Rubric v1 · 12 questions · 4 dimensions
How it scores
Each question is worth 0–10. A dimension is the mean of its three questions, scaled to 0–100. The final score is the weighted sum. Higher means more moat — buy it. Lower means wrapper — build it.
- Substance30%Is there anything under the prompt?
- Ownership25%Who accrues the value — you or them?
- Lock-in25%What breaks if you leave?
- Build delta20%What would it honestly cost you to build this yourself?
The two overrides
⚠️ TRAP
Substance < 40 and Lock-in < 40
You are paying wrapper prices for hostage terms. This is the worst quadrant on the grid: there is nothing here you could not build, and leaving will still cost you a year.
Underrated
Substance ≥ 70 and Build delta ≤ 40
This vendor has more substance than your build instinct is giving it credit for. You are proposing to rebuild something genuinely hard.
An override replaces the headline verdict but never the underlying score — the raw band is still shown.
Who is asking
After the verdict you can say whether you would be building it, buying it, or selling it. That changes the suggested next steps and nothing else.
The score is identical regardless of who is asking. Role changes only the recommended next steps. A test that told the engineer to build and the buyer to buy would be a horoscope. The number is the number, whoever ran it.
Bands
Prompt engineering with a price tag
Build it. You are renting something your own team could ship this quarter.
Thin — but you are not ready
Rent short. Build in parallel. Take it monthly. Do not sign multi-year for this.
Real product, bad deal
Buy it — after you fix the contract. The product stands up. The terms do not.
Genuine moat
Buy it. Rebuilding this would cost you more than it is worth.
Buy it and get closer
Treat as a strategic dependency. Rare. This is a partner, not a purchase.
Every question
1 · Substance
Swap their underlying model for a different frontier model tomorrow. How much still works?
- 0Almost nothing — it is prompts
- 4Most of it, but quality drops noticeably
- 8All of it — they abstract and route between models
- 10They train or fine-tune on proprietary data
Remedy
Model-independence warranty: the vendor commits to maintaining service levels across model changes, at no additional cost to you.
2 · Substance
Do they have data you could not get yourself?
- 0No — public data plus yours
- 4Licensed data you could also license
- 8Accumulated from their customer base, and it measurably improves output
- 10An exclusive corpus you have no path to
3 · Substance
Can they show error rates on your kind of task?
- 0No evals — just demos
- 3Their own benchmarks, nothing task-specific
- 7They will run an eval on your data during the POC
- 10A domain eval suite with published regression testing
Remedy
Acceptance criteria tied to a measured error rate on your own evaluation set, with the right to exit if it is not met.
4 · Ownership
Who owns the outputs and the corrections your team makes?
- 0The vendor — it feeds their model and you get nothing back
- 3Shared, or ambiguous in the contract
- 6You own them, but cannot export them usefully
- 10You own them and can export in a usable format
Remedy
Explicit customer ownership of all outputs, corrections and human feedback, with a machine-readable export right exercisable at any time.
5 · Ownership
Does your usage improve the product for you, or for everyone?
- 0For all their customers — including your competitors
- 3It does not learn from your usage at all
- 8It improves for your tenant specifically
- 10It improves for you, and you can take that with you
Remedy
Tenant isolation for any learning or fine-tuning derived from your data, with opt-out from cross-customer model improvement.
6 · Ownership
You leave in eighteen months. What do you walk out with?
- 0Nothing
- 4A raw data export with no structure
- 8Structured data plus the workflow definitions
- 10All of that, plus the labelled corrections and eval set
Remedy
Defined exit package: structured data, workflow definitions, labelled corrections and eval sets, delivered in an agreed format within 30 days of termination.
7 · Lock-in
How deep are they wired into your systems after a year?
- 10One system, at the edge
- 6A few, read-only
- 3Deep write access to systems of record
- 0They become the system of record
8 · Lock-in
Are they the workflow, or a step in it?
- 10A step you could swap out
- 6They own one whole workflow
- 3Workflow plus the UI your staff live in
- 0Workflow, UI and the data model
9 · Lock-in
Contract term and renewal mechanics?
- 10Monthly, no commitment
- 6Annual, with a capped renewal
- 3Multi-year with usage escalators
- 0Multi-year, uncapped, priced on volume you do not control
Remedy
Renewal price cap (CPI or a fixed percentage, whichever is lower) and a termination-for-convenience right with reasonable notice.
10 · Build delta
How long for your team to build the eighty-percent version?
- 0Under six weeks, with people we already have
- 4Three to six months
- 8Six to twelve months, and we would need to hire
- 10We genuinely cannot — it needs data, licences or approvals we lack
11 · Build delta
Annual licence versus the fully-loaded cost of the team who would build it?
- 0The licence costs more than a year of that team
- 3Roughly the same
- 7About half
- 10Under a quarter
12 · Build delta
Is this your differentiator, or your plumbing?
- 0Core to how we compete
- 4Adjacent to it
- 8Useful but undifferentiated
- 10Commodity plumbing