MVP Development Logo
Proof of Concept Development

Proof of Concept Development: Can It Actually Be Built?

We build the narrowest piece of working software that settles your hardest technical question, measure it against a pass mark you set, and hand you the answer in writing. Two to four weeks, fixed scope, and an honest no when the answer is no.

Get your proof of concept scoped

Tell us the technical question you need settled. We come back within 24 hours with a scope, a pass mark, and a fixed price.

Founders only. 24-hour response. No spam, ever.

Trusted By Founders

Admissions Angle - SaaS MVP development client
Hengcheng - SaaS MVP development client
Locus Digital - SaaS MVP development client

What Every Proof of Concept Includes

One Question, Answered
Fixed 2 To 4 Weeks
Fixed Scope, Fixed Price
Measured, Not Estimated
A Written Technical Verdict
Senior Engineers Only
Full Code & IP Ownership
A Costed Path To Production
An Honest No When It Is No
2-4Weeks end to end
1Question per engagement
100%Code and IP yours
24hScoping response
What this is

A proof of concept buys you an answer, not a product

Building the smallest piece of working software that settles one specific technical question: can this be done, with these tools, at this cost and this speed.

01

What it is not

Not a product, not a demo, and not meant to survive. It exists to replace an assumption with evidence while that evidence is still cheap, and the code is the instrument rather than the deliverable.

02

Why builds actually fail

Rarely because nobody wanted the product. Usually because something in the middle turned out to be impossible, or possible only at a cost that broke the business model, and nobody checked until the budget was committed.

03

Which engagement you need

The shortest test is to ask what you are actually unsure about. If it is whether people will use it, you want an MVP. If it is whether the thing can be made to work at all, you are in the right place. The long version of that distinction is in MVP vs proof of concept.

The whole point

It is deliberately narrow, deliberately short, and deliberately allowed to come back negative. A test that cannot fail proves nothing.

Most common

Model and AI feasibility

Whether a model reaches a usable accuracy on your data rather than on a public benchmark, and whether the remaining gap closes with more data or is structural.

Integration reality

Whether a third-party API, an ERP, a legacy database or a device actually exposes the data your product depends on, at the rate and in the shape you need it.

Performance and scale

Whether the critical path survives the concurrency you have promised a customer, measured on production-like infrastructure rather than argued about in a document.

Unit economics

What each expensive operation truly costs once it is instrumented, and whether the business model still stands at the volume you are forecasting.

What we are answering

A proof of concept answers one question, precisely

Every engagement starts by naming the single technical risk that decides whether the product is worth building. We build only enough software to answer it, then we tell you the answer in writing.

01

Can the model actually hit the accuracy we need?

What we build
A working pipeline run against a sample of your real data, not a benchmark set.
What you get back
Precision and recall on your data, the failure cases, and whether the gap closes with more data or not at all.
02

Does that integration give us the data we think it does?

What we build
A live connection to the API, warehouse, or device, pulling real records end to end.
What you get back
The fields that exist, the fields that do not, rate limits, and what the vendor documentation left out.
03

Will it hold at the load we are promising a customer?

What we build
The critical path only, deployed on production-like infrastructure and load tested.
What you get back
Throughput and p95 latency at your target concurrency, and the component that breaks first.
04

What does this cost per user once it is real?

What we build
Instrumented runs that measure the true cost of each expensive operation.
What you get back
Cost per request and per active user, extrapolated to your volume, with the drivers ranked.
05

Is this even possible with the tools available today?

What we build
The thinnest path from input to output, cutting everything that is not the hard part.
What you get back
A yes or a no, with the specific technical wall if it is a no, and the workaround if one exists.

If your question is not on this list, it is still probably a proof of concept. The test is whether a two-week build could produce evidence that changes your decision. If it could, this is the right engagement. If your risk is that nobody wants the product, that is a different job and you want an MVP instead.

AI proof of concept

Most AI demos are true and useless at the same time

AI feasibility is the largest category we are asked to settle, and it fails in a small number of predictable ways. Every one of them produces a demo that works, which is exactly why they get through.

01

The benchmark number

A model reports 94% on a public dataset, so the deck says 94%. Public benchmarks are clean, balanced and nothing like your data.

We measure instead: We benchmark on a sample of your own records, including the messy ones you would rather not send, and report the number that comes back.

02

The demo that only works forwards

Every AI demo works on the inputs chosen for it. The question is what happens on the twenty per cent that nobody rehearsed.

We measure instead: We build an adversarial set deliberately: the edge cases, the malformed inputs, the categories with almost no training examples, and we report accuracy separately for those.

03

Cost per call, discovered late

Inference looks free at demo volume. At ten thousand users a day it can quietly be the largest line in the P&L, and the architecture that made it cheap is not the one you built.

We measure instead: Every run is instrumented for token or compute cost, and the report extrapolates to your forecast volume with the cost drivers ranked.

04

Latency measured on an empty machine

A model that answers in 800ms alone can take four seconds under real concurrency, which is the difference between a product and a demo.

We measure instead: We measure p95 under realistic concurrency on production-equivalent infrastructure, not a single call on a warm machine.

05

Ignoring the failure mode

The important question is rarely whether the model is right. It is what happens when it is wrong, and whether the product can tell the difference.

We measure instead: We test whether confidence scores actually correlate with correctness, because a model that knows when it is unsure can ship behind a review queue and one that does not cannot.

The AI questions we are asked most

  • Can a model reach a usable accuracy on our documents, transcripts or images?
  • Is retrieval over our own corpus good enough to answer the questions users will really ask?
  • Can we run this on-device, or does it have to go to a provider?
  • Does a smaller open model get close enough to a frontier model to change the unit economics?
  • Will an agent complete this workflow reliably enough to run without a human watching?
  • Can we do this without our customers' data leaving our infrastructure?

All six are settled the same way: on your data, against a number agreed in advance, in two to four weeks.

The first deliverable

A proof of concept that cannot fail is not a test

The number gets agreed on day three, in writing, before any code exists. It is the single thing that separates an engagement that ends in a decision from one that ends in a demo and a shrug, and it is where almost every disappointing proof of concept actually went wrong.

How it usually arrives

See if the AI is good enough

Good enough for what, measured how, against whose data? Three people will read this three ways and all of them will be able to claim they were right afterwards.

What we write down instead

95% field-level accuracy across 500 of your own invoices, with low-confidence extractions flagged rather than silently wrong

How it usually arrives

Check the integration works

Works is not a state. Most integrations return something for most calls, and the failure is always in the one field or the one edge case that the whole product depends on.

What we write down instead

Read availability, write a booking and reconcile a cancellation end to end against a live tenant account, with every request logged

How it usually arrives

Make sure it scales

Scales to what, at what latency, at what cost? Without the three numbers this is a feeling, and feelings do not survive a board meeting.

What we write down instead

p95 under 400ms at 50,000 records and 200 concurrent users, at an infrastructure cost under $0.01 per operation

How it usually arrives

Prove the concept is viable

This is the request as it usually arrives, and it is the reason so many proofs of concept quietly become small products. Viable has no threshold, so the work has no end.

What we write down instead

One named technical unknown, one measurable threshold, one date. If we cannot write that sentence together, there is nothing here to build yet.

Two rules make the difference. The threshold has to be a number somebody outside the room could check, and it has to be set before anyone knows whether it is achievable. Agreeing a bar after seeing the result is not a test, it is a negotiation with yourself, and it is how teams talk themselves into building things that were never going to work.

It also has to be a bar the product genuinely needs. Ninety-five per cent accuracy sounds more serious than eighty, but if eighty with a human review queue ships a viable product and ninety-five does not exist at any price, then eighty was the right question and the engagement should have been scoped to it. Part of what you are buying on day three is somebody arguing with you about the number.

Know which one you need

Proof of concept, prototype, or MVP

These three get used interchangeably and they are not interchangeable. They answer different questions, for different audiences, and buying the wrong one is the most expensive mistake at this stage.

Proof of concept

This page
The question it answers
Can this be built at all?
Who it is for
Your engineers, your CTO, a technical investor
What you end up holding
A narrow working spike and a written verdict
Typical timeline
2 to 4 weeks
Does the code survive
Usually not, and that is fine. You are buying the answer, not the code.
MVP vs proof of concept

Prototype

The question it answers
What will it feel like to use?
Who it is for
Test users, a pitch audience, your design partner
What you end up holding
A clickable interface with no working backend
Typical timeline
1 to 3 weeks
Does the code survive
The design does. The prototype itself is throwaway by design.
MVP vs prototype

MVP

The question it answers
Will real people use it and pay?
Who it is for
Real users in the real market
What you end up holding
A working product, live, with one complete core flow
Typical timeline
from 7 days
Does the code survive
Yes. It is production code you keep building on.
What is an MVP
How it runs

Four weeks, end to end

Short enough that the answer still matters when it arrives. Most engagements land nearer three weeks; the scope is fixed on day three and does not move.

  1. Days 1 to 3

    Name the question

    A working session with your team to find the one technical risk worth spending the budget on. We write it down as a question with a pass mark attached, because a proof of concept with no pass mark cannot fail, and a test that cannot fail proves nothing.

    You see: A one-page scope: the question, the pass mark, what is explicitly out.

  2. Days 4 to 10

    Build the thinnest path

    We build only the part that carries the risk. No auth, no dashboard, no polish, no second use case. Everything that is not the hard part gets stubbed, because every hour spent on the easy parts is an hour not spent on the question.

    You see: A running spike you can watch execute on your own data.

  3. Days 11 to 16

    Measure it honestly

    The spike gets run against real inputs and instrumented properly: accuracy, latency, throughput, cost per operation. We look for the failure cases rather than the demo path, because the failure cases are what you are paying to discover.

    You see: Raw numbers against the pass mark, including the ones we did not want.

  4. Days 17 to 20

    Write the verdict

    You get the answer in writing: whether it clears the bar, what broke, what it would cost to take to production, and what we would build differently knowing what we now know. Delivered as a document you can hand to a board or a CTO.

    You see: The technical verdict, the numbers, and a costed path forward or a clear stop.

What we need from you

Every proof of concept that slips, slips here

Not in the code. In credentials that took three weeks to arrive, sample data that turned out to be synthetic, or nobody internally with the authority to act on the answer. All five are worth checking before you commit to a date.

1

A sample of real data

Representative rather than tidy. The awkward records are the ones that decide the number, so a clean export of the easy cases will produce an answer nobody can rely on.

If it is missing
We can work from synthetic data, but the report will say so and the confidence in the result drops accordingly.

2

Access to the systems involved

Sandbox or production credentials for any API, database or device the question depends on, plus whoever administers them.

If it is missing
This is the most common cause of delay. At a large company credentials can take three weeks, which is longer than the engagement, so we start this on day one rather than day five.

3

One person who can decide

Somebody with the authority to agree the pass mark and to act on the answer. Not a committee, and not somebody who has to take it away for approval.

If it is missing
Without it the engagement produces a result that sits unread, which is the most expensive outcome available.

4

The constraint that actually binds

The latency ceiling, the cost per user, the accuracy floor, the device that has to be supported. Whatever number the product genuinely cannot survive breaching.

If it is missing
If you do not know it yet, that is fine and it is what day one to three is for. What does not work is discovering it in week three.

5

Roughly four hours across the engagement

A kickoff to name the question and set the pass mark, a mid-point check, and a handover. That is the whole ask on your side.

If it is missing
If your calendar cannot absorb four hours across three weeks, the engagement will drift, and a free scoping call is the better use of the time you do have.

None of this is a prerequisite for talking to us. We work most of it out on the scoping call, and the point of listing it here is that you can start chasing credentials this week rather than in week two of an engagement you are paying for.

What you get

The code is the instrument. The answer is the deliverable.

Most engagements hand over software and leave you to work out what it means. A proof of concept is the other way around.

The software

Built to be measured, not to be shipped

  • A running spike that executes the risky path end to end
  • Real integrations wherever the risk lives in the integration
  • The instrumentation used to produce every number we report
  • A repository you own outright, with the steps to run it yourself
  • Infrastructure as code for the environment it was measured on

The answer

The technical verdict, in writing

  • The verdict against the pass mark we agreed on day three
  • Accuracy, latency, throughput and cost per operation, measured not estimated
  • The failure cases, including the ones that are inconvenient for us
  • The architecture we would use in production, and why it differs from the spike
  • A costed, scoped path to a real build, or a plain recommendation to stop
What one actually looks like

Four engagements, start to answer

These are the shapes rather than named projects: the question as the founder phrased it, the pass mark we agreed before starting, what got built, and the number that came back. Two of the four ended in a build, one ended in a rescope, and one ended in a stop.

Model accuracy

“Can we pull line items off scanned supplier invoices accurately enough to skip manual entry?”

The pass mark we agreed
95% field-level accuracy across 500 real invoices, with errors flagged rather than silently wrong
What we built
An ingestion pipeline against a sample of the client's own scans, two model approaches run side by side, and a scoring harness that compared every extracted field against a hand-keyed answer set.
What came back
91% on the first pass. The gap was almost entirely one supplier's carbon-copy forms, which no model handled. The recommendation was to ship with a review queue for low-confidence extractions rather than chase the last four points.
OutcomeBuild, with a narrower promise
Integration reality

“Does the incumbent booking system actually expose enough through its API for us to sit on top of it?”

The pass mark we agreed
Read availability, write a booking, and reconcile a cancellation, end to end, against a real tenant account
What we built
A thin client against the vendor sandbox, then the same calls against a live account the client already held, with every request and response logged.
What came back
Availability and booking worked. Cancellations came back without the identifier needed to reconcile them, and the vendor confirmed there was no roadmap for it. The whole product concept rested on that field.
OutcomeStop, for the price of two weeks
Performance and cost

“Does our matching engine still work when the marketplace has fifty thousand active listings rather than five hundred?”

The pass mark we agreed
p95 under 400ms at 50,000 listings and 200 concurrent searches, at an infrastructure cost under $0.01 per search
What we built
The matching path only, seeded with generated data at the target volume, deployed on production-equivalent infrastructure and load tested.
What came back
p95 held at 180ms. Cost per search came in an order of magnitude under budget. The bottleneck that did appear was in an unrelated ranking step nobody had suspected, and it was cheap to fix before the build rather than after.
OutcomeBuild, and one thing rescoped
Feasibility at all

“Can we do this on-device rather than in the cloud, so the data never leaves the phone?”

The pass mark we agreed
Run the full inference path on a mid-range Android handset in under two seconds, without draining the battery in normal use
What we built
The smallest viable model quantised and compiled for mobile, wrapped in a bare harness with no interface, measured on three real handsets rather than an emulator.
What came back
1.4 seconds on the mid-range target and well inside the power budget. The older handset in the test set could not run it at all, which turned into an explicit minimum-spec decision rather than a support problem discovered after launch.
OutcomeBuild, with the device floor set deliberately

The second one is the engagement people find hardest to justify in advance and are most grateful for afterwards. A missing field in somebody else’s API ended a product concept in two weeks rather than in month five, and the founder took the same idea to a different market segment where the integration was not required.

How this ends

One of these three, and we do not know which

If we told you in advance that the answer would be yes, we would be selling you a build with extra steps. A proof of concept is only worth buying if it is allowed to come back negative.

It clears the bar

Build it, and here is the number

The approach works and the measurements back it. You get the production architecture, a scoped quote, and a timeline. Roughly half of engagements end here, and the spike usually gets thrown away even so, because code written to answer a question is not code written to run a business.

It works, but not the way you assumed

Buildable, on different terms

The concept holds but something in the brief does not: the latency target is out of reach, the unit cost is three times the model, the vendor API is missing a field the whole flow depended on. You get the constraint in writing and the options around it, which is usually a narrower scope or a different component. This is the most common outcome and the most valuable one.

It does not work

Do not build it, and here is why

The wall is real and no amount of budget moves it this year. You get the specific technical reason, the evidence, and what would have to change for the answer to flip. You found this out for the cost of a few weeks instead of the cost of a product, which is the entire reason this engagement exists.

In the first two cases the natural next step is a build, and we scope it as rapid MVP development with the proof of concept behind it, which makes that build faster and cheaper than one scoped blind. In the third, you owe us nothing further and we will say so.

Not sure you need to build anything yet?

A proof of concept is worth paying for only when the question cannot be settled by talking. Plenty of them can. If an experienced engineer could answer yours in an hour with you on a call, take MVP consulting instead. It is free, no code gets written, and you leave with a scope, a budget range and a stack recommendation. We would rather send you there and be right than sell you a build you did not need.

What it costs, and why

Published proof of concept prices span 3,000 to 75,000 dollars

That range is real, and on its own it is useless. It is wide because the phrase covers three different engagements: a genuinely narrow spike, a spike bundled with a discovery phase, and an enterprise pilot. Here is what actually moves a quote, so you can estimate your own before speaking to anyone.

How narrow the question is

Largest effect

One question with one threshold is the cheapest engagement we run. Three questions is three engagements wearing one invoice, and it is the single most common reason a quote comes back higher than expected.

Whether the data already exists

Large effect

If you can hand us a representative sample on day one, week one is measurement. If the data has to be gathered, labelled or cleaned first, that work is real and it lands before anything can be tested.

How many systems it has to touch

Large effect

A spike against one API is straightforward. A spike that has to authenticate into a legacy system, wait on a sandbox account and reconcile against a second source carries the cost of every handoff.

Whether the infrastructure has to be realistic

Moderate effect

Accuracy questions can be answered on a laptop. Latency, throughput and cost questions cannot, and standing up production-equivalent infrastructure to measure them properly is a real line item.

Compliance and data handling

Situational effect

Regulated data, on-premise constraints or a security review before we can touch anything all add time before the work starts. It is rarely the biggest driver, but it is the one most often left out of the plan.

Two rules we hold to. The quote is fixed against a fixed scope and agreed before anything starts, so a proof of concept cannot quietly become a project. And it should always be a small fraction of the build it is protecting. If a proof of concept costs a third of the MVP, it has stopped being a test and become the first phase of the thing you were trying to de-risk.

If the number matters more than the question right now, the cost calculator gives you a build range in three minutes, and a scoping call costs nothing and will tell you whether you need this engagement at all.

Where these go wrong

Six ways a proof of concept wastes your money

If you have bought one before and got something you could not act on, it almost certainly failed in one of these six ways. They are worth knowing whoever you hire, because every one of them is visible in how the engagement is scoped before it starts.

It became a small product

How it shows up
Scope grew a login, then an interface, then a second use case. Week four arrives and there is software but no answer.

What prevents it
One question, one threshold, agreed in writing on day three. Anything that is not load-bearing for that question gets stubbed rather than built.

Nobody set a bar

How it shows up
The result is presented as a demo, everyone nods, and the decision it was meant to inform gets made on the same instincts as before.

What prevents it
A number somebody outside the room could check, set before anyone knows whether it is achievable.

It was tested on the wrong data

How it shows up
Excellent accuracy on a clean sample, then a very different picture once real records arrive in month two of the build.

What prevents it
Your own data, including the awkward records, plus an adversarial set built deliberately from edge cases.

The throwaway code shipped

How it shows up
The spike looked close enough to a product that it became the foundation, and the shortcuts taken to answer a question fast turned into technical debt on day one.

What prevents it
We say plainly which parts are instruments and which architecture we would use in production, and they are usually not the same.

The answer arrived too late to matter

How it shows up
A six-week feasibility study delivered after the board meeting it was meant to inform, or after the build had already started.

What prevents it
Two to four weeks, fixed. Past a month it stops being a test and becomes the project it was supposed to de-risk.

It could only ever say yes

How it shows up
The vendor building it also quotes for the build that follows, and no proof of concept they have ever run has recommended stopping.

What prevents it
A negative result is a legitimate outcome here, it is written up in full, and we will say do not proceed when that is what the numbers say.

The sixth is worth asking any vendor about directly, including us. Ask how many of their proofs of concept have ended in a recommendation not to build. If the answer is none, the engagement is a sales process with a technical costume on.

Honest fit

Is this right for you?

A great fit

You have one technical unknown that decides whether the product exists
An investor or a board wants evidence before releasing the build budget
You are betting on a model, an integration, or a vendor API you have not tested
An enterprise buyer needs a working demonstration before signing
Your internal team is split on whether the approach is viable

Probably not

Your risk is whether anyone wants it, which is a build-and-test job
You already know it can be built and just need it built
You want a polished demo for a pitch, that is a prototype
You expect the proof of concept code to become the product
The question cannot be settled in weeks, such as a multi-year research bet
Related MVP services

Explore our other MVP builds

Building something that spans categories? These related MVP development services share the same senior team, fixed timeline, and full code ownership.

Common Questions About Proof of Concept Development

What is proof of concept development?

How much does a proof of concept cost?

How long does a proof of concept take?

What is the difference between a proof of concept and an MVP?

What is the difference between a POC and a prototype?

How is this different from your MVP consulting?

Can the proof of concept code become the product?

What happens if the proof of concept fails?

Who owns the code and the IP?

Do you build AI proof of concepts?

What is the one thing you are not sure can be built?

Tell us the technical question and we come back within 24 hours with a scope, a pass mark, and a fixed price. If a proof of concept is the wrong engagement for you, we will say that instead.

Scope Your Proof of Concept