﻿# Build vs. buy, reloaded – and platform risk

> AI made the first version cheap, not the software: operations, integration, change and responsibility cost what they did. What you really buy (four layers, and you inherit the fourth), the seam test a demo never runs, the 90/10 trap, what a vendor learns from forty edge cases, platform risk from Exchange Basic auth to model deprecations – and why the line between build and buy is whether you can get out.

Source: https://laszlonemes.com/blog/build-vs-buy · Published: 2026-10-01 · Author: László Nemes

October 1, 2026 · 25 min read

`Architecture` `AI engineering` `Integration` `.NET`

Build or buy is one of the oldest questions in software, and every few years something arrives that is supposed to settle it. This time it is AI: if an agent can write the thing in a week, why would anyone pay a licence for it? I hear the question from both directions – from teams that want to replace a vendor with a weekend of prompting, and from customers of my own product who wonder whether they still need it.

The honest answer is that the question didn't change. One input did. The first working version got cheap – weeks became days. But the cost of software was never the first working version. It was everything after it.

> **The thesis:** build vs. buy didn't change – only the v1 got cheap. Operations, integration, change and responsibility cost exactly what they did five years ago. What did change is a risk that used to be quieter: the platform you build on is also your competitor. It sees the usage, it sees what works, and building it in costs it close to nothing.

> **Fig. 1** · Five years of owning software. The bar that AI shortened is the first one. The other four are where the money always went, and they are exactly as long as they were.
>
> *Diagram:* Five years of owning a piece of software, as two bars. Before AI, the first working version was roughly a fifth of the total. The rest was integration – login, file storage, backup; operations – on-call and incidents; maintenance – upgrades and security patches; and handover – documentation and the next person who has to understand it. With AI the first version shrinks to a sliver, and every other segment is exactly as long as it was.

## What you actually buy

When you buy software, you buy four separate things, and each needs its own price. There is the **hardware and infrastructure**. There is **the software** itself. There is the **service**: support, an SLA, operations, scaling. And there is the fourth, which the offer doesn't mention at all: **integration**. The vendor sells the first three. You inherit the fourth, and it never finishes.

> **Fig. 2** · Three layers on the invoice, one in your backlog. Identity mapping, storage, secrets, monitoring, backup and the runbook are work you do whether you build or buy – buying only decides who owns the rest.
>
> *Diagram:* What you buy when you buy software, as four layers. Hardware and infrastructure, the software itself, and the service – support, SLA, operations and scaling – are what the vendor sells and prices in the offer. The fourth layer, integration – identity, storage, secrets, monitoring, backup and the runbook – is not in the offer. You inherit it, and it never finishes.

That is why I say **buy is build too**. You just build the wrong half: the integration, the monitoring, the runbook, the permission mapping, the backup strategy – and then you pay rent on top. In an enterprise, a "buy" is typically 20–40% of the engineering a build would have taken, plus a fee that never stops. So the real comparison isn't build versus buy. It is **build-and-own versus integrate-and-rent**, and both columns have engineering in them.

### Where it runs is part of what you buy

The deployment model isn't a hosting detail. It decides who patches, who sees the logs and who answers the data-protection questions:

| Model | What you get | What you pay |
| --- | --- | --- |
| On-premises – in your infrastructure | You patch, you see the logs, the data never leaves. Responsibility for it is yours too. | Slower support: they can’t see inside your environment, so every ticket starts with "send us the logs". |
| SaaS – in theirs | Features arrive fast; nobody on your side runs anything. | The data leaves: a DPA, a sub-processor list, a region – and the least comfortable question of all, the exit. |
| BYOC / private cloud | Their software, your cloud account. Often the only option legal will sign. | Both sets of downsides at once: you run it, and they still need a way in. |

The licence model is a technical question as well, and it gets reviewed by the wrong people. Per-seat pricing is fine for forty analysts. Add five thousand read-only viewers who open a report twice a month, and the business case is dead – however good the software is. Whoever designs the solution should read the price list with the same care as the API documentation, because the price list decides the architecture as often as the API does.

> Buy is build. You build the wrong half – and pay rent on top.

## The seam test

In a demo you see the feature. What you are going to inherit is the **seams**: every place the product touches something of yours. A demo touches none of them, because it runs in the vendor's environment, against the vendor's storage, with the vendor's identity provider. Each seam has its own concrete, reproducible way of falling apart, and it is worth going through them in order.

> **Fig. 3** · The demo lives in the middle box. The installation lives in the eight around it. Ask about each one before signing; the exit, which everybody asks about last, is the one to ask about first.
>
> *Diagram:* The feature the demo shows sits in the middle. Around it are the eight seams you inherit: file storage, the database, secrets, auth, network and licence checks, deployment, observability, and the exit. The demo touches none of them; each one has its own way of failing on install day.

### File storage – the “S3-compatible” trap

The datasheet says S3-compatible, and it is – on AWS. Their client uses virtual-hosted addressing, where the bucket is part of the hostname. Your MinIO runs on-premises under an internal hostname, behind a port, and needs path-style addressing, where the bucket is part of the path. The demo on AWS was perfect; your first upload fails. It isn't a missing feature. It is one flag in the client – it just has to come out before you sign, not after.

```csharp
// "S3-compatible" – and it is, on AWS. Their client builds https://<bucket>.s3.amazonaws.com/<key>.
// Your MinIO lives at https://minio.internal:9000 and there is no wildcard DNS for bucket hostnames.
var s3 = new AmazonS3Client(credentials, new AmazonS3Config
{
    ServiceURL = "https://minio.internal:9000",
    ForcePathStyle = true,                  // https://minio.internal:9000/<bucket>/<key> – the whole difference
    AuthenticationRegion = "us-east-1",
});
```

### SharePoint and throttling

The customer uses SharePoint as a file store, and Microsoft throttles per tenant. If the product doesn't handle a `429` and its `Retry-After` header properly, everything works with fifty files and the job dies at night with fifty thousand, and nobody understands why. Worse, a client that retries on its own schedule extends the throttle it is fighting. You will never see this in a demo. You only see it under load.

```csharp
// SharePoint throttles per tenant and says for how long. A sync that ignores the answer works with
// 50 files and dies at night with 50,000 – after hammering the tenant into a longer throttle.
async Task<HttpResponseMessage> SendRespectingThrottle(Func<HttpRequestMessage> build, CancellationToken ct)
{
    for (var attempt = 1; ; attempt++)
    {
        var response = await http.SendAsync(build(), ct);   // a fresh request each time: upload bodies are streams
        if (response.StatusCode is not (HttpStatusCode.TooManyRequests or HttpStatusCode.ServiceUnavailable)
            || attempt == MaxAttempts)
            return response;

        var wait = RetryDelay(response, attempt);
        response.Dispose();
        await Task.Delay(wait, ct);
    }
}

static TimeSpan RetryDelay(HttpResponseMessage response, int attempt) => response.Headers.RetryAfter switch
{
    { Delta: { } delta } => delta,                                               // what the server asked for
    { Date: { } until } when until > DateTimeOffset.UtcNow => until - DateTimeOffset.UtcNow,
    _ => TimeSpan.FromSeconds(Math.Min(60, Math.Pow(2, attempt))),               // no header: back off anyway
};
```

### The database – version and privileges

"Postgres support" often means, in practice: tested on Postgres 14, in the `public` schema, by a user allowed to create extensions. You run a managed database where nobody is a superuser, and the installer's first migration asks for one. That comes out on install day, not in the proof of concept. And right behind it is the next question: is there a way back from the migrations, or only forward?

```sql
-- 001_init.sql, as shipped. "Postgres supported" meant: tested on 14, in public, connected as postgres.
CREATE EXTENSION IF NOT EXISTS pg_stat_statements;   -- for their diagnostics page; not a trusted extension
CREATE EVENT TRIGGER audit_ddl ON ddl_command_end    -- superuser only, on every Postgres there is
    EXECUTE FUNCTION public.log_ddl();
CREATE TABLE public.documents (                      -- your policy: one schema per application, nothing in public
    id   uuid PRIMARY KEY,
    body jsonb NOT NULL
);
-- 002 … 041: forward only. Nobody asked for a down step until install day.
```

### Secrets – Vault or plaintext

Your policy says service credentials don't sit anywhere in plain text. The product wants the connection string in a YAML file. That is not "adding a field" – it is their whole configuration and secret-handling layer. If it can't do it, you write the sidecar that injects the secret, which means you are building again, only now around the thing you bought. And there is a subtler trap behind the obvious one:

```yaml
# What their product wants: the password, in a file, forever.
database:
  connectionString: "Host=pg.internal;Database=docs;Username=docs_app;Password=S3cret!"
---
# What your policy allows: Vault issues a short-lived database user per pod, and a sidecar writes it to a file.
metadata:
  annotations:
    vault.hashicorp.com/agent-inject: "true"
    vault.hashicorp.com/role: "docs-app"
    vault.hashicorp.com/agent-inject-secret-db: "database/creds/docs-app"
    vault.hashicorp.com/agent-inject-template-db: |
      {{- with secret "database/creds/docs-app" -}}
      Host=pg.internal;Database=docs;Username={{ .Data.username }};Password={{ .Data.password }}
      {{- end }}
# The credential expires with its lease. A product that reads the connection string once, at start-up,
# works for exactly one TTL. That is the question to ask – not "do you support Vault?"
```

### Auth – where most of the hidden work lives

OIDC, or only SAML? Is there group-claim mapping, or do four hundred users get entered by hand? Is there SCIM provisioning? And the most important one: what does it do for a *service* account – key-pair, mTLS, or only a username and a password? In the environment I work in, Snowflake sits behind a proxy, and the service users authenticate with RSA key pairs, because password sign-in is being phased out. That isn't an exotic requirement. It is what will be mandatory everywhere next year.

```sql
-- A service user with no password at all, and two key slots, so rotation needs no downtime.
CREATE USER svc_reporting TYPE = SERVICE DEFAULT_ROLE = reporting_reader;
ALTER USER svc_reporting SET RSA_PUBLIC_KEY = 'MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8A...';

-- Rotation: the new key goes into the second slot, the clients switch, then the first slot is cleared.
ALTER USER svc_reporting SET RSA_PUBLIC_KEY_2 = 'MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8B...';
ALTER USER svc_reporting UNSET RSA_PUBLIC_KEY;
```

```csharp
// .NET connector: key-pair (JWT) auth through the corporate proxy. There is no password in the string
// because there is no password. The key and its passphrase come from the secret store, not from config.
var cs = "account=myorg-myaccount;user=svc_reporting;authenticator=snowflake_jwt;" +
         $"private_key_file={keyPath};private_key_pwd={keyPassphrase};" +
         "useProxy=true;proxyHost=proxy.internal;proxyPort=8080;" +
         "warehouse=REPORTING_WH;db=ANALYTICS;role=REPORTING_READER";

await using var connection = new SnowflakeDbConnection(cs);
await connection.OpenAsync(ct);
```

### Network, licence, deployment, observability

- **Network and licence.** Does the product need outbound traffic for telemetry or a licence check? What happens air-gapped? And the least pleasant one: when it can't reach the licence server, is there a grace period, or a hard failure at two in the morning?
- **Deployment.** If the vendor ships only a Helm chart, the vendor has also decided that you run Kubernetes. Now you operate a cluster for one application, with everything that means in operations and skills. That price is nowhere in the offer.
- **Observability.** Structured logs, a health endpoint, metrics – or only their own dashboard, with no API, that you can't wire into your monitoring?

### The exit – asked last, should be asked first

Is there a documented export, with the schema, or a pile of CSV files? The best question to put to a reference customer is not who moved *to* them. It is who moved **away from them**, and how long it took. And the export gets tested in the first week of onboarding, not on the day of the breakup – I come back to this below.

> You buy the feature. You inherit the seams.

## The 90/10 trap

This sentence is spoken in every one of these projects: "the vendor covers 90% of what we need, the remaining 10% is custom development". There are two problems with it.

The first: **the 90% is always self-reported, and it always comes from the demo.** Ask who measured it, when, on what data, and in whose environment. In my experience, after an integration test the number lands somewhere around 60–70. Not because the vendor lied, but because the difference is in the seams, and the demo doesn't touch the seams.

> **Fig. 4** · The same product, measured twice. The first bar is honest about the features. The second is honest about your environment, and the gap between them is the previous section.
>
> *Diagram:* Two bars of how much of the requirement a product covers. The vendor's figure, from the demo: ninety percent fit and ten percent custom development. The same product after an integration test in your own environment: sixty to seventy percent fit, and the gap is in the seams the demo never touched.

The second: **the remaining 10% is never evenly spread.** It sits exactly where you differ from everyone else – which is exactly where it is most expensive to touch. There are three ways to close it, and each one costs something:

- **Fork it, or customise it deeply.** From now on the merge conflicts are yours, there is no upgrade path, and the support contract became paper. The first security patch tells you what you signed up for.
- **Wrap it.** You run two systems plus the integration between them. The bugs will live in the gap, where nobody is on call and each side points at the other.
- **Shrink the requirement.** The least glamorous option and the most often right. A good part of the 10% isn't a business need but process inheritance. Ask "why do we do it this way?" and it often turns out that nobody will defend it.

And the main counter-argument when someone wants to build because of the 10%: **if you build your own because of the 10%, you will also run the 90%.** The licence fee isn't to be compared with the 10%. It is to be compared with five years of owning the whole system.

> The 90% always comes from the demo. After an integration test it's 60–70.

## AI made the v1 cheap – only the v1

The point in one sentence: **AI doesn't make software cheap. It makes the v1 cheap** – and the v1 is roughly a fifth of the total cost.

A vibe-coded app is genuinely good for a proof of concept. I'd say that is its best use today: in two days you put together a throwaway prototype that settles one question. But it isn't ready for production, and not because the code is bad. It is because **the information and the experience are missing behind it.** For a product to work properly you need customer feedback: what users actually do with it, which configurations really occur, which error message makes them call, and which feature a hundred people asked for and three use. That information can't be generated. You get it by serving customers, for years.

What AI didn't make cheaper: on-call, security patches for every vulnerability found in your dependencies, migrations, the upgrade path, backward compatibility, integration testing against a real identity provider and a real database rather than a mock, documentation for the next person – and the next person. Then there is the **bus factor**: whoever put it together in a week moves on within a year, and the repository has no owner. I wrote about the same gap from the engineer's side in [Engineering mindset in the age of AI](https://laszlonemes.com/blog/engineering-mindset); here it shows up on the balance sheet.

### Where AI actually helps

It is worth separating, because it isn't the same across the three phases:

> **Fig. 5** · AI's help is front-loaded: strongest when you are deciding, decent while you build the first version, thin once you operate. Ownership cost runs the other way.
>
> *Diagram:* How much AI helps in each phase of a build-or-buy decision. Planning and evaluation: the most – reading vendor documentation, building the integration test matrix, writing the procurement questions. Building: strong for the first version, weak by the third year – it writes the adapter fast but doesn't know the edge cases. Operating: the least – log triage and runbook drafts, but responsibility can't be delegated.

- **Planning and evaluation – the strongest.** Digesting vendor documentation, putting together an integration test matrix, writing the question list for procurement.
- **Building – strong for the v1, weak by year three.** It writes the adapter fast. It doesn't know the edge cases.
- **Operating – the least.** Log triage and a runbook draft, fine. But responsibility can't be delegated.

### The best use: disprove the demo

Two consequences that aren't obvious. The first: **the best build-vs-buy use of AI isn't writing the build. It is disproving the demo.** In two days you write a throwaway PoC that wires the product to your own Vault, your own Snowflake, your own MinIO, and you see whether what was promised actually works. That used to take three weeks, which is why nobody did it. Now there is no excuse.

```csharp
// The demo-disproving PoC: two days with an agent, deleted afterwards. Every test is one seam,
// run against the real thing in our network – not their sandbox, not a mock.
public class VendorSeamTests(VendorFixture vendor, OurInfra our)
    : IClassFixture<VendorFixture>, IClassFixture<OurInfra>
{
    [Fact]
    public Task Writes_to_our_MinIO_path_style() =>
        vendor.Files.UploadAsync(our.MinioBucket, "probe.txt", "hello");

    [Fact]
    public async Task Survives_a_Vault_lease_expiring()
    {
        await vendor.ConnectAsync(our.VaultRoleWithFiveMinuteTtl);
        await Task.Delay(TimeSpan.FromMinutes(6));
        Assert.True(await vendor.HealthAsync());          // the kind that reads YAML once fails here
    }

    [Fact]
    public Task Reads_Snowflake_with_a_key_pair_through_the_proxy() =>
        vendor.Sources.TestConnectionAsync(our.SnowflakeKeyPair);

    [Fact]
    public async Task Syncs_50_000_SharePoint_files_without_losing_one()
    {
        var run = await vendor.Sync.RunAsync(our.SharePointLoadTestSite);
        Assert.Equal(0, run.Failed);                      // the 429 lives here, not in a 50-file demo
    }
}
```

The second: **AI makes migration cheaper too.** Exporters, transformers and schema mappings are much faster to write now. If leaving is cheaper, lock-in is worth less, and with it the vendor's pricing power drops. AI doesn't only strengthen the build side – on the buy side it improves your negotiating position.

A prediction to finish with. The wall for internal tools put together with AI in 2025–26 will arrive in 2027–28. Not as a spectacular collapse, but as a slow maintenance cliff: repositories without owners, dependencies nobody updated, and the sentence "someone built this back then, but they're not here any more".

> Perfect for a PoC. Not for production – not because of the code, but because there's no customer behind it.

## The same thing from the vendor’s side

I sit on both sides. One of my own products is used by a lot of companies, and the customer runs it in their own infrastructure. Everything from the previous sections comes up there, and it is a mature product, so the list is long: four internal databases – SQLite, SQL Server, Postgres and MySQL; seven file stores – local disk, Azure Blob, S3, MinIO, Google Cloud Storage, Box and SharePoint; login with local users, Active Directory, or OAuth with Google, Okta, Microsoft or any OIDC provider; and eleven data-source connectors, with password, OAuth or RSA key-pair authentication. From that side the picture looks different.

### Why most vendors can't be generic

Because every adapter multiplies the test matrix and the support surface. Four databases, seven file stores and three kinds of login are eighty-four combinations before a single data source is connected. I test every one of them – and that is exactly the point. "Supported" has to mean tested; a combination that only exists on the datasheet is a promise nobody has checked. But the full matrix costs CI hours on every release, real SharePoint, Box and Microsoft test tenants because those don't come in a container, and the knowledge of what is worth asserting inside each cell. A vendor that sells to one kind of customer doesn't pay that bill, and has no reason to.

> **Fig. 6** · Twenty-eight cells per login type, eighty-four in total, every one of them a CI job. The next adapter isn't one more piece of code – it is a whole new column of jobs, on every release from then on.
>
> *Diagram:* A grid of four internal databases by seven file stores – local disk, Azure Blob, S3, MinIO, Google Cloud Storage, Box and SharePoint – twenty-eight combinations, times three login types: eighty-four CI jobs, all of them run on every release. An eighth file store would not be one more piece of code: it adds a whole column, four databases times three login types, twelve more jobs on every release from then on.

the matrix, as a CI workflow (sketch)

```yaml
# 4 internal databases × 7 file stores × 3 login types = 84 jobs, on every release. All of them.
strategy:
  fail-fast: false                # one red cell must not hide the other 83
  matrix:
    db: [sqlite, mssql, postgres, mysql]
    store: [local, azure-blob, s3, minio, gcs, box, sharepoint]
    login: [local, active-directory, oauth]     # oauth: Google, Okta, Microsoft or any OIDC provider
# Databases, MinIO and a directory server start as containers. SharePoint, Box and the Microsoft
# login don't come in a container: those jobs run against real test tenants – throttling included.
# The matrix is the cheap part. What each job asserts is a ticket some customer once filed.
```

So **the adapter layer isn't a technical decision. It is a product decision**, and it only pays off when you sell a boxed product, where the customer's infrastructure is a given, not a choice. Here is the file-store port from that product, trimmed. Two details answer seams from earlier: MinIO gets its own adapter, configured the way a MinIO install actually looks, and every adapter keeps its secret in a separate object from its plain settings.

```csharp
// The file-store port. One interface, seven adapters: local disk, Azure Blob, S3, MinIO,
// Google Cloud Storage, Box and SharePoint – chosen per installation, by configuration.
public interface IFileStorage
{
    FileStorageType Type { get; }
    int CurrentConfigVersion { get; }      // stored settings carry a version, so old ones can be migrated

    // Every adapter checks its own configuration – and the connection – before anything is saved.
    Task<FileStorageStatus> ValidateAndGetStatus(int configVersion, string settingsJson,
        string secureSettingsJson, AppConfig config);

    Task<IFileInfo> GetFileInfoAsync(Guid definitionId, int configVersion, string settingsJson,
        string secureSettingsJson, AppConfig config, PathType pathType, params string[] pathSegments);
    Task UpdateFileAsync(/* … */ Stream contents, UpdateFileMode mode, PathType pathType, params string[] pathSegments);
    Task DeleteAsync(/* … */ PathType pathType, params string[] pathSegments);
}

// One adapter. MinIO isn't "S3 with another URL" here: it is configured the way a MinIO install
// actually looks – a host and a port. And the secret lives in its own object, apart from the settings.
internal class MinioFileStorage : BaseFileStorage<MinioSettings, MinioSecureSettings>
{
    public override FileStorageType Type => FileStorageType.Minio;
    public override int CurrentConfigVersion => 1;

    protected override List<ValidationResult> validateModel(MinioSettings settings, MinioSecureSettings secure)
    {
        var errors = new List<ValidationResult>();
        if (string.IsNullOrEmpty(settings.HostAndPort)) errors.Add(new("'HostAndPort' is required"));
        if (string.IsNullOrEmpty(settings.AccessKey))   errors.Add(new("'AccessKey' is required"));
        if (string.IsNullOrEmpty(settings.BucketName))  errors.Add(new("'BucketName' is required"));
        if (string.IsNullOrEmpty(secure.SecretKey))     errors.Add(new("'SecretKey' is required"));
        return errors;
    }

    protected override IWriteableFileProvider getProvider(MinioSettings settings, MinioSecureSettings secure, AppConfig _) =>
        new WriteableMinioFileProvider(new()
        {
            HostAndPort = settings.HostAndPort,
            UseSSL = settings.UseSSL,
            BucketName = settings.BucketName,
            AccessKey = settings.AccessKey,
            SecretKey = secure.SecretKey,
        });
}
```

The data-source side has the same shape, and its auth enum is short enough to show whole. The general pattern – one port, many implementations, chosen by configuration – is in the [micromonolith post](https://laszlonemes.com/blog/micromonoliths).

```csharp
// The data-source side: eleven connectors – SQL Server, Postgres, MySQL, Oracle, Snowflake, BigQuery,
// Redshift, Teradata, SAP HANA and two ODBC variants – and one enum for how they log in.
public enum DatasourceAuthType
{
    Password,
    OAuth,
    RsaKeyPair,      // the Snowflake story from the seam section, seen from the vendor's side
}
```

It matters not to say this as a boast, because as a boast it is a weak argument. The precise claim is this: **the adapter isn't hard to write. The forty edge cases that make you rewrite it twice are hard to learn.** A few of them:

| Where | What bites |
| --- | --- |
| Oracle | Identifiers were capped at 30 bytes before 12.2. A generated index name that works everywhere else fails on the old one. |
| SQL Server | The default collation is case-insensitive. A unique key that is unique in Postgres isn’t unique here. |
| Postgres | `search_path`: the table exists – in a schema this connection doesn’t look at. |
| SQLite | One writer at a time. `SQLITE_BUSY` the moment a background job and a request write together. |
| S3-compatible stores | Path-style versus virtual-hosted addressing – see above. |
| SharePoint | `429` with `Retry-After`, counted per tenant, so another app’s traffic throttles yours. |
| LDAP / Active Directory | A search returns a referral to a domain controller the server can’t reach. Follow it and hang; ignore it and lose users. |
| Entra ID | Token lifetimes are randomised, and with continuous access evaluation a token can be revoked before it expires. A cached token is not a valid token. |

That knowledge comes from customers, over years. It isn't in the training data, which is why the code got faster and the experience didn't.

### What feature requests teach you

Custom development, feature requests, roadmap presentations – from the inside you learn which request is good and which is pointless. And the bad request usually isn't bad faith: the customer doesn't know exactly what they want either. The difference between the two is that **a good request describes the constraint, not the solution.** "Support Vault" is practically useless. "Our policy forbids credentials at rest in plain text, so we use dynamic database credentials" – that can be solved three different ways, and one of them might be ready tomorrow.

Which leads back to the build decision: **when you build for yourself, you have to write that constraint down yourself, and usually nobody does.** That is why the first version of an in-house solution is a faithful picture of its authors' own misunderstanding.

### Drawing the line

Because you can't do everything. Every integration is a permanent compatibility contract: versions, deprecations, and the other side's auth changes, which you don't decide. If one team needs system number X+1 for one job once a month, that is not something to integrate. The right answer is an escape hatch: a documented file format, a webhook, a CLI, an API. You don't integrate – you make integration possible, and the last mile is done by whoever needs it.

> Don't integrate everything. Give people an interface, and let whoever needs the last mile build it.

## Platform risk

### The AI version, briefly

The pattern is familiar. A startup builds something on an AI model, it takes off, and in doing so it proves there is demand. The model provider sees in its usage data that many users need this, and ships it as a built-in feature of its own product. That was the startup.

Told like that, the story is too fatalistic, and it needs handling with care. The example people usually reach for – an AI e-mail assistant whose obituary was written the week the model providers shipped mail connectors – mostly proves the opposite: it kept growing next to the built-ins. What the platform absorbed was the **feature**, not the **workflow**.

> **Fig. 7** · Same built-in, two outcomes. When a product is the feature, the platform release is its end. When a product owns the workflow and the data, the release mostly advertises the category.
>
> *Diagram:* Two products built on a model API, after the platform ships the same feature as a built-in. Left: a product that was a prompt and an API call; when the feature is absorbed, nothing remains. Right: a product that owns the workflow – state, permissions, audit trail, domain data and integrations; the feature is absorbed, the workflow remains, and the built-in grows the category and sends the serious users over.

So the real question isn't whether they will build it in, but **what is left when they do**:

- If your product is essentially a prompt and an API call, you don't have a product. You have a deadline.
- If your product owns the workflow and the data – state, permissions, an audit trail, domain data, integrations – the built-in feature tends to grow the category and send the serious users your way.

### The classic version – technically the more important one

Platform risk isn't an AI phenomenon. It is only louder now. A few cases, all of them the same pattern:

| Platform | What changed | What it cost the integrator |
| --- | --- | --- |
| Exchange Online | Basic authentication switched off for most protocols from October 2022; SMTP AUTH on its own, repeatedly moved timeline. | Every IMAP or POP integration with a username and password became an app registration, consent, an OAuth flow and narrower permissions – for everyone at once. |
| Docker Hub | Anonymous pulls rate-limited per IP from 2020, tightened again in 2025. | CI fleets behind one NAT shared one IP’s quota. Pipelines stopped worldwide at companies that thought this had nothing to do with them. |
| Snowflake | Single-factor password sign-in being phased out, service users included. | Every username-and-password integration became key-pair auth, with key storage and rotation – the snippet in the seam section. |
| Public APIs | Pricing changed with weeks of notice – Twitter and Reddit in 2023 are the famous ones. | Products that were viable in January weren’t in July. |
| Model APIs | Model versions retired on a published schedule. | Prompts and tuning fitted to one model became a forced migration with behaviour drift. Almost nobody prices this into a build. |

```yaml
# Anonymous Docker Hub pulls are rate-limited per IP – and a CI fleet behind one NAT is one IP.
# The fix is boring: authenticate, or pull through a registry you run yourself.
- uses: docker/login-action@v3
  with:
    username: ${{ vars.DOCKERHUB_USER }}
    password: ${{ secrets.DOCKERHUB_TOKEN }}
- run: docker pull registry.internal/mirror/library/node:22-alpine   # or: a pull-through cache
```

The sentence that holds it together: platform risk is also when your supplier's supplier changes its auth policy, and you are the one held to account.

> **Fig. 8** · The change starts two links away from you and lands on you, because you are the last link your users can call. Nobody in your company decided it, and nobody in your company can postpone it.
>
> *Diagram:* A chain of four: your users depend on you, you depend on your vendor, and your vendor depends on its platform. A change at the far end – passwords switched off, a rate limit, a new price, a retired model – travels back along the chain and lands on you, although nobody at your company decided it.

### Mitigation, technically

- **Have an eval set.** A model abstraction layer leaks: tool-calling formats, structured output, prompt caching and context windows all differ. An eval set, on the other hand, travels. With one, switching models is an engineering task. Without one, it is a matter of faith. I treat the provider as an adapter in the [shadow AI post](https://laszlonemes.com/blog/shadow-ai-local-ai); the eval set is what lets you actually swap it.
- **Keep the data and the state on your side**, not in the platform's conversation history.
- **Run the margin test.** What share of your gross margin is one vendor's API? What happens if they triple the price – or ship the same thing for free?
- **Look at where the big platforms don't go:** on-premises, air-gapped, sector-specific compliance, the local regulatory last mile, legacy integration. It is boring work, which is exactly why it lasts.

An eval case doesn't need a framework. It needs an input, an expectation you can check mechanically, and the reason it exists – the reason is what makes it worth keeping through the next three model migrations:

evals/invoices/017.json

```json
{
  "id": "invoice-extract-017",
  "input": "fixtures/invoices/017.pdf",
  "expect": {
    "total": "1234.50",
    "currency": "EUR",
    "vatId": { "present": true },
    "lineItems": { "count": 7 }
  },
  "check": "exact-fields",
  "why": "two VAT rates on one invoice – the case the March model got wrong"
}
```

> If your product is a prompt and an API call, you don't have a product – you have a deadline.

## Lock-in, procurement and the exit

**Lock-in isn't automatically bad.** It is one of the prices you can pay wisely: you buy speed and support with it. The trouble starts when you don't know how much you are paying. It helps to split it three ways:

- **Data lock-in** – does it come out, in a usable form?
- **Workflow lock-in** – how much process has been built on top of it?
- **Competence lock-in** – how many people's knowledge is now tied to one product?

The practical measure is simple: **how long would it take us to leave, and what stops while we do?** If nobody can answer, the lock-in isn't a decision. It is an accident.

### Onboarding is often bigger than the licence

Security review, a DPIA, a pen-test report, the sub-processor list, data residency, the contract, procurement – in an enterprise this takes from six months to a year and a half. From the vendor's side it is a good filter, by the way: a vendor who has been through it a hundred times has every document ready. If a vendor has to assemble them now, that is information in itself.

Many people draw the wrong conclusion from this: "in an enterprise it's faster to write it than to procure it". Often that is true. **But it isn't an argument for building. It is a criticism of the procurement process.** Taking on five years of operational responsibility to avoid a six-month process is a bad deal.

### The breakup

Think it through in the PoC phase, not when the relationship has already gone bad. What comes out as data, in what form, in how much time? Can you run both systems in parallel? Where do users keep their data in the meantime? Whoever hasn't tested the export doesn't have an exit option, only a hope.

```shell
# Week one, not the day of the breakup. Export, load it somewhere that isn't theirs, and count.
vendor export --format parquet --include-schema --out ./exit-test/      # whatever "documented export" means
duckdb -c "SELECT count(*) FROM './exit-test/documents/*.parquet'"
duckdb -c "SELECT count(*) FROM './exit-test/attachments/*.parquet'"
# Compare with what the product says it holds. Then time it: that number is your exit, not the contract clause.
```

> Test the export in the first week of onboarding, not on the day of the breakup.

## Build, buy, or both

### Build, when

- the integration *is* the product – the glue is the value, and no vendor will ever make it;
- the data can't leave the building, and there is no real on-premises or BYOC option;
- the pricing model runs against your usage curve;
- the scope is small and **stable** – a 500-line script often beats a SaaS at tens of thousands of dollars a year;
- you have the substrate and the skills already, and it extends an application you own.

### Buy, even if you could write it

- auth, cryptography, identity – never your own SSO;
- where the certification is the product: payroll, tax, e-invoicing, banking interfaces – a regulatory target that keeps moving;
- where you would compete with the features of a full-time R&D team;
- anything that will need fixing at two in the morning, when there is no on-call for it;
- the "it's just a CRUD app" trap: the CRUD is ten percent, the permission model is forty.

### Most often: both

The most common right answer is the two together: **buy the substrate** – database, identity, storage, model API, queue – and **build on it what makes you different.** The dividing line isn't whether something is "core". It is **whether you can get out of it.**

> **Fig. 9** · Build the top, buy the bottom, and draw the line by exits rather than by importance. A standard interface – the Postgres wire protocol, the S3 API, OIDC – means a second supplier exists. A closed platform means it doesn't.
>
> *Diagram:* The usual right answer is both. Build on top what makes you different – your workflow, your domain rules, the glue nobody sells. Buy underneath the substrate: a database behind the Postgres protocol, identity behind OIDC or SAML, storage behind the S3 API, a model API with your own eval set. Where there is a standard interface there is a way out, so buy freely. A closed platform with no way out deserves a second look, even when it is cheaper.

Where there is a standard interface and a way out – the Postgres wire protocol, the S3 API, OIDC – buy with a clear conscience. Where there is no way out, think twice, even when it looks cheaper. It is the same reasoning as every trade-off in the [engineering mindset post](https://laszlonemes.com/blog/engineering-mindset): not a verdict, a price you know.

> The question isn't whether it is core. It is whether you can get out of it.

## Objections, and the answers

| They say | You say |
| --- | --- |
| "With AI we can build it in a week. Why pay rent?" | You can build the v1 in a week. The rent pays for years two to five: on-call, security patches, upgrades, the next person. Compare the licence with five years of ownership, not with a week of prompting. |
| "The vendor covers 90%." | Who measured it, when, on what data, in whose environment? After an integration test it is usually 60–70, and the missing part sits exactly where you differ. |
| "Lock-in is always bad." | It is a price, and often a good one – speed and support. It is bad when you can’t say how long leaving would take and what would stop meanwhile. |
| "Procurement takes a year. Building is faster." | Often true. That is a criticism of procurement, not an argument for building. Five years of operating something is a high price for skipping six months of paperwork. |
| "Then never build on a model API." | No – own the data and the state, keep an eval set, and run the margin test. The risk isn’t the API. It is a product that is nothing but a prompt. |
| "Just abstract the model provider away." | The abstraction leaks: tool calling, structured output, caching and context windows all differ. The eval set is what travels between models. |

## Takeaways

1. **The question didn't change – only the price of the v1.** AI made the first version cheap; operations, integration, change and responsibility cost what they did.
2. **Buy is build too.** You buy three layers and inherit the fourth, integration. The real choice is build-and-own versus integrate-and-rent.
3. **Run the seam test before signing:** storage, database, secrets, auth, network and licence, deployment, observability – and the exit first.
4. **The 90% comes from the demo.** Measure it in your environment, and close the gap by shrinking the requirement before forking or wrapping.
5. **Use AI to disprove the demo**, not to write the build. A two-day PoC against your own Vault, database and storage is now cheap enough that skipping it is a choice.
6. **The adapter layer is a product decision.** The adapter is easy; the forty edge cases are years of customers.
7. **A good request describes the constraint, not the solution** – and when you build for yourself, nobody writes the constraint down.
8. **Platform risk is a supply chain.** Own the data and the state, keep an eval set, know how much of your margin one API is.
9. **Draw the line by exits.** Buy the substrate behind standard interfaces, build what makes you different, and test the export in week one.

Build or buy was never really the question. The question is what you will own in five years, and whether you chose it.

## Sources

- [Virtual hosting of buckets](https://docs.aws.amazon.com/AmazonS3/latest/userguide/VirtualHosting.html) – AWS on virtual-hosted versus path-style addressing.
- [Avoid getting throttled or blocked in SharePoint Online](https://learn.microsoft.com/en-us/sharepoint/dev/general-development/how-to-avoid-getting-throttled-or-blocked-in-sharepoint-online) – `429`, `503` and honouring `Retry-After`.
- [CREATE EVENT TRIGGER](https://www.postgresql.org/docs/current/sql-createeventtrigger.html) – superuser only, and [CREATE EXTENSION](https://www.postgresql.org/docs/current/sql-createextension.html) on trusted and untrusted extensions.
- [Vault database secrets engine](https://developer.hashicorp.com/vault/docs/secrets/databases) and the [Agent Injector annotations](https://developer.hashicorp.com/vault/docs/platform/k8s/injector/annotations) – dynamic credentials and their leases.
- [Snowflake key-pair authentication and rotation](https://docs.snowflake.com/en/user-guide/key-pair-auth) – `RSA_PUBLIC_KEY` and `RSA_PUBLIC_KEY_2`.
- [Deprecation of Basic authentication in Exchange Online](https://learn.microsoft.com/en-us/exchange/clients-and-mobile-in-exchange-online/deprecation-of-basic-authentication-exchange-online) – the timeline and what replaced it.
- [Docker Hub usage and limits](https://docs.docker.com/docker-hub/usage/) – pull limits for anonymous and authenticated users.
- [Continuous access evaluation](https://learn.microsoft.com/en-us/entra/identity/conditional-access/concept-continuous-access-evaluation) – why a token can stop being valid before it expires.
