Matthew Loebenstein
Analysis AI · Sovereignty · 2026

Frontier AI for less than your CRM

A self-hosted platform running open Chinese models covers ninety percent of real business workloads, costs less than fifty seats of Salesforce on auditable AWS infrastructure you fully control, and keeps your data, and your IP, out of a competitor's hands. Here is the case, with the numbers.

// Context

This started as a cost-modelling exercise and became something more opinionated. I set out to compare what it costs to self-host frontier-class open-source models from Chinese labs (DeepSeek V4 Flash, GLM-5.2) against a run-of-the-mill SaaS subscription, and ended up making the case that the real risk to a company's intellectual property is not a foreign lab but the American vendor whose API it already pays for.

All pricing is drawn from public cloud rate cards and vendor pricing pages as of mid-2026, and all model claims reference published benchmarks. Where I use my own client work as evidence, company and product details are stripped out.

The part nobody prices in: your IP

In 2021, Jasper built one of the fastest-growing AI companies on record. A $1.5 billion valuation, $125 million raised, roughly $120 million in annual revenue, all on the back of a simple proposition: a well-designed interface over OpenAI's models. Then, in November 2022, OpenAI shipped ChatGPT. Free. The product Jasper sold was now a feature of the thing Jasper rented. Revenue fell by more than half within a year. The layoffs came in rounds. Both co-founders eventually stepped down.

Jasper is the cleanest documented case, but it is not an outlier. It is the pattern. Every time a US AI lab ships a feature (PDF upload, web browsing, code execution, memory), a category of startups built on that lab's API dies overnight. Not through competition. Through absorption. One founder on OpenAI's own forum, having watched his startup's core idea get baked into the platform twice, put it plainly: imagine you sell hand-painted sneakers and Walmart clones them and floods every store. You cannot out-distribute a company worth half a trillion dollars that just read your product in its own request logs.

That is the structural point, and it deserves to be stated without euphemism. If your product runs on someone else's model, you are renting your company from a landlord who is also your competitor, and who can read every request you send. Your prompts, your documents, your customers' data, your product roadmap as it forms in real time: all of it flows through infrastructure you do not control, owned by a company with every commercial incentive to build what you have just proven is worth building.

The honest version of this argument needs a caveat. The US enterprise APIs, OpenAI and Anthropic both, contractually commit to not training on your data. They offer retention windows measured in days and zero-retention endpoints for qualifying customers. That is real. But it is a policy, not an architecture. It rests on trust, and that trust is thinner than it looks. Consumer accounts play by different rules, and your employees use them for work every day. Retention policies can change. And even where the letter of the policy holds, "we promise not to look" is a weaker guarantee than "the server is in your rack."

Then there is the risk that has nothing to do with intent. Access to a US model is a policy decision away from disappearing. When the US government issued an export-control directive prohibiting foreign nationals from using Anthropic's latest model, the company took the model offline entirely. If your product's engine depends on that API, an arbitrary decision made in Washington becomes your problem, wherever in the world you are. Being subject to American domestic and foreign policy is a real and growing constraint on any entrepreneur trying to build AI capability globally. Pricing has a similar effect: opaque, unilateral, and revised without a seat at the table for the companies paying the invoices. Models and their capabilities can feel as though they change at random, and safety features, however well-intentioned, can stand in the way of legitimate work and genuine innovation.

Here is the reframe this article is built on. We have been trained to worry about Chinese models seeing our data. We should be at least as worried about the company that already does: the American vendor whose invoice you pay every month, whose terms of service you clicked through, and whose product team ships features that look, increasingly, like your roadmap. The threat to your intellectual property is not a lab in Hangzhou you have never used. It is the landlord you already have.

The bill you already pay without thinking

Fifty seats of Salesforce, on the Enterprise edition most mid-size companies land on, costs $105,000 a year. That is before a single AI feature, before CPQ, before the implementation partner. Add those and a realistic first year runs to $150,000 or more, settling to roughly $150,000 annually ongoing. For software that manages your pipeline.

I pick Salesforce deliberately, not because it is bad value, but because it is the most ordinary software purchase a growing company makes. Nobody gets fired for buying Salesforce. It is the default, the line item that sails through budget season without a fight. That ordinariness is exactly what makes it the right baseline. This is what "normal" costs. When people say self-hosted AI is expensive, this is the number they are unconsciously comparing it against, and it is worth holding both in view at once.

What do you get for that $150,000? A database with a polished interface, a deep ecosystem of integrations, and a workflow engine your sales team lives in. Genuinely useful. But it does not think. It does not read a contract, draft a proposal, summarise a call, or reason over your data. It records. The intelligence layer, the part that actually does work, is sold separately.

And it is sold separately at a price that makes the point for me. Salesforce's own AI tier, Agentforce, runs to $550 per seat per month. Those same fifty seats, with the AI your CRM vendor will happily sell you, come to $330,000 a year. The intelligence costs more than twice the system it sits on. Keep that figure in mind as we turn to what frontier-class AI actually costs when you own the machine it runs on.

The alternative: own the machine

"Self-hosted AI" sounds like a research project. It is not. In practice it means renting a GPU server from a cloud provider, exactly as you already rent web servers, downloading a set of open model weights once, and serving them to your team through the same kind of API you would call at OpenAI. The weights are MIT-licensed, free to use commercially, and free to modify. Your prompts, your documents, and your customers' data stay inside your own cloud perimeter. Nobody trains on them, nobody retains them, and nobody reads your product roadmap out of a request log. The models in question are not toys. They are the ones currently trading blows with GPT-5.5 on the coding and agentic benchmarks.

Here is the comparison this article exists to make. I have priced the self-hosted column on AWS, the same kind of auditable, SOC 2-attested infrastructure your CRM runs on, because the whole point of this article is control over your data, and you do not get that from an anonymous budget cloud. One column is a CRM; the other is a frontier-class AI you own outright, on a dedicated instance inside your own cloud perimeter.

// Annual cost, 50 seats vs self-hosted frontier AI (AWS, 8× A100)

Salesforce Enterprise (50 seats) . . . . . . . ~$105,000 / yr
Salesforce + CPQ & implementation . . . . ~$150,000 / yr
Salesforce Agentforce AI (50 seats) . . . ~$330,000 / yr

DeepSeek V4 Flash, on-demand . . . . . . . . . ~$192,000 / yr
DeepSeek V4 Flash, 3-yr reserved . . . . . . . . ~$82,000 / yr
GLM-5.2, on-demand . . . . . . . . . . . . . . . ~$192,000 / yr
GLM-5.2, 3-yr reserved . . . . . . . . . . . . . ~$82,000 / yr

Read the reserved rows first. Committed for three years, either model, GLM-5.2 trading blows with GPT-5.5, or DeepSeek V4 Flash beating its larger sibling across the benchmarks its publisher reports, comes in under the base Salesforce bill, and at roughly a quarter of the Agentforce AI tier. Even on-demand, with no commitment at all, both sit comfortably under the loaded Salesforce figure and well under half of what your CRM vendor wants for its own AI. The intelligence is not the expensive line item. The record-keeping software is.

One honest caveat, because the table is only useful if it survives scrutiny. These figures are the compute on AWS, auditable and attested, and they do not include the engineering time to stand the system up and keep it healthy. Add a part-time engineer and the all-in figure rises, but it still lands under the Agentforce bill. And if auditability matters less for a given workload, moving the same model to a budget GPU cloud cuts the compute cost by more than half again; the $82,000 figure is the conservative, fully-auditable ceiling, not the floor. The American APIs sit between the columns on cost and carry every one of the IP risks from the first section. The shape of the comparison holds across all of those adjustments: owning the machine is cheaper than the CRM's own AI, and it is not close.

And the self-hosted column buys things no SaaS contract offers. No per-seat pricing, so the cost does not grow as your team does. No rate limits. The freedom to fine-tune the model on your own data, so it gets better at your specific work the longer you run it. The vendor cannot deprecate your model, change your pricing, or take your capability offline with a policy decision. You are not renting the engine. You own it.

What "good enough for 90% of the work" actually means

The claim that an open model covers ninety percent of use-cases only means something if you are honest about what that work is. For most businesses it is not frontier research. It is summarising a hundred-page contract, drafting the proposal, extracting structured data from messy documents, answering questions over your own knowledge base, writing and reviewing code, and running agents that chain tools together to get a job done. That is the workload. And on exactly those tasks, summarisation, drafting, extraction, code, agentic tool-use, the open Chinese models are not behind the frontier. On several of them they set it.

This is not a hedge. It is a claim with numbers behind it. DeepSeek's V4 Flash, the smaller of its two current models, outperforms its own larger sibling on all nine of the benchmarks DeepSeek publishes. GLM-5.2 trades blows with GPT-5.5 on long-horizon coding and agentic work. These are the same tasks a business actually buys AI to do. "Near-frontier" is not a polite way of saying second-rate; on the work that matters, the gap has effectively closed. (For readers who think in API terms: GLM-5.2's hosted version costs $1.40 per million input tokens. Self-hosted, you stop counting tokens altogether.)

I am not asking you to take this on faith, because I have built it. For a document-intelligence product I designed and shipped a multi-agent pipeline that reads, reasons over, and extracts from complex documents, running entirely on self-hosted open models. The all-in cost came to $0.0034 per scan. That is not a projection or a vendor's benchmark; it is a production system I controlled end to end, with no third party touching the data. The reason I am confident a self-hosted platform covers ninety percent of real workloads is that I have watched one do it.

The honest ten percent exists, and naming it makes the ninety credible. At the absolute frontier of reasoning, and in some multimodal work, the largest proprietary US models still lead. If your product depends on that specific edge, you may need them, for now, for that slice. But most businesses are not operating at that edge. They are operating in the broad middle where the open models have already won on price, on control, and increasingly on capability. The question is not whether self-hosted AI can do the work. It is whether you are comfortable continuing to pay a premium, in money and in IP, for a margin of capability you will never use.

Who is actually copying whom

There is a comfortable story the industry tells itself: American labs invent, Chinese labs copy. It was perhaps true once, in other industries, a generation ago. In AI today the engineering record says the opposite, and you only have to read the papers rather than the headlines to see it. The techniques that make the current generation of powerful, affordable models possible did not come out of San Francisco. They came out of Hangzhou and Beijing, published openly, and are now quietly used by everyone.

Three examples, all from DeepSeek, all now embedded in how the field builds models. Multi-head Latent Attention compresses the key-value cache that long-context inference depends on by more than ninety percent, which is precisely what makes a million-token context window economically viable; it is now the reference approach serious long-context systems are built on. GRPO, the reinforcement-learning method behind DeepSeek's reasoning models, was published openly and has become the default recipe other labs use for reasoning post-training. And the quantization-aware training recipes, the methods for training models to run at low precision without losing capability, are the entire reason the frontier-class weights in the previous section are cheap enough to run on hardware you can actually rent. None of these were minor tweaks. Each one moved the economics of the whole field.

The adoption numbers say the quiet part out loud. Qwen, Alibaba's open model family, is now the most-downloaded open model family in the world. The world is not downloading copies of American innovation. It is building, at scale, on Chinese research.

The deeper difference is not technical. It is cultural. The US labs that once published their methods have gone quiet; the architectures, the training recipes, the scaling insights are now trade secrets. The Chinese labs are doing the opposite: publishing the architecture, the training method, and the weights, under licences that let anyone use, study, and build on them. One culture is accumulating private advantage and selling access to it by the token. The other is contributing to a commons that makes everyone who touches it more capable. That difference is not a footnote to this argument. It is the argument, and it leads directly to the question of what we should want this industry to be.

What we should want this industry to be

Listen to how the leaders of the major US labs talk about artificial intelligence and a pattern emerges. It is a race. It must be won. It is framed as a geopolitical contest, an arms race, a contest for dominance in which the prize is so large that almost anything is justified in pursuit of it. The behaviour follows the rhetoric: the weights are closed, the methods are secret, the pricing extracts as much as the market will bear, and the safety language so often serves the competitive position. I want to be careful here, because these are not stupid people and the stakes they describe are real. But when winning and money become the explicit goals, moving the science forward becomes a means rather than the end, and it is hard to escape the sense that the people building the most powerful technology of our generation would trade almost anything for an edge.

Software, at its best, has never worked that way. No other industry has anything like open source. Nowhere else do the most skilled practitioners in a field take the product of their labour, the thing they could most easily sell, and give it away, so that a stranger on another continent can use it, learn from it, and build something better on top of it. That culture is not a quirk of the software industry. It is the reason the industry is extraordinary. The internet runs on it. Every major technology of the last thirty years was built on it. It is the clearest proof that the people who make software have always believed the craft matters more than the margin, that the thing itself, made well and shared widely, is the point.

I say "we" deliberately. Most of us who build technology came into it because we cared about the product and the craft, about making something genuinely good, more than we cared about the money. That is not naivety; it is the culture that produced everything worth using. When I look at the Chinese labs publishing their architectures and giving away frontier-class weights, and at the American labs closing their methods and metering access by the token, I know which of the two feels like the software industry I recognise. Choosing open, self-hosted models is not only the pragmatic choice, on cost and on control of your IP. It is a choice about which set of values you want to build your company on.

Because underneath the pricing and the benchmarks, these are two different visions of what technology is for. In one, intelligence is a resource to be enclosed: gated behind an opaque API, priced per token, watching what you build, optimised to extract. It treats the people and companies who use it as metered endpoints, and there is something faintly anti-human in that, a reduction of judgement, creativity, and work to usage on someone else's meter. In the other vision, intelligence is a capability to be shared: weights you can hold, run, study, and improve, that get better the more openly they are used. I know which one I want to build on.

What I would build for you

This is the part of the article where I stop analysing and make the offer, so let me be plain about what it is. I build self-hosted AI platforms for companies that want frontier-class capability without handing their intellectual property to a competitor. A platform that is close enough to the frontier to cover ninety percent of real workloads, that costs less to run than the SaaS you already pay for without thinking, and inside which your data never leaves your own perimeter. Not a proof of concept. A system your business runs on.

In practice that means standing up the serving infrastructure on rented GPU capacity, wiring the model into the workflows where it actually earns its keep, and, where it helps, fine-tuning on your own data so the system gets better at your specific work the longer it runs. Then I hand it over. It is yours: the infrastructure, the model, the data, the capability. No per-seat pricing, no rate limits, no vendor who can deprecate the thing you built on or read your roadmap out of a request log.

I am not describing an aspiration. I have built exactly this kind of system: a multi-agent pipeline that reads and reasons over complex documents, running entirely on self-hosted open models, at a cost of $0.0034 per scan. The platform I am proposing is a repeat of work I have shipped, not a bet on something I have only read about.

I will also tell you honestly that it is not for everyone. If your product genuinely depends on the absolute bleeding edge of proprietary reasoning, or your workloads are too light to justify a dedicated platform, the API is a reasonable choice, for now, for that slice. But here is the counterweight that I think tips it for most companies, and it is only getting heavier. At a moment when everyone is chasing the AI edge, and bending the normal rules of business to get there, being able to offer an auditable guarantee that your customer data is contained within your own systems stops being a defensive precaution and becomes a competitive advantage. Especially in regulated industries, the ability to prove, not merely promise, that nothing leaves your perimeter is something your competitors cannot easily copy and your customers will increasingly pay for.

Which brings it back to where we started. The question was never really China against America, or cheap against expensive. It is whether you own the engine your business runs on, or rent it from a landlord who competes with you. And it is more than that: it is a question of how you see technology, and humanity's place in it. Whether intelligence is something to be enclosed, metered, and sold back to you by the token, or something to be shared, studied, and built on together. I know which future I want to help build.

If this is a conversation worth having for your company, I'd like to hear from you.

· End ·


Home