TheoryAugust 26, 20266 min read

AI with a "proprietary model": the question that unmasks it

More and more providers sell "we have our own AI model" as a differentiator. In most cases it is a third party model with a prompt on top. Here are the three real levels and the question that tells them apart.

Pol

Fundador de AutoBoost

Applied AIAI models
AI with a "proprietary model": the question that unmasks it

AI "with a proprietary model" is the phrase that comes up most in sales calls with AI vendors this year, and the one explained least. It sounds like a serious differentiator: "we don't use ChatGPT like everyone else, we have our own model." In most cases, when you ask what is actually behind it, the answer falls apart. Let's put the three real levels on the table, and the exact question that tells them apart.

What the vendor is selling when they say "proprietary model"

"Proprietary model" is a label that sounds good and costs nothing to say. The problem is that three very different things fit inside that label, with completely different cost, time and meaning, and almost nobody separates them when selling:

  1. Training a language model from scratch. Building the neural network and training it on proprietary data, without starting from any existing model.
  2. Fine-tuning an already trained base model with proprietary data, so it performs better in a specific domain.
  3. Putting a third party base model to work on proprietary data, with a good information retrieval system (RAG), tools and a well built prompt, without touching the model's weights at all.

All three get presented as "we have our own AI." Only the first one is a proprietary model in the strict sense. The second is a modified third party model. The third, by far the most common, is an unmodified third party model given access to proprietary data. None of the three is bad on its own; the problem is selling all three under the same name, because the client pays for the label believing they are buying the first one.

The three levels, without the marketing

LevelWhat it really isWhat it requiresWho it makes sense for
Trained from scratchA new neural network, trained end to end on proprietary dataMassive data volumes, a specialized team, months of workPractically no small or medium business; not even most large ones
Fine-tunedAn existing base model, partially retrained on proprietary examplesA curated, sufficiently large dataset, and knowing when it beats other optionsVery specific cases: a language or format the base model doesn't handle well
Base model + proprietary data (RAG, tools, agent)The model itself doesn't change; what changes is what information and actions it has access toGetting the data in order, building retrieval and tools, integrating with the real systemThe vast majority of businesses that want AI applied to their operation

The table makes clear something that sales conversations tend to blur on purpose: the level almost any business actually needs is the third one, not the first. And the third one doesn't require training anything: it requires having the data in order and building the integration well. That's exactly the angle we work from at AutoBoost: get the data in order first, then apply AI on top, not the other way around.

The question that unmasks it

You don't need to be technical to find out which level your vendor is really at. One two part question is enough, and watching whether they answer both halves:

"What data did you train the model on, and what base model runs underneath?"

  • If they answer with proprietary dataset names, training size and compute time: level 1 or 2, genuinely proprietary or fine-tuned.
  • If they answer "that's confidential about our technology" with no concrete detail: high suspicion, most likely level 3 sold as level 1.
  • If they answer with the name of a known commercial model (or actively dodge the question): level 3. Nothing wrong with that, as long as it's sold for what it is.

This is the decision rule that actually matters: if your vendor cannot tell you what data the model was trained or fine-tuned on, nor which model runs underneath, they don't have a proprietary model: they have a third party model with a well built system on top. That can be an excellent solution. What it can't be is priced as if it were the other thing.

Why this matters to whoever runs the business, not just the engineer

This distinction isn't an engineering nuance, it's a purchasing decision. It changes three things:

  • What price makes sense to pay. Training a model from scratch is an investment of an entirely different order than integrating an existing model with your data. If you're charged as if it were the first and it's actually the second, you're overpaying for a label.
  • The real lock-in you take on. A genuinely proprietary model ties you to the vendor who trained it. A base model with your own data on top ties you far less: the data and the integration are yours, and the underlying model can be swapped without redoing the whole project.
  • What you can demand if something goes wrong. If the problem is that the model "doesn't know" something about your business, the fix at level 3 is improving the data it accesses; at level 1 or 2, it means retraining, which is far slower and far more expensive.

When we built a sovereign AI platform with RAG over technical manuals for an industrial company, the language model runs in Spain, never leaving their infrastructure, and that is a serious commitment: it isn't a "proprietary model" because nothing was trained from scratch, it's tailor made integration with data sovereignty, and that's exactly how we describe it, without inflating the label.

Checklist before paying for a "proprietary model"

Before signing with a vendor who uses that phrase, ask this and write down the answers:

  • What data was the model trained or fine-tuned on? (dataset name, source, approximate size)
  • What base model runs underneath, if any?
  • What part is training, and what part is retrieval over your documents?
  • If you switch base model providers tomorrow, what do you lose from your investment, and what stays with you (the data, the integration)?
  • Does the price reflect training a model, or integrating an existing one with your system?

If more than two answers come back without real detail, you now know which level you're actually negotiating.

The label isn't the problem, the confusion is

There's nothing wrong with building on a well integrated base model with your own data on top: it's what almost any business needs, and when done right, it's where the real value sits. The problem is paying for the "proprietary model" label without knowing which of the three levels you're actually buying. At AutoBoost we work at the level that genuinely moves the needle for a small or medium business: getting the data in order and integrating AI into the real system, tailor made software, not generic templates. If you want us to review what you're being sold, or what you actually need, let's talk at /contacto.

Share article
AI "proprietary model": what it really means | AutoBoost