AI "with a proprietary model" is the phrase that comes up most in sales calls with AI vendors this year, and the one explained least. It sounds like a serious differentiator: "we don't use ChatGPT like everyone else, we have our own model." In most cases, when you ask what is actually behind it, the answer falls apart. Let's put the three real levels on the table, and the exact question that tells them apart.
What the vendor is selling when they say "proprietary model"
"Proprietary model" is a label that sounds good and costs nothing to say. The problem is that three very different things fit inside that label, with completely different cost, time and meaning, and almost nobody separates them when selling:
- Training a language model from scratch. Building the neural network and training it on proprietary data, without starting from any existing model.
- Fine-tuning an already trained base model with proprietary data, so it performs better in a specific domain.
- Putting a third party base model to work on proprietary data, with a good information retrieval system (RAG), tools and a well built prompt, without touching the model's weights at all.
All three get presented as "we have our own AI." Only the first one is a proprietary model in the strict sense. The second is a modified third party model. The third, by far the most common, is an unmodified third party model given access to proprietary data. None of the three is bad on its own; the problem is selling all three under the same name, because the client pays for the label believing they are buying the first one.
The three levels, without the marketing
| Level | What it really is | What it requires | Who it makes sense for |
|---|---|---|---|
| Trained from scratch | A new neural network, trained end to end on proprietary data | Massive data volumes, a specialized team, months of work | Practically no small or medium business; not even most large ones |
| Fine-tuned | An existing base model, partially retrained on proprietary examples | A curated, sufficiently large dataset, and knowing when it beats other options | Very specific cases: a language or format the base model doesn't handle well |
| Base model + proprietary data (RAG, tools, agent) | The model itself doesn't change; what changes is what information and actions it has access to | Getting the data in order, building retrieval and tools, integrating with the real system | The vast majority of businesses that want AI applied to their operation |
The table makes clear something that sales conversations tend to blur on purpose: the level almost any business actually needs is the third one, not the first. And the third one doesn't require training anything: it requires having the data in order and building the integration well. That's exactly the angle we work from at AutoBoost: get the data in order first, then apply AI on top, not the other way around.
The question that unmasks it
You don't need to be technical to find out which level your vendor is really at. One two part question is enough, and watching whether they answer both halves:
"What data did you train the model on, and what base model runs underneath?"
- If they answer with proprietary dataset names, training size and compute time: level 1 or 2, genuinely proprietary or fine-tuned.
- If they answer "that's confidential about our technology" with no concrete detail: high suspicion, most likely level 3 sold as level 1.
- If they answer with the name of a known commercial model (or actively dodge the question): level 3. Nothing wrong with that, as long as it's sold for what it is.
This is the decision rule that actually matters: if your vendor cannot tell you what data the model was trained or fine-tuned on, nor which model runs underneath, they don't have a proprietary model: they have a third party model with a well built system on top. That can be an excellent solution. What it can't be is priced as if it were the other thing.
Why this matters to whoever runs the business, not just the engineer
This distinction isn't an engineering nuance, it's a purchasing decision. It changes three things:
- What price makes sense to pay. Training a model from scratch is an investment of an entirely different order than integrating an existing model with your data. If you're charged as if it were the first and it's actually the second, you're overpaying for a label.
- The real lock-in you take on. A genuinely proprietary model ties you to the vendor who trained it. A base model with your own data on top ties you far less: the data and the integration are yours, and the underlying model can be swapped without redoing the whole project.
- What you can demand if something goes wrong. If the problem is that the model "doesn't know" something about your business, the fix at level 3 is improving the data it accesses; at level 1 or 2, it means retraining, which is far slower and far more expensive.
When we built a sovereign AI platform with RAG over technical manuals for an industrial company, the language model runs in Spain, never leaving their infrastructure, and that is a serious commitment: it isn't a "proprietary model" because nothing was trained from scratch, it's tailor made integration with data sovereignty, and that's exactly how we describe it, without inflating the label.
Checklist before paying for a "proprietary model"
Before signing with a vendor who uses that phrase, ask this and write down the answers:
- What data was the model trained or fine-tuned on? (dataset name, source, approximate size)
- What base model runs underneath, if any?
- What part is training, and what part is retrieval over your documents?
- If you switch base model providers tomorrow, what do you lose from your investment, and what stays with you (the data, the integration)?
- Does the price reflect training a model, or integrating an existing one with your system?
If more than two answers come back without real detail, you now know which level you're actually negotiating.
The label isn't the problem, the confusion is
There's nothing wrong with building on a well integrated base model with your own data on top: it's what almost any business needs, and when done right, it's where the real value sits. The problem is paying for the "proprietary model" label without knowing which of the three levels you're actually buying. At AutoBoost we work at the level that genuinely moves the needle for a small or medium business: getting the data in order and integrating AI into the real system, tailor made software, not generic templates. If you want us to review what you're being sold, or what you actually need, let's talk at /contacto.

