Why Thomson Reuters Actually Needed This
Thomson Reuters sits on an enormous, unique dataset that nobody else has. Decades of legal research content, tax and accounting guidance, case law analysis, built up over generations and available nowhere else in that combination. A general-purpose AI model, no matter how capable, cannot replicate the depth and specificity of that proprietary content because it was never trained on it. For a company whose entire competitive advantage is that proprietary content, using an off-the-shelf model for AI features means giving up the one thing that makes them different, or spending enormous effort trying to feed proprietary data into someone else’s model in a way that still would not be as effective as training directly on it.
That is the actual bar for when building your own model makes sense: you have a genuinely unique, defensible dataset, large enough and valuable enough that a general model cannot replicate what you know, and the AI capability built on that data is core to what makes your product worth paying for.
Why This Almost Never Applies to a Founder-Sized Business
Most businesses, even good, successful ones, do not clear that bar, and that is not a criticism, it is just math. Building and training a model, even from an open-source base as Thomson Reuters did, requires the kind of compute and specialized talent investment that only makes sense when the proprietary data advantage is large enough to justify it. Thomson Reuters spent 40 million dollars on this. That number alone should clarify who this is actually for.
If your business does not have decades of unique, proprietary data at genuinely large scale, unavailable anywhere else, you are very likely better served using an existing frontier model and focusing your effort on how you apply it, your prompts, your workflows, your systems, rather than trying to build the underlying model yourself. The advantage in almost every founder-sized business is not going to come from a better base model. It is going to come from a smarter system built on top of a model someone else already trained.
The Real Lesson Is About Data, Not Models
What Thomson Reuters actually demonstrates is not “build your own AI,” it is “know what your genuinely proprietary asset is, and protect it accordingly.” For Thomson Reuters, that asset is decades of legal and financial content. For your business, it might be a specific process you have refined over years, a dataset of customer interactions and outcomes, or institutional knowledge about your specific market that a generic model would never have access to.
The question worth asking is not whether you should build your own model, for almost everyone reading this the answer is no. The question is whether you have identified the specific data or knowledge that makes your business defensible, and whether your current AI usage is actually leveraging that asset or just using a generic model in a generic way that any competitor could replicate with the same tool.
What to Actually Do With This
Look at whatever proprietary knowledge, data, or process actually makes your business different from a competitor. Then ask honestly whether your current AI tools and workflows are built around that asset, feeding it in as context, using it to fine-tune prompts, building systems specifically shaped by what you know that others do not, or whether you are just using AI in a generic way disconnected from what actually makes you valuable.
You do not need Thomson Reuters’ budget to apply the same principle at your scale. You need clarity about what your version of “decades of Westlaw content” actually is, and a deliberate decision to build your AI usage around leveraging that asset rather than treating AI as a generic tool doing generic work.
Also read: What Growth Strategy Actually Looks Like When AI Does the Execution
Neon Aliens exists at the intersection of AI and the ideas most brands are not paying attention to yet.
We work with a small number of founders at a time. See if you qualify.
See If We’re a Fit