Access, Not Ownership
There is an attractive idea circulating among software companies: train an AI on your own codebase and you get a programmer that understands your business better than any off-the-shelf model ever could. Thousands, perhaps millions, of lines of code; your developers already understand the architecture and the rules; feed it all into a private model and the result should be an expert in your software.
It is an appealing proposition. It is also frequently the wrong one.
The distinction that matters is between giving an existing model access to your code and training a new model on your code. Those are radically different undertakings, and for most organisations the first is considerably more useful than the second.
Five Things People Mean by “Our Own AI”
When a company says it wants its own AI, that can mean running an open-weight model on its own servers, connecting a commercial model to its private repositories, retrieving relevant code via RAG, fine-tuning an existing model, or actually pre-training a new foundation model from scratch. These are not equivalent, and the last one is the spectacularly expensive one.
Pre-training a foundation model means competing, on infrastructure, data, and engineering talent, with companies whose entire business is building foundation models. An NVIDIA H100 SXM carries a manufacturer-rated thermal design power of up to 700W. A Lawrence Berkeley National Laboratory study that physically measured an eight-GPU H100 training node under load recorded a sustained draw of 8.4kW — and that's one node, before cooling, networking, or the wider production system around it. The IEA's April 2026 update puts global data-centre electricity demand up 17% in 2025 alone, on course to double again by 2030.
That is the infrastructure bill before a single line of training code has run. The company isn't buying a few GPUs. It's taking on infrastructure, model engineering, evaluation, security, and the opportunity cost of its best engineers working on the model instead of the product.
You Don’t Need to Retrain the Brain
Hiring an exceptionally capable engineer doesn't normally involve handing them your codebase and saying study this for six months, then we'll let you write something. You hire someone who already knows how to program, then give them the documentation, the issue tracker, the tests, and colleagues to ask.
The same principle increasingly applies to coding agents. OpenAI's Codex is explicitly designed to connect to a company's repositories and fix bugs, write tests, and build features without the model being trained on that company's code at all — OpenAI's own enterprise documentation states plainly that Codex does no training on enterprise data. Anthropic's Claude Code follows the same agentic approach: navigate the codebase, edit multiple files, run commands, test the result.
The model doesn’t need to memorise your codebase. It needs to be able to interrogate it. That is a much easier problem.
Retrieval-Augmented Generation is the mechanism that makes this work. Instead of modifying the model's weights, a system indexes the company's code and documentation; a developer asks why the authentication service rejects a token, and the system retrieves the relevant files, configuration, and history for the model to reason over. The underlying model stays general-purpose. The context is private. And the codebase can change tomorrow without the model needing to be retrained — which matters, because a repository is not a static textbook. Dependencies get upgraded, APIs get deprecated, bugs get fixed, conventions change, every single day.
A 2025 study examining over 160,000 internal C++ files at Tencent, run jointly with researchers at CUHK and Harbin Institute of Technology, compared RAG against fine-tuning for industrial code completion directly. RAG, implemented with proper retrieval, achieved higher accuracy than fine-tuning alone, the two approaches combined for a further improvement, and RAG scaled better as the codebase grew. The obvious-sounding answer to how do we make the AI understand our enormous codebase may not be to train the AI on the enormous codebase. It may simply be to give the AI good access to it.
Your Code Is Not a Training Dataset
Your company's code is not a collection of best practices. It's simply the code your company currently has: obsolete APIs, duplicated logic, inconsistent naming, undocumented workarounds, historical mistakes nobody has had time to fix. Train a model primarily on that material and you've built a feedback loop — the organisation asks the AI to improve the code, the AI has learned from the existing code, the existing code contains the mistakes. The model has to distinguish between how your organisation currently does things and how things should actually be done, and that distinction doesn't come for free.
There's a real caveat here: fine-tuning doesn't have to mean blindly copying the repository. A well-designed process can include curated examples, expert-written patches, security findings, and tests as training signal, which is exactly how frontier coding models are themselves trained. The problem isn't that learning from bad code inevitably produces bad code.
A company’s raw repository is not, by itself, a sufficiently good training dataset for creating a superior software engineer.
The Frontier Has a Head Start
Companies like OpenAI, Anthropic, and Google have spent enormous resources on general-purpose models, trained on far broader datasets with heavy investment in reasoning, tool use, and safety. A smaller company buying a rack of GPUs doesn't inherit any of that by proximity.
On SWE-bench Verified, the most widely cited real-world coding benchmark, Claude Opus 4.5 became the first model to break 80%, scoring 80.9%; GPT-5.2 followed close behind around 80.0%, with Gemini 3 Pro at roughly 76.2%. Worth knowing: OpenAI's own internal audit found that frontier models could reproduce verbatim gold patches for some SWE-bench Verified tasks, and OpenAI deprecated the benchmark in February 2026 over exactly that contamination concern. Impressive scores on a benchmark the model may have partially memorised are not the same thing as a demonstrated skill.
CCBench was built to address precisely that gap — small, real-world repositories specifically excluded from training data, testing whether an agent can actually work on something unfamiliar rather than recall something it's seen. In its February 2026 results, GPT-5.2-Codex scored 75.4% and Claude Code with Opus 4.6 scored 72.7%. The model doesn't need to have trained on your source code to work effectively with it. It needs to be good at software engineering, on code it has never seen before.
Where Private Models Actually Win
None of this means commercial models are automatically superior at everything. Research into OpenAPI code completion found a fine-tuned Code Llama model — 7 billion parameters, roughly 25 times smaller than the commercial system it was measured against — achieved a 55.2% correctness improvement over GitHub Copilot on that specific, narrow task. That's not evidence that small models generally win. It's evidence that specialisation can beat generalisation when the problem is sufficiently narrow: a specific internal configuration language, repetitive boilerplate, a very particular coding convention.
The question isn’t “can our model beat ChatGPT?” The question is “can our system perform this particular task well enough to justify its total cost of ownership?”
The Economics Are Backwards
Building your own model means GPU infrastructure, a training pipeline, data collection and cleaning, evaluation, inference infrastructure, ongoing maintenance, retraining every time the codebase moves, and security on top of all of it — compared against a model a frontier lab has already spent billions refining. Using an existing model means taking a strong coding model, giving it controlled repository access, good retrieval, the test suite, the compiler, your coding standards, and human review, then measuring and improving the system around it rather than attempting to recreate the foundation model underneath. For most organisations, the second path is vastly easier to justify, and the real-world productivity data backs the direction even if the specific numbers vary by study: GitHub's randomised trial with Accenture (450 developers against a 200-person control group) found measurable gains in pull-request volume and merge rate, and an earlier controlled experiment found Copilot users completed a coding task roughly 56% faster than a control group. These are vendor-sponsored studies and should be read as such, not as neutral proof of universal productivity gains — but they're a long way from nothing.
The Middle Ground Nobody Mentions
The choice isn't between training a gigantic model from scratch and sending all your source code to a third party. An organisation can take an open-weight model and run it entirely inside its own infrastructure — real advantages around data sovereignty, privacy, and customisation, and potentially cost, once usage volume is high enough. If a company is generating billions of tokens a day, paying frontier-model prices forever can become expensive enough to justify operating that infrastructure itself.
Notice the distinction, though: it may make sense to self-host somebody else's model. That does not mean it makes sense to train your own. Those are completely different economic propositions, and conflating them is how companies end up building an AI research team when what they actually needed was better retrieval.
The Real Advantage Is the Orchestration Layer
The valuable “company AI” of the future is probably not a company-trained LLM at all. It's an orchestration layer around a very capable general model — one with access to the entire repository, architectural documentation, git history, issue trackers, CI/CD, test suites, static analysis, vulnerability scanners, dependency databases, and the company's accumulated engineering knowledge. The model itself stays general-purpose. The competitive advantage comes from what it's allowed to access and do, which is a far harder thing for an outside competitor to copy than a model checkpoint, and it can be extended continuously: add a repository, update the documentation, add a scanner, add a test suite. Nothing needs to be retrained.
The strongest agent architectures also close the loop between generation and execution rather than just predicting the next token: the AI writes code, the compiler complains, the AI fixes it, the tests fail, the AI investigates, the scanner finds a vulnerability, the AI changes the implementation, a human reviews the result. The objective was never to generate code that looks like the existing code. It's to generate code that works, and that distinction is what turns software development into an environment where the AI gets feedback from reality rather than from its own training data.
So, Are Private Coding LLMs Worth It?
Sometimes. But the question tends to get asked in the wrong order. Before spending millions on GPUs and standing up an internal AI research team, work out what problem is actually being solved. Want an AI that understands your code? Start with retrieval and an agent that can inspect the repository. Can't let source code leave your infrastructure? Look at privately deploying an open-weight model. Have millions of repetitive coding tasks and genuinely enormous inference volume? Self-hosting may pay for itself. Need a very specific internal format or behaviour enforced consistently? Fine-tuning earns its place.
Want to build a model better at software engineering than the companies that build frontier models for a living?
Then the first question should probably be: why? Because that is no longer a software-development project. It is an AI research company.
The future of private AI coding is unlikely to be thousands of companies independently training miniature versions of Claude, GPT, or Gemini on their own repositories. It's far more likely to be frontier or open-weight models, plus private data, plus retrieval, plus tools, tests, and agentic workflows, plus human oversight. The model provides the general intelligence. The company's infrastructure provides the context. The test suite provides the reality check. That is a considerably more powerful proposition than feeding yesterday's code into tomorrow's model.
Sources
- Wang et al., “RAG or Fine-tuning? A Comparative Study on LCMs-based Code Completion in Industry,” FSE 2025
- CCBench coding-agent leaderboard, February 2026 results
- Petryshyn & Lukoševičius, “Optimizing Large Language Models for OpenAPI Code Completion,” 2024
- GitHub & Accenture, randomised controlled trial on Copilot adoption
- OpenAI Codex — Enterprise admin documentation
- NVIDIA H100 Tensor Core GPU datasheet
- Latif et al., “Single-Node Power Demand During AI Training: Measurements on an 8-GPU NVIDIA H100 System,” IEEE Access, 2025
- IEA, “Key Questions on Energy and AI,” April 2026 update