Free whitepaper
AI Engine under your own control.
Inference and model hosting in an environment you control, and why the bottleneck is not the model. Eight pages by Andreas Hankel, CTO onehundred, for IT decision-makers, CTOs and cloud leads.
Andreas Hankel
CTO onehundred · Former CTO idealo · CIO of the Year 2014
„With an AI engine, choosing the model is the easiest decision in the project. The bottleneck is almost always operations: GPU capacity, data pipelines, an operating model for continuous running, not the model itself.“
- Former CTO idealo: Nearly nine years, 2016 to 2025.
- CIO of the Year 2014: Mid-market category, awarded by IDG, CIO-Magazin and Computerwoche. As VP Technology at ImmobilienScout24 he brought server operations back in-house, built a private cloud and switched the virtualisation platform while in live operation.
- Over 30 years in IT: Since 1990, including 18 years at Fiducia IT, then ImmobilienScout24 and idealo.
„What has been created is exemplary and shows the way for designing the systems of the future.“
Manfred Broy, jury CIO of the Year 2014







Request the whitepaper
You will receive the PDF by email at your business address.
The situation
42 to over 80 % never reach production: why the bottleneck is operations, not the model.
The legal position is the clearer part of the decision, operations the harder one. The figures behind that are documented, not asserted.
42 to over 80 %
of AI initiatives, depending on the study and definition, never reach production. The cause lies almost consistently in the environment before and after the model, GPU capacity planning, data pipelines and an operating model for continuous running, not the model itself.
Quelle: S&P Global, IDC/Lenovo, RAND Corporation
2 December 2027
From this date the EU AI Act's high-risk obligations apply to standalone AI systems, to systems embedded in products from August 2028, following the Digital Omnibus agreement of June 2026.
Quelle: European Commission, Council of the EU
12 months
The horizon over which the whitepaper's self-check asks about expected request volume, the basis of the make-or-buy calculation of request volume, run time and evenness of load.
Quelle: onehundred, AI Engine
What's inside
Does your own inference pay off, or does the API stay cheaper?
The whitepaper covers three propositions: why sovereignty decides where inference runs, why the bottleneck is operations rather than the model, and why make or buy is a calculation about run time. Each proposition comes with three questions to place your own position within minutes.
01
Not every request may leave through an external API
When external model APIs are used, every request leaves your own environment, including the data contained in the prompt. For a portion of use cases that is unproblematic. For requests involving personal data, trade secrets or regulated content it is a question that has to be settled before the pilot, not after it.
02
Choosing the model is the easiest decision in the project
The studies on failed AI initiatives are inconsistent in magnitude, a 42 % abandonment rate according to S&P Global, 88 % of proofs of concept never rolled out according to IDC/Lenovo, more than 80 % failed projects according to RAND, but they largely agree on the cause: not the model, but GPU capacity planning, data pipelines and an operating model for continuous running.
03
Make or buy is a question of run time
External APIs are billed per request and usable without lead time, unbeatable at low or fluctuating volume. Your own inference infrastructure has high fixed costs and low marginal costs; it pays off above a certain constant level of utilisation. The decision is therefore not a matter of principle but a calculation across three quantities: expected request volume, the run time over which it will be used, and how evenly the load is distributed, supplemented by the cases in which the sovereignty requirement overrides the calculation.
Further chapters
- Technical detail: three interchangeable building blocks
- Benefits and economics
- What happens if you do nothing
- What onehundred takes on here
- Frequently asked questions
- Conclusion: where do you stand?
- Sources
What onehundred takes on
Consulting, transition, operation.
onehundred builds and operates AI infrastructure for inference and model hosting in an environment you control or a sovereign environment.
Consulting
Technical feasibility and sovereignty assessment, make-or-buy calculation, capacity planning
Transition
Building the inference and hosting infrastructure, including GPU orchestration
Operation
Managed operation of the AI engine with defined availability and capacity control
Two limits stated openly. AI strategy on the business side, which use cases are worthwhile, does not sit with onehundred but with your business functions or with specialist consultancies. And: the team is confident it can build this, but has not yet demonstrated it in a reference project in this form. We say so in the first conversation rather than talking around it.
Questions
Frequently asked questions.
Is AI Engine a finished product we can simply order?
Does AI Engine cover all our AI consulting needs?
Do we have to replace our existing AI tools?
Does its own AI infrastructure pay off for an organisation of our size?
Conclusion
Where do you stand?
The AI engine is an operations decision, not a model decision. Sovereignty determines where inference may run, utilisation and run time determine whether running it yourself pays off, and continuous operation determines whether anything reaches production at all. The nine questions in the whitepaper settle those three points in the order in which they also arise in the project.
- Where inference may run: the sovereignty question
- Whether running it yourself pays off: utilisation and run time
- Whether the application reaches production: continuous operation
Why now
Running the feasibility assessment now costs comparatively little and creates the ability to act, regardless of whether an AI engine of your own is built in the end.
