ONEHUNDRED

Free whitepaper

AI Engine under your own control.

Inference and model hosting in an environment you control, and why the bottleneck is not the model. Eight pages by Andreas Hankel, CTO onehundred, for IT decision-makers, CTOs and cloud leads.

Andreas Hankel, CTO onehundred

Andreas Hankel

CTO onehundred · Former CTO idealo · CIO of the Year 2014

„With an AI engine, choosing the model is the easiest decision in the project. The bottleneck is almost always operations: GPU capacity, data pipelines, an operating model for continuous running, not the model itself.“

Andreas Hankel, Autor des Whitepapers
  • Former CTO idealo: Nearly nine years, 2016 to 2025.
  • CIO of the Year 2014: Mid-market category, awarded by IDG, CIO-Magazin and Computerwoche. As VP Technology at ImmobilienScout24 he brought server operations back in-house, built a private cloud and switched the virtualisation platform while in live operation.
  • Over 30 years in IT: Since 1990, including 18 years at Fiducia IT, then ImmobilienScout24 and idealo.

„What has been created is exemplary and shows the way for designing the systems of the future.“

Manfred Broy, jury CIO of the Year 2014

Request the whitepaper

You will receive the PDF by email at your business address.

Why now

In your inbox within minutes. With the nine questions you will know today whether your bottleneck is sovereignty, utilisation or continuous operation.

The situation

42 to over 80 % never reach production: why the bottleneck is operations, not the model.

The legal position is the clearer part of the decision, operations the harder one. The figures behind that are documented, not asserted.

42 to over 80 %

of AI initiatives, depending on the study and definition, never reach production. The cause lies almost consistently in the environment before and after the model, GPU capacity planning, data pipelines and an operating model for continuous running, not the model itself.

Quelle: S&P Global, IDC/Lenovo, RAND Corporation

2 December 2027

From this date the EU AI Act's high-risk obligations apply to standalone AI systems, to systems embedded in products from August 2028, following the Digital Omnibus agreement of June 2026.

Quelle: European Commission, Council of the EU

12 months

The horizon over which the whitepaper's self-check asks about expected request volume, the basis of the make-or-buy calculation of request volume, run time and evenness of load.

Quelle: onehundred, AI Engine

What's inside

Does your own inference pay off, or does the API stay cheaper?

The whitepaper covers three propositions: why sovereignty decides where inference runs, why the bottleneck is operations rather than the model, and why make or buy is a calculation about run time. Each proposition comes with three questions to place your own position within minutes.

01

Not every request may leave through an external API

When external model APIs are used, every request leaves your own environment, including the data contained in the prompt. For a portion of use cases that is unproblematic. For requests involving personal data, trade secrets or regulated content it is a question that has to be settled before the pilot, not after it.

02

Choosing the model is the easiest decision in the project

The studies on failed AI initiatives are inconsistent in magnitude, a 42 % abandonment rate according to S&P Global, 88 % of proofs of concept never rolled out according to IDC/Lenovo, more than 80 % failed projects according to RAND, but they largely agree on the cause: not the model, but GPU capacity planning, data pipelines and an operating model for continuous running.

03

Make or buy is a question of run time

External APIs are billed per request and usable without lead time, unbeatable at low or fluctuating volume. Your own inference infrastructure has high fixed costs and low marginal costs; it pays off above a certain constant level of utilisation. The decision is therefore not a matter of principle but a calculation across three quantities: expected request volume, the run time over which it will be used, and how evenly the load is distributed, supplemented by the cases in which the sovereignty requirement overrides the calculation.

Further chapters

  • Technical detail: three interchangeable building blocks
  • Benefits and economics
  • What happens if you do nothing
  • What onehundred takes on here
  • Frequently asked questions
  • Conclusion: where do you stand?
  • Sources

What onehundred takes on

Consulting, transition, operation.

onehundred builds and operates AI infrastructure for inference and model hosting in an environment you control or a sovereign environment.

Consulting

Technical feasibility and sovereignty assessment, make-or-buy calculation, capacity planning

Transition

Building the inference and hosting infrastructure, including GPU orchestration

Operation

Managed operation of the AI engine with defined availability and capacity control

Two limits stated openly. AI strategy on the business side, which use cases are worthwhile, does not sit with onehundred but with your business functions or with specialist consultancies. And: the team is confident it can build this, but has not yet demonstrated it in a reference project in this form. We say so in the first conversation rather than talking around it.

Questions

Frequently asked questions.

Is AI Engine a finished product we can simply order?
No, and we say so deliberately. AI Engine is an offering under construction with a sound technical basis but without a completed reference project. Every engagement therefore begins with a feasibility assessment rather than an order.
Does AI Engine cover all our AI consulting needs?
No. The largest demand in the AI field is advice and the integration of existing tools, AI Engine is not intended for that. It specifically addresses running your own, self-controlled inference infrastructure.
Do we have to replace our existing AI tools?
Not necessarily. In practice a hybrid model is often sensible: uncritical use cases continue through standard tools, sensitive or customer-related processing runs through your own AI engine. Which split makes sense is settled in the consulting phase.
Does its own AI infrastructure pay off for an organisation of our size?
That cannot be answered in general terms and depends on data volume, the sensitivity of the use cases and customer requirements. At low volume an API-based solution may remain cheaper. That calculation is exactly what we deliver transparently in the feasibility assessment.

Conclusion

Where do you stand?

The AI engine is an operations decision, not a model decision. Sovereignty determines where inference may run, utilisation and run time determine whether running it yourself pays off, and continuous operation determines whether anything reaches production at all. The nine questions in the whitepaper settle those three points in the order in which they also arise in the project.

  1. Where inference may run: the sovereignty question
  2. Whether running it yourself pays off: utilisation and run time
  3. Whether the application reaches production: continuous operation
Get the whitepaper

Why now

Running the feasibility assessment now costs comparatively little and creates the ability to act, regardless of whether an AI engine of your own is built in the end.