01 Solution · AI infrastructure
A dedicated GPU server is the easy part. You need AI infrastructure that survives operations.
We build and run AI infrastructure for inference and model hosting in an environment you control: in your own data centre or in a European sovereign data centre. The entry point is a readiness assessment, not an order form. Consulting, build and operations come from one team.
- Entry point
- AI readiness assessment: data infrastructure, compute capacity, sovereignty
- Where it runs
- On premises or in a European sovereign data centre
- Models
- Open models from the Llama and Mistral families, you control the versions
- Operations
- 24/7 with on-call duty, patch and escalation management
02 Starting point
The bottleneck is operations, not the model
Depending on the study and the definition, between 42 per cent (S&P Global, "2025 Voice of the Enterprise") and more than 80 per cent (RAND Corporation) of AI initiatives never reach production. The causes rarely sit in the model.
The pilot ran on one machine
A proof of concept needs a data export and a GPU. Production needs orchestration, tenant separation, monitoring and recovery. That step is rarely planned as a project of its own, and afterwards nobody owns the pager when the application stops answering at three in the morning.
Every request leaves your organisation
With external model APIs, every request leaves your environment together with whatever sits in the prompt, and it is rarely documented what that is. Standard terms of use also allow data to be reused for model improvement: towards your own customers you have to be able to prove the opposite.
GPU capacity gets procured, not planned
Buying or renting a dedicated GPU server case by case leaves you without capacity planning and without a defensible make-or-buy calculation. Without an expected request volume over twelve months and without a statement about how even the load is, the investment decision is a gut feeling.
Then there is the evidence question. For high-risk systems the EU AI Act requires a conformity assessment and documented evidence. Nobody who cannot say where their inference happens can produce any of it.
03 The solution
Inference runs in your environment, not in a black box
The AI Engine is built from three deliberately interchangeable building blocks, so a new model version, a different hardware backend or a change of hosting environment does not force a rewrite. Before any of that comes the question of whether it pays off for you at all.
AI readiness assessment
A technical assessment of the prerequisites: data infrastructure, available compute capacity, GPU demand and sovereignty requirements. This is a platform question only, not an evaluation of your use cases.
Feasibility and make or buy
We quantify model size, expected load and maintenance effort, then calculate make or buy across three variables: expected request volume, how long the use case will run, and how evenly the load is distributed. The calculation is open in both directions.
Hosting layer
The layer for open or licensed models, typically LLMs, extended by image or embedding models. You keep control over model versions, so a vendor update does not change how your model behaves unasked.
Inference runtime
The runtime on GPU or specialised accelerator hardware. Sized against the model size and the load established in the feasibility step, not against a datasheet.
Orchestration layer
Scaling, load distribution and access control, plus tenant separation and recovery. This is the layer where pilots fail on their way into production, whether the capacity is your own or bought as GPU as a service.
Managed operations
Running the AI Engine with defined availability and capacity management: monitoring of latency, cost and model quality, update management for models, scaling with load and a monthly compliance and cost report.
04 How it works
Four steps, three exit points
Every engagement starts with an assessment, not with an order. At each transition there is a go or no-go decision.
Step 1
Readiness
We look at what already exists: data, compute capacity, platform, sovereignty requirements. The result is a list of what is missing for productive AI workloads.
Step 2
Feasibility and make or buy
Target architecture, capacity planning and the lifetime calculation against API usage. First go or no-go decision. "The API stays cheaper" is a permitted outcome.
Step 3
Build and pilot
Hosting layer, inference runtime and GPU orchestration are built and put under real requests while your existing tools keep running. Second go or no-go decision.
Step 4
Operations
Handover into managed operations with monitoring, capacity management and update management. Early customer projects get particularly close technical support.
Business-level AI strategy, meaning the question of which use cases are worth doing, belongs in your business units or with specialised consultancies. We are accountable for the platform and for running it.
05 Proof
What is proven: the platform underneath
The AI Engine is being built. There is no completed reference project, and we claim none. What is proven is the layer an inference environment sits on: containers, storage, network and operations that hold under load.
06 Read on and check
Free whitepaper · 8 pages
AI Engine
By Andreas Hankel, CTO onehundred. Download in exchange for your e-mail address, no sales call.
07 To be honest
What you are right to ask at this point
"Is the AI Engine a finished product we can simply order?"
No. There is no configurator and no one-click replacement for cloud APIs. Every engagement starts with a feasibility and sovereignty assessment, and its outcome may well be that you should stay on the API.
"Does self-hosted inference pay off for an organisation our size?"
Possibly not. Self-hosted inference has high fixed costs and low marginal costs, so at small or strongly fluctuating volumes external model APIs still beat it more often than not. Where the sovereignty requirement overrides the cost calculation, it is the decision, and the numbers are not.
"Do we have to replace the AI tools we already use?"
No, a hybrid model is usually the better one. Uncritical use cases keep running on standard tools, while sensitive or customer-related processing moves to your own AI Engine. Which requests must not leave your environment is settled in the consulting phase.
"Will you also advise us on where AI is worth using?"
No, and that is deliberate. Business AI strategy, use case evaluation and the integration of existing SaaS tools belong to your business units or to specialised consultancies. Nor do we provide the legal assessment under the EU AI Act or the GDPR: we are accountable for an architecture that can technically produce the evidence.
08 Fasttrack analysis
Check whether it pays off before you buy hardware
The initial call commits you to nothing. You describe the use case, we name the questions that have to be answered before an investment decision and the variables that belong in the make-or-buy calculation.
15 minutes · not a sales call · book directly in the calendar
09 Frequently asked questions
Frequently asked questions about AI infrastructure and LLM hosting
Self-hosted LLM or model API?
When does self-hosted LLM inference pay off?
Which GPU do we need for LLM inference?
What do AI infrastructure companies actually deliver?
How can Llama or Mistral be hosted in line with the GDPR?
Who runs MLOps, including monitoring and drift?
When do the high-risk obligations of the EU AI Act apply?
Is GPU as a service enough, or do we need our own hardware?
10 Insights · AI infrastructure
Further reading
The pillar article explains what AI infrastructure covers and why sovereign AI is an infrastructure question rather than a model question.

Book a call