ONEHUNDRED

01 Solution · AI infrastructure

A dedicated GPU server is the easy part. You need AI infrastructure that survives operations.

We build and run AI infrastructure for inference and model hosting in an environment you control: in your own data centre or in a European sovereign data centre. The entry point is a readiness assessment, not an order form. Consulting, build and operations come from one team.

Entry point
AI readiness assessment: data infrastructure, compute capacity, sovereignty
Where it runs
On premises or in a European sovereign data centre
Models
Open models from the Llama and Mistral families, you control the versions
Operations
24/7 with on-call duty, patch and escalation management

02 Starting point

The bottleneck is operations, not the model

Depending on the study and the definition, between 42 per cent (S&P Global, "2025 Voice of the Enterprise") and more than 80 per cent (RAND Corporation) of AI initiatives never reach production. The causes rarely sit in the model.

01

The pilot ran on one machine

A proof of concept needs a data export and a GPU. Production needs orchestration, tenant separation, monitoring and recovery. That step is rarely planned as a project of its own, and afterwards nobody owns the pager when the application stops answering at three in the morning.

02

Every request leaves your organisation

With external model APIs, every request leaves your environment together with whatever sits in the prompt, and it is rarely documented what that is. Standard terms of use also allow data to be reused for model improvement: towards your own customers you have to be able to prove the opposite.

03

GPU capacity gets procured, not planned

Buying or renting a dedicated GPU server case by case leaves you without capacity planning and without a defensible make-or-buy calculation. Without an expected request volume over twelve months and without a statement about how even the load is, the investment decision is a gut feeling.

Then there is the evidence question. For high-risk systems the EU AI Act requires a conformity assessment and documented evidence. Nobody who cannot say where their inference happens can produce any of it.

03 The solution

Inference runs in your environment, not in a black box

The AI Engine is built from three deliberately interchangeable building blocks, so a new model version, a different hardware backend or a change of hosting environment does not force a rewrite. Before any of that comes the question of whether it pays off for you at all.

01

AI readiness assessment

A technical assessment of the prerequisites: data infrastructure, available compute capacity, GPU demand and sovereignty requirements. This is a platform question only, not an evaluation of your use cases.

02

Feasibility and make or buy

We quantify model size, expected load and maintenance effort, then calculate make or buy across three variables: expected request volume, how long the use case will run, and how evenly the load is distributed. The calculation is open in both directions.

03

Hosting layer

The layer for open or licensed models, typically LLMs, extended by image or embedding models. You keep control over model versions, so a vendor update does not change how your model behaves unasked.

04

Inference runtime

The runtime on GPU or specialised accelerator hardware. Sized against the model size and the load established in the feasibility step, not against a datasheet.

05

Orchestration layer

Scaling, load distribution and access control, plus tenant separation and recovery. This is the layer where pilots fail on their way into production, whether the capacity is your own or bought as GPU as a service.

06

Managed operations

Running the AI Engine with defined availability and capacity management: monitoring of latency, cost and model quality, update management for models, scaling with load and a monthly compliance and cost report.

04 How it works

Four steps, three exit points

Every engagement starts with an assessment, not with an order. At each transition there is a go or no-go decision.

Step 1

Readiness

We look at what already exists: data, compute capacity, platform, sovereignty requirements. The result is a list of what is missing for productive AI workloads.

Step 2

Feasibility and make or buy

Target architecture, capacity planning and the lifetime calculation against API usage. First go or no-go decision. "The API stays cheaper" is a permitted outcome.

Step 3

Build and pilot

Hosting layer, inference runtime and GPU orchestration are built and put under real requests while your existing tools keep running. Second go or no-go decision.

Step 4

Operations

Handover into managed operations with monitoring, capacity management and update management. Early customer projects get particularly close technical support.

Business-level AI strategy, meaning the question of which use cases are worth doing, belongs in your business units or with specialised consultancies. We are accountable for the platform and for running it.

05 Proof

What is proven: the platform underneath

The AI Engine is being built. There is no completed reference project, and we claim none. What is proven is the layer an inference environment sits on: containers, storage, network and operations that hold under load.

GeoMobile Digitalisation partner for public transport. The entire application moved from virtual machines to containers, including separated environments and training for the in-house dev team. Availability above 99.9 per cent for years, with much shorter deployment cycles. Kubernetes · Istio · Velero · CephFS · PostgreSQL
konversionsKRAFT Digital marketing agency with around 80 staff. A SaaS platform handling terabit data volumes, consolidated onto a highly available private cloud with a redundant database cluster. Proxmox · Percona · GitLab
Aerosoft Simulation software with brutal load peaks at product launches and on Black Friday. Availability permanently above 99.9 per cent, the full stack from network through virtualisation to database from one provider. Proxmox · HAProxy · Percona · Redis

06 Read on and check

Cover of the whitepaper AI Engine

Free whitepaper · 8 pages

AI Engine

By Andreas Hankel, CTO onehundred. Download in exchange for your e-mail address, no sales call.

07 To be honest

What you are right to ask at this point

"Is the AI Engine a finished product we can simply order?"

No. There is no configurator and no one-click replacement for cloud APIs. Every engagement starts with a feasibility and sovereignty assessment, and its outcome may well be that you should stay on the API.

"Does self-hosted inference pay off for an organisation our size?"

Possibly not. Self-hosted inference has high fixed costs and low marginal costs, so at small or strongly fluctuating volumes external model APIs still beat it more often than not. Where the sovereignty requirement overrides the cost calculation, it is the decision, and the numbers are not.

"Do we have to replace the AI tools we already use?"

No, a hybrid model is usually the better one. Uncritical use cases keep running on standard tools, while sensitive or customer-related processing moves to your own AI Engine. Which requests must not leave your environment is settled in the consulting phase.

"Will you also advise us on where AI is worth using?"

No, and that is deliberate. Business AI strategy, use case evaluation and the integration of existing SaaS tools belong to your business units or to specialised consultancies. Nor do we provide the legal assessment under the EU AI Act or the GDPR: we are accountable for an architecture that can technically produce the evidence.

08 Fasttrack analysis

Nils Hornke
Nils HornkeCEO, onehundred

Check whether it pays off before you buy hardware

The initial call commits you to nothing. You describe the use case, we name the questions that have to be answered before an investment decision and the variables that belong in the make-or-buy calculation.

15 minutes · not a sales call · book directly in the calendar

09 Frequently asked questions

Frequently asked questions about AI infrastructure and LLM hosting

Self-hosted LLM or model API?
A lifetime calculation across three variables: expected request volume, how long the use case runs, and how evenly the load is distributed. Self-hosted inference has high fixed costs and low marginal costs, while APIs need no lead time and are hard to beat at low or fluctuating volumes.
When does self-hosted LLM inference pay off?
At a growing and reasonably constant request volume over a longer period. And always when requests carry personal data, trade secrets or regulated content that must not leave your environment.
Which GPU do we need for LLM inference?
From model size, expected load and latency requirement, not from a datasheet. We quantify those three variables in the feasibility step and derive the capacity plan from them.
What do AI infrastructure companies actually deliver?
In our case: assessment, build and operation of the inference platform, meaning hosting layer, inference runtime, orchestration and managed operations. Model training, fine-tuning and business AI strategy are not part of it.
How can Llama or Mistral be hosted in line with the GDPR?
By running inference in an environment you control: on premises or in a European sovereign data centre. Prompts, uploaded reference material and generated results stay there and are not passed to third parties for model improvement.
Who runs MLOps, including monitoring and drift?
That is the operations phase: defined availability, capacity management, monitoring of latency, cost and model quality, and update management for models. Variable GPU and compute costs are not included.
When do the high-risk obligations of the EU AI Act apply?
Following the Digital Omnibus agreement in June 2026, they apply to standalone AI systems from 2 December 2027 and to systems embedded in products from August 2028. Anyone still planning against the original date of August 2026 has more time than assumed.
Is GPU as a service enough, or do we need our own hardware?
Rented GPU capacity solves the compute question and leaves the operating question open: orchestration, tenant separation, monitoring and recovery are needed either way. Which model fits follows from the make-or-buy calculation and your sovereignty requirements.