Cloud & platform engineering · AI in production

From notebook to production.

Versioning, deployment, monitoring and retraining — so a model doesn't get stuck as a prototype. Fully within your own infrastructure if you want.

WHY MODELS DIE AS PROTOTYPES

A model in a notebook is not a product. It lacks what every other piece of software takes for granted: versioning, a path to production, monitoring and someone accountable. This is exactly where most AI initiatives fail — not on the model, but on operations.

The lifecycle a model needs

Four steps that form a circle. Only once the circle closes does a model stay useful over time.

1

Version

Model, training data, hyperparameters and code are versioned together. Every result is reproducible — in a year's time and by someone else.

2

Deploy

Models go to production through the same pipeline as your applications: tested, approved, reversible. As an API, a batch job or embedded directly.

3

Monitor

Not just latency and error rate, but the quality of predictions too. Drift is detected before anyone in the business notices results getting worse.

4

Retrain

New data flows into a new model that competes against the old one. Only the better one goes live — automated, without a gut call.

Private AI: models at your site, data in-house

If your data must not leave the building, the AI runs at your site instead. Technically that is no longer a compromise.

dns

Models in your infrastructure

Open models run on your GPUs — in your own data centre or in the Swiss cloud. No prompt, no document and no customer record leaves your environment.

lock

Role-based access

The model only sees what the person asking is allowed to see. Permissions from your directory apply to the AI as well — not retroactively.

receipt_long

Traceable & auditable

Every request, answer and source logged. You can prove afterwards how a decision came about — a requirement in regulated areas.

payments

Predictable cost

Your own hardware instead of per-token billing. With steady load that is predictable — and beyond a certain volume clearly cheaper.

What we actually do

inventory_2

Model registry & pipelines

One central registry for all models with state, metrics and approval status. Training and deployment pipelines as code, not manual work.

speed

Inference that scales

Model serving on Kubernetes with GPU sharing and autoscaling. Expensive cards stay utilised instead of belonging to individual teams.

rule

Evaluation & guardrails

Automated tests for quality, hallucinations and unwanted output — on every model change, before it goes live.

WHAT THIS ACTUALLY DELIVERS
Reproducible

instead of "it worked back then" — every result can be rebuilt.

In-house

your data stays, if that is what you need — without giving up AI.

None

model lock-in: open formats and standards, you can switch provider.

Technologies

Chosen vendor-neutrally, to fit the problem.

MLflowKubeflowKServevLLMOllamaOpenShift AIDVCRay

Related

UW
YOUR DIRECT CONTACT
Uwe Winnwa
Sales · cloud37 AG

30-min intro call — free.
No standard pitch. I listen, ask the right questions and give an honest assessment — even if we are not the right partner.