Skip to content
Cloud
Blog
Cloud8 min

Hosting an LLM in Switzerland: sovereign AI on GPU

Mattia Eleuteri11 August 2026

Generative AI has entered the enterprise, but one question holds back serious projects in Switzerland: where does the data go? Sending internal documents, client files or proprietary code to an API hosted in the United States is rarely acceptable for regulated sectors. The answer is two words: sovereign AI, meaning running models on GPUs in Switzerland with data that never leaves.

Why sovereignty changes everything for AI

An AI project handles, by nature, what a company holds most sensitive: its documents, its conversations, its customer data. Yet most hyperscaler GPU offerings do not guarantee data residency for AI workloads, and remain subject to extraterritorial laws such as the US CLOUD Act. For a bank, an insurer, a healthcare player or a public body in Switzerland, that is a dealbreaker.

Running the model on sovereign infrastructure solves the problem at the root: context data, model weights and prompts stay physically in Switzerland, on locally operated infrastructure.

What Hikube provides

Hikube offers GPU as a Service with NVIDIA H100, A100 and L40S cards, available two ways:

  • from dedicated VMs, for a direct GPU-server experience;
  • from Kubernetes pods via the NVIDIA device plugin, to industrialize training and serving at scale.

Around the GPU, the platform brings the bricks a "GPU only" offer often lacks: S3-compatible object storage for datasets and checkpoints, managed databases, and observability. All open source, replicated across three Swiss datacenters.

Three concrete use cases

Inference plus RAG. The most common case: connect an open-source LLM to your internal documents (Retrieval Augmented Generation) to answer business questions without exposing the corpus to a third party. Fast to set up, high impact.

Fine-tuning. Specialize a model on your domain, vocabulary or tone. Useful when generic inference is not enough, more GPU-intensive, best reserved for cases where the gain is measurable.

Training and scientific computing. For R&D teams training their own models or running intensive compute, with H100 and A100 sized to the workload.

A telling example from French-speaking Switzerland: Novel-T uses the platform to launch complex training runs and validate models at scale, with no infrastructure constraint and without data leaving Switzerland.

Where to start

There is no need to aim for training a giant model on day one. The pragmatic path starts with an inference plus RAG use case on a limited scope, measures the value, then expands. GPU sizing (which card, how many, VM or Kubernetes) is decided based on the model and target throughput, not the other way around.

If you have an AI project stuck on the data-sovereignty question, let us talk: our engineers in Geneva frame the use case, size the GPU and operate the platform. Explore Hikube or talk to an engineer.

Mattia Eleuteri

Cloud Product Manager

Cloud Product Manager at Hidora. Specialist in Kubernetes, IaC and observability.

CKACKSElastic Certified

Does this article resonate?

Hidora can support you on this topic.

Need support?

Let's talk about your project. 30 minutes, no strings attached.