Generative AI has entered the enterprise, but one question holds back serious projects in Switzerland: where does the data go? Sending internal documents, client files or proprietary code to an API hosted in the United States is rarely acceptable for regulated sectors. The answer is two words: sovereign AI, meaning running models on GPUs in Switzerland with data that never leaves.
Why sovereignty changes everything for AI
An AI project handles, by nature, what a company holds most sensitive: its documents, its conversations, its customer data. Yet most hyperscaler GPU offerings do not guarantee data residency for AI workloads, and remain subject to extraterritorial laws such as the US CLOUD Act. For a bank, an insurer, a healthcare player or a public body in Switzerland, that is a dealbreaker.
Running the model on sovereign infrastructure solves the problem at the root: context data, model weights and prompts stay physically in Switzerland, on locally operated infrastructure.
What Hikube provides
Hikube offers GPU as a Service with NVIDIA H100, A100 and L40S cards, available two ways:
- from dedicated VMs, for a direct GPU-server experience;
- from Kubernetes pods via the NVIDIA device plugin, to industrialize training and serving at scale.
Around the GPU, the platform brings the bricks a "GPU only" offer often lacks: S3-compatible object storage for datasets and checkpoints, managed databases, and observability. All open source, replicated across three Swiss datacenters.
Three concrete use cases
Inference plus RAG. The most common case: connect an open-source LLM to your internal documents (Retrieval Augmented Generation) to answer business questions without exposing the corpus to a third party. Fast to set up, high impact.
Fine-tuning. Specialize a model on your domain, vocabulary or tone. Useful when generic inference is not enough, more GPU-intensive, best reserved for cases where the gain is measurable.
Training and scientific computing. For R&D teams training their own models or running intensive compute, with H100 and A100 sized to the workload.
A telling example from French-speaking Switzerland: Novel-T uses the platform to launch complex training runs and validate models at scale, with no infrastructure constraint and without data leaving Switzerland.
Where to start
There is no need to aim for training a giant model on day one. The pragmatic path starts with an inference plus RAG use case on a limited scope, measures the value, then expands. GPU sizing (which card, how many, VM or Kubernetes) is decided based on the model and target throughput, not the other way around.
If you have an AI project stuck on the data-sovereignty question, let us talk: our engineers in Geneva frame the use case, size the GPU and operate the platform. Explore Hikube or talk to an engineer.

Cloud Product Manager
Cloud Product Manager at Hidora. Specialist in Kubernetes, IaC and observability.



