Vianordis
--:--:-- UTCVianordis / ED. 02 / 2026
§ 07 — App support

Inference

Run open AI models on your own dedicated GPUs in Europe.

Inference gives organizations a private vLLM endpoint for chat, speech-to-text, and text-to-speech on EU-hosted GPU capacity. It speaks the OpenAI API, supports mTLS access, integrates with the GSC AI Hub, and keeps prompts, documents, and audio inside European infrastructure.

01Problem

Most private AI still depends on someone else's model cloud.

Teams want assistants, RAG, voice agents, document intelligence, and workflow automation, but many inference paths still send prompts, documents, audio, and embeddings to public hyperscale APIs. Self-hosting avoids that dependency, but GPU procurement, drivers, runtimes, model swaps, monitoring, and upgrades quickly become operational work.

  • Public AI APIs can move sensitive prompts, documents, or audio outside the jurisdiction the organization is trying to protect.
  • Shared token-bucket capacity makes latency, throughput, and model availability difficult to explain to business owners.
  • Self-managed GPU servers require driver maintenance, CUDA compatibility, runtime tuning, model caching, monitoring, backup, and incident response.
  • Existing applications often depend on OpenAI-compatible SDKs, so switching providers can become a code migration instead of an endpoint change.
  • Voice workloads add more moving parts when chat, speech-to-text, and text-to-speech are operated as separate services.
02Solution

Inference provides a managed EU GPU endpoint instead of another public model account.

Inference runs customer-scoped vLLM and voice services on dedicated European GPU capacity. Applications connect through an OpenAI-compatible API, while Vianordis handles runtime operation, model deployment, GPU monitoring, backups, and integration with the wider platform.

  • Small-tier production capacity runs on a dedicated 20 GB GPU for 8B-14B-class quantized models and shared STT/TTS workloads.
  • Medium and Cluster tiers are scoped for larger models, higher throughput, and multi-GPU workloads that exceed single 20 GB capacity.
  • The API pattern lets existing OpenAI SDKs, agents, AI Hub routes, and internal products move to private inference with minimal integration change.
  • Access follows the platform's zero-trust posture through mTLS, internal gateway controls, monitoring, and region-aware deployment.
03Benefits

Get the control of self-hosting without taking over the GPU stack.

Inference is designed for teams that need private AI capacity, predictable operations, and a straightforward path from prototype to production.

Keep inference in Europe

Run prompts, documents, and audio on EU infrastructure rather than sending them to third-party model providers outside the European environment.

Use familiar APIs

Connect existing OpenAI-compatible SDKs, agent frameworks, and application integrations to a private endpoint instead of rewriting the application layer.

Reserve real capacity

Dedicated GPU tiers give workloads a clearer capacity model than public shared inference pools.

Serve chat and voice together

Run LLM chat, faster-whisper speech-to-text, and Kokoro text-to-speech behind the same private inference offering.

Swap open models deliberately

Use validated open-weight models such as Ministral, Mistral, Llama, Qwen, Mixtral, DeepSeek, or private fine-tunes according to tier capacity.

Reduce operations burden

Vianordis handles GPU drivers, vLLM runtime operation, model cache, monitoring, backup hooks, and lifecycle management.

04How it works

From model requirement to private endpoint in three steps.

  1. 01

    Scope the workload

    Choose the model family, context needs, voice requirements, throughput target, region, latency expectations, and isolation posture.

  2. 02

    Provision the endpoint

    Vianordis prepares the GPU tier, deploys the vLLM and voice runtime, configures access, connects monitoring, and exposes an OpenAI-compatible endpoint.

  3. 03

    Integrate and operate

    Applications call the private endpoint directly or through AI Hub, while runtime health, GPU usage, model behavior, and scaling needs are monitored during production use.

05Features

What Inference provides.

OpenAI-compatible endpoint

Expose familiar routes such as chat completions, model listing, and speech-oriented endpoints so existing clients can move with minimal code change.

Dedicated vLLM runtime

Serve open-weight chat models through vLLM with model names that can remain stable even when the backing model is swapped.

Voice AI services

Support speech-to-text with faster-whisper and text-to-speech with Kokoro for private voice agents and audio workflows.

Three capacity tiers

Small starts with 20 GB VRAM, Medium targets 96 GB single-GPU workloads, and Cluster targets multi-GPU or 96 GB+ models.

EU region placement

Offer deployment across European locations such as Helsinki, Falkenstein, and Nuremberg, with region rollout matched to capacity availability.

Managed operations

GPU monitoring, runtime supervision, model storage, backup hooks, lifecycle management, and operational handover are part of the service.

06Architecture

Architecture and operating model.

Inference sits behind the GSC AI Hub and internal gateway pattern so customer applications get a familiar API while the GPU runtime remains isolated and managed.

  • Customer applications, SaaS products, or agents call an OpenAI-compatible API for chat, speech-to-text, text-to-speech, and model discovery.
  • AI Hub can route tenant workloads to the private inference endpoint with platform-level policy, gateway controls, and integration consistency.
  • Internal gateway and mesh controls isolate traffic to the GPU runtime rather than exposing raw node services directly.
  • vLLM serves chat models under a stable model name so the backing model can be changed without forcing downstream configuration changes.
  • Voice services run alongside chat inference for STT and TTS workloads that need the same EU-resident posture.
  • Small tier is backed by a 20 GB NVIDIA RTX 4000 SFF Ada class node; Medium and Cluster capacity are scoped for larger GPU memory and parallel serving.
  • Monitoring tracks GPU utilization, memory, temperature, power, service health, and operational state for support and capacity planning.
07Use cases

Where Inference fits.

Private enterprise assistant

Run a company assistant, Bicameral-style agent, or internal copilot against EU-resident model capacity.

RAG over controlled data

Keep document retrieval, prompt construction, and answer generation inside the same sovereign infrastructure boundary.

Voice agent

Combine speech-to-text, LLM reasoning, and text-to-speech for a private voice workflow without public speech APIs.

Document and ticket automation

Classify, summarize, extract, translate, and route support tickets, records, or documents through dedicated model capacity.

Product AI backend

Power Vianordis products or customer applications that need predictable private inference instead of shared public model calls.

Private fine-tunes and large models

Scope Medium or Cluster tiers for larger open models, high-QPS production workloads, or private model variants.

08Integrations

Connects model runtime to the Vianordis platform.

Inference is designed to be consumed directly by applications or indirectly through the AI Hub so model serving remains a platform capability rather than a one-off server.

GSC AI Hub

Shows how this service fits into the wider Vianordis environment instead of standing alone.

OpenAI-compatible API

Lets approved systems connect without ad hoc exports or manual copy-paste.

vLLM

Powers AI features while keeping model execution inside the controlled environment.

faster-whisper

Shows how this service fits into the wider Vianordis environment instead of standing alone.

Kokoro TTS

Shows how this service fits into the wider Vianordis environment instead of standing alone.

mTLS

Shows how this service fits into the wider Vianordis environment instead of standing alone.

Internal gateway

Shows how this service fits into the wider Vianordis environment instead of standing alone.

Istio authorization

Shows how this service fits into the wider Vianordis environment instead of standing alone.

Envoy

Shows how this service fits into the wider Vianordis environment instead of standing alone.

Zabbix monitoring

Shows how this service fits into the wider Vianordis environment instead of standing alone.

Katello lifecycle

Shows how this service fits into the wider Vianordis environment instead of standing alone.

Podman quadlets

Shows how this service fits into the wider Vianordis environment instead of standing alone.

Model cache

Powers AI features while keeping model execution inside the controlled environment.

RAG pipelines

Shows how this service fits into the wider Vianordis environment instead of standing alone.

Agent frameworks

Shows how this service fits into the wider Vianordis environment instead of standing alone.

Vianordis apps

Shows how this service fits into the wider Vianordis environment instead of standing alone.

09Trust

Data residency and operational trust.

Inference is built for organizations that need a clear answer to where model execution happens, who operates the runtime, how access is controlled, and what capacity is actually available.

EU residency

Inference capacity is designed for European GPU locations, with no customer prompt, document, or audio data sent to non-EU model providers.

Dedicated capacity

Tiers are scoped as customer-facing GPU capacity rather than a generic public shared model pool.

Access control

Customer-facing access follows the platform posture with mTLS and gateway isolation rather than raw node exposure.

Status clarity

Small-tier production capacity is live; Medium, Cluster, Helsinki, and Nuremberg expansion capacity are offered and scoped during onboarding.

Runtime visibility

GPU utilization, memory, temperature, power, service health, and operational lifecycle are monitored.

Model governance

Open models can be selected, validated, swapped, or fine-tuned according to the workload and tier capacity.

10 — Pricing

Start with Small from EUR 549 per month or scope larger GPU capacity.

Small provides dedicated 20 GB GPU capacity for 8B-14B-class models, assistants, RAG, GenUI, and voice workloads. Medium targets 96 GB single-GPU workloads, while Cluster is scoped for multi-GPU, high-throughput, or very large open models.

11FAQ

Questions technical buyers usually ask.

Is Inference a public shared model API?

No. Inference is sold as dedicated GPU-backed capacity with a private OpenAI-compatible endpoint, not as a generic public token pool.

Which tier is live today?

The Small tier is live on 20 GB GPU capacity. Medium, Cluster, and additional regional GPU capacity are offered and scoped during onboarding.

Can existing OpenAI SDKs connect to it?

Yes. The service exposes an OpenAI-compatible API so existing clients, agent frameworks, and applications can usually move by changing endpoint and credentials.

Does it support voice?

Yes. Inference includes chat plus speech-to-text and text-to-speech service patterns for private voice agents and audio workflows.

What models can run on Small?

Small is designed for 8B-14B-class quantized models and voice workloads. Larger models or higher concurrency should be scoped against Medium or Cluster capacity.

Does customer data leave the EU?

No. The product is designed so prompts, documents, and audio are processed on European infrastructure rather than sent to third-party model providers outside the EU.

12 — Next step

Move model execution from public APIs to dedicated European GPUs.

See how Inference can give your agents, applications, RAG systems, and voice workflows a private OpenAI-compatible endpoint on managed EU GPU capacity.