Skip to content

AI

Our LLMaaS Gateway (Large Language Models as a Service) provides high-performance access to a curated selection of current open-weight language models. Inference runs entirely on our Swiss-hosted GPU infrastructure — your prompts, embeddings, and generated responses never leave Switzerland.

Available Models

The following models are currently available in production via the gateway. The list is fetched live from the API.

*Loading models…*

All models are addressed using the provider/model format (e.g. ew/glm-5.2), so switching models is typically a one-line change. More top-tier models are in the evaluation phase and will be added soon.

OpenAI-Compatible API

The gateway exposes an OpenAI-compatible REST interface — existing code using the OpenAI SDK (Python, Node, Go, …) can be pointed at our endpoints with no changes to application logic:

  • POST /v1/chat/completions — chat and reasoning requests, including streaming and tool calling
  • POST /v1/embeddings — vector embeddings for RAG, semantic search, classification
  • POST /v1/rerank — re-ranking of search results for higher hit quality
  • GET /v1/models — list of all currently available models

→ Full interface specification in the API Reference.

→ Per-request options — response caching and content-logging control — are described in Request Options.

Virtual Keys & Governance

The gateway supports virtual keys (prefix sk-bf-...) for fine-grained access control, model routing, and per-team / per-project / per-use-case usage tracking. You can create and manage your virtual keys yourself in the Cloud Services Portal under AI.

→ See AI Management for creating AI Projects and generating API Keys.

Typical Use Cases

  • RAG pipelines — document search with embeddings + rerank, context-aware answer generation
  • Code assistance — internal developer tooling, code review, and refactoring suggestions
  • Classification & extraction — structured data extraction from emails, reports, tickets
  • Agents & automation — tool-calling-enabled workflows with controlled write access
  • Multilingual content — translation and localisation with a focus on German-speaking markets

Limited Access

Selected customers can generate their API key directly in the Cloud Services Portal under AI. Access to this feature is granted gradually — if your portal does not show the AI section yet, your account is not yet enabled.