Skip to content

AI

Our LLMaaS Gateway (Large Language Models as a Service) provides high-performance access to a curated selection of current open-weight language models. Inference runs entirely on our Swiss-hosted GPU infrastructure — your prompts, embeddings, and generated responses never leave Switzerland.

Available Models

The following models are currently available in production via the gateway. The list is fetched live from the API.

*Loading models…*

All models are addressed using the provider/model format (e.g. ew/glm-5.3-flash), so switching models is typically a one-line change. More top-tier models are in the evaluation phase and will be added soon.

OpenAI-Compatible API

The gateway exposes an OpenAI-compatible REST interface — existing code using the OpenAI SDK (Python, Node, Go, …) can be pointed at our endpoints with no changes to application logic:

  • POST /v1/chat/completions — chat and reasoning requests, including streaming and tool calling
  • POST /v1/embeddings — vector embeddings for RAG, semantic search, classification
  • POST /v1/rerank — re-ranking of search results for higher hit quality
  • GET /v1/models — list of all currently available models

→ Full interface specification in the API Reference.

→ Per-request options — response caching and content-logging control — are described in Request Options.

Systemone-Compatible API

As an alternative to the OpenAI-compatible interface, the gateway supports the Systemone API (POST /v1/systemone) with the model ew/dex. Instead of generating free text, you submit a state together with any number of typed questions and receive exactly one typed answer per question — including probabilities:

  • noul — yes/no question, answered with the probability that the answer is "yes"
  • choice — one of up to 255 named options, answered with the winning option and per-option probabilities
  • score — rating on an ordered scale of up to 10 levels, answered with the expected level and its distribution

Answers are computed by scoring the candidate answer labels directly on the model — no text is generated, nothing has to be parsed, and the response is a single JSON document (streaming is not supported; output_tokens is always 0). This makes the API well suited wherever a fixed decision matters more than prose:

  • Ticket routing & triage — assign incoming requests to billing, technical, or sales and decide urgency with noul
  • Sentiment & escalation scoring — grade customer messages on a defined scale before handing them to a person
  • Classification with thresholds — act directly on the returned probabilities, e.g. allow/block with a confidence cutoff
  • Guardrails for agents — cheap yes/no checks that gate tool calls or escalations

The API follows the published System One specification — existing System One clients and SDKs (e.g. the TypeSafe SDK) work by pointing their base URL at the gateway and setting the model to ew/dex.

→ Full interface specification in the API Reference.

Virtual Keys & Governance

The gateway supports virtual keys (prefix sk-bf-...) for fine-grained access control, model routing, and per-team / per-project / per-use-case usage tracking. You can create and manage your virtual keys yourself in the Cloud Services Portal under AI.

→ See AI Management for creating AI Projects and generating API Keys.

Typical Use Cases

  • RAG pipelines — document search with embeddings + rerank, context-aware answer generation
  • Code assistance — internal developer tooling, code review, and refactoring suggestions
  • Classification & extraction — structured data extraction from emails, reports, tickets
  • Agents & automation — tool-calling-enabled workflows with controlled write access
  • Multilingual content — translation and localisation with a focus on German-speaking markets

Limited Access

Selected customers can generate their API key directly in the Cloud Services Portal under AI. Access to this feature is granted gradually — if your portal does not show the AI section yet, your account is not yet enabled.