Skip to content
Myra AI Workspace Administration Manual Updated · 01 Sep 2026

Technical description

This section describes the technical structure of Myra AI Workspace, the processing of a request, the interfaces, and the error handling.

Functional overview

Myra AI Workspace is a multi-tenant reverse proxy between the customer's applications and the programming interfaces of the AI providers. The structure knows two levels. A tenant is the top-level account — usually a company, an application, or a team — and a gateway is a named installation within a tenant, for example production or staging. Every gateway carries its own provider keys, its own authentication tokens, its own rate limits, and its own routing rules, fully separated from every other gateway.

Every inference request carries the tenant and the gateway in the URL path, either as /v1/{tenant}/{gateway}/compat/chat/completions for the unified endpoint or as /v1/{tenant}/{gateway}/{provider}/chat/completions for the native path of a provider. The gateway resolves both slugs per request and loads the matching gateway configuration. A change to the configuration takes effect within seconds.

Each request passes three phases. Each phase either stops and answers the request immediately, or it hands the request on to the next step:

  • ■ Access phase: Runs before the body of the request is read, which makes a rejection inexpensive. The access phase covers the authentication against the stored token hash, the sliding window rate limit of the gateway and of the token, the quota check against the budgets of the token, the tenant, and the gateway, and the IP allowlist.
  • ■ Content phase: Runs with the complete body available. The content phase covers the cache lookup, the two-tier guardrail chain on the outgoing request, the evaluation of the routing rules, the call to the provider, the guardrail chain on the incoming response, the cost calculation, and the cache store.
  • ■ Logging phase: Runs after the response has been sent, so an error here no longer affects the caller. The logging phase writes the structured log record and the metrics.

The routing rules are evaluated in their configured order, and the first matching rule wins. A rule overrides the provider, the model, and the fallback chain. If no rule matches, the provider and the model of the request itself apply:

  • ■ On a native endpoint both come from the URL path.
  • ■ On the unified endpoint the gateway derives the provider from the field model over the prefix of the model name, falling back to OpenRouter.

If the provider answers with an error of class 5xx, the gateway retries the call up to retry_count times — the default value 2 yields three attempts in total — and then works through the fallback chain. Responses of class 4xx are returned to the caller immediately and without a retry.

System architecture

The gateway runs as a single OpenResty process — nginx with LuaJIT — that carries the complete policy chain inside the worker itself. There is no sidecar and no additional network hop for authentication, rate limiting, caching, or the tier 1 guardrails. The product consists of the following components:

  • ■ Gateway process (nginx/OpenResty, Lua): The multi-tenant reverse proxy. The reverse proxy hooks into the three nginx phases access, content, and log, and runs an ordered middleware chain in which every step either answers the request immediately or enriches a per-request context object for the next step. One process serves many tenants and gateways with hard separation between them.
  • ■ Web interface (React, TypeScript, Vite): The surface for daily work and for administration. The navigation bar has two states. In the workspace, the navigation bar carries the Workspace group with the Chat, Projects, Workflows, Agents, and Playground entries, and the Overview group with the Dashboard and Recent chats entries. Through the user menu, the navigation bar changes to one of the four Settings, User Management, Account, and My access areas. The navigation bar then shows only the entries of that area and the Back to workspace entry. Which entries appear depends on the permissions of the account. The navigation bar only hides entries. Every view and every interface checks the permission again on the server.
  • ■ Configuration and log database (MySQL/MariaDB): Holds the durable configuration — tenants, gateways, provider keys, authentication tokens, routing rules, users, roles, and model prices — and the structured request log. The configuration of a gateway is read from the database and held in shared memory for the time of config_cache_ttl, so a request needs no database read of its own.
  • ■ Shared memory (nginx shared dictionaries): The hot state of the process — response cache, sliding window counters of the rate limit, gateway configuration, decrypted provider keys with a lifetime of 60 seconds, and the counters for the metrics.
  • ■ myrapod (LiteLLM proxy in front of the Myra GPU cluster): Serves every request to the provider myra and hosts the tier 2 services — the Presidio analyzer and anonymizer for the detection of personal data, the Llama Guard 3 classification for Prompt Guard, and the document conversion. Every model operated by Myra is reached exclusively over this path.
  • ■ Blob store: A file store per owner on the host, into which attachments, knowledge files, uploads, and generated images are moved out of the database.
  • ■ Supporting services: An SMTP relay for the one-time codes of the login, the push services for the mobile clients, a search interface for the web search, an OpenTelemetry collector for the tracing, and the SIEM target of the customer.
  • ■ Upstream providers: The connected AI providers, called over HTTPS during the content phase.

The production environment is reachable at the following addresses:

Address Meaning
ai.myra.eu Web interface
ai-api.myra.eu Inference interface under /v1/
ai-api-admin.myra.eu Administration interface under /admin/
ai-docs.myra.eu Documentation

Technical details

Interfaces. The product exposes two separate interfaces. The inference interface under /v1/ carries the requests to the AI providers and authenticates with an opaque token per gateway. The administration interface under /admin/v1/ manages tenants, gateways, users, tokens, provider keys, routing rules, model prices, statistics, logs, and guardrail events, and authenticates with a browser session held in the signed JWT cookie aig_admin.

Authentication. The gateway accepts a token in three headers and takes the first one present:

  • ■ x-aig-token
  • ■ Authorization: Bearer <token>
  • ■ x-api-key for compatibility with the SDK of Anthropic

Through this order the gateway takes the place of the interface of OpenAI without any change to an existing client. A token is hashed with SHA-256 before it is stored, is valid for exactly one gateway, and carries an optional expiry, an own rate limit, an own budget, and a binding to an account. The role of that account decides what the token may do.

Roles. The platform knows the roles admin, tenant_admin, ki_manager, member, finance, viewer, and demouser. The role is enforced during authentication, before the body of the request is read. Deleting a user invalidates every authentication token of that user immediately.

Login. The web interface authenticates without a password over a six-digit one-time code sent by e-mail. After five failed attempts per address the check answers with 429. A tenant additionally sets up OpenID Connect or SAML 2.0. Both methods sign in existing accounts only and never create one. A temporary fault never invalidates a session and never leads to a login — the check always fails to the safe side with 503.

Request headers. Payload logging is switched off per gateway with log_payloads: false and per request with x-aig-collect-log-payload: false. The header x-aig-collect-log: false drops the log record entirely. The header x-aig-byok-alias selects a specific provider key. On an unknown alias the gateway never falls back to the default key.

Response headers. A cached response carries X-AIG-Cache: HIT. An error response carries the persistent code in X-AIG-Error. Every response carries X-RateLimit-Limit and X-RateLimit-Remaining. A rejection additionally carries Retry-After.

Data separation. The URL prefix {tenant_slug}/{gateway_slug} is resolved per request, and an unknown slug is answered with 404 tenant_not_found. The whole internal state is named separately per tenant and gateway, all data is strictly assigned to the owning tenant on the storage layer, and rate limits and budgets are kept per gateway and never shared.

Streaming. With "stream": true the gateway forwards the chunks of the response as they arrive. If an error occurs after the stream has been opened, the HTTP status can no longer be changed. The response carries the status 200 on the wire, and the error arrives as an event in the stream. A streaming client therefore evaluates both paths — the HTTP status before the first event and the field error in the events after it. The log records the actual status, not the 200 of the wire.

Cryptography and key handling. A provider key is encrypted with AES-256-CBC and stored as a pair of initialization vector and ciphertext. The key is decrypted on first use and held in shared memory for 60 seconds, then injected into the call to the provider. The master key comes from the environment of the process. The session of the web interface is a stateless JSON web token with HMAC-SHA256, held in the cookie aig_admin with the attributes HttpOnly and SameSite=Strict and a lifetime of eight hours by default.

Rate limit and budget. The rate limit is a sliding window over two counters. The counter of the previous window is weighted by the share of the current window that has not yet elapsed and added to the counter of the current window. Costs are held as micro-dollars, checked before the request against the limit of the token, of the tenant, and of the gateway, and increased after the response of the provider.

Roles and permissions. A role is a set of permissions, and every gate resolves against a permission instead of a role name. The seven roles named above are held as system roles maintained by Myra. A tenant additionally defines its own roles from the same set of permissions. A role of a tenant never carries a permission of the platform class, which the data model itself enforces. Single sign-on assigns the default role member and never sets the platform role. Groups of the identity provider are mapped onto internal groups over an allowlist per tenant.

Response cache. Beside the exact match over a checksum of provider, model, and the canonical body of the request, a gateway optionally switches on a semantic cache. On a miss the prompt is embedded and compared with the stored embeddings of the same gateway and the same model. Above the configured similarity — 0.95 by default — the stored response is returned with the headers X-AIG-Cache: SEMANTIC_HIT and X-AIG-Similarity. Responses transferred as a stream are not cached semantically.

Error handling

Every error response of the gateway carries an HTTP status and a persistent error code. The code names the cause precisely and does not change when the wording of the message is revised. A calling application therefore evaluates error.code and never the text error.message. The code arrives both in the JSON body and in the header X-AIG-Error. The web interface shows a translated message for every code. The errors fall into four groups.

Limits and billing

Error Cause and remedy
429 rate_limited The sliding window limit is exhausted. Wait for the period stated in Retry-After and raise the limit of the gateway or of the token.
429 quota_exceeded A budget of the token, the tenant, or the gateway is reached. Raise the budget or reset the counter.
402 subscription_inactive The subscription is not active. Reactivate the subscription or store a valid payment method.
402 trial_expired The trial period has expired. Reactivate the subscription or store a valid payment method.

Access and security

Error Cause and remedy
401 unauthorized The token is missing, unknown, expired, revoked, or scoped to a different gateway. The wording of the message distinguishes the three cases.
403 forbidden The source address is not in the IP allowlist, or the role does not permit the action.
403 data_residency_blocked The selected model would run over a provider or a region outside the EU. Select an EU model or a gateway without the armed floor.
400 guardrail_blocked A guardrail rejected the content of the response. Rephrase the content, or adjust the pattern or the action of the guardrail in case of a false positive. A block in the request phase carries a different status, see the section Shape of a blocked response.
403 processing_restricted Processing is restricted for the account, for example after a request under Article 18 GDPR.

Request and model

Error Cause and remedy
400 invalid_request The body is not valid JSON, or a mandatory field is missing.
400 context_length_exceeded The prompt exceeds the context window. Shorten the input, switch on context compaction, or move to a model with a larger context window.
413 context_overflow The prompt exceeds the context window. The response carries estimated, limit, suggested_model, and can_compact_now.
413 request_too_large No provider accepted the request, usually because of oversized attachments.
400 model_capability_mismatch The selected model does not support the requested tools.
400 web_search_not_supported The selected model does not support the web search.
404 tenant_not_found The tenant or the gateway in the path does not exist.

Provider and operations

Error Cause and remedy
502 provider_error The provider answered with a server error after every retry.
502 all_providers_failed Every provider of the routing chain failed. Check the request log, the status pages of the providers, and the keys of every provider in the chain.
424 provider_key_missing No key is stored for the resolved provider on this gateway.
504 request_timeout The request did not finish within the time budget. Simplify the request or switch off the web search.
503 service_starting The gateway is starting up. Repeat the request after the time stated in Retry-After.
500 configuration_error To be resolved by Myra Security. Support needs the request ID and the gateway slug from the request log.
500 internal_error To be resolved by Myra Security. Support needs the request ID and the gateway slug from the request log.

Circuit breaker. If the circuit breaker is switched on and a provider reaches the configured failure threshold within the configured window, the gateway skips that provider and moves to the next entry of the chain. After the cool-down the circuit breaker allows one probe request. Only an attempt that actually returns a response closes it. The state per provider is read with GET /admin/v1/gateways/{id}/circuit-breaker.

Shape of a blocked response

A guardrail blocks in two phases, and the response differs per phase:

  • ■ Request phase: The gateway returns a synthetic response in the wire format of the request, with the status 200 and the refusal text in place of the model answer. The calling application therefore receives a validly shaped response instead of an error. With "stream": true, the block arrives as the event aig_status: "guardrail_blocked" in the stream.
  • ■ Response phase: A call through the inference interface receives the error 400 guardrail_blocked.

A calling application therefore evaluates both paths.

Guardrail service unreachable

If a tier 2 service cannot be reached, the setting Fail-open of the affected guardrail decides the outcome:

  • ■ Switched on — the default — the request passes unchecked.
  • ■ Switched off, the request is blocked.

The request log shows whether a tier 2 verdict is missing on a request that passed.

Use cases

A public sector body and a regulated enterprise face the same conflict. The most capable models — OpenAI, Anthropic, Google Gemini, AWS Bedrock — are operated by companies from the United States, are subject to the law there, and stand in data centers there. Sending confidential information to those providers directly raises questions under the GDPR and under the requirements of the respective industry:

  • ■ Who accesses that information?
  • ■ Under which law?
  • ■ Under whose supervision?

At the same time, teams that are denied AI start using consumer services on their own, which is exactly the outflow the organization wanted to prevent.

Myra AI Workspace places the certified EU infrastructure of Myra Security between the organization and every one of those providers. All requests and responses run through the network of Myra, and a provider from the United States only ever receives what the policies of the customer explicitly permit. The services that inspect the content and detect personal data run inside that certified infrastructure, so confidential content is checked, masked, or replaced by placeholders inside the EU.

Three typical scenarios illustrate the use of the product:

  • ■ The legal department of an insurer works with contracts that contain personal data. The department uses a project of the access level PII protection required. Personal data is replaced by reversible placeholders before the request leaves the EU and restored in the answer. The department budget is capped at 2,000 USD per month.
  • ■ A development team integrates a customer-facing chatbot. The application authenticates with its own token against the gateway production and never sees a provider key. The token carries a budget of 50 USD per month and its own rate limit. A fallback chain and the circuit breaker keep the chatbot answering when a provider fails.
  • ■ The compliance officer answers the question of an auditor after a security incident: what did the model receive, and what did it answer? The request log holds the identity, the routing, the token counts, the cost, the latency per phase, the verdict of every guardrail, and the residency zone of every segment, sealed against change by a chain of checksums.