llmd provider sends OpenAI-compatible requests to an
llm-d Router and adds
trusted scheduling metadata for the Endpoint Picker (EPP). Use this provider
when GoModel is the authenticated API gateway in front of llm-d.
Prerequisites
- A running llm-d deployment with an HTTPRoute that reaches the Router/EPP.
- The model name exposed by that route.
- Network access from GoModel to the Router service.
Configure with environment variables
LLMD_BASE_URL is required and should include /v1. LLMD_API_KEY is
optional; set it when the Gateway in front of llm-d requires bearer
authentication.
When LLMD_INFERENCE_OBJECTIVE is omitted, GoModel does not send an inference
objective header. Set it when the llm-d deployment defines multiple inference
objectives and this GoModel provider should select one of them.
Declare LLMD_MODELS when your route does not serve GET /v1/models. GoModel
uses that list as the provider’s model inventory.
Configure with YAML
fairness_from_user_path defaults to true. GoModel derives the fairness ID
from the effective user_path, after authentication
and key policy have been applied. Set it to false if another trusted layer
sets fairness metadata.
A raw
X-GoModel-User-Path supplied by a client is not enough to set the
fairness ID. Bind the path to a managed auth key or an extension identity so
GoModel can treat it as an effective authenticated path.x-llm-d-*
headers and the deprecated x-gateway-* aliases with the same trusted values.
This supports stable llm-d 0.8 deployments while using the canonical names
recognized by llm-d 0.9 and later.
When EPP flow control returns a 429 on a translated inference route, GoModel
preserves the x-llm-d-request-dropped-reason response header so clients can
distinguish rejection from post-dispatch eviction.
Verify
First create a managed GoModel API key in the dashboard, bind itsuser_path
to /team/alpha, and assign the returned value to TEAM_ALPHA_GOMODEL_KEY.
The key binding, rather than a client-asserted header, establishes the fairness
ID used by this request.
llmd/ routing qualifier before sending the model
name upstream. Slash-shaped Hugging Face model IDs remain intact.
Supported routes
The provider implements chat completions, Responses, embeddings, and model listing through the OpenAI-compatible API. Passthrough is enabled by default for the other routes in the llm-d HTTP API reference, including OpenAI completions, Anthropic Messages, and vLLM Generate:Multiple llm-d routes
Use suffixed variables to register independent Router services:llmd-prod and llmd-dev. Select them with model
names such as llmd-dev/Qwen/Qwen2.5-0.5B-Instruct.