> ## Documentation Index
> Fetch the complete documentation index at: https://dripart-codex-api-first-result.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Use Gemini 3.8 Flash with Comfy Router

> Call vertexai/gemini-3.8-flash through Comfy Router: endpoint, request shape and the response Router returns.

API Reference for `vertexai/gemini-3.8-flash`, served by Comfy Router from Google.

## Quick start

Create a key in [your Comfy workspace](https://platform.comfy.org/profile/api-keys) and export it as `COMFY_API_KEY`. The Python and TypeScript snippets use the Comfy SDKs (`pip install comfy-sdk`, `npm install @comfyorg/sdk`); the cURL snippet is the same call over raw HTTP.

**Model ID:** `vertexai/gemini-3.8-flash`

**Endpoint:** `POST https://api.comfy.org/v2/models/vertexai/gemini-3.8-flash`

<Tabs>
  <Tab title="Wait for the result">
    <CodeGroup>
      ```python Python theme={null}
      from comfy_sdk import Comfy

      # Reads COMFY_API_KEY from the environment.
      # The SDK automatically creates an idempotency key and reuses it for automatic retries.
      with Comfy() as client:
          result = client.models.run(
              "vertexai/gemini-3.8-flash",
              {
                  "contents": [
                      {
                          "parts": [
                              {
                                  "text": "Describe a robot learning to paint, in two sentences.",
                              },
                          ],
                          "role": "user",
                      },
                  ],
              },
          )

      print(result)
      ```

      ```typescript TypeScript theme={null}
      import { comfy } from "@comfyorg/sdk";

      // Reads COMFY_API_KEY from the environment.
      // The SDK automatically creates an idempotency key and reuses it for automatic retries.
      const { data } = await comfy.models.run("vertexai/gemini-3.8-flash", {
        contents: [
          {
            parts: [
              {
                text: "Describe a robot learning to paint, in two sentences.",
              },
            ],
            role: "user",
          },
        ],
      });

      console.log(data);
      ```

      ```bash cURL theme={null}
      curl https://api.comfy.org/v2/models/vertexai/gemini-3.8-flash \
        -H "X-API-Key: $COMFY_API_KEY" \
        -H "Idempotency-Key: $(uuidgen)" \
        -H "Content-Type: application/json" \
        -d "{\"contents\": [{\"parts\":[{\"text\":\"Describe a robot learning to paint, in two sentences.\"}],\"role\":\"user\"}]}"
      ```
    </CodeGroup>
  </Tab>

  <Tab title="Queue and collect later">
    <Note>
      Queued delivery is rolling out per workspace. Until yours is enabled, the submit route answers `403` with `X-Comfy-Error-Type: not_enabled`. Nothing about the request is wrong, and the same body works through the synchronous route in the meantime.
    </Note>

    The same body, sent to `POST https://api.comfy.org/v2/models/vertexai/gemini-3.8-flash/requests`. Router answers `201` with a `request_id` as soon as the run is admitted, and the result is collected once it is ready, from this process or another one. [Queued delivery](/development/comfy-router/queue) walks through status, cancellation and collection.

    <CodeGroup>
      ```python Python theme={null}
      from comfy_sdk import Comfy

      # Reads COMFY_API_KEY from the environment.
      # Each submit() call mints its own Idempotency-Key and reuses it for automatic retries.
      with Comfy() as client:
          handle = client.models.submit(
              "vertexai/gemini-3.8-flash",
              {
                  "contents": [
                      {
                          "parts": [
                              {
                                  "text": "Describe a robot learning to paint, in two sentences.",
                              },
                          ],
                          "role": "user",
                      },
                  ],
              },
          )
          print("request_id:", handle.request_id)  # with the model ID, all another process needs

          # Poll until the request completes, waiting the Retry-After the server names.
          for update in handle.iter_events():
              print(update.status, update.queue_position)

          # The provider's own payload, the same value models.run() returns.
          # A request that failed or was cancelled raises the typed Router error here.
          result = handle.get()

      print(result)
      ```

      ```typescript TypeScript theme={null}
      import { comfy } from "@comfyorg/sdk";

      // Reads COMFY_API_KEY from the environment.
      // Each submit() call mints its own Idempotency-Key and reuses it for automatic retries.
      const handle = await comfy.models.submit("vertexai/gemini-3.8-flash", {
        contents: [
          {
            parts: [
              {
                text: "Describe a robot learning to paint, in two sentences.",
              },
            ],
            role: "user",
          },
        ],
      });
      console.log("requestId:", handle.requestId); // with the model ID, all another process needs

      // Poll until the request completes, waiting the Retry-After the server names.
      for await (const update of handle.events()) {
        console.log(update.status, update.queuePosition);
      }

      // The same result models.run() returns. A request that failed or was cancelled rejects here.
      const result = await handle.get();

      console.log(result.data);
      ```

      ```bash cURL theme={null}
      # 1. Submit. Router answers 201 with request_id, status_url, response_url and cancel_url.
      curl https://api.comfy.org/v2/models/vertexai/gemini-3.8-flash/requests \
        -H "X-API-Key: $COMFY_API_KEY" \
        -H "Idempotency-Key: $(uuidgen)" \
        -H "Content-Type: application/json" \
        -d "{\"contents\": [{\"parts\":[{\"text\":\"Describe a robot learning to paint, in two sentences.\"}],\"role\":\"user\"}]}"

      # 2. Poll until status is COMPLETED, waiting the Retry-After seconds each response names.
      REQUEST_ID="<request_id from the 201 body>"
      curl -i https://api.comfy.org/v2/models/vertexai/gemini-3.8-flash/requests/$REQUEST_ID/status \
        -H "X-API-Key: $COMFY_API_KEY"

      # 3. Collect. 200 with the model's native output, 202 with the status body while it is still running.
      curl https://api.comfy.org/v2/models/vertexai/gemini-3.8-flash/requests/$REQUEST_ID \
        -H "X-API-Key: $COMFY_API_KEY"
      ```
    </CodeGroup>
  </Tab>
</Tabs>

## Schema

### Input

<ParamField body="contents" type="object[]" required>
  The content of the current conversation with the model. For single-turn queries, this is a single instance. For multi-turn queries, this is a repeated field that contains conversation history and the latest request.
</ParamField>

<ParamField body="contents[].parts" type="object[]" required />

<ParamField body="contents[].parts[].fileData" type="object">
  URI based data.
</ParamField>

<ParamField body="contents[].parts[].fileData.fileUri" type="string">
  URI
</ParamField>

<ParamField body="contents[].parts[].fileData.mimeType" type="string">
  The media type of the file specified in the data or fileUri fields. Acceptable values include the following. For gemini-2.0-flash-lite and gemini-2.0-flash, the maximum length of an audio file is 8.4 hours and the maximum length of a video file (without audio) is one hour. For more information, see Gemini audio and video requirements. Text files must be UTF-8 encoded. The contents of the text file count toward the token limit. There is no limit on image resolution.

  Possible values: `application/pdf`, `audio/mpeg`, `audio/mp3`, `audio/wav`, `image/png`, `image/jpeg`, `image/webp`, `text/plain`, `video/mov`, `video/mpeg`, `video/mp4`, `video/mpg`, `video/avi`, `video/wmv`, `video/mpegps`, `video/flv`, `image/heic`, `image/heif`, `audio/flac`, `video/webm`
</ParamField>

<ParamField body="contents[].parts[].inlineData" type="object">
  Inline data in raw bytes. For gemini-2.0-flash-lite and gemini-2.0-flash, you can specify up to 3000 images by using inlineData.
</ParamField>

<ParamField body="contents[].parts[].inlineData.data" type="string (byte)">
  The base64 encoding of the image, PDF, or video to include inline in the prompt. When including media inline, you must also specify the media type (mimeType) of the data. Size limit: 20MB

  Format: `byte`
</ParamField>

<ParamField body="contents[].parts[].inlineData.mimeType" type="string">
  The media type of the file specified in the data or fileUri fields. Acceptable values include the following. For gemini-2.0-flash-lite and gemini-2.0-flash, the maximum length of an audio file is 8.4 hours and the maximum length of a video file (without audio) is one hour. For more information, see Gemini audio and video requirements. Text files must be UTF-8 encoded. The contents of the text file count toward the token limit. There is no limit on image resolution.

  Possible values: `application/pdf`, `audio/mpeg`, `audio/mp3`, `audio/wav`, `image/png`, `image/jpeg`, `image/webp`, `text/plain`, `video/mov`, `video/mpeg`, `video/mp4`, `video/mpg`, `video/avi`, `video/wmv`, `video/mpegps`, `video/flv`, `image/heic`, `image/heif`, `audio/flac`, `video/webm`
</ParamField>

<ParamField body="contents[].parts[].mediaProcessing" type="string">
  How the model reads this part's video. Set "AGENTIC" to let the model decide which segments to inspect, instead of fixed-rate frame sampling. Omit for the default fixed-rate sampling. Supported on gemini-3.7-flash and newer Flash models.
</ParamField>

<ParamField body="contents[].parts[].text" type="string">
  A text prompt or code snippet.
</ParamField>

<ParamField body="contents[].parts[].thought" type="boolean">
  Indicates this part is a thinking/reasoning step from the model.
</ParamField>

<ParamField body="contents[].role" type="string">
  Possible values: `user`, `model`
</ParamField>

<ParamField body="generationConfig" type="object">
  Sampling, length and output settings for the generation. Every field is optional: the fields below that declare a `default` apply it when omitted, and the rest fall back to the model's own behaviour.
</ParamField>

<ParamField body="generationConfig.imageConfig" type="object">
  Configuration for image generation
</ParamField>

<ParamField body="generationConfig.imageConfig.aspectRatio" type="string">
  Aspect ratio for generated images
</ParamField>

<ParamField body="generationConfig.imageConfig.imageOutputOptions" type="object">
  Optional. The image output format for generated images.
</ParamField>

<ParamField body="generationConfig.imageConfig.imageOutputOptions.compressionQuality" type="integer">
  Optional. The compression quality of the output image.
</ParamField>

<ParamField body="generationConfig.imageConfig.imageOutputOptions.mimeType" type="string">
  Optional. The image format that the output should be saved as.
</ParamField>

<ParamField body="generationConfig.imageConfig.imageSize" type="string">
  Optional. Specifies the size of generated images. Supported values are 1K, 2K, 4K. If not specified, the model will use default value 1K.
</ParamField>

<ParamField body="generationConfig.maxOutputTokens" type="integer">
  Maximum number of tokens that can be generated in the response. A token is approximately 4 characters. 100 tokens correspond to roughly 60-80 words.

  Range: `16` to `65536`
</ParamField>

<ParamField body="generationConfig.responseModalities" type="`TEXT`, `IMAGE`[]" />

<ParamField body="generationConfig.seed" type="integer">
  When seed is fixed to a specific value, the model makes a best effort to provide the same response for repeated requests. Deterministic output isn't guaranteed. Also, changing the model or parameter settings, such as the temperature, can cause variations in the response even when you use the same seed value. By default, a random seed value is used. Available for the following models:, gemini-2.5-flash, gemini-2.5-pro, gemini-2.5-flash-preview-04-1, gemini-2.5-pro-preview-05-0, gemini-2.0-flash-lite-00, gemini-2.0-flash-001
</ParamField>

<ParamField body="generationConfig.stopSequences" type="string[]" />

<ParamField body="generationConfig.temperature" type="number" default="1">
  The temperature is used for sampling during response generation, which occurs when topP and topK are applied. Temperature controls the degree of randomness in token selection. Lower temperatures are good for prompts that require a less open-ended or creative response, while higher temperatures can lead to more diverse or creative results. A temperature of 0 means that the highest probability tokens are always selected. In this case, responses for a given prompt are mostly deterministic, but a small amount of variation is still possible. If the model returns a response that's too generic, too short, or the model gives a fallback response, try increasing the temperature

  Range: `0` to `2`

  Format: `float`
</ParamField>

<ParamField body="generationConfig.thinkingConfig" type="object">
  Optional. Configuration for thinking features. Thinking is a process where the model breaks down a complex task into smaller steps to generate a higher-quality response.
</ParamField>

<ParamField body="generationConfig.thinkingConfig.includeThoughts" type="boolean">
  Optional. If true, the model will include its thoughts in the response.
</ParamField>

<ParamField body="generationConfig.thinkingConfig.thinkingBudget" type="integer">
  Optional. The token budget for the model's thinking process. The model will make a best effort to stay within this budget.
</ParamField>

<ParamField body="generationConfig.thinkingConfig.thinkingLevel" type="string">
  Optional. The thinking level for the model.

  Possible values: `THINKING_LEVEL_UNSPECIFIED`, `LOW`, `MEDIUM`, `HIGH`, `MINIMAL`
</ParamField>

<ParamField body="generationConfig.topK" type="integer" default="40">
  Top-K changes how the model selects tokens for output. A top-K of 1 means the next selected token is the most probable among all tokens in the model's vocabulary. A top-K of 3 means that the next token is selected from among the 3 most probable tokens by using temperature.

  Range: `1` to `…`
</ParamField>

<ParamField body="generationConfig.topP" type="number" default="0.95">
  If specified, nucleus sampling is used.
  Top-P changes how the model selects tokens for output. Tokens are selected from the most (see top-K) to least probable until the sum of their probabilities equals the top-P value. For example, if tokens A, B, and C have a probability of 0.3, 0.2, and 0.1 and the top-P value is 0.5, then the model will select either A or B as the next token by using temperature and excludes C as a candidate.
  Specify a lower value for less random responses and a higher value for more random responses.

  Range: `0` to `1`

  Format: `float`
</ParamField>

<ParamField body="safetySettings" type="object[]">
  Per request settings for blocking unsafe content. Enforced on GenerateContentResponse.candidates.
</ParamField>

<ParamField body="safetySettings[].category" type="string" required>
  Possible values: `HARM_CATEGORY_SEXUALLY_EXPLICIT`, `HARM_CATEGORY_HATE_SPEECH`, `HARM_CATEGORY_HARASSMENT`, `HARM_CATEGORY_DANGEROUS_CONTENT`
</ParamField>

<ParamField body="safetySettings[].threshold" type="string" required>
  Possible values: `OFF`, `BLOCK_NONE`, `BLOCK_LOW_AND_ABOVE`, `BLOCK_MEDIUM_AND_ABOVE`, `BLOCK_ONLY_HIGH`
</ParamField>

<ParamField body="systemInstruction" type="object">
  Instructions for the model to steer it toward better performance. For example, "Answer as concisely as possible" or "Don't use technical terms in your response". The text strings count toward the token limit. The role field of systemInstruction is ignored and doesn't affect the performance of the model. Note: Only text should be used in parts and content in each part should be in a separate paragraph.
</ParamField>

<ParamField body="systemInstruction.parts" type="object[]" required>
  A list of ordered parts that make up a single message. Different parts may have different IANA MIME types. For limits on the inputs, such as the maximum number of tokens or the number of images, see the model specifications on the Google models page.
</ParamField>

<ParamField body="systemInstruction.parts[].text" type="string">
  A text prompt or code snippet.
</ParamField>

<ParamField body="systemInstruction.role" type="string">
  The identity of the entity that creates the message. The following values are supported: user: This indicates that the message is sent by a real person, typically a user-generated message. model: This indicates that the message is generated by the model. The model value is used to insert messages from the model into the conversation during multi-turn conversations. For non-multi-turn conversations, this field can be left blank or unset.

  Possible values: `user`, `model`
</ParamField>

<ParamField body="tools" type="object[]">
  A piece of code that enables the system to interact with external systems to perform an action, or set of actions, outside of knowledge and scope of the model. See Function calling.
</ParamField>

<ParamField body="tools[].functionDeclarations" type="object[]" />

<ParamField body="tools[].functionDeclarations[].description" type="string" />

<ParamField body="tools[].functionDeclarations[].name" type="string" required />

<ParamField body="tools[].functionDeclarations[].parameters" type="object">
  JSON schema for the function parameters
</ParamField>

<ParamField body="uploadImagesToStorage" type="boolean">
  If true, generated images will be uploaded to cloud storage and returned as signed URLs instead of inline base64 data. The URLs expire after 24 hours.
</ParamField>

<ParamField body="videoMetadata" type="object">
  For video input, the start and end offset of the video in Duration format. For example, to specify a 10 second clip starting at 1:00, set "startOffset": \{ "seconds": 60 } and "endOffset": \{ "seconds": 70 }. The metadata should only be specified while the video data is presented in inlineData or fileData.
</ParamField>

<ParamField body="videoMetadata.endOffset" type="object">
  Represents a duration offset for video timeline positions.
</ParamField>

<ParamField body="videoMetadata.endOffset.nanos" type="integer">
  Signed fractions of a second at nanosecond resolution. Negative second values with fractions must still have non-negative nanos values.

  Range: `0` to `999999999`
</ParamField>

<ParamField body="videoMetadata.endOffset.seconds" type="integer">
  Signed seconds of the span of time. Must be from -315,576,000,000 to +315,576,000,000 inclusive.

  Range: `-315576000000` to `315576000000`
</ParamField>

<ParamField body="videoMetadata.startOffset" type="object">
  Represents a duration offset for video timeline positions.
</ParamField>

<ParamField body="videoMetadata.startOffset.nanos" type="integer">
  Signed fractions of a second at nanosecond resolution. Negative second values with fractions must still have non-negative nanos values.

  Range: `0` to `999999999`
</ParamField>

<ParamField body="videoMetadata.startOffset.seconds" type="integer">
  Signed seconds of the span of time. Must be from -315,576,000,000 to +315,576,000,000 inclusive.

  Range: `-315576000000` to `315576000000`
</ParamField>

Generated from the schema Router serves at `GET /v2/models/vertexai/gemini-3.8-flash/openapi.json`, the same document it validates a call against before the request reaches the provider.

### Output

<ResponseField name="candidates" type="object[]" />

<ResponseField name="candidates[].citationMetadata" type="object" />

<ResponseField name="candidates[].citationMetadata.citations" type="object[]" />

<ResponseField name="candidates[].citationMetadata.citations[].authors" type="string[]" />

<ResponseField name="candidates[].citationMetadata.citations[].endIndex" type="integer" />

<ResponseField name="candidates[].citationMetadata.citations[].license" type="string" />

<ResponseField name="candidates[].citationMetadata.citations[].publicationDate" type="string (date)">
  Format: `date`
</ResponseField>

<ResponseField name="candidates[].citationMetadata.citations[].startIndex" type="integer" />

<ResponseField name="candidates[].citationMetadata.citations[].title" type="string" />

<ResponseField name="candidates[].citationMetadata.citations[].uri" type="string" />

<ResponseField name="candidates[].content" type="object">
  The content of the current conversation with the model. For single-turn queries, this is a single instance. For multi-turn queries, this is a repeated field that contains conversation history and the latest request.
</ResponseField>

<ResponseField name="candidates[].content.parts" type="object[]" required />

<ResponseField name="candidates[].content.parts[].fileData" type="object">
  URI based data.
</ResponseField>

<ResponseField name="candidates[].content.parts[].fileData.fileUri" type="string">
  URI
</ResponseField>

<ResponseField name="candidates[].content.parts[].fileData.mimeType" type="string">
  The media type of the file specified in the data or fileUri fields. Acceptable values include the following. For gemini-2.0-flash-lite and gemini-2.0-flash, the maximum length of an audio file is 8.4 hours and the maximum length of a video file (without audio) is one hour. For more information, see Gemini audio and video requirements. Text files must be UTF-8 encoded. The contents of the text file count toward the token limit. There is no limit on image resolution.

  Possible values: `application/pdf`, `audio/mpeg`, `audio/mp3`, `audio/wav`, `image/png`, `image/jpeg`, `image/webp`, `text/plain`, `video/mov`, `video/mpeg`, `video/mp4`, `video/mpg`, `video/avi`, `video/wmv`, `video/mpegps`, `video/flv`, `image/heic`, `image/heif`, `audio/flac`, `video/webm`
</ResponseField>

<ResponseField name="candidates[].content.parts[].inlineData" type="object">
  Inline data in raw bytes. For gemini-2.0-flash-lite and gemini-2.0-flash, you can specify up to 3000 images by using inlineData.
</ResponseField>

<ResponseField name="candidates[].content.parts[].inlineData.data" type="string (byte)">
  The base64 encoding of the image, PDF, or video to include inline in the prompt. When including media inline, you must also specify the media type (mimeType) of the data. Size limit: 20MB

  Format: `byte`
</ResponseField>

<ResponseField name="candidates[].content.parts[].inlineData.mimeType" type="string">
  The media type of the file specified in the data or fileUri fields. Acceptable values include the following. For gemini-2.0-flash-lite and gemini-2.0-flash, the maximum length of an audio file is 8.4 hours and the maximum length of a video file (without audio) is one hour. For more information, see Gemini audio and video requirements. Text files must be UTF-8 encoded. The contents of the text file count toward the token limit. There is no limit on image resolution.

  Possible values: `application/pdf`, `audio/mpeg`, `audio/mp3`, `audio/wav`, `image/png`, `image/jpeg`, `image/webp`, `text/plain`, `video/mov`, `video/mpeg`, `video/mp4`, `video/mpg`, `video/avi`, `video/wmv`, `video/mpegps`, `video/flv`, `image/heic`, `image/heif`, `audio/flac`, `video/webm`
</ResponseField>

<ResponseField name="candidates[].content.parts[].mediaProcessing" type="string">
  How the model reads this part's video. Set "AGENTIC" to let the model decide which segments to inspect, instead of fixed-rate frame sampling. Omit for the default fixed-rate sampling. Supported on gemini-3.7-flash and newer Flash models.
</ResponseField>

<ResponseField name="candidates[].content.parts[].text" type="string">
  A text prompt or code snippet.
</ResponseField>

<ResponseField name="candidates[].content.parts[].thought" type="boolean">
  Indicates this part is a thinking/reasoning step from the model.
</ResponseField>

<ResponseField name="candidates[].content.role" type="string">
  Possible values: `user`, `model`
</ResponseField>

<ResponseField name="candidates[].finishReason" type="string" />

<ResponseField name="candidates[].safetyRatings" type="object[]" />

<ResponseField name="candidates[].safetyRatings[].category" type="string">
  Possible values: `HARM_CATEGORY_SEXUALLY_EXPLICIT`, `HARM_CATEGORY_HATE_SPEECH`, `HARM_CATEGORY_HARASSMENT`, `HARM_CATEGORY_DANGEROUS_CONTENT`
</ResponseField>

<ResponseField name="candidates[].safetyRatings[].probability" type="string">
  The probability that the content violates the specified safety category

  Possible values: `NEGLIGIBLE`, `LOW`, `MEDIUM`, `HIGH`, `UNKNOWN`
</ResponseField>

<ResponseField name="createTime" type="string">
  Timestamp when the response was created.
</ResponseField>

<ResponseField name="modelVersion" type="string">
  The model version used to generate the response.
</ResponseField>

<ResponseField name="promptFeedback" type="object" />

<ResponseField name="promptFeedback.blockReason" type="string" />

<ResponseField name="promptFeedback.blockReasonMessage" type="string" />

<ResponseField name="promptFeedback.safetyRatings" type="object[]" />

<ResponseField name="promptFeedback.safetyRatings[].category" type="string">
  Possible values: `HARM_CATEGORY_SEXUALLY_EXPLICIT`, `HARM_CATEGORY_HATE_SPEECH`, `HARM_CATEGORY_HARASSMENT`, `HARM_CATEGORY_DANGEROUS_CONTENT`
</ResponseField>

<ResponseField name="promptFeedback.safetyRatings[].probability" type="string">
  The probability that the content violates the specified safety category

  Possible values: `NEGLIGIBLE`, `LOW`, `MEDIUM`, `HIGH`, `UNKNOWN`
</ResponseField>

<ResponseField name="responseId" type="string">
  Unique identifier for the response.
</ResponseField>

<ResponseField name="usageMetadata" type="object" />

<ResponseField name="usageMetadata.cachedContentTokenCount" type="integer">
  Output only. Number of tokens in the cached part in the input (the cached content).
</ResponseField>

<ResponseField name="usageMetadata.candidatesTokenCount" type="integer">
  Number of tokens in the response(s).
</ResponseField>

<ResponseField name="usageMetadata.candidatesTokensDetails" type="object[]">
  Breakdown of candidate tokens by modality.
</ResponseField>

<ResponseField name="usageMetadata.candidatesTokensDetails[].modality" type="string">
  Type of input or output content modality.

  Possible values: `MODALITY_UNSPECIFIED`, `TEXT`, `IMAGE`, `VIDEO`, `AUDIO`, `DOCUMENT`
</ResponseField>

<ResponseField name="usageMetadata.candidatesTokensDetails[].tokenCount" type="integer">
  Number of tokens for the given modality.
</ResponseField>

<ResponseField name="usageMetadata.promptTokenCount" type="integer">
  Number of tokens in the request. When cachedContent is set, this is still the total effective prompt size meaning this includes the number of tokens in the cached content.
</ResponseField>

<ResponseField name="usageMetadata.promptTokensDetails" type="object[]">
  Breakdown of prompt tokens by modality.
</ResponseField>

<ResponseField name="usageMetadata.promptTokensDetails[].modality" type="string">
  Type of input or output content modality.

  Possible values: `MODALITY_UNSPECIFIED`, `TEXT`, `IMAGE`, `VIDEO`, `AUDIO`, `DOCUMENT`
</ResponseField>

<ResponseField name="usageMetadata.promptTokensDetails[].tokenCount" type="integer">
  Number of tokens for the given modality.
</ResponseField>

<ResponseField name="usageMetadata.thoughtsTokenCount" type="integer">
  Number of tokens present in thoughts output.
</ResponseField>

<ResponseField name="usageMetadata.toolUsePromptTokenCount" type="integer">
  Number of tokens present in tool-use prompt(s).
</ResponseField>

<ResponseField name="usageMetadata.toolUsePromptTokensDetails" type="object[]">
  Breakdown of tool-use prompt tokens by modality.
</ResponseField>

<ResponseField name="usageMetadata.toolUsePromptTokensDetails[].modality" type="string">
  Type of input or output content modality.

  Possible values: `MODALITY_UNSPECIFIED`, `TEXT`, `IMAGE`, `VIDEO`, `AUDIO`, `DOCUMENT`
</ResponseField>

<ResponseField name="usageMetadata.toolUsePromptTokensDetails[].tokenCount" type="integer">
  Number of tokens for the given modality.
</ResponseField>

<ResponseField name="usageMetadata.totalTokenCount" type="integer">
  Total number of tokens (prompt + candidates).
</ResponseField>

<ResponseField name="usageMetadata.trafficType" type="string">
  Traffic type used for the request (e.g., PROVISIONED\_THROUGHPUT).
</ResponseField>

## Examples

### Input

```json theme={null}
{
  "contents": [
    {
      "parts": [
        {
          "text": "Describe a robot learning to paint, in two sentences."
        }
      ],
      "role": "user"
    }
  ]
}
```

### Output

```json theme={null}
{
  "candidates": [
    {
      "content": {
        "parts": [
          {
            "text": "A lighthouse stands at the edge of the harbour, its lamp still turning as the sun comes up."
          }
        ],
        "role": "model"
      },
      "finishReason": "STOP"
    }
  ],
  "modelVersion": "gemini-3.8-flash",
  "responseId": "0d1f2a3b-4c5d-6e7f-8a9b-0c1d2e3f4a5b",
  "usageMetadata": {
    "candidatesTokenCount": 21,
    "promptTokenCount": 12,
    "totalTokenCount": 33
  }
}
```

## Before you ship

The SDKs create an `Idempotency-Key` and reuse it for automatic retries. For manual retries, reuse the original key. Router can hold the connection for up to 10 minutes.

When a request fails, Router sends an `X-Comfy-Error-Type` response header explaining why. A `422` means Router rejected the input before calling the provider. Download generated assets promptly because [result URLs can expire](/development/comfy-router/reference#result-assets).

<CardGroup cols={3}>
  <Card title="Headers" icon="list" href="/development/comfy-router/headers">
    Authentication, idempotency, request IDs, error buckets, retry pacing, spend limits.
  </Card>

  <Card title="Using the Router API" icon="code" href="/development/comfy-router/api">
    Model discovery, validation errors, retries, and billing.
  </Card>

  <Card title="Limitations" icon="triangle-exclamation" href="/development/comfy-router/limitations">
    What Router does not do today, and what to use instead.
  </Card>
</CardGroup>
