CGU AI Gateway V7

LLM API Guide

This page explains how to call the CGU LLM Gateway with OpenAI-compatible APIs. With one Base URL and one API key you can use text generation, streaming, embeddings, image generation and editing, speech-to-text, speech translation, text-to-speech, Realtime WebSocket and usage queries.

Never put your API key in public front-end code, HTML, Git, screenshots or any file users can download. The examples below use the placeholder YOUR_CGU_API_KEY; in production, read the key from an environment variable on your back end and make the call there.

Quick start

Base URLhttps://air.cgu.edu.tw/cgullmapi/v1
AuthorizationAuthorization: Bearer YOUR_CGU_API_KEY
JSON requestsText, embeddings, image generation and TTS mostly use Content-Type: application/json.
File uploadsAudio uploads and image edits use multipart/form-data (-F in the curl examples).
WebSocketRealtime uses wss://air.cgu.edu.tw/cgullmapi/v1/..., not a regular https:// POST.
Model listSee the model catalog below for the available models and what they are for. Call /v1/models to see exactly which models your account can use.
curl https://air.cgu.edu.tw/cgullmapi/v1/models \
  -H "Authorization: Bearer YOUR_CGU_API_KEY"

Model catalog

These models are currently available to everyone. Put the model name from the tables into the model field of your request; /v1/models shows exactly which ones your account can use.

Models use two separate quotas: the school compute pool (Spark) and local models use your “local model token quota” and cost no USD; OpenAI models use your “OpenAI cost quota” (USD). When your OpenAI quota runs out, you can still use the school compute pool models. Check your remaining quota on the API Key page.

School compute pool (Spark) models | local model token quota

The university's own GPU servers. They support /v1/chat/completions and /v1/responses, and you can also use them for coding with opencode.

ModelGood forInput
gpt-oss-120bGeneral Q&A, reasoning and analysis, coding; the default model for opencode.Text
gemma-4-26b-a4b-itGeneral Q&A, summarizing and translation; long documents; understanding images (screenshots, charts, photos).Text, images
mistral-small-4-119b-2603Coding, multi-step work that calls tools (agents), organizing data; long documents; understanding images.Text, images
nemotron-3-super-120b-a12bProblems that need more thinking: reasoning, math, debugging code; long documents.Text

Text example

curl https://air.cgu.edu.tw/cgullmapi/v1/chat/completions \
  -H "Authorization: Bearer YOUR_CGU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-oss-120b",
    "messages": [
      {"role": "user", "content": "Introduce Chang Gung University in one sentence."}
    ]
  }'

Image example (Python)

import base64
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_CGU_API_KEY",
    base_url="https://air.cgu.edu.tw/cgullmapi/v1",
)

with open("photo.png", "rb") as f:
    image = base64.b64encode(f.read()).decode()

response = client.chat.completions.create(
    model="gemma-4-26b-a4b-it",
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "Describe this image."},
        {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image}"}},
    ]}],
)
print(response.choices[0].message.content)

OpenAI models | OpenAI cost quota (USD)

UseModelsEndpoint
Text, reasoning, Codexgpt-6-luna (recommended; half the price of gpt-5.6-luna), gpt-5.6-luna/v1/responses, /v1/chat/completions
Embeddingstext-embedding-3-small, text-embedding-3-large/v1/embeddings
Image generation and editinggpt-image-2, gpt-image-1.5, gpt-image-1, gpt-image-1-mini, chatgpt-image-latest/v1/images/generations, /v1/images/edits
Speech-to-text, speech translationwhisper-1/v1/audio/transcriptions, /v1/audio/translations
Text-to-speechgpt-4o-mini-tts, tts-1, tts-1-hd/v1/audio/speech
Realtime voice (WebSocket)gpt-realtime-2.1, gpt-realtime-2.1-mini, gpt-realtime-2, gpt-realtime-mini, gpt-realtime-translate, gpt-realtime-whisperSee Realtime below

Other, more expensive OpenAI models are currently locked; see the API Key page for the list.

Local models (Ollama) | local model token quota

ModelUseEndpointNotes
gpt-oss:20bText chat/v1/chat/completionsSimple Q&A and short text tasks.
bge-m3:latestEmbeddings (multilingual)/v1/embeddings1024-dimensional vectors.
translategemma:latestTranslation/v1/chat/completionsPut the text to translate in the message, e.g. “Translate into English: 長庚大學”.

If a model has been idle for a while, the first call has to load it first and takes longer. For images, use the school compute pool models gemma-4-26b-a4b-it or mistral-small-4-119b-2603 above.

Which API should I use?

Text, tools, Codex or agents

POST /v1/responses

Recommended for new projects. Good for Q&A, summaries, reasoning, tool calling, code generation, streaming and multi-step workflows.

Legacy chat compatibility

POST /v1/chat/completions

For older SDKs, Open WebUI, existing systems and third-party tools that only support Chat Completions.

Search, RAG, similarity

POST /v1/embeddings

Turns text into vectors for database search, document comparison, recommendations and pre-processing for classification.

Image generation and editing

POST /v1/images/generationsPOST /v1/images/edits

Use generations to create an image from text; use edits to modify, redraw or restyle an uploaded image.

Audio

POST /v1/audio/transcriptionsPOST /v1/audio/translationsPOST /v1/audio/speech

transcriptions turns speech into text in the original language; translations turns speech into English text; speech turns text into an audio file.

Realtime voice and low-latency interaction

WS /v1/realtimeWS /v1/realtime/translationsWS /v1/realtime/transcription_sessions

Requires WebSocket. Good for realtime voice assistants, live translation and live transcription.

Responses API

Recommended for new projects, Codex and agent workflows. The Responses API takes the user's input in input; with the SDK, the reply text is in response.output_text.

When to useQ&A, summaries, coding help, data extraction, streaming and tool-based apps.
Required fieldsmodel, input
Common fieldsmax_output_tokens limits the output length; stream: true turns on streaming.
NoteNot every model supports both Responses and Chat Completions; check /v1/models and test.
curl https://air.cgu.edu.tw/cgullmapi/v1/responses \
  -H "Authorization: Bearer YOUR_CGU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-luna",
    "input": "Introduce the CGU AI Center in one sentence.",
    "max_output_tokens": 128
  }'
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_CGU_API_KEY",
    base_url="https://air.cgu.edu.tw/cgullmapi/v1",
)

response = client.responses.create(
    model="gpt-5.6-luna",
    input="List three key points for testing an API.",
    max_output_tokens=160,
)
print(response.output_text)

Streaming

Streaming is good for showing text as it is generated, for long answers, or for cutting the perceived wait. When testing with curl, add -N so curl doesn't buffer the output.

curl -N https://air.cgu.edu.tw/cgullmapi/v1/responses \
  -H "Authorization: Bearer YOUR_CGU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-luna",
    "input": "Explain what a reverse proxy is for, in three points.",
    "stream": true,
    "max_output_tokens": 180
  }'

Chat Completions (compatibility)

Older SDKs, Open WebUI and tools that only support Chat Completions can use this route. It takes a messages array; the usual roles are system, user and assistant.

When to useExisting OpenAI chat code, Open WebUI, older LangChain flows and tools that only accept the messages format.
Required fieldsmodel, messages
Common fieldsmax_completion_tokens, temperature, stream
NoteCodex-type models and newer agent models may only support /v1/responses, not the chat endpoint.
curl https://air.cgu.edu.tw/cgullmapi/v1/chat/completions \
  -H "Authorization: Bearer YOUR_CGU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-luna",
    "messages": [
      {"role": "system", "content": "You are a concise teaching assistant."},
      {"role": "user", "content": "Reply OK and confirm the API works."}
    ],
    "max_completion_tokens": 128
  }'

Embeddings

Embeddings turn text into numeric vectors. The vectors aren't meant to be read; databases and programs use them for similarity search, RAG document retrieval, duplicate detection, classification and recommendations.

When to useIndex documents in chunks, find the passages relevant to a question, semantic search or similarity comparison.
Required fieldsmodel, input
Input formatA single string or an array of strings. Send large amounts of data in batches.
Common modelstext-embedding-3-small, text-embedding-3-large (OpenAI cost quota); local bge-m3:latest (local model token quota).
curl https://air.cgu.edu.tw/cgullmapi/v1/embeddings \
  -H "Authorization: Bearer YOUR_CGU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "text-embedding-3-small",
    "input": [
      "CGU LLM Gateway test text",
      "This text will be turned into a vector"
    ]
  }'

Images: generation and editing

There are two image functions: /images/generations creates an image from a text prompt; /images/edits modifies, redraws, restyles or extends an uploaded image. chatgpt-image-latest, gpt-image-1, gpt-image-1-mini, gpt-image-1.5 and gpt-image-2 have all passed generation and editing tests; /v1/models shows which ones your account can use.

Generate an image from text

Give only a text prompt and the model creates a new image. Good for icons, illustrations, poster drafts and concept art.

curl https://air.cgu.edu.tw/cgullmapi/v1/images/generations \
  -H "Authorization: Bearer YOUR_CGU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "chatgpt-image-latest",
    "prompt": "A clean blue square icon with CGU letters on a white background",
    "size": "1024x1024",
    "quality": "low",
    "n": 1
  }'

Edit an image

Upload an image and describe the change in the prompt. Use multipart; the upload field must be named image.

curl https://air.cgu.edu.tw/cgullmapi/v1/images/edits \
  -H "Authorization: Bearer YOUR_CGU_API_KEY" \
  -F model=chatgpt-image-latest \
  -F image=@input.png \
  -F prompt="Keep the main composition and turn it into a clean blue-and-white tech-style icon" \
  -F size=1024x1024
Image APIs usually return JSON containing the image data or an image URL. Decode or download it on your back end before showing it on the front end.

Audio: speech-to-text, translation, text-to-speech

Speech-to-text (STT)

Use /audio/transcriptions: upload an audio file and get a transcript in the original language. Good for meeting notes, interviews, lecture recordings and voice input.

curl https://air.cgu.edu.tw/cgullmapi/v1/audio/transcriptions \
  -H "Authorization: Bearer YOUR_CGU_API_KEY" \
  -F model=whisper-1 \
  -F language=en \
  -F response_format=json \
  -F file=@meeting.wav

Speech translation

Use /audio/translations: upload non-English audio and get English text back. To keep the original language, use transcriptions.

curl https://air.cgu.edu.tw/cgullmapi/v1/audio/translations \
  -H "Authorization: Bearer YOUR_CGU_API_KEY" \
  -F model=whisper-1 \
  -F response_format=json \
  -F file=@speech.wav

Text-to-speech (TTS)

Use /audio/speech: send text and get the audio bytes back. Call it from your back end and pass the audio to the front end to play or download.

curl https://air.cgu.edu.tw/cgullmapi/v1/audio/speech \
  -H "Authorization: Bearer YOUR_CGU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini-tts",
    "voice": "alloy",
    "input": "Today is a great day to finish your API integration.",
    "instructions": "Read in a natural, clear tone.",
    "response_format": "mp3"
  }' \
  --output speech.mp3

Realtime WebSocket

Realtime needs WebSocket and is good for low-latency voice, live conversation, live translation or live transcription. If you go through an external reverse proxy, make sure it forwards Upgrade: websocket, Connection: Upgrade and the related WebSocket headers.

/realtimeGeneral realtime multimodal conversation, e.g. gpt-realtime-2.
/realtime/translationsLive translation workflows, e.g. gpt-realtime-translate.
/realtime/transcription_sessionsLive transcription workflows, e.g. gpt-realtime-whisper.
Testing tipsHTTP 101 Switching Protocols means the WebSocket handshake succeeded; a full session test should also wait for server events and send audio or text events.
const ws = new WebSocket(
  "wss://air.cgu.edu.tw/cgullmapi/v1/realtime?model=gpt-realtime-2",
  [],
);

ws.addEventListener("open", () => {
  ws.send(JSON.stringify({
    type: "session.update",
    session: { instructions: "Reply in English." }
  }));
});

ws.addEventListener("message", (event) => {
  console.log(JSON.parse(event.data));
});
A browser's built-in WebSocket can't easily add an Authorization header. In production, open the Realtime connection from your back end, or have your back end issue a short-lived token or session for the front end.

Usage and billing

/v1/me/usage returns the local and OpenAI usage, token quota, OpenAI USD quota and remaining amount for the account bound to your API key. Check it before and after heavy testing to make sure your requests are counted correctly.

curl https://air.cgu.edu.tw/cgullmapi/v1/me/usage \
  -H "Authorization: Bearer YOUR_CGU_API_KEY"

Troubleshooting

401The API key is missing, not sent as a Bearer token, or doesn't exist.
403The model isn't allowed, streaming is disabled, or your token / USD quota is used up.
404The route isn't supported, or the model isn't available on this endpoint.
413The request is too large or over the per-request token limit.
415 / 422Usually a multipart field problem, e.g. the audio field isn't named file, or the image-edit field isn't named image.
426A Realtime endpoint was called over plain HTTP; use WS/WSS.
Model doesn't support the endpointA model doesn't necessarily support every endpoint. For example, Codex-type models usually use Responses, image models use Images, and audio models use Audio.
Brotli / brThe gateway sends Accept-Encoding: identity upstream and removes Content-Encoding and Content-Length from responses, so clients never receive already-decompressed content still labeled br.

Codex guide

Keep your CGU API Key safe. It usually starts with sk-. Don't publish it, share screenshots of it or paste it on public websites.

1. Quit the Codex app completely

If you have used Codex before, sign out of your account in the Codex app first, then close the app.

Windows: also checkLook in the notification area at the bottom right of the taskbar for a Codex icon. If it's there, right-click it and choose Exit so Codex quits completely.

2. Activate and copy your CGU API Key

Open the CGU API Key pageOpen the link below to activate and copy your CGU API Key:
https://air.cgu.edu.tw/workspace4/LLMAPI/index.php
API Key formatThe API Key usually starts with sk-, for example sk-xxxxxxxxxxxxxxxxxxxxxxxx

3. Find the Codex settings folder

Windows pathReplace YourUserName with your Windows user name.
C:\Users\YourUserName\.codex\
macOS pathReplace YourUserName with your macOS user name. In Finder, press Command + Shift + G and enter the path.
/Users/YourUserName/.codex/

4. Edit config.toml

In the Codex settings folder, find or create config.toml and add the following:

model_provider = "custom"
model = "gpt-5.6-luna"
model_reasoning_effort = "medium"

[model_providers.custom]
name = "CGU API"
base_url = "https://air.cgu.edu.tw/cgullmapi/v1"
requires_openai_auth = true
wire_api = "responses"

Keeping your existing project settings

If config.toml already has a [projects] section, replace only the content above [projects] and keep your existing project settings below it. If you don't need your old settings, you can replace the whole config.toml with the content above.

5. Restart the Codex app

6. Enter your CGU API Key when prompted

opencode guide (school compute pool models)

opencode is an open-source AI coding assistant that runs in your terminal: it reads code, edits files and runs commands for you. Connected to this service, it can use the school compute pool (Spark) models, which only count against your local model token quota and cost no USD.

GitHubhttps://github.com/anomalyco/opencode
Website and docshttps://opencode.ai/docs
opencode reads and writes the files in your project and runs commands. Start it inside the project folder you want it to work on, and back up with Git before important changes. Don't use /share: it uploads the conversation as a public link.

1. Install opencode

WindowsInstall Node.js (LTS) first, then run npm install -g opencode-ai in PowerShell or Command Prompt. You can also use choco install opencode or scoop install opencode.
macOSbrew install anomalyco/tap/opencode, or curl -fsSL https://opencode.ai/install | bash
Linuxcurl -fsSL https://opencode.ai/install | bash

2. Create opencode.json in your project folder

In your project folder (the one you want opencode to work on), create a file named opencode.json with the content below. It contains no key, so it's safe to commit to Git.

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "cgu": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "CGU API",
      "options": {
        "baseURL": "https://air.cgu.edu.tw/cgullmapi/v1"
      },
      "models": {
        "gpt-oss-120b": { "name": "gpt-oss-120b (CGU compute pool)", "limit": { "context": 131072, "output": 32768 } },
        "gemma-4-26b-a4b-it": { "name": "gemma-4-26b-a4b-it (CGU compute pool)", "limit": { "context": 262144, "output": 32768 } },
        "mistral-small-4-119b-2603": { "name": "mistral-small-4-119b-2603 (CGU compute pool)", "limit": { "context": 262144, "output": 32768 } },
        "nemotron-3-super-120b-a12b": { "name": "nemotron-3-super-120b-a12b (CGU compute pool)", "limit": { "context": 262144, "output": 32768 } }
      }
    }
  },
  "model": "cgu/gpt-oss-120b"
}

To use this setup in every project on macOS / Linux, you can also put it in ~/.config/opencode/opencode.json.

3. Enter your CGU API Key (one time only)

Step 1Open a terminal in your project folder and run opencode.
Step 2Type /connect, scroll to the bottom of the provider list and choose Other.
Step 3Enter cgu as the provider ID (it must match "cgu" in opencode.json).
Step 4Paste your CGU API Key (it starts with sk-). The key is stored in opencode's own settings, not in opencode.json.

4. Start using it

Just describe what you want, e.g. “Show me the structure of this project” or “How do I fix this error?”. The default model is gpt-oss-120b; type /models to switch. The first time in a project, you can type /init to have it write a project overview (AGENTS.md). If something goes wrong, /undo reverts the last step.

gpt-oss-120bRecommended default: good for reading, writing and debugging code in general.
gemma-4-26b-a4b-itLarge projects where it needs to read many files at once; it can also read screenshots.
mistral-small-4-119b-2603Changes that take several steps, e.g. editing a few files and then running the tests; it can also read screenshots.
nemotron-3-super-120b-a12bHarder debugging, algorithms or logic problems.
Small local models (such as gpt-oss:20b) don't yet support the tool calling opencode needs (reading files, running commands); please use the models above. To code with OpenAI models, use Codex (above).