LLM API Guide
This page explains how to call the CGU LLM Gateway with OpenAI-compatible APIs. With one Base URL and one API key you can use text generation, streaming, embeddings, image generation and editing, speech-to-text, speech translation, text-to-speech, Realtime WebSocket and usage queries.
YOUR_CGU_API_KEY; in production, read the key from an environment variable on your back end and make the call there.Quick start
| Base URL | https://air.cgu.edu.tw/cgullmapi/v1 |
|---|---|
| Authorization | Authorization: Bearer YOUR_CGU_API_KEY |
| JSON requests | Text, embeddings, image generation and TTS mostly use Content-Type: application/json. |
| File uploads | Audio uploads and image edits use multipart/form-data (-F in the curl examples). |
| WebSocket | Realtime uses wss://air.cgu.edu.tw/cgullmapi/v1/..., not a regular https:// POST. |
| Model list | See the model catalog below for the available models and what they are for. Call /v1/models to see exactly which models your account can use. |
curl https://air.cgu.edu.tw/cgullmapi/v1/models \
-H "Authorization: Bearer YOUR_CGU_API_KEY"Model catalog
These models are currently available to everyone. Put the model name from the tables into the model field of your request; /v1/models shows exactly which ones your account can use.
School compute pool (Spark) models | local model token quota
The university's own GPU servers. They support /v1/chat/completions and /v1/responses, and you can also use them for coding with opencode.
| Model | Good for | Input |
|---|---|---|
gpt-oss-120b | General Q&A, reasoning and analysis, coding; the default model for opencode. | Text |
gemma-4-26b-a4b-it | General Q&A, summarizing and translation; long documents; understanding images (screenshots, charts, photos). | Text, images |
mistral-small-4-119b-2603 | Coding, multi-step work that calls tools (agents), organizing data; long documents; understanding images. | Text, images |
nemotron-3-super-120b-a12b | Problems that need more thinking: reasoning, math, debugging code; long documents. | Text |
Text example
curl https://air.cgu.edu.tw/cgullmapi/v1/chat/completions \
-H "Authorization: Bearer YOUR_CGU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-oss-120b",
"messages": [
{"role": "user", "content": "Introduce Chang Gung University in one sentence."}
]
}'Image example (Python)
import base64
from openai import OpenAI
client = OpenAI(
api_key="YOUR_CGU_API_KEY",
base_url="https://air.cgu.edu.tw/cgullmapi/v1",
)
with open("photo.png", "rb") as f:
image = base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="gemma-4-26b-a4b-it",
messages=[{"role": "user", "content": [
{"type": "text", "text": "Describe this image."},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image}"}},
]}],
)
print(response.choices[0].message.content)OpenAI models | OpenAI cost quota (USD)
| Use | Models | Endpoint |
|---|---|---|
| Text, reasoning, Codex | gpt-6-luna (recommended; half the price of gpt-5.6-luna), gpt-5.6-luna | /v1/responses, /v1/chat/completions |
| Embeddings | text-embedding-3-small, text-embedding-3-large | /v1/embeddings |
| Image generation and editing | gpt-image-2, gpt-image-1.5, gpt-image-1, gpt-image-1-mini, chatgpt-image-latest | /v1/images/generations, /v1/images/edits |
| Speech-to-text, speech translation | whisper-1 | /v1/audio/transcriptions, /v1/audio/translations |
| Text-to-speech | gpt-4o-mini-tts, tts-1, tts-1-hd | /v1/audio/speech |
| Realtime voice (WebSocket) | gpt-realtime-2.1, gpt-realtime-2.1-mini, gpt-realtime-2, gpt-realtime-mini, gpt-realtime-translate, gpt-realtime-whisper | See Realtime below |
Other, more expensive OpenAI models are currently locked; see the API Key page for the list.
Local models (Ollama) | local model token quota
| Model | Use | Endpoint | Notes |
|---|---|---|---|
gpt-oss:20b | Text chat | /v1/chat/completions | Simple Q&A and short text tasks. |
bge-m3:latest | Embeddings (multilingual) | /v1/embeddings | 1024-dimensional vectors. |
translategemma:latest | Translation | /v1/chat/completions | Put the text to translate in the message, e.g. “Translate into English: 長庚大學”. |
If a model has been idle for a while, the first call has to load it first and takes longer. For images, use the school compute pool models gemma-4-26b-a4b-it or mistral-small-4-119b-2603 above.
Which API should I use?
Text, tools, Codex or agents
POST /v1/responses
Recommended for new projects. Good for Q&A, summaries, reasoning, tool calling, code generation, streaming and multi-step workflows.
Legacy chat compatibility
POST /v1/chat/completions
For older SDKs, Open WebUI, existing systems and third-party tools that only support Chat Completions.
Search, RAG, similarity
POST /v1/embeddings
Turns text into vectors for database search, document comparison, recommendations and pre-processing for classification.
Image generation and editing
POST /v1/images/generationsPOST /v1/images/edits
Use generations to create an image from text; use edits to modify, redraw or restyle an uploaded image.
Audio
POST /v1/audio/transcriptionsPOST /v1/audio/translationsPOST /v1/audio/speech
transcriptions turns speech into text in the original language; translations turns speech into English text; speech turns text into an audio file.
Realtime voice and low-latency interaction
WS /v1/realtimeWS /v1/realtime/translationsWS /v1/realtime/transcription_sessions
Requires WebSocket. Good for realtime voice assistants, live translation and live transcription.
Responses API
Recommended for new projects, Codex and agent workflows. The Responses API takes the user's input in input; with the SDK, the reply text is in response.output_text.
| When to use | Q&A, summaries, coding help, data extraction, streaming and tool-based apps. |
|---|---|
| Required fields | model, input |
| Common fields | max_output_tokens limits the output length; stream: true turns on streaming. |
| Note | Not every model supports both Responses and Chat Completions; check /v1/models and test. |
curl https://air.cgu.edu.tw/cgullmapi/v1/responses \
-H "Authorization: Bearer YOUR_CGU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-luna",
"input": "Introduce the CGU AI Center in one sentence.",
"max_output_tokens": 128
}'from openai import OpenAI
client = OpenAI(
api_key="YOUR_CGU_API_KEY",
base_url="https://air.cgu.edu.tw/cgullmapi/v1",
)
response = client.responses.create(
model="gpt-5.6-luna",
input="List three key points for testing an API.",
max_output_tokens=160,
)
print(response.output_text)Streaming
Streaming is good for showing text as it is generated, for long answers, or for cutting the perceived wait. When testing with curl, add -N so curl doesn't buffer the output.
curl -N https://air.cgu.edu.tw/cgullmapi/v1/responses \
-H "Authorization: Bearer YOUR_CGU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-luna",
"input": "Explain what a reverse proxy is for, in three points.",
"stream": true,
"max_output_tokens": 180
}'Chat Completions (compatibility)
Older SDKs, Open WebUI and tools that only support Chat Completions can use this route. It takes a messages array; the usual roles are system, user and assistant.
| When to use | Existing OpenAI chat code, Open WebUI, older LangChain flows and tools that only accept the messages format. |
|---|---|
| Required fields | model, messages |
| Common fields | max_completion_tokens, temperature, stream |
| Note | Codex-type models and newer agent models may only support /v1/responses, not the chat endpoint. |
curl https://air.cgu.edu.tw/cgullmapi/v1/chat/completions \
-H "Authorization: Bearer YOUR_CGU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-luna",
"messages": [
{"role": "system", "content": "You are a concise teaching assistant."},
{"role": "user", "content": "Reply OK and confirm the API works."}
],
"max_completion_tokens": 128
}'Embeddings
Embeddings turn text into numeric vectors. The vectors aren't meant to be read; databases and programs use them for similarity search, RAG document retrieval, duplicate detection, classification and recommendations.
| When to use | Index documents in chunks, find the passages relevant to a question, semantic search or similarity comparison. |
|---|---|
| Required fields | model, input |
| Input format | A single string or an array of strings. Send large amounts of data in batches. |
| Common models | text-embedding-3-small, text-embedding-3-large (OpenAI cost quota); local bge-m3:latest (local model token quota). |
curl https://air.cgu.edu.tw/cgullmapi/v1/embeddings \
-H "Authorization: Bearer YOUR_CGU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "text-embedding-3-small",
"input": [
"CGU LLM Gateway test text",
"This text will be turned into a vector"
]
}'Images: generation and editing
There are two image functions: /images/generations creates an image from a text prompt; /images/edits modifies, redraws, restyles or extends an uploaded image. chatgpt-image-latest, gpt-image-1, gpt-image-1-mini, gpt-image-1.5 and gpt-image-2 have all passed generation and editing tests; /v1/models shows which ones your account can use.
Generate an image from text
Give only a text prompt and the model creates a new image. Good for icons, illustrations, poster drafts and concept art.
curl https://air.cgu.edu.tw/cgullmapi/v1/images/generations \
-H "Authorization: Bearer YOUR_CGU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "chatgpt-image-latest",
"prompt": "A clean blue square icon with CGU letters on a white background",
"size": "1024x1024",
"quality": "low",
"n": 1
}'Edit an image
Upload an image and describe the change in the prompt. Use multipart; the upload field must be named image.
curl https://air.cgu.edu.tw/cgullmapi/v1/images/edits \
-H "Authorization: Bearer YOUR_CGU_API_KEY" \
-F model=chatgpt-image-latest \
-F image=@input.png \
-F prompt="Keep the main composition and turn it into a clean blue-and-white tech-style icon" \
-F size=1024x1024Audio: speech-to-text, translation, text-to-speech
Speech-to-text (STT)
Use /audio/transcriptions: upload an audio file and get a transcript in the original language. Good for meeting notes, interviews, lecture recordings and voice input.
curl https://air.cgu.edu.tw/cgullmapi/v1/audio/transcriptions \
-H "Authorization: Bearer YOUR_CGU_API_KEY" \
-F model=whisper-1 \
-F language=en \
-F response_format=json \
-F file=@meeting.wavSpeech translation
Use /audio/translations: upload non-English audio and get English text back. To keep the original language, use transcriptions.
curl https://air.cgu.edu.tw/cgullmapi/v1/audio/translations \
-H "Authorization: Bearer YOUR_CGU_API_KEY" \
-F model=whisper-1 \
-F response_format=json \
-F file=@speech.wavText-to-speech (TTS)
Use /audio/speech: send text and get the audio bytes back. Call it from your back end and pass the audio to the front end to play or download.
curl https://air.cgu.edu.tw/cgullmapi/v1/audio/speech \
-H "Authorization: Bearer YOUR_CGU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini-tts",
"voice": "alloy",
"input": "Today is a great day to finish your API integration.",
"instructions": "Read in a natural, clear tone.",
"response_format": "mp3"
}' \
--output speech.mp3Realtime WebSocket
Realtime needs WebSocket and is good for low-latency voice, live conversation, live translation or live transcription. If you go through an external reverse proxy, make sure it forwards Upgrade: websocket, Connection: Upgrade and the related WebSocket headers.
/realtime | General realtime multimodal conversation, e.g. gpt-realtime-2. |
|---|---|
/realtime/translations | Live translation workflows, e.g. gpt-realtime-translate. |
/realtime/transcription_sessions | Live transcription workflows, e.g. gpt-realtime-whisper. |
| Testing tips | HTTP 101 Switching Protocols means the WebSocket handshake succeeded; a full session test should also wait for server events and send audio or text events. |
const ws = new WebSocket(
"wss://air.cgu.edu.tw/cgullmapi/v1/realtime?model=gpt-realtime-2",
[],
);
ws.addEventListener("open", () => {
ws.send(JSON.stringify({
type: "session.update",
session: { instructions: "Reply in English." }
}));
});
ws.addEventListener("message", (event) => {
console.log(JSON.parse(event.data));
});Authorization header. In production, open the Realtime connection from your back end, or have your back end issue a short-lived token or session for the front end.Usage and billing
/v1/me/usage returns the local and OpenAI usage, token quota, OpenAI USD quota and remaining amount for the account bound to your API key. Check it before and after heavy testing to make sure your requests are counted correctly.
curl https://air.cgu.edu.tw/cgullmapi/v1/me/usage \
-H "Authorization: Bearer YOUR_CGU_API_KEY"Troubleshooting
| 401 | The API key is missing, not sent as a Bearer token, or doesn't exist. |
|---|---|
| 403 | The model isn't allowed, streaming is disabled, or your token / USD quota is used up. |
| 404 | The route isn't supported, or the model isn't available on this endpoint. |
| 413 | The request is too large or over the per-request token limit. |
| 415 / 422 | Usually a multipart field problem, e.g. the audio field isn't named file, or the image-edit field isn't named image. |
| 426 | A Realtime endpoint was called over plain HTTP; use WS/WSS. |
| Model doesn't support the endpoint | A model doesn't necessarily support every endpoint. For example, Codex-type models usually use Responses, image models use Images, and audio models use Audio. |
| Brotli / br | The gateway sends Accept-Encoding: identity upstream and removes Content-Encoding and Content-Length from responses, so clients never receive already-decompressed content still labeled br. |
Codex guide
sk-. Don't publish it, share screenshots of it or paste it on public websites.1. Quit the Codex app completely
If you have used Codex before, sign out of your account in the Codex app first, then close the app.
| Windows: also check | Look in the notification area at the bottom right of the taskbar for a Codex icon. If it's there, right-click it and choose Exit so Codex quits completely. |
|---|
2. Activate and copy your CGU API Key
| Open the CGU API Key page | Open the link below to activate and copy your CGU API Key: https://air.cgu.edu.tw/workspace4/LLMAPI/index.php |
|---|---|
| API Key format | The API Key usually starts with sk-, for example sk-xxxxxxxxxxxxxxxxxxxxxxxx |
3. Find the Codex settings folder
| Windows path | Replace YourUserName with your Windows user name.C:\Users\YourUserName\.codex\ |
|---|---|
| macOS path | Replace YourUserName with your macOS user name. In Finder, press Command + Shift + G and enter the path./Users/YourUserName/.codex/ |
4. Edit config.toml
In the Codex settings folder, find or create config.toml and add the following:
model_provider = "custom"
model = "gpt-5.6-luna"
model_reasoning_effort = "medium"
[model_providers.custom]
name = "CGU API"
base_url = "https://air.cgu.edu.tw/cgullmapi/v1"
requires_openai_auth = true
wire_api = "responses"Keeping your existing project settings
config.toml already has a [projects] section, replace only the content above [projects] and keep your existing project settings below it. If you don't need your old settings, you can replace the whole config.toml with the content above.5. Restart the Codex app
6. Enter your CGU API Key when prompted
opencode guide (school compute pool models)
opencode is an open-source AI coding assistant that runs in your terminal: it reads code, edits files and runs commands for you. Connected to this service, it can use the school compute pool (Spark) models, which only count against your local model token quota and cost no USD.
| GitHub | https://github.com/anomalyco/opencode |
|---|---|
| Website and docs | https://opencode.ai/docs |
/share: it uploads the conversation as a public link.1. Install opencode
| Windows | Install Node.js (LTS) first, then run npm install -g opencode-ai in PowerShell or Command Prompt. You can also use choco install opencode or scoop install opencode. |
|---|---|
| macOS | brew install anomalyco/tap/opencode, or curl -fsSL https://opencode.ai/install | bash |
| Linux | curl -fsSL https://opencode.ai/install | bash |
2. Create opencode.json in your project folder
In your project folder (the one you want opencode to work on), create a file named opencode.json with the content below. It contains no key, so it's safe to commit to Git.
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"cgu": {
"npm": "@ai-sdk/openai-compatible",
"name": "CGU API",
"options": {
"baseURL": "https://air.cgu.edu.tw/cgullmapi/v1"
},
"models": {
"gpt-oss-120b": { "name": "gpt-oss-120b (CGU compute pool)", "limit": { "context": 131072, "output": 32768 } },
"gemma-4-26b-a4b-it": { "name": "gemma-4-26b-a4b-it (CGU compute pool)", "limit": { "context": 262144, "output": 32768 } },
"mistral-small-4-119b-2603": { "name": "mistral-small-4-119b-2603 (CGU compute pool)", "limit": { "context": 262144, "output": 32768 } },
"nemotron-3-super-120b-a12b": { "name": "nemotron-3-super-120b-a12b (CGU compute pool)", "limit": { "context": 262144, "output": 32768 } }
}
}
},
"model": "cgu/gpt-oss-120b"
}To use this setup in every project on macOS / Linux, you can also put it in ~/.config/opencode/opencode.json.
3. Enter your CGU API Key (one time only)
| Step 1 | Open a terminal in your project folder and run opencode. |
|---|---|
| Step 2 | Type /connect, scroll to the bottom of the provider list and choose Other. |
| Step 3 | Enter cgu as the provider ID (it must match "cgu" in opencode.json). |
| Step 4 | Paste your CGU API Key (it starts with sk-). The key is stored in opencode's own settings, not in opencode.json. |
4. Start using it
Just describe what you want, e.g. “Show me the structure of this project” or “How do I fix this error?”. The default model is gpt-oss-120b; type /models to switch. The first time in a project, you can type /init to have it write a project overview (AGENTS.md). If something goes wrong, /undo reverts the last step.
gpt-oss-120b | Recommended default: good for reading, writing and debugging code in general. |
|---|---|
gemma-4-26b-a4b-it | Large projects where it needs to read many files at once; it can also read screenshots. |
mistral-small-4-119b-2603 | Changes that take several steps, e.g. editing a few files and then running the tests; it can also read screenshots. |
nemotron-3-super-120b-a12b | Harder debugging, algorithms or logic problems. |
gpt-oss:20b) don't yet support the tool calling opencode needs (reading files, running commands); please use the models above. To code with OpenAI models, use Codex (above).