Generation
Four endpoints, one request shape. Every call takes a prompt, an optional
provider and model, and a parameters object whose
contents depend on the modality.
Omit provider and the account default for that modality is used. Omit
model and the provider's default model is used.
Text
POST /v1/generate/text — synchronous.
| Field | Type | Required | Notes |
|---|---|---|---|
| prompt | string | yes | Max 10,000 characters |
| system_prompt | string | no | Max 5,000 characters |
| provider | string | no | anthropic, openai |
| model | string | no | See models |
| template_id | integer | no | Apply a saved prompt template |
| template_variables | object | no | Values for the template's placeholders |
| parameters.max_tokens | integer | no | 1–4096 |
| parameters.temperature | number | no | 0.0–2.0 |
{
"prompt": "Write a professional product description for a time-tracking app.",
"system_prompt": "You are a senior SaaS copywriter.",
"provider": "anthropic",
"model": "claude-sonnet-5",
"parameters": { "max_tokens": 800, "temperature": 0.7 }
}
Image
POST /v1/generate/image — synchronous.
| Field | Type | Notes |
|---|---|---|
| prompt | string | Required. Max 5,000 characters |
| negative_prompt | string | What to avoid. Max 2,000 characters |
| parameters.size | string | 256x256, 512x512, 1024x1024, 1024x1792, 1792x1024 |
| parameters.quality | string | standard or hd. hd costs more |
| parameters.style | string | vivid or natural |
| parameters.n | integer | 1–4. Billed per image |
{
"prompt": "Isometric illustration of a server rack, muted palette",
"negative_prompt": "blurry, low quality, text",
"provider": "openai",
"model": "gpt-image-1",
"parameters": { "size": "1024x1024", "quality": "hd", "n": 1 }
}
The response result is a URL to the stored image.
Video
POST /v1/generate/video — asynchronous. Returns immediately with status: "processing".
| Field | Type | Notes |
|---|---|---|
| prompt | string | Required |
| parameters.duration | integer | Seconds. Billed per second |
| parameters.ratio | string | e.g. 16:9, 9:16 |
Poll for completion:
curl https://api.genzoai.com/v1/generate/$UUID \ -H "Authorization: Bearer $GENZO_API_KEY" # status: pending -> processing -> completed | failed
Poll every 5–10 seconds. Video jobs typically finish in 1–4 minutes. Credits are released automatically if the job fails.
Speech
POST /v1/generate/audio — asynchronous. Text to speech.
| Field | Type | Notes |
|---|---|---|
| prompt | string | Required. The text to speak |
| parameters.voice | string | Voice identifier |
| parameters.speed | number | Playback rate multiplier |
| parameters.format | string | mp3, wav |
Billed per 1,000 characters of input text.
Retrieving results
| Endpoint | Returns |
|---|---|
| GET /v1/generate/{'{'}uuid{'}'} | One generation with full result and metadata |
| GET /v1/generate/history | Paginated list. Filter with ?type=, ?status=, ?per_page= |
Status values
pending— accepted, queued.processing— running upstream.completed—resultpopulated.failed—errorpopulated; credits released.