Chat Completions & SSE Streaming
Primary endpoint for text and multimodal generation:
POST https://api.modelmart.io.vn/v1/chat/completions
1. Standard Request (Non-Streaming)
curl https://api.modelmart.io.vn/v1/chat/completions \ -H "Authorization: Bearer $MODELMART_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-sonnet-5", "messages": [ {"role": "system", "content": "You are a concise tech assistant."}, {"role": "user", "content": "Summarize ACID properties in databases."} ], "temperature": 0.3, "max_tokens": 500 }'
2. Real-time Streaming (Server-Sent Events)
Pass "stream": true to receive incremental token chunks in real-time.
Python Streaming
from openai import OpenAI import os client = OpenAI( api_key=os.environ["MODELMART_API_KEY"], base_url="https://api.modelmart.io.vn/v1" ) stream = client.chat.completions.create( model="gpt-5.6-sol", messages=[{"role": "user", "content": "Write a short poem about code compilation."}], stream=True ) for chunk in stream: content = chunk.choices[0].delta.content or "" print(content, end="", flush=True) print()
TypeScript Streaming
import OpenAI from 'openai'; const client = new OpenAI({ apiKey: process.env.MODELMART_API_KEY, baseURL: 'https://api.modelmart.io.vn/v1', }); async function runStream() { const stream = await client.chat.completions.create({ model: 'claude-sonnet-5', messages: [{ role: 'user', content: 'Explain Kubernetes Pod lifecycle.' }], stream: true, }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content || ''); } } runStream();
3. Parameter Reference
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Target model ID (e.g., claude-sonnet-5, gpt-5.6-sol). |
messages | array | Yes | Array of message objects (role, content). |
temperature | number | No | Sampling temperature (0.0 to 2.0). Default 1.0. |
top_p | number | No | Nucleus sampling probability threshold. Default 1.0. |
max_tokens | integer | No | Maximum token limit for output completion. |
stream | boolean | No | Stream response chunks via SSE. Default false. |
stop | string / array | No | Sequence where the model will stop generating tokens. |
seed | integer | No | Deterministic sampling seed for reproducible outputs. |