Chat Completions & SSE Streaming

Primary endpoint for text and multimodal generation: POST https://api.modelmart.io.vn/v1/chat/completions


1. Standard Request (Non-Streaming)

curl https://api.modelmart.io.vn/v1/chat/completions \
  -H "Authorization: Bearer $MODELMART_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5",
    "messages": [
      {"role": "system", "content": "You are a concise tech assistant."},
      {"role": "user", "content": "Summarize ACID properties in databases."}
    ],
    "temperature": 0.3,
    "max_tokens": 500
  }'

2. Real-time Streaming (Server-Sent Events)

Pass "stream": true to receive incremental token chunks in real-time.

Python Streaming

from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ["MODELMART_API_KEY"],
    base_url="https://api.modelmart.io.vn/v1"
)

stream = client.chat.completions.create(
    model="gpt-5.6-sol",
    messages=[{"role": "user", "content": "Write a short poem about code compilation."}],
    stream=True
)

for chunk in stream:
    content = chunk.choices[0].delta.content or ""
    print(content, end="", flush=True)
print()

TypeScript Streaming

import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: process.env.MODELMART_API_KEY,
  baseURL: 'https://api.modelmart.io.vn/v1',
});

async function runStream() {
  const stream = await client.chat.completions.create({
    model: 'claude-sonnet-5',
    messages: [{ role: 'user', content: 'Explain Kubernetes Pod lifecycle.' }],
    stream: true,
  });

  for await (const chunk of stream) {
    process.stdout.write(chunk.choices[0]?.delta?.content || '');
  }
}

runStream();

3. Parameter Reference

ParameterTypeRequiredDescription
modelstringYesTarget model ID (e.g., claude-sonnet-5, gpt-5.6-sol).
messagesarrayYesArray of message objects (role, content).
temperaturenumberNoSampling temperature (0.0 to 2.0). Default 1.0.
top_pnumberNoNucleus sampling probability threshold. Default 1.0.
max_tokensintegerNoMaximum token limit for output completion.
streambooleanNoStream response chunks via SSE. Default false.
stopstring / arrayNoSequence where the model will stop generating tokens.
seedintegerNoDeterministic sampling seed for reproducible outputs.