Error Handling & Retry Strategies
ModelMart returns standard OpenAI-compatible error objects:
{ "error": { "type": "invalid_request_error", "code": "balance_too_low", "message": "Insufficient account balance to reserve request." } }
1. HTTP Status & Error Codes
| Status | Code | Cause | Resolution |
|---|---|---|---|
| 400 | invalid_request | Malformed JSON or invalid parameter. | Validate schema and messages array. |
| 401 | invalid_api_key | Missing or invalid API key. | Check MODELMART_API_KEY in environment. |
| 402 | balance_too_low | Account balance exhausted. | Top up balance via Dashboard. |
| 404 | model_not_found | Invalid model ID. | Check model ID in catalog. |
| 413 | request_too_large | Token limit exceeded. | Shorten prompt or history. |
| 429 | rate_limit_exceeded | RPM/TPM limit hit. | Implement exponential backoff retry. |
| 502 | upstream_error | Upstream provider error. | Automatic retry handled by router. |
| 504 | upstream_timeout | Request processing timeout. | Reduce max_tokens. |
2. Exponential Backoff Pattern (Python)
import time import random from openai import OpenAI, RateLimitError, APIConnectionError, InternalServerError client = OpenAI( api_key="mm_live_xxxxxxxxxxxx", base_url="https://api.modelmart.io.vn/v1" ) def safe_completion(messages, model="claude-sonnet-5", retries=3): for i in range(retries): try: return client.chat.completions.create(model=model, messages=messages) except (RateLimitError, APIConnectionError, InternalServerError) as e: if i == retries - 1: raise e wait = (1.0 * (2 ** i)) + random.uniform(0, 0.5) time.sleep(wait)