Skip to main content
This is the core endpoint for all AI-powered features — summarization, extraction, analysis, drafting.
Endpoint
Response

Parameters

Required

Optional

Messages

Each message in the messages array:

System prompts

Set the AI’s behavior with a system message:

Multi-turn conversations

Include previous messages to maintain context:

Streaming

Get responses token-by-token as they’re generated:

Streaming HTTP failures

When an upstream provider rejects a streaming request before streaming begins, Case.dev returns a non-2xx JSON error instead of opening an SSE stream. The response preserves the upstream HTTP status and includes data.code: "UPSTREAM_PROVIDER_ERROR" with a generic message. Check this machine-readable code before interpreting the HTTP status. In particular, a 401 or 403 carrying UPSTREAM_PROVIDER_ERROR describes an upstream failure, not an invalid Case.dev API key. Do not automatically revoke or rotate the caller’s key based on that status alone. Provider response bodies are not returned. This describes pre-stream failures; failures after SSE has started follow the stream’s error handling.

Vision

Send images to models that support vision (GPT and Gemini):
Typescript

Usage and costs

Every response includes token counts and cost:
Response
Reduce costs: Use temperature: 0 for factual extraction. Try cheaper models like openai/gpt-6-luna or google/gemini-3.1-flash-lite for simpler tasks.

Common patterns

Deposition summary

Contract clause extraction

Medical record review