OpenAI's latest realtime speech-to-text model, built for low-latency use — it streams incremental transcripts as audio arrives, ideal for live captions, dictation, and voice interfaces. Supports automatic transcription across 57 languages, with keyword/context hints and a tunable latency-vs-accuracy tradeoff.
Pricing
$0.017 per minute (same as OpenAI's official rate, no markup), charged by the audio duration transcribed in real time. Usage is settled to the second, rounded up to the next whole second.
Input Modalities
- Text
Output Modalities
- Text