# OpenAI introduces next-generation audio models in the API

> OpenAI announced new speech-to-text and text-to-speech models in its API for voice-agent applications. The release described improvements to transcription, voice generation, and customization, and separated those models from later Realtime API availability updates. Product access, pricing, and terms remain specific to the stated release channels.

Canonical: https://hypler.com/ai-news/openai-audio-api-models-2025
Hypler publication: 2026-10-06
Original source date: 2025-03-20
Publisher: OpenAI
Article kind: source-briefing
Category: Developer tools
Tags: audio, voice, API

## Engineering relevance

Voice workflows add consent, retention, accessibility, and disclosure requirements. Hypler should define recording policy, human escalation, source attribution, and evaluation criteria before introducing spoken interactions.

## Original source

- [OpenAI](https://openai.com/index/introducing-our-next-generation-audio-models/)

## Image provenance

### Primary image

![Generated editorial illustration. Speech recognition and synthesis illustrated by a microphone, transcript ribbon and speaker.](https://hypler.com/assets/news/openai-audio-api-models-2025.webp)

generated illustration. Credit: Hypler / AI-generated editorial illustration. Hypler editorial asset; not vendor or event photography.
- [Image source](https://hypler.com/ai-news/media.md)

### Supporting image

![Generated illustration of microphone, camera optics, and compute interfaces for multimodal inputs.](https://hypler.com/assets/news/multimodal-inputs.webp)

generated illustration. Credit: Hypler / AI-generated editorial illustration. Hypler editorial asset; not vendor or event photography.
- [Image source](https://hypler.com/ai-news/media.md)

## Reading boundary

Historical source briefing, not current availability, compliance advice, partnership, or proof of Hypler deployment. Generated artwork is not event photography. Public reading grants no private access or execution authority.
