Public APIs
Convert Text to Speech API favicon

Convert Text to Speech API

Text Analysis

This API converts any text to speech automatically.

Convert Text to Speech API's website screenshot

About Convert Text to Speech API

Text to Speech API converts plain text or document content into MP3 audio. It accepts a text string, an uploaded document file, or a document URL, and generates speech using a selected voice gender of male or female. Text input is limited to 75,000 characters per request, and file uploads are capped at 100MB.

The API offers six POST endpoints covering the combinations of input type (uploaded file, document URL, or direct text) and output format (binary MP3 file returned directly, or a JSON response containing a signed download link). This lets it fit into workflows requiring either immediate file delivery or asynchronous download handling.

It is suited for developers building read-aloud features, accessibility-focused content playback, internal training libraries, or document narration into applications without implementing their own speech synthesis. The service is hosted on ApyHub and versioned at 1.0.0.

Key features

  • Converts plain text or document files to MP3 audio
  • Accepts text strings, uploaded document files, or document URLs as input
  • Choose between male and female voice gender
  • Returns audio as direct binary MP3 or as a JSON response with a signed download link
  • Handles up to 75,000 characters of text input
  • 6 POST endpoints covering file-to-speech and text-to-speech, each with binary or link output

Frequently asked questions

What input formats does the API accept?

It accepts plain text strings, uploaded document files, or document URLs, and converts the content into MP3 audio.

What output formats are available?

Depending on the endpoint, the API returns either a binary MP3 file directly or a JSON response with a signed download link.

Is there a limit on text length?

Yes, text input is limited to a maximum of 75,000 characters.

What voice options are supported?

You can choose a voice gender of male or female for the generated speech.

Who is this API useful for?

It suits developers building read-aloud modes, accessibility content playback, internal training libraries, or document narration features without building their own speech synthesis logic.

Advertise here

Featured products

  • SerpApi - Search API favicon
  • Screenshot Scout favicon
  • TalorData favicon
  • CoreClaw favicon

Show your product to thousands of developers

· 100k monthly pageviews
· 7k newsletter subscribers

Advertise your product