Speech API
for Voice-Enabled Experiences

Turn conversations into text and give your content a voice. Integrate speech recognition and voice generation into your products, support workflows, and multilingual experiences.

ChatBucket Speech API - ConsoleRequest example
Node.js

POST /stt/transcribe

const form = new FormData();
form.append("audio", audioFile);
form.append("language", "auto");

const response = await fetch(
  `${process.env.CB_API_BASE_URL}/stt/transcribe`,
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.CB_ACCESS_TOKEN}`,
    },
    body: form,
  }
);

if (!response.ok) throw new Error("Transcription failed");
const transcript = await response.json();

From audio to usable text

Upload a recording or capture speech in the workspace. Choose a language or use automatic detection, then review the transcript.

  1. 1Audio recording
  2. 2Language selection
  3. 3Transcript review
Open speech to text

Confirm your production endpoint, credentials, and supported formats with our team before integrating.

API Capabilities

Speech Tools Built for Developers

Audio to text. Text to voice. Practical controls for your next voice-enabled workflow.

Speech to Text
Text to Speech
Language Selection
Voice Controls
Developer-First Design
Authenticated Requests

Transcribe recordings and generate spoken content with language selection, voice settings, and authenticated server requests.

SPEECH AND LANGUAGE

Two Speech Workflows. One Workspace.

Move between transcription and voice generation, with languages and voices you can explore before integrating.

Audio to text

Record or upload audio, choose the source language or automatic detection, and review your transcript.

Text to voice

Choose a voice and language for your content, generate speech, and preview or download the audio.

Your voice settings

Adjust pitch and speaking rate. Explore available voice styles and languages in your workspace.

Speech to TextText to SpeechLanguage DetectionVoice SelectionPitch ControlSpeaking Rate

Use Cases

Built for Every Speech Workflow

Customer Support

  • Turn support recordings into readable transcripts
  • Generate spoken greetings and customer updates
  • Review conversations for recurring questions
Explore audio transcription

PRICING

Start with Your Workspace. Plan for Production.

Explore your account options and speak to our team about your speech requirements. View available plans.

WORKSPACE

Get started

Explore speech tools

Test speech with your own content.

  • Record or upload audio
  • Generate and preview speech
  • Language and voice selection
  • Dashboard usage tracking
Create your account

PRODUCTION

Custom

Discuss your requirements

Plan your integration with our team.

  • Confirm API access and credentials
  • Review supported audio formats
  • Discuss languages and voice options
  • Plan usage and production volume
Talk to our team

Questions, Answered

Yes, both are usage-based. Speech-to-Text costs ₹0.74/min ($0.0078) for streaming and ₹0.49/min ($0.0052) for pre-recorded audio. Text-to-Speech costs ₹0.95 per 1k characters ($0.01) for streaming and ₹2.85 per 1k characters ($0.03) for non-streaming. There are no minimums or commitments.

GET STARTED

Bring Speech to Your Next Experience

Try transcription and voice generation in your workspace, then connect the speech workflow that fits your product.

Smarter conversations

start here Early access. New features. Updates

background

Powered by

© 2026 ChatBucket. All rights reserved