openai-tts

10
1
Source

OpenAI Text-to-Speech API for high-quality speech synthesis. Use for generating natural-sounding audio from text with customizable voices and tones.

Install

mkdir -p .claude/skills/openai-tts && curl -L -o skill.zip "https://mcp.directory/api/skills/download/5813" && unzip -o skill.zip -d .claude/skills/openai-tts && rm skill.zip

Installs to .claude/skills/openai-tts

About this skill

OpenAI Text-to-Speech

Generate high-quality spoken audio from text using OpenAI's TTS API.

Authentication

The API key is available as environment variable:

OPENAI_API_KEY

Models

  • gpt-4o-mini-tts - Newest, most reliable. Supports tone/style instructions.
  • tts-1 - Lower latency, lower quality
  • tts-1-hd - Higher quality, higher latency

Voice Options

Built-in voices (English optimized):

  • alloy, ash, ballad, coral, echo, fable
  • nova, onyx, sage, shimmer, verse
  • marin, cedar - Recommended for best quality

Note: tts-1 and tts-1-hd only support: alloy, ash, coral, echo, fable, onyx, nova, sage, shimmer.

Python Example

from pathlib import Path
from openai import OpenAI

client = OpenAI()  # Uses OPENAI_API_KEY env var

# Basic usage
with client.audio.speech.with_streaming_response.create(
    model="gpt-4o-mini-tts",
    voice="coral",
    input="Hello, world!",
) as response:
    response.stream_to_file("output.mp3")

# With tone instructions (gpt-4o-mini-tts only)
with client.audio.speech.with_streaming_response.create(
    model="gpt-4o-mini-tts",
    voice="coral",
    input="Today is a wonderful day!",
    instructions="Speak in a cheerful and positive tone.",
) as response:
    response.stream_to_file("output.mp3")

Handling Long Text

For long documents, split into chunks and concatenate:

from openai import OpenAI
from pydub import AudioSegment
import tempfile
import re
import os

client = OpenAI()

def chunk_text(text, max_chars=4000):
    """Split text into chunks at sentence boundaries."""
    sentences = re.split(r'(?<=[.!?])\s+', text)
    chunks = []
    current_chunk = ""

    for sentence in sentences:
        if len(current_chunk) + len(sentence) < max_chars:
            current_chunk += sentence + " "
        else:
            if current_chunk:
                chunks.append(current_chunk.strip())
            current_chunk = sentence + " "

    if current_chunk:
        chunks.append(current_chunk.strip())

    return chunks

def text_to_audiobook(text, output_path):
    """Convert long text to audio file."""
    chunks = chunk_text(text)
    audio_segments = []

    for chunk in chunks:
        with tempfile.NamedTemporaryFile(suffix='.mp3', delete=False) as tmp:
            tmp_path = tmp.name

        with client.audio.speech.with_streaming_response.create(
            model="gpt-4o-mini-tts",
            voice="coral",
            input=chunk,
        ) as response:
            response.stream_to_file(tmp_path)

        segment = AudioSegment.from_mp3(tmp_path)
        audio_segments.append(segment)
        os.unlink(tmp_path)

    # Concatenate all segments
    combined = audio_segments[0]
    for segment in audio_segments[1:]:
        combined += segment

    combined.export(output_path, format="mp3")

Output Formats

  • mp3 - Default, general use
  • opus - Low latency streaming
  • aac - Digital compression (YouTube, iOS)
  • flac - Lossless compression
  • wav - Uncompressed, low latency
  • pcm - Raw samples (24kHz, 16-bit)
with client.audio.speech.with_streaming_response.create(
    model="gpt-4o-mini-tts",
    voice="coral",
    input="Hello!",
    response_format="wav",  # Specify format
) as response:
    response.stream_to_file("output.wav")

Best Practices

  • Use marin or cedar voices for best quality
  • Split text at sentence boundaries for long content
  • Use wav or pcm for lowest latency
  • Add instructions parameter to control tone/style (gpt-4o-mini-tts only)

latex-writing

benchflow-ai

Guide LaTeX document authoring following best practices and proper semantic markup. Use proactively when: (1) writing or editing .tex files, (2) writing or editing .nw literate programming files, (3) literate-programming skill is active and working with .nw files, (4) user mentions LaTeX, BibTeX, or document formatting, (5) reviewing LaTeX code quality. Ensures proper use of semantic environments (description vs itemize), csquotes (\enquote{} not ``...''), and cleveref (\cref{} not \S\ref{}).

569441

polyglot-rust-c

benchflow-ai

Guidance for creating polyglot source files that compile and run correctly as both Rust and C or C++. This skill applies when tasks require writing code that is valid in multiple languages simultaneously, exploiting comment syntax differences or preprocessor directives to create dual-language source files. Use this skill for polyglot programming challenges, CTF tasks, or educational exercises involving multi-language source compatibility.

135427

python-programming

benchflow-ai

Master Python fundamentals, OOP, data structures, async programming, and production-grade scripting for data engineering

6734

pytorch

benchflow-ai

Building and training neural networks with PyTorch. Use when implementing deep learning models, training loops, data pipelines, model optimization with torch.compile, distributed training, or deploying PyTorch models.

6931

search-flights

benchflow-ai

Search flights by origin, destination, and departure date using the bundled flights dataset. Use this skill when proposing flight options or checking whether a route/date combination exists.

19628

marker

benchflow-ai

Convert PDF documents to Markdown using marker_single. Use when Claude needs to extract text content from PDFs while preserving LaTeX formulas, equations, and document structure. Ideal for academic papers and technical documents containing mathematical notation.

6726

You might also like

nano-banana-pro

garg-aayush

Generate and edit images using Google's Nano Banana Pro (Gemini 3 Pro Image) API. Use when the user asks to generate, create, edit, modify, change, alter, or update images. Also use when user references an existing image file and asks to modify it in any way (e.g., "modify this image", "change the background", "replace X with Y"). Supports both text-to-image generation and image-to-image editing with configurable resolution (1K default, 2K, or 4K for high resolution). DO NOT read the image file first - use this skill directly with the --input-image parameter.

2,2683,903

ui-ux-pro-max

nextlevelbuilder

"UI/UX design intelligence. 50 styles, 21 palettes, 50 font pairings, 20 charts, 8 stacks (React, Next.js, Vue, Svelte, SwiftUI, React Native, Flutter, Tailwind). Actions: plan, build, create, design, implement, review, fix, improve, optimize, enhance, refactor, check UI/UX code. Projects: website, landing page, dashboard, admin panel, e-commerce, SaaS, portfolio, blog, mobile app, .html, .tsx, .vue, .svelte. Elements: button, modal, navbar, sidebar, card, table, form, chart. Styles: glassmorphism, claymorphism, minimalism, brutalism, neumorphism, bento grid, dark mode, responsive, skeuomorphism, flat design. Topics: color palette, accessibility, animation, layout, typography, font pairing, spacing, hover, shadow, gradient."

3,6603,342

pdf-to-markdown

aliceisjustplaying

Convert entire PDF documents to clean, structured Markdown for full context loading. Use this skill when the user wants to extract ALL text from a PDF into context (not grep/search), when discussing or analyzing PDF content in full, when the user mentions "load the whole PDF", "bring the PDF into context", "read the entire PDF", or when partial extraction/grepping would miss important context. This is the preferred method for PDF text extraction over page-by-page or grep approaches.

4,7072,234

flutter-development

aj-geddes

Build beautiful cross-platform mobile apps with Flutter and Dart. Covers widgets, state management with Provider/BLoC, navigation, API integration, and material design.

2,2851,724

drawio-diagrams-enhanced

jgtolentino

Create professional draw.io (diagrams.net) diagrams in XML format (.drawio files) with integrated PMP/PMBOK methodologies, extensive visual asset libraries, and industry-standard professional templates. Use this skill when users ask to create flowcharts, swimlane diagrams, cross-functional flowcharts, org charts, network diagrams, UML diagrams, BPMN, project management diagrams (WBS, Gantt, PERT, RACI), risk matrices, stakeholder maps, or any other visual diagram in draw.io format. This skill includes access to custom shape libraries for icons, clipart, and professional symbols.

2,4881,553

youtube-transcribe-skill

feiskyer

Extract subtitles/transcripts from a YouTube video URL and save as a local file. Use when you need to extract subtitles from a YouTube video.

3681,441