Getting started
Make your first API request in minutes.
Output reference
Learn the JSON schema and fields.
API overview
Langcraft Speech API provides pronunciation assessment and prosody analysis in a single call. It’s designed for EdTech, speech therapy, and linguistic analysis, and returns structured JSON you can use directly in your app. Key capabilities:- Word and phoneme alignment with millisecond timing
- Phoneme‑level scoring and error detection
- Automatic transcription with word‑level timestamps when no reference text is provided
- Multilingual support (36 languages)
- Model selection: use
model=aurora-1for adult speech andmodel=nova-1for early-childhood speech;model=defaultand an omitted model remain supported for automatic adult analysis
Highlights
- Per‑phoneme scores with timestamps
- Per‑word rollups and summaries
- Pitch and stress contours at the phoneme and word levels
- Alignment metadata to connect audio, phones, and text
Inputs
You can analyze speech with any of:- A reference text (
reference_text) plus language code (lang) — the API runs grapheme‑to‑phoneme generation to derive canonical phones - A direct IPA phone sequence (
reference_phones) — bypasses G2P, useful for pronunciation contrast tests - Audio only — the API runs automatic transcription and uses the transcript as the reference
reference_text accepts the alias text. reference_phones accepts the alias ipa.
Model selection
For adult speech in any supported language, including English, setmodel=aurora-1. The API uses the same request and response schema across all
supported languages.
You can also send model=default or omit model. Both remain supported and
automatically select adult analysis for the request language.
For early-childhood speech, use the separate public model:
nova-1 is designed for speakers from approximately six months to eight years
old. See Early-childhood speech analysis
for supported languages and request requirements.
