> ## Documentation Index
> Fetch the complete documentation index at: https://docs.langcraft.world/llms.txt
> Use this file to discover all available pages before exploring further.

# Changelog

> New features, improvements, and fixes to the Langcraft Speech API.

export const Change = ({type = "new", tags = [], title, children}) => {
  const labels = {
    new: "New",
    improved: "Improved",
    fixed: "Fixed",
    breaking: "Breaking change"
  };
  return <div className={`lc-change lc-change--${type}`}>
      <div className="lc-change__meta">
        <span className={`lc-badge lc-badge--${type}`}>{labels[type] || type}</span>
        {tags.map(tag => <span key={tag} className="lc-chip">
            {tag}
          </span>)}
      </div>
      {title ? <div className="lc-change__title">{title}</div> : null}
      <div className="lc-change__body">{children}</div>
    </div>;
};

<div className="lc-changelog-page" />

## September 2026

<Update label="Sep 1" tags={["Breaking change"]}>
  <Change type="breaking" tags={["v1/speech/analyze", "transcription_mode"]} title="Verbatim is now the default transcription mode">
    When the API transcribes audio automatically and the request does not set
    `transcription_mode`, the transcript is now **verbatim**: it keeps fillers,
    repeated words, interrupted fragments, and vocal sounds, and fills the
    `events` array. Requests that send `reference_text` or `reference_phones`
    are not transcribed and are not affected.

    **Why it matters:** apps that display the transcript now see what the
    speaker actually said, including hesitations.

    **Action:** send `transcription_mode=intended` to keep the previous
    clean-transcript behavior. [Transcription modes →](/transcription-modes)
  </Change>
</Update>

## August 2026

<Update label="Aug 31" tags={["Fixed"]}>
  <Change type="fixed" tags={["Transcription"]} title="Automatic transcription of short recordings">
    Fixed an issue where automatic transcription of recordings shorter than 30
    seconds could fail with `transcription_unavailable`.
  </Change>
</Update>

<Update label="Aug 26" tags={["Improved"]}>
  <Change type="improved" tags={["aurora-1", "Models"]} title="One adult model for every language">
    `model=aurora-1` is now the recommended choice for adult speech in every
    supported language, including English.

    [Model selection →](/quickstart#model-selection)
  </Change>

  <Change type="improved" tags={["Timestamps"]} title="More accurate utterance start and end times">
    For phrases and sentences in languages other than English, the start of
    the first scored phone and the end of the last scored phone now line up
    more closely with when speech begins and ends.

    **Why it matters:** better trimming, playback highlighting, and duration
    measurements for each attempt.
  </Change>
</Update>

<Update label="Aug 25" tags={["Improved", "Fixed"]}>
  <Change type="improved" tags={["v1/speech/analyze", "model"]} title="Unrecognized model values are rejected">
    The `model` field now accepts only the documented values (`aurora-1`,
    `nova-1`, and `default`). Any other value returns
    `400 model_unavailable` instead of being passed through.

    **Action:** if you send any other `model` value, switch to a documented one
    or omit the field. [Model selection →](/quickstart#model-selection)
  </Change>

  <Change type="fixed" tags={["Japanese", "Korean", "Phones"]} title="More precise Japanese and Korean phones">
    Japanese palatalized consonants keep their palatalization (for example
    `kʲ`, `ɾʲ`), Japanese *u* is reported as `ɯ`, and Korean aspiration is
    attached to the consonant it belongs to.
  </Change>

  <Change type="fixed" tags={["v1/speech/analyze"]} title="Consistent English analysis when no language is sent">
    Requests that include `reference_text` but neither `lang` nor `dialect`
    are now consistently analyzed as US English (`en-us`), as documented.
  </Change>

  <Change type="improved" tags={["Status page"]} title="Status page history and incident timelines">
    Daily uptime history now shows each day's uptime more precisely, you can
    select a day to see its incidents, and each incident shows a timeline from
    detection to resolution.
  </Change>
</Update>

<Update label="Aug 13" tags={["New"]}>
  <Change type="new" tags={["Docs", "Benchmarks"]} title="Published pronunciation benchmarks">
    We published benchmark results for adult and second-language speech, young
    learners, and French, Spanish, and German adult speech, together with the
    full methodology and comparisons with other pronunciation assessment
    providers.

    [Benchmarks →](/benchmarks)
  </Change>
</Update>

<Update label="Aug 10" tags={["Fixed"]}>
  <Change type="fixed" tags={["Transcription"]} title="Verbatim words link to the right scored words">
    In verbatim transcripts, each transcribed word now links reliably to the
    matching scored word in `word_groups`, including when the transcript
    contains fillers or repetitions.
  </Change>
</Update>

## July 2026

<Update label="Jul 30" tags={["Improved"]}>
  <Change type="improved" tags={["Lexical stress", "English"]} title="More accurate English lexical stress">
    Stress scores for English words of two or more syllables are now more
    accurate. [Lexical stress →](/lexical-stress)
  </Change>
</Update>

<Update label="Jul 29" tags={["New", "Improved"]}>
  <Change type="new" tags={["v1/speech/analyze", "recording_quality"]} title="Recording quality diagnostics (preview)">
    Responses can now include a `recording_quality` object with the estimated
    signal-to-noise ratio, reverberation risk, clipping, an overall status, and
    a recommended action such as asking the learner to move closer to the
    microphone. It does not change scores or grades.

    **Why it matters:** your app can tell a real pronunciation problem apart
    from a recording that was too noisy or echoey to judge fairly.

    [Recording quality →](/recording-quality)
  </Change>

  <Change type="improved" tags={["Reliability", "Errors"]} title="No more stalled requests while a model warms up">
    If the service is not ready to take a request, the API now returns
    `503 model_warming_up` with a `Retry-After` header within a few seconds,
    instead of holding the request open. Unrecognized
    `dialect` values now return `400 invalid_dialect`.

    **Action:** retry `model_warming_up` responses after the `Retry-After`
    interval.
  </Change>
</Update>

<Update label="Jul 28" tags={["New", "Fixed"]}>
  <Change type="new" tags={["nova-1", "Models"]} title="Nova 1 for early-childhood speech">
    Early-childhood analysis is now available as the public model Nova 1
    (`model=nova-1`), for young speakers from about six months to eight years
    old. It supports English, French, German, Portuguese, and Spanish, uses
    the same request and response format, and requires reference text or
    reference phones. `speaker_profile=early_childhood` also selects it.

    [Early-childhood speech →](/early-childhood-speaker-profile)
  </Change>

  <Change type="new" tags={["Transcription", "Errors"]} title="Language mismatch detection">
    When automatic transcription finds no usable text and the recording is
    confidently identified as a different language from the one requested, the
    API now returns `422 language_mismatch` with `selected_language`,
    `detected_language`, and `confidence`, instead of `transcription_empty`.
  </Change>

  <Change type="fixed" tags={["rate_metrics"]} title="Correct word counts in speaking-rate metrics">
    Fixed `rate_metrics.word_count`, `speech_rate_wpm`, and
    `articulation_rate_wpm` undercounting words when a timed transcript entry
    contained more than one word.
  </Change>

  <Change type="new" tags={["Dashboard"]} title="API status in the dashboard">
    The developer dashboard now shows live API status and recent availability
    in its header.
  </Change>
</Update>

<Update label="Jul 27" tags={["New", "Improved"]}>
  <Change type="new" tags={["Lexical stress", "English"]} title="Lexical stress (preview)">
    For English, scoreable syllables in `word_groups[].syllables[]` now include
    a `stress` object with a stress score and, where known, whether the
    syllable is expected to carry primary stress. No request field is needed.

    **Why it matters:** you can flag words stressed on the wrong syllable, such
    as *baNAna* said as *BAnana*.

    [Lexical stress →](/lexical-stress)
  </Change>

  <Change type="improved" tags={["Syllables", "English"]} title="Better English syllable boundaries">
    English syllable boundaries are more accurate, fixing several words whose
    syllables were split in the wrong place.
  </Change>

  <Change type="improved" tags={["Transcription"]} title="More accurate verbatim phone timing">
    In verbatim mode, timestamped phones in `transcription` are assigned more
    accurately to words, fillers, and cut-off words.
  </Change>

  <Change type="improved" tags={["Early-childhood"]} title="Better single-word accuracy in German and Portuguese">
    Early-childhood analysis of single words in German and Portuguese is more
    accurate.
  </Change>
</Update>

<Update label="Jul 24" tags={["Improved"]}>
  <Change type="improved" tags={["Early-childhood"]} title="Early-childhood analysis for phrases and sentences">
    The early-childhood profile now accepts phrases and sentences as well as
    single words. [Early-childhood speech →](/early-childhood-speaker-profile)
  </Change>
</Update>

<Update label="Jul 23" tags={["New", "Fixed"]}>
  <Change type="new" tags={["Early-childhood"]} title="Early-childhood speaker profile (preview)">
    Added `speaker_profile=early_childhood` for analyzing young children's
    speech: single-word requests with reference text in English, French,
    German, Portuguese, and Spanish. Invalid values return
    `400 invalid_speaker_profile`.

    [Early-childhood speech →](/early-childhood-speaker-profile)
  </Change>

  <Change type="fixed" tags={["Alignment", "Errors"]} title="Clear errors when audio cannot be aligned">
    Requests whose audio cannot be aligned to the reference, for example when
    the reference is much longer than the recording, now return
    `422 audio_reference_mismatch` instead of a server error. Some compound
    words, such as French *ceux-là*, are now analyzed correctly and returned
    as one word.
  </Change>
</Update>

<Update label="Jul 22" tags={["New"]}>
  <Change type="new" tags={["v1/speech/analyze", "transcription_mode"]} title="Transcription modes">
    Added `transcription_mode` for requests the API transcribes automatically.
    `intended` (the default at launch) returns a clean, readable transcript;
    `verbatim` returns what the speaker actually said, including fillers,
    repetitions, interrupted words, and vocal sounds, plus a normalized
    `events` array. Transcripts now also include timestamped surface phones.
    Invalid values return `400 invalid_transcription_mode`.

    **Why it matters:** fluency and speech-therapy apps can see and count
    disfluencies instead of having them cleaned away.

    [Transcription modes →](/transcription-modes)
  </Change>
</Update>

<Update label="Jul 6" tags={["New", "Fixed"]}>
  <Change type="new" tags={["Syllables"]} title="Syllables in word results">
    Each entry in `word_groups` now includes `syllables`: the written syllable
    text, the reference phones it covers, and a per-syllable `score_pct`.
  </Change>

  <Change type="fixed" tags={["Transcription"]} title="Long recordings are transcribed in full">
    Fixed automatic transcription of long recordings, which could stop after
    the first 30 seconds.
  </Change>
</Update>

## June 2026

<Update label="Jun 25" tags={["New"]}>
  <Change type="new" tags={["v1/speech/analyze", "functional_load"]} title="Functional load">
    Substituted phones can now include `functional_load` metadata: whether the
    error risks changing the word's meaning, a priority score and category,
    and example minimal pairs.

    **Why it matters:** you can show learners the errors that matter most
    first, such as *ship* heard as *sheep*.

    [Functional load →](/functional-load)
  </Change>
</Update>

<Update label="Jun 22" tags={["Improved"]}>
  <Change type="improved" tags={["Scoring"]} title="Better scoring for single words">
    Scoring of one-word requests is more accurate in English, Arabic, Chinese,
    French, German, Hindi, Italian, Portuguese, Russian, Spanish, and Turkish.
  </Change>
</Update>

<Update label="Jun 20" tags={["Improved"]}>
  <Change type="improved" tags={["Scoring", "Alignment"]} title="More reliable error detection in short phrases">
    For references of one to five words, substitutions, insertions, and
    deletions are identified more reliably.
  </Change>

  <Change type="improved" tags={["Scoring", "Languages"]} title="Fewer false errors in Japanese, Mandarin, Thai, Korean, and Vietnamese">
    Pronunciation scoring in these languages reports fewer false
    substitutions.
  </Change>
</Update>

<Update label="Jun 18" tags={["Improved", "Fixed"]}>
  <Change type="improved" tags={["aurora-1", "Languages"]} title="Aurora becomes the default for more languages">
    When `model` is omitted or set to `default`, requests in languages other
    than English, French, and Spanish now use `aurora-1`. `model=default` now
    follows the same language-based choice as omitting `model`.

    **Action:** set `model` explicitly if you need to keep the previous model
    for a language.
  </Change>

  <Change type="fixed" tags={["Mandarin"]} title="Mandarin word segmentation">
    Mandarin text written without spaces is now split into words, so
    `word_groups` has one entry per word instead of one for the whole
    sentence.
  </Change>
</Update>

<Update label="Jun 17" tags={["Improved"]}>
  <Change type="improved" tags={["Audio formats"]} title="M4A is officially supported">
    M4A joins WAV, FLAC, MP3, OGG, and WEBM as a documented audio format. For
    raw uploads, send it with `Content-Type: audio/mp4`.
    [Supported formats →](/quickstart#supported-formats)
  </Change>
</Update>

<Update label="Jun 10" tags={["Improved"]}>
  <Change type="improved" tags={["Errors"]} title="Consistent error responses">
    Errors now return a JSON body with a stable machine-readable `code` and a
    plain-language `error` message, for example `missing_audio`,
    `audio_too_short`, `audio_decode_failed`, `invalid_match_score_mode`, and
    `model_unavailable`.

    **Action:** branch on `code` rather than the message text.
  </Change>
</Update>

<Update label="Jun 9" tags={["New", "Fixed"]}>
  <Change type="new" tags={["Status page"]} title="End-to-end status monitoring">
    The status page now reflects end-to-end checks of the public API, so it
    matches what your integration experiences.
  </Change>

  <Change type="fixed" tags={["Arabic"]} title="Arabic short vowels">
    Arabic text written without diacritics now produces reference phones that
    include the short vowels.
  </Change>
</Update>

<Update label="Jun 3" tags={["New"]}>
  <Change type="new" tags={["aurora-1", "Models"]} title="Model selection and Aurora 1">
    The `model` request field is now documented, and Aurora 1
    (`model=aurora-1`) is available as an experimental multilingual model,
    recommended for German, French, and Spanish. Its predicted phones use the
    same IPA symbols as reference phones. Omit `model` for automatic
    selection.

    [Model selection →](/quickstart#model-selection)
  </Change>
</Update>

<Update label="Jun 2" tags={["Improved"]}>
  <Change type="improved" tags={["Phones"]} title="Cleaner predicted phone sequences">
    Adjacent duplicate phones are collapsed in predicted phone output, and the
    matching repeated-phone insertions are no longer reported in
    `word_groups`.
  </Change>
</Update>

## May 2026

<Update label="May 30" tags={["Improved"]}>
  <Change type="improved" tags={["Scoring"]} title="Fewer false errors on similar consonants">
    When you send a reference, the API now avoids two common false errors: a
    final /k/ heard as /t/, and, in English, an /s/ heard as /z/.
  </Change>
</Update>

<Update label="May 28" tags={["Improved"]}>
  <Change type="improved" tags={["Scoring", "English"]} title="More consistent partial credit in English">
    In English, close substitutions earn more consistent partial credit in
    `lenient` mode.
  </Change>
</Update>

<Update label="May 27" tags={["Improved"]}>
  <Change type="improved" tags={["Audio input", "Errors"]} title="More robust audio uploads">
    The audio format is now detected from the file itself when the
    `Content-Type` is missing or generic (for example
    `application/octet-stream`). Audio-only requests with no detectable speech
    return `422` with a clear message.
  </Change>
</Update>

<Update label="May 9" tags={["Fixed"]}>
  <Change type="fixed" tags={["Audio formats"]} title="iPhone M4A recordings">
    Fixed M4A recordings from iOS, such as Voice Memos, which could fail with a
    server error. Empty or header-only uploads now return a clear `400` error.
  </Change>
</Update>

<Update label="May 6" tags={["Fixed", "New"]}>
  <Change type="fixed" tags={["v1/speech/analyze"]} title="Requests with reference_text are scored against it">
    Requests that sent `reference_text` or `reference_phones` were treated as
    having no reference and were transcribed instead of scored. They are now
    scored against your reference.

    **Action:** if you sent `reference_text` before May 6 and saw a transcript
    in the response's `reference_text`, rerun those requests.
  </Change>

  <Change type="new" tags={["v1/speech/analyze"]} title="New request field names">
    `reference_text` and `reference_phones` (a space-separated IPA reference)
    are now the documented request fields. `text` and `ipa` keep working.
    [Request fields →](/quickstart#request-fields)
  </Change>
</Update>

## April 2026

<Update label="Apr 29" tags={["New"]}>
  <Change type="new" tags={["Website", "Playground"]} title="Redesigned website and playground">
    langcraft.world and the interactive playground were redesigned.
  </Change>
</Update>

<Update label="Apr 20" tags={["Improved"]}>
  <Change type="improved" tags={["Phones"]} title="Plain IPA phone tokens">
    Phone tokens in responses are now plain IPA symbols without tone or stress
    digits (for example, `i` instead of `i5`).
  </Change>
</Update>

<Update label="Apr 10" tags={["Improved"]}>
  <Change type="improved" tags={["Scoring"]} title="Equivalent phone spellings are not errors">
    Phones that are written differently but sound the same (for example `kʰ`
    and `k`, or `ɡ` and `g`) are no longer counted as substitutions.
  </Change>
</Update>

## March 2026

<Update label="Mar 27" tags={["Improved"]}>
  <Change type="improved" tags={["Timestamps"]} title="More accurate phoneme timestamps">
    Phone start and end times now follow the recognized sounds more closely,
    and adjacent phones share a boundary. In our tests this roughly halved the
    average boundary error.
  </Change>
</Update>

<Update label="Mar 9" tags={["New"]}>
  <Change type="new" tags={["v1/speech/analyze", "Scoring"]} title="Strict and lenient scoring">
    Added `match_score_mode`. `lenient` (the default) gives near-matches
    partial credit; `strict` scores substitutions `0` and accepted matches
    `100`.

    [Request fields →](/quickstart#request-fields)
  </Change>

  <Change type="new" tags={["Status page"]} title="Status page">
    Launched a public status page for the Speech API.
  </Change>
</Update>

<Update label="Mar 5" tags={["Improved"]}>
  <Change type="improved" tags={["Dashboard", "Billing"]} title="Redesigned dashboard">
    The developer dashboard moved to a new design and now shows your remaining
    free-trial days and audio minutes, with a link to the billing portal.
  </Change>
</Update>

<Update label="Mar 3" tags={["New", "Improved"]}>
  <Change type="new" tags={["v1/speech/analyze", "Fluency"]} title="Pause and speaking-rate metrics">
    Responses now include pause metrics (`pauses`, `pause_metrics`,
    `pause_threshold_sec`), speaking-rate metrics (`speech_rate_wpm`,
    `articulation_rate_wpm`, `rate_metrics`), and `audio_duration_sec`.

    [Output reference →](/output)
  </Change>

  <Change type="improved" tags={["Billing"]} title="Separate pricing for scripted and unscripted requests">
    Usage is now metered separately for scripted requests (with a reference)
    and unscripted requests (transcribed automatically).
  </Change>

  <Change type="improved" tags={["Website"]} title="New website">
    langcraft.world moved to a new design.
  </Change>
</Update>

## February 2026

<Update label="Feb 14" tags={["Breaking change"]}>
  <Change type="breaking" tags={["v1/speech/analyze"]} title="Response format finalized">
    Over the first weeks after launch, the response format settled into its
    current shape:

    * `reference_ipa` and `predicted_ipa` became `reference_phones` and
      `predicted_phones`, returned as arrays of phone tokens.
    * `confidence_pct` and `avg_confidence_pct` became `match_score_pct` and
      `avg_match_score_pct`.
    * Summary counts became `num_insertions`, `num_deletions`, and
      `num_substitutions`.
    * `grade` uses `match`, `substitute`, and `delete`.
    * `utterance_id`, `audio.path`, and the flat top-level `phones`, `edits`,
      and `insertions` lists were removed; phones and insertions are reported
      per word in `word_groups`.
    * `enable_prosody` was renamed `enable_prosody_contours`. The old name
      still works.

    **Action:** integrations built before February 14 should update these
    field names. [Output reference →](/output)
  </Change>
</Update>

<Update label="Feb 6" tags={["New"]}>
  <Change type="new" tags={["proficiency_metrics"]} title="Proficiency metrics">
    Set `enable_proficiency_metrics=true` to get holistic fluency, prosody, and
    intelligibility scores from 0 to 100, each with a short explanation.

    [Proficiency metrics →](/output#proficiency-metrics)
  </Change>

  <Change type="new" tags={["Audio formats"]} title="WEBM support">
    WEBM audio, as recorded by most browsers, is now supported.
  </Change>
</Update>

<Update label="Feb 5" tags={["New"]}>
  <Change type="new" tags={["Languages"]} title="Supported languages reference">
    Published the list of 36 supported languages with their `lang` codes and
    regional `dialect` codes.

    [Supported languages →](/dialects)
  </Change>
</Update>

<Update label="Feb 4" tags={["New"]}>
  <Change type="new" tags={["v1/speech/analyze"]} title="Versioned endpoint">
    The API is now served at `POST /v1/speech/analyze`. The original
    `/analyze` path continues to work.
  </Change>
</Update>

<Update label="Feb 2" tags={["New", "Improved"]}>
  <Change type="new" tags={["Audio input"]} title="More ways to send audio">
    Besides multipart uploads, audio can be sent as a raw request body, as
    base64 in `audio_b64`, or as a link in `audio_url`.
    [Audio upload options →](/quickstart#audio-upload-options)
  </Change>

  <Change type="new" tags={["Prosody"]} title="Prosody contours">
    Word-level pitch and stress contours, sampled every 10 ms, are available
    on request (now `enable_prosody_contours=true`).
    [Prosody →](/output#prosody)
  </Change>

  <Change type="new" tags={["topk"]} title="Alternative phone guesses">
    Each phone now includes `topk`: the model's top alternative guesses for
    that phone, each with its own score.
  </Change>

  <Change type="improved" tags={["Billing"]} title="Usage metered by audio minutes">
    Paid plans are metered by audio minutes processed, with usage and invoices
    in the dashboard.
  </Change>
</Update>

## January 2026

<Update label="Jan 30" tags={["New"]}>
  <Change type="new" tags={["Dashboard", "Authentication", "Billing"]} title="Developer dashboard, API keys, and plans">
    Every request is authenticated with an API key in the `x-api-key` header.
    In the developer dashboard you can create, name, reveal, and delete keys,
    review your API requests and usage, start a free trial with an included
    audio-minute allowance, and subscribe to a paid plan. Requests after the
    trial allowance ends return `402`.

    [Getting started →](/quickstart)
  </Change>
</Update>

<Update label="Jan 28" tags={["New"]}>
  <Change type="new" tags={["Speech API"]} title="Langcraft Speech API launch">
    The Langcraft Speech API is live at `api.langcraft.world` with this
    documentation site. One request returns phoneme recognition, alignment
    against a reference with millisecond timestamps, per-phone and per-word
    scores, per-phone pitch and stress, and an utterance summary. Provide
    reference text, an IPA phone sequence, or audio only for automatic
    transcription.

    [Getting started →](/quickstart)
  </Change>
</Update>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.