Speech recognition device, speech recognition method, and program

The E2E-ASR model with a dynamic vocabulary and bias encoder/decoder addresses the challenge of recognizing infrequent words by enhancing accuracy through easy registration, thereby improving speech recognition.

JP2025175589APending Publication Date: 2025-12-03HONDA MOTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024081768
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-12-03

AI Technical Summary

Technical Problem

Conventional End-to-End (E2E) speech recognition models struggle with accurately recognizing infrequent words, phrases, and sentences due to the lack of dictionaries, necessitating full model retraining for registration.

Method used

The E2E-ASR model incorporates a dynamic vocabulary with a bias encoder and decoder that registers pre-defined words, phrases, and sentences as bias tokens, allowing easy integration and improving recognition accuracy.

Benefits of technology

The model enhances speech recognition accuracy by enabling easy registration of infrequent terms, reducing errors and improving overall recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025175589000001_ABST
    Figure 2025175589000001_ABST
Patent Text Reader

Abstract

To provide a speech recognition device, speech recognition method, and program that can further improve speech recognition accuracy by utilizing an E2E-ASR model capable of simply registering low-frequency words, phrases, and sentences.SOLUTION: A speech recognition device comprises an acquisition unit that acquires audio data of speech, and a speech recognition unit that generates text from the audio data using an automatic speech recognition model. The automatic speech recognition model includes: an audio encoder that converts the audio data into features; a bias encoder that converts registered bias tokens into features; and a bias decoder, extended to correspond to the bias tokens, for estimating the next token based on the features output by the audio encoder, the features output by the bias encoder, and the previously estimated token sequence.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Voice Recognition System

    JP2021501376A