Speech recognition device, speech recognition method, and program
The E2E-ASR model with a dynamic vocabulary and bias encoder/decoder addresses the challenge of recognizing infrequent words by enhancing accuracy through easy registration, thereby improving speech recognition.
JP2025175589APending Publication Date: 2025-12-03HONDA MOTOR CO LTD
Patent Information
- Application Number
- JP2024081768
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-12-03
AI Technical Summary
Technical Problem
Conventional End-to-End (E2E) speech recognition models struggle with accurately recognizing infrequent words, phrases, and sentences due to the lack of dictionaries, necessitating full model retraining for registration.
Method used
The E2E-ASR model incorporates a dynamic vocabulary with a bias encoder and decoder that registers pre-defined words, phrases, and sentences as bias tokens, allowing easy integration and improving recognition accuracy.
Benefits of technology
The model enhances speech recognition accuracy by enabling easy registration of infrequent terms, reducing errors and improving overall recognition performance.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure 2025175589000001_ABST
Abstract
To provide a speech recognition device, speech recognition method, and program that can further improve speech recognition accuracy by utilizing an E2E-ASR model capable of simply registering low-frequency words, phrases, and sentences.SOLUTION: A speech recognition device comprises an acquisition unit that acquires audio data of speech, and a speech recognition unit that generates text from the audio data using an automatic speech recognition model. The automatic speech recognition model includes: an audio encoder that converts the audio data into features; a bias encoder that converts registered bias tokens into features; and a bias decoder, extended to correspond to the bias tokens, for estimating the next token based on the features output by the audio encoder, the features output by the bias encoder, and the previously estimated token sequence.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
Voice Recognition System
JP2021501376A