Speech Recognition Editing via Repeated Utterance Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in improving accuracy when appropriate learning has not been performed, leading to inconvenient speech recognition services.
Innovation Solution
An information processing apparatus and method that includes a recognition unit for identifying desired words in speech recognition results, a generating unit for acquiring and processing repeated speech information to generate speech information for editing, and a speech recognition unit for performing speech recognition on the edited information, allowing for the correction of errors in speech recognition results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If word replacement is performed based on past learning results, then speech recognition accuracy can be improved when learning has been performed, but the system fails to replace words as expected when appropriate learning has not been performed, lowering convenience
Solution Approach 1:
The system performs preliminary speech recognition to generate a first speech recognition result, then uses a language model to identify candidate words for replacement before the user interacts with the system. This preliminary processing prepares the system to handle both cases where learning has been performed and where it has not, by pre-identifying potential corrections.
Solution Approach 2:
The system displays the first speech recognition result to the user and accepts input indicating whether to replace identified words. This feedback mechanism allows the user to confirm or reject automatic replacements, ensuring convenience is maintained even when learning results are insufficient, while still enabling accuracy improvements when learning has been performed.
2Measurement precision
If automatic word replacement is performed without user confirmation, then speech recognition accuracy may be improved, but user control and convenience are reduced
Solution Approach 1:
The system dynamically adjusts its behavior based on user input. It can operate in automatic replacement mode when the language model is confident about corrections, or switch to manual confirmation mode when user control is needed. This dynamic adaptability resolves the contradiction by allowing both automatic accuracy improvement and user control depending on the situation.
Solution Approach 2:
The system changes the parameter of replacement execution from always automatic to conditionally automatic based on user confirmation. By introducing a user confirmation step, the system maintains flexibility to achieve accuracy improvements while preserving user control, adapting the level of automation based on user preference and system confidence.
Data Source
AI summary
There is provided an information processing apparatus, an information processing method, and a program capable of providing a more convenient speech recognition service. The processing of recognizing, as an edited portion, a desired word configuring a sentence presented to a user as a speech recognition result, acquiring speech information repeatedly uttered for editing a word of the edited portion, and connecting speech information other than a repeated utterance to the speech information is performed, and speech information for speech recognition for editing is generated. Then, speech recognition is performed on the generated speech information for speech recognition for editing.


