Speech Recognition Editing via Repeated Utterance Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in improving accuracy when appropriate learning has not been performed, leading to inconvenient speech recognition services.

Innovation Solution

An information processing apparatus and method that includes a recognition unit for identifying desired words in speech recognition results, a generating unit for acquiring and processing repeated speech information to generate speech information for editing, and a speech recognition unit for performing speech recognition on the edited information, allowing for the correction of errors in speech recognition results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If word replacement is performed based on past learning results, then speech recognition accuracy can be improved when learning has been performed, but the system fails to replace words as expected when appropriate learning has not been performed, lowering convenience

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidconvenience of speech recognition service
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs preliminary speech recognition to generate a first speech recognition result, then uses a language model to identify candidate words for replacement before the user interacts with the system. This preliminary processing prepares the system to handle both cases where learning has been performed and where it has not, by pre-identifying potential corrections.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system displays the first speech recognition result to the user and accepts input indicating whether to replace identified words. This feedback mechanism allows the user to confirm or reject automatic replacements, ensuring convenience is maintained even when learning results are insufficient, while still enabling accuracy improvements when learning has been performed.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If automatic word replacement is performed without user confirmation, then speech recognition accuracy may be improved, but user control and convenience are reduced

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiduser control over recognition results
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically adjusts its behavior based on user input. It can operate in automatic replacement mode when the language model is confident about corrections, or switch to manual confirmation mode when user control is needed. This dynamic adaptability resolves the contradiction by allowing both automatic accuracy improvement and user control depending on the situation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter of replacement execution from always automatic to conditionally automatic based on user confirmation. By introducing a user confirmation step, the system maintains flexibility to achieve accuracy improvements while preserving user control, adapting the level of automation based on user preference and system confidence.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11308951B2Information processing apparatus, information processing method, and program
Publication Date: 2022.04.19 SONY GROUP CORP
  • US11308951B2 patent drawing
  • US11308951B2 patent drawing
  • US11308951B2 patent drawing

AI summary

There is provided an information processing apparatus, an information processing method, and a program capable of providing a more convenient speech recognition service. The processing of recognizing, as an edited portion, a desired word configuring a sentence presented to a user as a speech recognition result, acquiring speech information repeatedly uttered for editing a word of the edited portion, and connecting speech information other than a repeated utterance to the speech information is performed, and speech information for speech recognition for editing is generated. Then, speech recognition is performed on the generated speech information for speech recognition for editing.