Voice Recognition Confidence Update via Interaction Consistency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems often misrecognize user speech, leading to incorrect system responses and requiring user intervention for correction, which disrupts the interaction and may result in missed correction opportunities if the user is not immediately aware of the misrecognition.
Innovation Solution
An information processing apparatus and method that includes a voice recognition section and a learning processing section, which execute a voice recognition process and update the confidence levels of recognition results based on user interactions, allowing the system to detect and correct misrecognition without requiring special user operations or speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the system displays candidate data sets and requires user selection for correction, then the accuracy of speech recognition can be improved, but the ease of operation deteriorates because additional user actions are required
Solution Approach 1:
The system automatically detects misrecognition by monitoring interaction consistency between user speech and system responses, eliminating the need for users to manually select from candidate data sets. The system serves itself by autonomously identifying and correcting recognition errors through contextual analysis of the interaction flow.
Solution Approach 2:
The system incorporates feedback mechanisms where the learning processing section analyzes the interaction between user speech and system responses to automatically adjust confidence levels. This feedback loop enables the system to correct misrecognition based on contextual inconsistencies without requiring explicit user input for correction.
2Reliability
If the system requires user intervention to correct misrecognition, then the reliability of information processing can be improved, but the loss of time increases due to interruption of interaction flow
Solution Approach 1:
The system performs preliminary detection of potential misrecognition by analyzing interaction consistency in advance, before the user becomes aware of the error. The learning processing section continuously monitors the interaction flow and identifies inconsistencies that indicate misrecognition, enabling timely automatic correction without waiting for user perception or intervention.
Solution Approach 2:
The system autonomously detects and corrects misrecognition by monitoring interaction consistency, eliminating the need for user intervention. The learning processing section automatically adjusts confidence levels and corrects errors based on contextual analysis, maintaining interaction flow without time loss.
3Ease of operation
If the system uses automatic confidence level updates based on interaction consistency, then the ease of operation is improved by eliminating user intervention, but the device complexity increases due to additional learning processing
Solution Approach 1:
The learning processing section serves multiple functions: it updates confidence levels, detects misrecognition, analyzes interaction consistency, and corrects errors. By consolidating these functions into a single multi-functional component, the system achieves ease of operation without proportionally increasing overall device complexity.
Data Source
AI summary
Provided is an apparatus that includes a voice recognition section that executes a voice recognition process on a user speech and a learning processing section that executes a process of updating a degree of confidence on the basis of an interaction made between a user and the information processing apparatus after the user speech. The degree of confidence is an evaluation value indicating the reliability of a voice recognition result of the user speech. The voice recognition section generates data on degrees of confidence in recognition of the user speech in which data plural user speech candidates based on the voice recognition result of the user speech are associated with the degrees of confidence which are evaluation values each indicating reliability of the corresponding user speech candidate.


