Voice Input Correction via Signal Similarity Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems fail to accurately determine when a second voice input is intended to correct a previously misrecognized audio signal, leading to inefficient user interaction and incorrect responses.
Innovation Solution
A method and device that analyze the similarity between first and second audio signals, identify vocal characteristics, and use natural language processing models to determine if the second signal is for correction, thereby obtaining and processing corrected words or syllables to provide accurate responses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech recognition is used to convert user voice to text, then user interaction convenience is improved, but misrecognition occurs leading to incorrect responses
Solution Approach 1:
The system analyzes the user's second voice input to determine whether it is a correction of the first misrecognized input. By providing feedback about the detected correction intention and presenting correction options to the user, the system enables the user to confirm or modify the correction, thereby improving recognition accuracy while maintaining ease of operation
Solution Approach 2:
The system performs preliminary analysis of the second voice input to detect correction intention before finalizing the recognition result. By identifying correction patterns and presenting potential corrections in advance, the system prepares the correction data structure and allows the user to confirm the intended meaning before execution
2Measurement precision
If the system requests user re-input for correction, then recognition accuracy is improved, but interaction time increases
Solution Approach 1:
Instead of requesting complete re-input from the user, the system performs partial analysis by detecting correction intention in the second voice input and automatically generating correction options. This partial action approach reduces the user's burden while maintaining high recognition accuracy
Solution Approach 2:
The system automatically detects correction patterns in the user's second voice input and generates correction suggestions without requiring explicit user commands. The system serves itself by identifying and processing corrections autonomously based on voice pattern analysis
3Measurement precision
If the system analyzes voice patterns and vocal characteristics to detect correction intention, then correction accuracy is improved, but processing complexity increases
Solution Approach 1:
The system segments the voice analysis process into distinct components: extracting vocal characteristics, analyzing voice patterns, detecting correction intention, and generating correction options. This segmentation allows each component to be optimized independently while maintaining overall system manageability
Data Source
AI summary
A method, performed by an electronic device, of processing a voice input of a user. The method includes obtaining a first audio signal from a first user voice input, obtaining a second audio signal from a second user voice input that is obtained subsequent to the first audio signal, identifying whether the second audio signal is an audio signal for correcting the obtained first audio signal, when the obtained second audio signal is an audio signal for correcting the obtained first audio signal, obtaining, from the obtained second audio signal, at least one of one or more corrected words or one or more corrected syllables, based on the at least one of the one or more corrected words or the one or more corrected syllables, identifying at least one corrected audio signal for the obtained first audio signal, and processing the at least one corrected audio signal.


