Voice Recognition Accuracy via Contextual Data Disambiguation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice recognition systems often suffer from low recognition accuracy due to the need for standard pronunciation, failing to account for variations in voice environments and contextual cues, which can lead to incorrect interpretations of homonyms and non-standard pronunciations.
Innovation Solution
A method and device that obtain audio data and contextual data characterizing the voice environment, using this information to improve recognition accuracy by selecting the most probable recognition result based on the context, such as location or historical data, thereby enhancing the correct rate of voice recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional voice recognition requires standard pronunciation, then the system structure remains simple, but recognition accuracy deteriorates due to inability to handle non-standard pronunciation and homonyms
Solution Approach 1:
The system performs preliminary acquisition of contextual data (location, time, historical information) before voice recognition occurs. This preliminary action enables the recognition system to have reference information ready, improving accuracy for non-standard pronunciation and homonym disambiguation without requiring complex real-time processing
Solution Approach 2:
Contextual data serves as an intermediary element between the audio input and recognition result. The contextual information (location, historical data, time) acts as a mediator that helps disambiguate homonyms and correct non-standard pronunciations, resolving the contradiction by adding an intermediate processing layer rather than directly complicating the core recognition engine
2Measurement precision
If contextual data is acquired to improve recognition accuracy, then recognition precision improves, but information processing time increases
Solution Approach 1:
Contextual data such as location, time, and historical information is acquired and prepared in advance before the voice recognition process. This preliminary preparation reduces the need for complex real-time analysis, thereby improving recognition accuracy while minimizing additional processing time during actual voice input
Solution Approach 2:
The system selectively acquires only the most relevant contextual data needed for recognition improvement rather than gathering all possible information. By focusing on key contextual elements (location, historical data, time) that most directly impact recognition accuracy, the system achieves high precision without excessive processing overhead
Data Source
AI summary
An information processing method and an electronic device are provided. The method includes: obtaining audio data collected by a slave device; obtaining contextual data corresponding to the slave device; and obtaining a recognition result of recognizing the audio data based on the contextual data. The contextual data characterizes a voice environment of the audio data collected by the slave device.


