Voice recognition method and device, storage medium, and air conditioner
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current far-field voice recognition technologies face challenges in accuracy and noise reduction due to limitations in microphone arrays and deep learning methods, particularly in complex environments, where location and direction data of sound sources cannot be simultaneously sensed, leading to low recognition rates and long response times.
Innovation Solution
The proposed solution involves using a microwave radar module to accurately locate sound sources and adjust the state of a microphone array based on this information, combined with a deep learning algorithm like LSTM for training a far-field voice recognition model, allowing for enhanced voice data collection and processing in complex environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If deep learning methods or microphone array methods are used to remove reverberation and noise, then noise reduction effect is improved, but far-field voice recognition rate remains low and response time is long
Solution Approach 1:
The system performs preliminary sound source localization using microwave radar before voice recognition processing. By determining the location and direction of the sound source in advance, the microphone array can be pre-adjusted to focus on the correct direction, ensuring that the subsequent deep learning-based noise reduction operates on already-filtered directional audio signals. This preliminary action significantly improves far-field voice recognition rates while maintaining noise reduction effectiveness.
Solution Approach 2:
The system segments the voice processing pipeline into distinct functional modules: microwave radar-based sound source localization, microphone array directional filtering, deep learning-based noise reduction, and voice recognition. Each module handles a specific aspect of the processing chain, allowing optimization of each component independently. This segmentation enables the system to achieve high recognition rates by ensuring that each stage contributes effectively to the overall performance.
2Device complexity
If general methods are used to process voice data, then device complexity is reduced, but far-field voice recognition rate is low
Solution Approach 1:
The system merges microwave radar technology with traditional microphone array processing to create a hybrid solution. The microwave radar provides accurate sound source localization and direction information, which is then integrated with the microphone array's directional filtering capabilities. This combination achieves high far-field voice recognition rates without requiring overly complex processing methods, as the radar pre-processing simplifies the subsequent audio analysis tasks.
Solution Approach 2:
The system introduces an intermediary component: the microwave radar system that acts as a mediator between the environment and the voice recognition system. The radar provides intermediate information about sound source location and direction, which bridges the gap between simple microphone arrays and complex deep learning models. This intermediary enables the system to achieve high recognition rates while keeping the overall processing complexity manageable.
3Device complexity
If microphone array state is not adjusted, then device complexity is reduced, but voice data collection accuracy in complex environments is poor
Solution Approach 1:
The system implements dynamic adjustment of the microphone array state based on real-time sound source location information from the microwave radar. The microphone array transitions from a static configuration to a dynamic one, where the beam direction and focus are continuously adjusted according to the detected sound source position. This dynamic adaptation significantly improves voice data collection accuracy in complex environments while maintaining manageable device complexity through automated control.
Solution Approach 2:
The system establishes a feedback loop where the microwave radar continuously monitors sound source position and feeds this information back to the microphone array controller. The microphone array then adjusts its state based on this feedback, creating a closed-loop system that automatically optimizes voice data collection. This feedback mechanism improves measurement precision without requiring complex manual intervention, as the system self-adjusts based on real-time conditions.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach improves the accuracy and reliability of far-field voice recognition by accurately determining sound source location and adjusting microphone array settings, resulting in higher recognition rates and reduced noise, effectively addressing the limitations of existing technologies.
Implementation Method 1
using a microwave radar module to accurately locate sound sources
Implementation Method 2
adjust the state of a microphone array based on this information, combined with a deep learning algorithm like LSTM for training a far-field voice recognition model
Data Source
Figure 1
Figure 2~3
Figure 4~6
AI summary
Provided is a voice recognition method and apparatus, a storage medium, and an air conditioner. The method includes: acquiring first voice data (S110); adjusting, according to the first voice data, a collection state of second voice data to obtain an adjusted collection state, and acquiring the second voice data based on the adjusted collection state (S120); and performing far-field voice recognition on the second voice data using a preset far-field voice recognition model so as to obtain semantic information corresponding to the acquired second voice data (S130). The application can solve the problem in which far-field voice recognition performance is poor when a deep learning method or a microphone array method is used to remove reverberation and noise from far-field voice data, thereby enhancing far-field voice recognition performance.