Voice Recognition Model Direction Interval Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition technologies perform poorly in low SNR scenarios, such as far-field environments, due to signal attenuation and noise interference, leading to unstable recognition accuracy.
Innovation Solution
A method and apparatus that utilize a pre-trained voice recognition model with multiple processing layers, including convolutional layers and Fourier transform networks, to improve voice recognition accuracy by training the model on voice samples in specific direction intervals and determining voice recognition results based on initial text outputs from these networks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional voice recognition models are used, then the system is simple and easy to implement, but the recognition accuracy deteriorates in low SNR scenarios such as far-field environments
Solution Approach 1:
The voice recognition model is segmented into multiple specialized recognition networks, each trained on voice samples from specific direction intervals. This segmentation allows each network to specialize in recognizing voices from particular directions, improving overall accuracy in far-field environments without requiring a single overly complex model.
Solution Approach 2:
Different recognition networks are assigned different local expertise by training them on voice samples from specific direction intervals. Each network develops local quality specialized for its designated direction range, allowing the system to maintain high recognition accuracy across various directions while managing complexity through modular specialization.
2Measurement precision
If multiple recognition networks with direction interval training are used, then voice recognition accuracy in preset direction intervals is improved, but the computational complexity and processing time increase
Solution Approach 1:
Voice samples from different direction intervals are used to pre-train specialized recognition networks before deployment. This preliminary action during the training phase enables the networks to quickly and accurately recognize voices from their designated directions during inference, reducing processing time while maintaining high accuracy.
Solution Approach 2:
The system dynamically selects and weights the outputs of different recognition networks based on the detected direction of the incoming voice signal. This dynamic adaptation allows the system to efficiently process voices from different directions by activating the most relevant networks and adjusting their confidence weights, optimizing processing time while maintaining accuracy.
Data Source
AI summary
A method and an apparatus for recognizing a voice are provided. The method may include: inputting a target voice into a pre-trained voice recognition model to obtain an initial text output by at least one recognition network in the voice recognition model, the recognition network including a plurality of preset types of processing layers, and at least one type of processing layer of the recognition network being obtained by training based on a voice sample in a preset direction interval; and determining a voice recognition result of the target voice, based on the initial text.


