Speech Recognition Mode Selection via Silent Section Length
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems face a tradeoff between improving performance and response due to limited processing resources, where high-performance audio processing for better recognition performance leads to slower processing speed and delayed response, while prioritizing response results in compromised recognition performance.
Innovation Solution
A speech recognition method that determines a criteria value to assess the length of silent sections within processing frames, allowing for the selection of appropriate processing modes to adjust performance and response by executing audio processing only on sections of interest, including silent sections, thereby optimizing processing time and load.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-performance audio processing is used to improve speech recognition performance, then recognition accuracy is improved, but processing time increases and response speed deteriorates
Solution Approach 1:
The patent applies dynamics by making the processing mode adjustable and changeable based on silent section length. The system dynamically switches between first processing mode (higher performance, longer time) and second processing mode (lower performance, shorter time) according to the detected silent section characteristics, allowing optimal balance between recognition performance and response speed
Solution Approach 2:
The patent applies local quality by differentiating processing intensity based on local characteristics of the audio signal. Silent sections are identified and processed differently from speech sections, with the processing mode specifically adapted to the silent section length to reduce unnecessary processing time while maintaining recognition accuracy where needed
2Measurement precision
If audio processing is performed on entire processing sections including silent sections, then recognition performance is maintained, but processing load and time increase
Solution Approach 1:
The patent applies taking out by extracting and separately handling silent sections from the audio signal. The silent section detection unit identifies silent portions, and the processing mode determination unit uses this information to adjust processing, effectively removing unnecessary processing from non-speech portions while maintaining necessary processing for speech recognition
Solution Approach 2:
The patent applies partial action by performing audio processing selectively based on silent section length. When silent sections are detected, the system uses a second processing mode with reduced processing intensity, applying only the necessary level of processing rather than full processing across the entire audio section, thereby improving efficiency
Data Source
Figure 1
Figure 2~3
Figure 4
AI summary
In a speech recognition method, a criteria value is determined to determine the length of a silent section included in a processing section, and a processing mode to use is determined in accordance with the criteria value. The criteria value is used to obtain audio information of the processing section. Audio processing is executed on the audio information in the processing section, using the processing mode that has been determined. Speech recognition processing is executed on the audio information in the processing section that has been subjected to audio processing.