Sound Signal Processing Device for Accurate Speech Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech detection methods face challenges in accurately distinguishing speech segments from noise and multiple speech sources, leading to incorrect recognition and inefficient sound source extraction, particularly in environments with mixed signals.
Innovation Solution
A sound signal processing device that employs a directional point detecting unit to identify the direction of arrival of sound signals by generating null beam and directionality patterns, and a dynamic threshold calculation to differentiate between speech and non-speech segments, enhancing the accuracy of speech detection by classifying directional characteristics patterns and adjusting thresholds based on 'speech likeliness' determination.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional speech detection methods are used, then the processing is simpler, but the accuracy of distinguishing speech from noise and multiple speech sources deteriorates
Solution Approach 1:
The sound signal is divided into multiple frequency bands through spectral decomposition. The speech detection device separately processes each frequency band to identify directional points, then integrates results across bands. This segmentation allows accurate detection in complex mixed signals by analyzing individual frequency components rather than treating the entire spectrum as a single signal.
Solution Approach 2:
The invention adds a spatial dimension to speech detection by estimating direction of arrival (DOA) for each frequency band. Instead of only temporal analysis, the system creates a three-dimensional detection space combining frequency, time, and spatial dimensions. This enables differentiation of speech sources based on their directional characteristics, significantly improving accuracy in separating speech from noise and multiple speakers.
2Reliability
If simple threshold-based detection is used, then the processing speed is faster, but the reliability of speech segment identification deteriorates in mixed signal environments
Solution Approach 1:
The system dynamically adjusts detection parameters including frequency band selection, threshold values, and directional point connection criteria based on signal characteristics. Instead of using fixed thresholds, the device adapts parameters to match the specific acoustic environment, improving reliability in varying conditions while maintaining efficient processing through automated parameter optimization.
Solution Approach 2:
The speech detection device incorporates feedback mechanisms where detection results from one time frame influence parameter selection in subsequent frames. The system continuously refines its detection criteria based on accumulated information about signal patterns, environmental noise characteristics, and speech source behavior, thereby improving reliability without requiring exponentially increased processing power.
3Measurement precision
If directional point tracking is implemented, then the accuracy of speech segment detection is improved, but the computational load increases
Solution Approach 1:
The system performs preliminary processing by pre-defining frequency bands and calculating expected directional patterns before actual speech detection. Reference models of directional characteristics are prepared in advance for different speech scenarios. This preliminary action reduces real-time computational requirements while maintaining high detection accuracy, as the system only needs to match incoming signals against pre-computed reference data rather than performing full spectral analysis from scratch.
Data Source
AI summary
A device and a method for determining a speech segment with a high degree of accuracy from a sound signal in which different sounds coexist are provided. Directional points indicating the direction of arrival of the sound signal are connected in the temporal direction, and a speech segment is detected. In this configuration, pattern classification is performed in accordance with directional characteristics with respect to the direction of arrival, and a directionality pattern and a null beam pattern are generated from the classification results. Also, an average null beam pattern is also generated by calculating the average of the null beam patterns at a time when a non-speech-like signal is input. Further, a threshold that is set at a slightly lower value than the average null beam pattern is calculated as the threshold to be used in detecting the local minimum point corresponding to the direction of arrival from each null beam pattern, and a local minimum point equal to or lower than the threshold is determined to be the point corresponding to the direction of arrival.


