Stereo Audio Speech Detection via Center-Surround Energy Ratios
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice activity detection methods are prone to errors in mixed media contents with noise and struggle to accurately distinguish speech from music and sound effects, requiring extensive data and computational resources for training and feature extraction.
Innovation Solution
A method and apparatus that utilize inter-channel relation information from stereo audio signals to separate center and surround channel elements, calculate energy ratio values, and determine speech and non-speech sections without prior training, using energy ratios between center and surround channel signals and a mono signal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional threshold-based voice activity detection methods are used, then the detection process is simple, but detection accuracy deteriorates in mixed media contents with noise and music
Solution Approach 1:
The patent segments the audio signal processing into multiple stages: inter-channel relation analysis, center/surround channel separation, energy ratio calculation, and speech/non-speech determination. This segmentation allows accurate speech detection in mixed media by analyzing specific spatial characteristics rather than treating the audio signal as a whole, resolving the contradiction between accuracy and complexity.
Solution Approach 2:
The patent introduces a new dimension of analysis by utilizing inter-channel spatial relationships in stereo audio signals. Instead of analyzing only temporal or spectral features, the method adds spatial dimension through center-surround channel separation and energy ratio calculation, enabling accurate speech detection in complex mixed media contents.
2Measurement precision
If statistical feature extraction and classifier training methods are used for voice/music classification, then classification performance improves, but data preparation time and computational resources increase significantly
Solution Approach 1:
The patent employs self-service by utilizing inherent spatial characteristics of speech signals in stereo audio without requiring external training data or classifier learning. The inter-channel relation-based energy ratio method automatically adapts to different audio contents, eliminating the time-consuming data preparation and training phases while maintaining high classification accuracy.
3Measurement precision
If simple threshold-based detection is used, then computational resources and memory usage are minimal, but detection accuracy deteriorates in the presence of noise and mixed media
Solution Approach 1:
The patent applies local quality by focusing analysis on specific spatial regions through center and surround channel separation. Instead of processing the entire audio signal uniformly, the method calculates energy ratios for specific spatial components, achieving accurate speech detection with reduced computational burden compared to comprehensive signal analysis.
Data Source
AI summary
Provided is an apparatus for detecting a speech/non-speech section. The apparatus includes an acquisition unit which obtains inter-channel relation information of a stereo audio signal, a separation unit which separates each element of the stereo audio signal into a center channel element and a surround element on the basis of the inter-channel relation information, a calculation unit which calculates an energy ratio value between a center channel signal composed of center channel elements and a surround channel signal composed of surround elements, for each frame, and an energy ratio value between the stereo audio signal and a mono signal generated on the basis of the stereo audio signal, and a judgment unit which determines a speech section and a non-speech section from the stereo audio signal by comparing the energy ratio values.


