Stereo Audio Speech Detection via Center-Surround Energy Ratios

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice activity detection methods are prone to errors in mixed media contents with noise and struggle to accurately distinguish speech from music and sound effects, requiring extensive data and computational resources for training and feature extraction.

Innovation Solution

A method and apparatus that utilize inter-channel relation information from stereo audio signals to separate center and surround channel elements, calculate energy ratio values, and determine speech and non-speech sections without prior training, using energy ratios between center and surround channel signals and a mono signal.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional threshold-based voice activity detection methods are used, then the detection process is simple, but detection accuracy deteriorates in mixed media contents with noise and music

Engineering Contradiction:
Improvespeech section detection accuracyVSAvoiddetection method complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio signal processing into multiple stages: inter-channel relation analysis, center/surround channel separation, energy ratio calculation, and speech/non-speech determination. This segmentation allows accurate speech detection in mixed media by analyzing specific spatial characteristics rather than treating the audio signal as a whole, resolving the contradiction between accuracy and complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of analysis by utilizing inter-channel spatial relationships in stereo audio signals. Instead of analyzing only temporal or spectral features, the method adds spatial dimension through center-surround channel separation and energy ratio calculation, enabling accurate speech detection in complex mixed media contents.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If statistical feature extraction and classifier training methods are used for voice/music classification, then classification performance improves, but data preparation time and computational resources increase significantly

Engineering Contradiction:
Improvevoice/music classification accuracyVSAvoiddata preparation and training time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent employs self-service by utilizing inherent spatial characteristics of speech signals in stereo audio without requiring external training data or classifier learning. The inter-channel relation-based energy ratio method automatically adapts to different audio contents, eliminating the time-consuming data preparation and training phases while maintaining high classification accuracy.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If simple threshold-based detection is used, then computational resources and memory usage are minimal, but detection accuracy deteriorates in the presence of noise and mixed media

Engineering Contradiction:
Improvespeech section detection accuracyVSAvoidcalculation and memory usage
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies local quality by focusing analysis on specific spatial regions through center and surround channel separation. Instead of processing the entire audio signal uniformly, the method calculates energy ratios for specific spatial components, achieving accurate speech detection with reduced computational burden compared to comprehensive signal analysis.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9336796B2Method and apparatus for detecting speech/non-speech section
Publication Date: 2016.05.10 ELECTRONICS & TELECOMM RES INST
  • US9336796B2 patent drawing
  • US9336796B2 patent drawing
  • US9336796B2 patent drawing

AI summary

Provided is an apparatus for detecting a speech/non-speech section. The apparatus includes an acquisition unit which obtains inter-channel relation information of a stereo audio signal, a separation unit which separates each element of the stereo audio signal into a center channel element and a surround element on the basis of the inter-channel relation information, a calculation unit which calculates an energy ratio value between a center channel signal composed of center channel elements and a surround channel signal composed of surround elements, for each frame, and an energy ratio value between the stereo audio signal and a mono signal generated on the basis of the stereo audio signal, and a judgment unit which determines a speech section and a non-speech section from the stereo audio signal by comparing the energy ratio values.