Sound Signal Processing Device for Accurate Speech Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech detection methods face challenges in accurately distinguishing speech segments from noise and multiple speech sources, leading to incorrect recognition and inefficient sound source extraction, particularly in environments with mixed signals.

Innovation Solution

A sound signal processing device that employs a directional point detecting unit to identify the direction of arrival of sound signals by generating null beam and directionality patterns, and a dynamic threshold calculation to differentiate between speech and non-speech segments, enhancing the accuracy of speech detection by classifying directional characteristics patterns and adjusting thresholds based on 'speech likeliness' determination.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech detection methods are used, then the processing is simpler, but the accuracy of distinguishing speech from noise and multiple speech sources deteriorates

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The sound signal is divided into multiple frequency bands through spectral decomposition. The speech detection device separately processes each frequency band to identify directional points, then integrates results across bands. This segmentation allows accurate detection in complex mixed signals by analyzing individual frequency components rather than treating the entire spectrum as a single signal.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention adds a spatial dimension to speech detection by estimating direction of arrival (DOA) for each frequency band. Instead of only temporal analysis, the system creates a three-dimensional detection space combining frequency, time, and spatial dimensions. This enables differentiation of speech sources based on their directional characteristics, significantly improving accuracy in separating speech from noise and multiple speakers.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If simple threshold-based detection is used, then the processing speed is faster, but the reliability of speech segment identification deteriorates in mixed signal environments

Engineering Contradiction:
Improvespeech segment identification reliabilityVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system dynamically adjusts detection parameters including frequency band selection, threshold values, and directional point connection criteria based on signal characteristics. Instead of using fixed thresholds, the device adapts parameters to match the specific acoustic environment, improving reliability in varying conditions while maintaining efficient processing through automated parameter optimization.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The speech detection device incorporates feedback mechanisms where detection results from one time frame influence parameter selection in subsequent frames. The system continuously refines its detection criteria based on accumulated information about signal patterns, environmental noise characteristics, and speech source behavior, thereby improving reliability without requiring exponentially increased processing power.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If directional point tracking is implemented, then the accuracy of speech segment detection is improved, but the computational load increases

Engineering Contradiction:
Improvedirectional point detection accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary processing by pre-defining frequency bands and calculating expected directional patterns before actual speech detection. Reference models of directional characteristics are prepared in advance for different speech scenarios. This preliminary action reduces real-time computational requirements while maintaining high detection accuracy, as the system only needs to match incoming signals against pre-computed reference data rather than performing full spectral analysis from scratch.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10013998B2Sound signal processing device and sound signal processing method
Publication Date: 2018.07.03 SONY GROUP CORP
  • US10013998B2 patent drawing
  • US10013998B2 patent drawing
  • US10013998B2 patent drawing

AI summary

A device and a method for determining a speech segment with a high degree of accuracy from a sound signal in which different sounds coexist are provided. Directional points indicating the direction of arrival of the sound signal are connected in the temporal direction, and a speech segment is detected. In this configuration, pattern classification is performed in accordance with directional characteristics with respect to the direction of arrival, and a directionality pattern and a null beam pattern are generated from the classification results. Also, an average null beam pattern is also generated by calculating the average of the null beam patterns at a time when a non-speech-like signal is input. Further, a threshold that is set at a slightly lower value than the average null beam pattern is calculated as the threshold to be used in detecting the local minimum point corresponding to the direction of arrival from each null beam pattern, and a local minimum point equal to or lower than the threshold is determined to be the point corresponding to the direction of arrival.