Sound Detection via Spatial Spectrum Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in accurately detecting overlapping speech in scenarios with multiple speakers, leading to low accuracy in speaker diarization and speech recognition.

Innovation Solution

A sound detection method that involves obtaining an initial sound signal and its spatial distribution spectrum, segmenting the signal to identify target sound segments, and inputting these segments into a sound detection model to determine the presence of multiple speakers, thereby improving the accuracy of overlapping speech detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If overlapping speech detection technology is used to identify multiple speakers, then speaker diarization and speech recognition can be performed, but the detection accuracy is relatively low

Engineering Contradiction:
Improveoverlapping speech detection accuracyVSAvoiddetection system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio signal into multiple frames and processes each frame independently through the neural network model. This segmentation approach allows the system to detect overlapping speech in localized time windows, improving overall detection accuracy while maintaining manageable computational complexity through parallel processing of segments

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces spatial distribution spectrum as an additional dimension beyond traditional temporal audio analysis. By incorporating spatial information from multiple microphone channels and analyzing the spectral distribution across different frequencies and time frames, the system achieves more accurate overlapping speech detection without proportionally increasing system complexity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If traditional speech recognition systems are used in multi-speaker scenarios, then speech can be recognized, but the accuracy deteriorates due to overlapping speech

Engineering Contradiction:
Improvespeech recognition reliabilityVSAvoidspeaker identification accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary detection of overlapping speech segments before conducting speaker diarization and speech recognition. By identifying and flagging overlapping regions in advance using the trained neural network model, the system can apply specialized processing strategies to these segments, improving overall reliability and accuracy of speaker identification in multi-speaker scenarios

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12170098B2Sound detection method
Publication Date: 2024.12.17 ALIBABA DAMO (HANGZHOU) TECH CO LTD
  • US12170098B2 patent drawing
  • US12170098B2 patent drawing
  • US12170098B2 patent drawing

AI summary

The present disclosure discloses a sound detection method. The method includes: obtaining an initial sound signal and a spatial distribution spectrum of the initial sound signal; segmenting the initial sound signal, to obtain a target sound segment, and obtaining a timestamp corresponding to the target sound segment, the target sound segment including a speech of at least one object, and the timestamp being used for indicating a start time of the target sound segment and an end time of the target sound segment; segmenting the spatial distribution spectrum by using the timestamp, to obtain a spatial distribution spectrum segment corresponding to the target sound segment; and inputting the target sound segment and the spatial distribution spectrum segment into a sound detection model, to obtain a first sound detection result, the first sound detection result being used for describing whether sound of multiple objects exists in the initial sound signal.