Speech Zone Detection Using Sound Source Localization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech signal-processing apparatuses face challenges in accurately detecting speech zones, leading to degraded speech recognition accuracy due to the inclusion of non-speech zones in the separation process.
Innovation Solution
A speech-processing apparatus and method that utilize sound source localization and multiple threshold values for precise speech zone detection, incorporating sound source localization, clustering, and event information to improve accuracy and reduce insertion errors and discontinuities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If speech zone detection is performed using a single threshold value on the speech signal, then the detection process is simple, but speech recognition accuracy is degraded due to inclusion of non-speech zones
Solution Approach 1:
The speech zone detection process is segmented into multiple stages: first detecting sound source candidates using a first threshold value, then performing clustering on these candidates, and finally detecting speech zones using a second threshold value for each cluster. This segmentation allows the system to achieve high detection accuracy without excessive complexity by dividing the detection task into manageable steps.
Solution Approach 2:
The invention introduces a new dimension to speech zone detection by incorporating spatial information through sound source localization. Instead of detecting speech zones based solely on temporal speech signal characteristics, the system adds spatial dimension by localizing sound sources and using their positions to define speech zones, thereby improving detection accuracy.
2Measurement precision
If multiple threshold values and clustering processes are used for speech zone detection, then speech recognition accuracy is improved, but the detection process becomes more complex
Solution Approach 1:
The complex detection process is segmented into distinct stages: sound source candidate detection using a first threshold, clustering of candidates, and speech zone detection using a second threshold for each cluster. This segmentation makes the complex process more manageable and systematic, improving accuracy while controlling complexity through structured processing.
Solution Approach 2:
The system performs preliminary sound source localization and candidate detection before final speech zone detection. By pre-processing the speech signal to identify sound source candidates and their spatial positions, the system prepares the data in advance, making the subsequent speech zone detection more accurate and efficient.
3Measurement precision
If sound source localization is performed before speech zone detection, then speech zone detection accuracy is improved, but the overall processing complexity increases
Solution Approach 1:
The sound source localization unit serves multiple functions: it localizes sound sources in space, identifies sound source candidates, and provides spatial information for speech zone detection. This multi-functionality reduces overall processing complexity by using a single unit for multiple purposes rather than requiring separate dedicated units for each function.
Solution Approach 2:
Sound source localization is performed as a preliminary step that provides essential spatial information for subsequent speech zone detection. By localizing sound sources first, the system prepares spatial data that simplifies the speech zone detection process, making the overall system more efficient despite the added processing step.
Data Source
AI summary
A speech-processing apparatus includes: a sound source localization unit that localizes a sound source based on an acquired speech signal; and a speech zone detection unit that performs speech zone detection based on localization information localized by the sound source localization unit.


