Speech Processing Device Crosstalk Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech processing devices struggle to effectively suppress acoustic crosstalk between multiple speakers in a closed space, particularly when microphones are biased towards one speaker, leading to poor sound quality due to uneven sound pressure ratios and difficulty in learning filter coefficients for adaptive filtering.

Innovation Solution

A speech processing device and method that utilize multiple microphones to detect single-talk states and estimate mixing rates of speech signals, determining the necessity of crosstalk suppression based on sound pressure ratios to adaptively suppress crosstalk components and improve sound quality by adjusting filter coefficients and directionality in real-time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single microphone is disposed in front of the driver to collect driver speech, then the driver speech can be collected at high sound pressure, but the passenger speech cannot be collected at high sound pressure due to distance and position bias

Engineering Contradiction:
Improvesound pressure ratioVSAvoidspeech collection capability for multiple speakers
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent divides the speech collection function into multiple microphones positioned at different locations (driver-side microphone and passenger-side microphone). Each microphone is optimized to collect speech from its respective side, enabling high sound pressure collection for both driver and passenger speech simultaneously through segmented positioning.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by assigning different microphones to different spatial zones (driver zone and passenger zone). The driver-side microphone is optimized for collecting driver speech with appropriate directional characteristics, while the passenger-side microphone is optimized for passenger speech, allowing each microphone to have specialized local optimization for its target speaker.

Inventive Principle:
Principle #3Local quality

2Reliability

If the microphone position is biased toward the driver, then the driver speech is collected effectively, but the passenger speech cannot be collected at high sound pressure making it difficult to learn filter coefficients for adaptive filtering

Engineering Contradiction:
Improvespeech signal qualityVSAvoidfilter coefficient learning difficulty
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the adaptive filtering task into two independent filter learning processes: one filter for suppressing passenger speech in the driver speech channel, and another filter for suppressing driver speech in the passenger speech channel. This segmentation allows each filter to learn from sufficient speech data of the target speaker without being biased by microphone positioning.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the system parameters by introducing multiple microphones with different positional parameters and directional characteristics. This parameter change enables the system to collect sufficient speech data from both speakers, making it possible to learn accurate filter coefficients for adaptive filtering of crosstalk from either speaker.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If a single microphone is used, then the device structure is simple, but the sound quality deteriorates when both driver and passenger speak simultaneously due to inability to suppress crosstalk effectively

Engineering Contradiction:
Improvemicrophone configurationVSAvoidspeech signal clarity
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the speech processing function across multiple microphones and multiple adaptive filters. Each microphone captures a specific spatial zone, and each adaptive filter processes crosstalk from a specific speaker. This segmentation enables effective crosstalk suppression while maintaining clear speech quality for the intended speaker.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges multiple speech collection channels (driver-side channel and passenger-side channel) into a unified speech processing system. By combining the outputs of multiple microphones and coordinating multiple adaptive filters, the system achieves improved speech clarity and crosstalk suppression that neither single channel could achieve alone.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12039993B2Speech processing device and speech processing method
Publication Date: 2024.07.16 PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
  • US12039993B2 patent drawing
  • US12039993B2 patent drawing
  • US12039993B2 patent drawing

AI summary

A speech processing device includes a processor. The processor performs operations including: detecting a single-talk state based on a speech signal collected by each of microphones, the single-talk state in which any one of persons speaks; estimating a mixing rate indicating a ratio of a speech signal of the main speaking person to a speech signal of another person based on a sound pressure ratio of the speech signals collected by the microphones in the single-talk state of the main speaking person and a sound pressure ratio of the speech signals collected by the plurality of microphones in the single-talk state of the another person; and determining whether suppression of a crosstalk component due to speaking of the another person contained in the speech signal of the main speaking person is necessary based on an estimation result of the mixing rate.