Voice Extraction Unit Channel Shift for Microphone Array Position Changes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

When the positions of microphones in a microphone array change, the learned noise characteristic information becomes obsolete, leading to deteriorated voice extraction performance, as sufficient learning time cannot be secured, and noise suppression performance is affected.

Innovation Solution

The signal processing device employs a channel shift method, where the signals from the microphones are treated as signals from other microphones, allowing voice extraction to continue without interruption, even when the microphone positions change, by using noise characteristics learned previously and adjusting the degree of reflection based on time and positional errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the positions of microphones are changed to face the speech direction, then the device can track and focus on the speaker, but the learned noise characteristic information becomes obsolete and voice extraction performance deteriorates

Engineering Contradiction:
Improveability to track speech directionVSAvoidvoice extraction performance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system performs preliminary noise characteristic learning during periods when microphones are stationary and noise characteristics are stable. By accumulating noise data in advance during stable periods, the system prepares noise suppression models before position changes occur, ensuring that when the device rotates to track speech, the pre-learned noise characteristics can still be applied without immediate relearning delays.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts the noise suppression processing based on the current operational state. When microphones are stationary, the system emphasizes using learned noise characteristics; when position changes occur, the system transitions to relying more on spatial processing and beamforming. This dynamic adaptation allows the system to maintain voice extraction performance across different operational phases.

Inventive Principle:
Principle #15Dynamics

2Speed

If the device rotates to face the speech direction immediately after detecting speech, then the response time is reduced, but sufficient learning time cannot be secured and noise suppression performance deteriorates

Engineering Contradiction:
Improveresponse speed to speechVSAvoidnoise characteristic learning time
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The system continuously performs noise characteristic learning in the background during stable operational periods, accumulating noise data before speech detection occurs. This preliminary learning ensures that when speech is detected and the device needs to rotate quickly, the noise suppression models are already prepared and do not require immediate relearning, thus preserving both fast response and adequate learning time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The noise characteristic learning process operates continuously during stable periods rather than being interrupted by speech detection events. This continuous learning approach ensures that noise characteristics are constantly updated and refined over time, maintaining noise suppression performance while allowing the system to respond quickly to speech events without sacrificing learning opportunities.

Inventive Principle:
Principle #20Continuity of useful action

3Reliability

If noise characteristic learning is performed continuously, then the noise suppression performance is improved, but when microphone positions change, the learned information becomes inaccurate and must be relearned

Engineering Contradiction:
Improvenoise suppression performanceVSAvoidadaptability to position changes
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system dynamically switches between two operational modes: during stable periods, it performs continuous noise characteristic learning to improve noise suppression; when position changes are detected, it transitions to a mode that relies on spatial processing and beamforming while suspending noise characteristic updates. This dynamic mode switching allows the system to optimize for noise suppression during stable operation while adapting quickly to position changes without carrying forward obsolete learned information.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11386904B2Signal processing device, signal processing method, and program
Publication Date: 2022.07.12 SONY GROUP CORP
  • US11386904B2 patent drawing
  • US11386904B2 patent drawing
  • US11386904B2 patent drawing

AI summary

Deterioration of voice extraction performance when positions of a plurality of microphones are changed is prevented.A signal processing device according to an embodiment of the present technology includes a voice extraction unit that performs voice extraction from signals of a plurality of microphones, in which the voice extraction unit uses, when respective positions of the plurality of microphones are changed to positions where other microphones have been present, respective signals of the plurality of microphones as signals of the other microphones. Thus, it is possible to cancel the effect of changing the positions of respective microphones on the voice extraction.