Voice Processing Device Switching Input and Separation Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conferencing systems face challenges in accurately recognizing multiple speakers' voices due to wraparound speech and signal distortion during sound source separation, which affects the accuracy of voice recognition.

Innovation Solution

A voice processing device that receives input signals from multiple microphones, separates them by sound sources, and switches between input and separation signals based on the number of active sound sources to suppress wraparound speech and avoid signal distortion, thereby enhancing voice recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If sound source separation is performed to recognize multiple speakers' voices, then voice recognition capability is improved, but wraparound speech and signal distortion occur which worsen recognition accuracy

Engineering Contradiction:
Improvevoice recognition capabilityVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system dynamically switches between two operational modes: using separation signals when multiple sound sources are detected, and using original input signals when a single sound source is detected. This dynamic adaptation resolves the contradiction by selecting the optimal signal processing approach based on real-time acoustic conditions, thereby maintaining high recognition accuracy while preserving the capability to handle multiple speakers.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter of signal selection based on the number of detected sound sources. When the number of sound sources changes from one to multiple, the system transitions from using original input signals to using separation signals. This parameter-based switching resolves the technical contradiction by adapting the signal processing strategy to the specific acoustic scenario.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If separation signals are always used to handle multiple speakers, then speaker differentiation is improved, but signal distortion occurs which worsens overall recognition quality

Engineering Contradiction:
Improvespeaker differentiation accuracyVSAvoidsignal quality
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The system employs a conditional approach where separation processing is applied only when necessary (multiple speakers detected), rather than continuously. This selective application avoids the signal distortion that would result from always applying separation processing, while still achieving speaker differentiation when needed. The separation processing is essentially used 'short-lived' only in the specific condition when multiple speakers are present.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Solution Approach 2:

The system dynamically adjusts signal processing based on the number of active speakers. When only one speaker is detected, the original input signal is used to avoid distortion. When multiple speakers are detected, the system switches to using separation signals to achieve proper speaker differentiation. This dynamic switching resolves the contradiction between speaker differentiation and signal quality.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If the system switches between input and separation signals based on sound source count, then recognition accuracy is improved, but system complexity increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The voice processing device is designed to perform multiple functions: it can process both single-speaker and multi-speaker scenarios using a unified architecture. The switching mechanism between input signals and separation signals is integrated into the existing signal processing pipeline, allowing the system to maintain high recognition accuracy across different scenarios without requiring entirely separate processing systems, thus limiting the increase in complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The sound source separation processing is performed in advance on all input signals, creating separation signals that are then selectively used based on the number of detected speakers. This preliminary action allows the switching mechanism to simply select between pre-computed signals rather than performing complex real-time separation only when needed, thereby reducing the computational complexity of the switching decision itself.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10504523B2Voice processing device, voice processing method, and computer program product
Publication Date: 2019.12.10 KK TOSHIBA
  • US10504523B2 patent drawing
  • US10504523B2 patent drawing
  • US10504523B2 patent drawing

AI summary

According to an embodiment, a voice processing device includes a receiver, a separator, and an output controller. The receiver is configured to receive n input signals input into n voice input devices respectively corresponding to n sound sources, where n is an integer of 2 or more. The separator is configured to separate the input signals by the sound sources to produce n separation signals. The output controller is configured to, according to the number of sound sources having uttered voice sounds, switch between an output signal produced based on the input signal and an output signal produced based on the separation signal, and output the output signal.