Signal Processing Device for Speech Removal via Similarity Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional technologies are ineffective in removing speech signals from acoustic signals under Conditions 3 and 4, where background sounds are equal or unequal between channels, leading to difficulties in distinguishing and separating speech from monaural signals.

Innovation Solution

A signal processing device that acquires and processes acoustic signals from two channels, calculates a background sound signal by removing speech, generates a reference signal, calculates similarity between feature data of the background sound signals, and computes a weighted sum based on this similarity to enhance the separation of background sounds, effectively addressing the limitations of existing technologies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If conventional speech removal technology is applied to acoustic signals of Conditions 3 and 4, then the processing complexity is reduced, but the speech removal effectiveness deteriorates significantly

Engineering Contradiction:
Improveprocessing complexityVSAvoidspeech removal effectiveness
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent dynamically adapts the speech removal processing based on the input signal characteristics. The system automatically determines whether the input signal is a stereo signal or monaural signal and adjusts the processing method accordingly. For monaural signals (Conditions 3 and 4), the system uses a different processing path that involves generating artificial stereo separation, while for stereo signals (Conditions 1 and 2), it uses conventional processing. This dynamic adaptation resolves the contradiction by maintaining speech removal effectiveness across different signal types without significantly increasing overall system complexity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the processing parameters based on the input signal type. When detecting monaural signals, the system changes the processing approach by using magnitude and phase extraction, generating artificial left and right channel signals through specific mathematical operations, and applying weighted summation based on similarity calculations. This parameter change allows effective speech removal from monaural signals while keeping the system architecture relatively simple.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If conventional speech removal technology is used for stereo signals, then the speech removal effectiveness is maintained, but the adaptability to different signal conditions deteriorates

Engineering Contradiction:
Improvespeech removal effectivenessVSAvoidadaptability to different signal conditions
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal speech removal system that can handle multiple signal types (stereo and monaural) through a unified architecture. The system includes a signal type determination module that automatically identifies whether the input is stereo or monaural, and then routes to appropriate processing paths. This multi-functionality allows the same system to effectively process both stereo signals (maintaining conventional effectiveness) and monaural signals (achieving new capability), thereby resolving the contradiction between maintaining effectiveness and improving adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent segments the speech removal process into distinct modules: signal type determination, magnitude/phase extraction, artificial stereo generation (for monaural inputs), similarity calculation, and weighted summation. This segmentation allows the system to apply different processing strategies for different signal types while maintaining a unified overall structure, thus achieving both effectiveness for stereo signals and adaptability to monaural signals.

Inventive Principle:
Principle #1Segmentation

3Device complexity

If speech signals are removed from monaural signals using conventional methods, then the processing simplicity is maintained, but the background sound extraction quality deteriorates

Engineering Contradiction:
Improveprocessing simplicityVSAvoidbackground sound extraction quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent introduces intermediate processing steps for monaural signals: magnitude extraction, phase extraction, and artificial stereo signal generation. These intermediaries transform the monaural signal into a form that can be processed using speech removal techniques. The system then uses similarity calculation as another intermediary to determine the appropriate weighting for combining the processed signals. This chain of intermediaries improves background sound extraction quality from monaural signals while keeping the overall processing approach relatively simple and systematic.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9412391B2Signal processing device, signal processing method, and computer program product
Publication Date: 2016.08.09 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9412391B2 patent drawing
  • US9412391B2 patent drawing
  • US9412391B2 patent drawing

AI summary

According to an embodiment, a signal processing device includes a background calculator, a signal generator, an extractor, a similarity calculator, and a mixer. The background calculator is configured to calculate a first background signal in which a speech signal is removed, based on the acoustic signals. The signal generator is configured to generate a reference signal from at least one of the acoustic signals. The extractor is configured to extract a second background signal by removing a speech signal from the reference signal. The similarity calculator is configured to calculate a similarity between feature data of the background signals. The mixer is configured to calculate a weighted sum of the background signals in such a way that a greater weight is given to the first background signal as the similarity is higher and a greater weight is given to the second background signal as the similarity is lower.