Signal Processing Device for Speech Removal via Similarity Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional technologies are ineffective in removing speech signals from acoustic signals under Conditions 3 and 4, where background sounds are equal or unequal between channels, leading to difficulties in distinguishing and separating speech from monaural signals.
Innovation Solution
A signal processing device that acquires and processes acoustic signals from two channels, calculates a background sound signal by removing speech, generates a reference signal, calculates similarity between feature data of the background sound signals, and computes a weighted sum based on this similarity to enhance the separation of background sounds, effectively addressing the limitations of existing technologies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If conventional speech removal technology is applied to acoustic signals of Conditions 3 and 4, then the processing complexity is reduced, but the speech removal effectiveness deteriorates significantly
Solution Approach 1:
The patent dynamically adapts the speech removal processing based on the input signal characteristics. The system automatically determines whether the input signal is a stereo signal or monaural signal and adjusts the processing method accordingly. For monaural signals (Conditions 3 and 4), the system uses a different processing path that involves generating artificial stereo separation, while for stereo signals (Conditions 1 and 2), it uses conventional processing. This dynamic adaptation resolves the contradiction by maintaining speech removal effectiveness across different signal types without significantly increasing overall system complexity.
Solution Approach 2:
The patent changes the processing parameters based on the input signal type. When detecting monaural signals, the system changes the processing approach by using magnitude and phase extraction, generating artificial left and right channel signals through specific mathematical operations, and applying weighted summation based on similarity calculations. This parameter change allows effective speech removal from monaural signals while keeping the system architecture relatively simple.
2Reliability
If conventional speech removal technology is used for stereo signals, then the speech removal effectiveness is maintained, but the adaptability to different signal conditions deteriorates
Solution Approach 1:
The patent creates a universal speech removal system that can handle multiple signal types (stereo and monaural) through a unified architecture. The system includes a signal type determination module that automatically identifies whether the input is stereo or monaural, and then routes to appropriate processing paths. This multi-functionality allows the same system to effectively process both stereo signals (maintaining conventional effectiveness) and monaural signals (achieving new capability), thereby resolving the contradiction between maintaining effectiveness and improving adaptability.
Solution Approach 2:
The patent segments the speech removal process into distinct modules: signal type determination, magnitude/phase extraction, artificial stereo generation (for monaural inputs), similarity calculation, and weighted summation. This segmentation allows the system to apply different processing strategies for different signal types while maintaining a unified overall structure, thus achieving both effectiveness for stereo signals and adaptability to monaural signals.
3Device complexity
If speech signals are removed from monaural signals using conventional methods, then the processing simplicity is maintained, but the background sound extraction quality deteriorates
Solution Approach 1:
The patent introduces intermediate processing steps for monaural signals: magnitude extraction, phase extraction, and artificial stereo signal generation. These intermediaries transform the monaural signal into a form that can be processed using speech removal techniques. The system then uses similarity calculation as another intermediary to determine the appropriate weighting for combining the processed signals. This chain of intermediaries improves background sound extraction quality from monaural signals while keeping the overall processing approach relatively simple and systematic.
Data Source
AI summary
According to an embodiment, a signal processing device includes a background calculator, a signal generator, an extractor, a similarity calculator, and a mixer. The background calculator is configured to calculate a first background signal in which a speech signal is removed, based on the acoustic signals. The signal generator is configured to generate a reference signal from at least one of the acoustic signals. The extractor is configured to extract a second background signal by removing a speech signal from the reference signal. The similarity calculator is configured to calculate a similarity between feature data of the background signals. The mixer is configured to calculate a weighted sum of the background signals in such a way that a greater weight is given to the first background signal as the similarity is higher and a greater weight is given to the second background signal as the similarity is lower.


