Crosstalk Reduction in Automatic Speech Translation Headsets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automatic speech translation systems face challenges in reducing crosstalk from other voices in noisy environments, particularly in two-way translation scenarios where separating voices with similar frequency domains is difficult.
Innovation Solution
A method and device that utilize a headset with both in-ear and out-ear microphones to receive and process voice signals, employing voice activity detection and pattern matching to isolate and remove crosstalk from the signal received by the out-ear microphone, allowing for improved noise reduction and accurate speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition is performed in noisy environments with multiple speakers, then automatic speech translation can be provided in conference rooms and on streets, but crosstalk from other people's voices degrades recognition accuracy
Solution Approach 1:
The patent segments the audio signal into different frequency bands using filter banks. By dividing the frequency spectrum into multiple bands and processing each band separately, the system can identify and remove crosstalk components more effectively while preserving the target speaker's voice, thus maintaining recognition accuracy in noisy environments
Solution Approach 2:
The patent introduces an auxiliary microphone as an intermediary device that captures crosstalk signals from other speakers. This auxiliary signal serves as a reference that is combined with the main microphone signal through adaptive filtering to cancel out crosstalk components, thereby improving speech recognition accuracy in conference rooms and street environments
2Device complexity
If conventional speech recognition processes all frequency components equally, then processing is simpler, but crosstalk and environmental noise cannot be effectively separated when they have similar frequency domains
Solution Approach 1:
The patent applies filter banks to segment the audio signal into multiple frequency bands. This segmentation allows the system to analyze and process different frequency components separately, enabling effective separation of crosstalk from environmental noise even when they overlap in certain bands, thereby improving reliability without excessive complexity
Solution Approach 2:
The patent applies different processing strategies to different frequency bands based on local characteristics. By adapting the noise reduction and crosstalk removal parameters to each frequency band's specific conditions, the system achieves better noise separation capability while maintaining reasonable processing complexity
Data Source
AI summary
Disclosed are a method, a device, and a computer-readable storage medium for reducing crosstalk when performing automatic speech translation between at least two users speaking different languages. The method for reducing crosstalk includes receiving a signal inputted to an out-ear microphone of a first user, wherein the first user is wearing a headset equipped with an in-ear microphone and the out-ear microphone and the signal includes a voice signal A of the first user and a voice signal b of a second user, receiving a voice signal Binear inputted to an in-ear microphone of the second user, wherein the second user is wearing a headset equipped with the in-ear microphone and an out-ear microphone, and removing the voice signal b of the second user from the signal A+b inputted to the out-ear microphone of the first user, based on the voice signal Binear inputted to the in-ear microphone of the second user.


