Voice Processing System Beamforming and Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice processing systems face challenges in separating speech signals from multiple speakers, especially when speakers are not at fixed positions, leading to poor sound quality and crosstalk, and are not scalable for simultaneous translation in a single room with reduced power efficiency.
Innovation Solution
A voice processing system comprising microphone units and a central unit that use beam forming and metadata generation to isolate and identify individual speakers, reduce crosstalk, and optimize power consumption by selectively streaming data packages based on speech detection, allowing for improved transcription and simultaneous translation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple microphones at fixed positions are used to separate speech signals, then speech separation quality is improved, but the system becomes non-scalable and cannot handle users at moving positions
Solution Approach 1:
The patent implements dynamic beamforming where the beam direction is continuously adapted based on the detected position of the target speaker. The system calculates arrival time differences across microphone arrays and adjusts beamforming weights in real-time to track speakers as they move, transforming the static fixed-position system into a dynamic tracking system that maintains speech separation quality regardless of user position.
Solution Approach 2:
The system employs feedback mechanisms by continuously monitoring the audio signals from multiple microphones, detecting speaker position based on arrival time differences, and using this information to adjust the beamforming control signals. This closed-loop feedback enables the system to adapt to moving speakers and maintain optimal speech separation.
2Measurement precision
If continuous streaming of audio data is implemented, then speech detection accuracy is improved, but power consumption increases
Solution Approach 1:
The patent implements periodic or event-driven streaming where audio data is streamed to the central unit only when speech is detected or at scheduled intervals, rather than continuously. The microphone unit analyzes local audio signals and selectively transmits data packages based on speech activity detection, reducing unnecessary transmissions and power consumption while maintaining speech detection accuracy.
Solution Approach 2:
The microphone unit performs self-service by locally processing audio signals to detect speech presence and autonomously deciding when to stream data to the central unit. This distributed intelligence reduces the need for continuous centralized processing and enables power-efficient selective streaming based on local conditions.
3Object-generated harmful factors
If beam forming is applied to isolate individual speakers, then crosstalk is reduced, but system complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the audio processing into distinct functional blocks: microphone array for signal capture, beamforming processing for spatial separation, metadata generation for speaker identification, and selective streaming for data transmission. This modular segmentation reduces crosstalk through systematic signal separation while managing complexity through organized functional decomposition.
Solution Approach 2:
The system introduces metadata as an intermediary element that carries information about speaker position and identity alongside the audio signal. This metadata acts as a mediator that enables the central unit to correctly associate audio segments with specific speakers, reducing crosstalk effects and enabling accurate speech separation without requiring overly complex processing.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The system achieves improved sound quality for individual speakers, reduces crosstalk, and enhances power efficiency, enabling effective simultaneous translation of multiple conversations in a single room with flexible user positioning.
Implementation Method 1
The unit derives from a group of Y consecutive samples of the source localisation signal a beam form control signal. Under control of the beam form control signal the unit generates a group of Y consecutive samples of a beam formed audio signal from the N microphone signals
Implementation Method 2
The unit determines from the N microphone signals a source localisation signal... a field having a value derived from a group of Y consecutive samples of the source localisation signal
Data Source
Figure 1
Figure 2~3
AI summary
Methods for a voice processing system comprising P microphone units (102A…102D) and a central unit (104) are disclosed. Each microphone unit is linked to a person and derives from N microphone signals a source localisation signal. The source localisation signal is used to control an adaptive beam form process to obtain a beam formed audio signal. The microphone unit is further configured to derive metadata from for N microphone signals, such direction the sound is coming from. Packages with the metadata and beam formed audio signal are transmitted to the central unit. The central unit processes the metadata to determine which parts of the P beam formed audio signal comprises speech from a person that is linked to another microphone unit. By removing said parts from the audio signals before transcription, the quality of the transcription is improved. The transcriptions are displayed on a remote device.