Voice Capturing System Buffer Fusion for Direction Change Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice capturing systems in electrical devices face delays and distortions when the direction of a target speaker changes, as they require collecting at least N seconds of voice data to obtain sufficient speaker information for processing.
Innovation Solution
A voice capturing method and system that store voice data from multiple microphones, determine the presence and direction of a target speaker, insert a voice segment from the previous tracking direction into the current position in the data to generate fusion voice data, perform voice enhancement based on the current direction, and shorten the enhanced data to eliminate delays.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the voice capturing system waits to collect N seconds of voice data before processing, then sufficient target speaker information can be obtained for spatial enhancement, but voice data becomes delayed and distorted when the target speaker direction changes
Solution Approach 1:
The system performs preliminary voice data collection and stores it in a buffer before processing is needed. When direction change is detected, the pre-collected voice data is immediately available for fusion processing, eliminating the N-second waiting delay while ensuring sufficient data is available for accurate spatial enhancement
Solution Approach 2:
The voice data processing is divided into segments: pre-collection segment stored in buffer, fusion segment combining previous and current direction data, and output segment after enhancement. This segmentation allows the system to process direction changes immediately by working with segmented data portions rather than waiting for complete N-second intervals
2Reliability
If the voice capturing system waits N seconds to collect sufficient voice data, then adequate information is available for spatial enhancement, but the voice data becomes distorted when target speaker direction changes
Solution Approach 1:
Voice data is preliminarily collected and stored in a buffer before processing, ensuring sufficient data is available when needed. This preliminary action maintains data quality by avoiding delayed collection that would capture direction changes, while still providing adequate data quantity for reliable spatial enhancement processing
Solution Approach 2:
The system dynamically adjusts the voice data processing approach based on direction change detection. When direction changes are detected, the system switches to fusion processing that combines pre-collected data with current data, maintaining reliability by adapting to dynamic speaker movement rather than relying on fixed N-second collection intervals
Data Source
AI summary
A voice capturing method includes following operations: storing, by a buffer, voice data from a plurality of microphones; determining, by a processor, whether a target speaker exists and whether a direction of the target speaker changes according to the voice data and target speaker information; inserting a voice segment corresponding to a previous tracking direction into a current position in the voice data to generate fusion voice data when the target speaker exists and the direction of the target speaker changes from the previous tracking direction to a current tracking direction; performing, by the processor, a voice enhancement process on the fusion voice data according to the current tracking direction to generate enhanced voice data; performing, by the processor, a voice shortening process on the enhanced voice data to generate voice output data; and playing, by a playing circuit, the voice output data.


