Video Capture Device Sound Source Localization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for controlling video capture devices during videoconferencing or online teaching suffer from low accuracy in locating sound sources distant from the device, leading to inaccurate device control due to the installation of microphones on the device itself.
Innovation Solution
The method involves using microphones disposed outside the video capture device to capture voice signals and calculate voice power values based on coordinate information, identifying the position with the maximum voice power value as the sound source, and adjusting the device's focal length accordingly to accurately direct at the sound source.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If the microphone is installed on the video capture device, then the device structure is simple and easy to implement, but the sound source localization accuracy deteriorates for distant sound sources
Solution Approach 1:
The system divides the sound source localization function from the video capture device by using separate microphones positioned around the scene. These microphones are spatially segmented to create a microphone array that covers a wider area, enabling accurate localization of distant sound sources while the video capture device focuses on visual capture.
Solution Approach 2:
The patent introduces an intermediate processing system that receives signals from multiple microphones, calculates voice power values for different preset positions based on coordinate information, and determines the sound source location. This intermediary processing layer enables accurate distant sound source localization without requiring the video capture device itself to have integrated microphones.
2Device complexity
If the microphone is installed on the video capture device, then the integration is high and the system is compact, but the coverage area for sound capture is limited
Solution Approach 1:
The system transitions from a single-point microphone installation on the video capture device to a distributed microphone array positioned around the scene. By adding spatial dimensionality with microphones at multiple locations, the coverage area for sound capture is significantly expanded while maintaining manageable system complexity through modular deployment.
Solution Approach 2:
The separate microphone array system serves multiple functions: it captures sound from the entire scene, enables accurate sound source localization, and provides voice power value calculations for preset positions. This multi-functional approach expands coverage area while the video capture device maintains its primary visual capture function.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach significantly improves the accuracy of sound source localization and device control by allowing the capture of sound from the entire scene, enabling precise targeting of the speaker regardless of distance.
Implementation Method 1
acquiring a voice signal captured by each of the microphones
Implementation Method 2
adjusting a focal length of the video capture device to the target focal length
Data Source
Figure 1~2
Figure 2(b)~3
Figure 4~5
AI summary
Embodiments of the present application provide a method, apparatus and system for controlling a device. The method is applicable to a video capture device (610) in a sound source localization system. The sound source localization system further includes microphones (620) disposed outside the video capture device (610). The method includes: acquiring a voice signal captured by each of the microphones (620), and acquiring coordinate information of each of the microphones (620) and coordinate information of each of preset positions (S101), calculating a voice power value corresponding to each of the preset positions based on the voice signal captured by each of the microphones (620), the coordinate information of each of the microphones (620) and the coordinate information of each of the preset positions (S102), identifying a preset position with a maximum voice power value, determining the preset position as a position of the sound source, and controlling the video capture device to direct at the position of the sound source (S103). In the present application, microphones (620) are disposed outside the video capture device (610), and the sound in the scene to be captured can all be captured by the microphones (620). Therefore, the accuracy of the sound source localization can be improved, thereby improving the accuracy of the device control.