Image Sound Pickup Device Audio Video Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound pickup systems require high directionality to identify sound source locations, which is challenging in large rooms with many participants, and they rely solely on audio data, lacking effective methods to enhance specific sound sources.
Innovation Solution
An image and sound pickup device that combines audio signal phase differences with face recognition technology to generate location information and display estimated utterer images, allowing for the enhancement of specific sound sources by controlling audio output based on user selection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sound source location is calculated based only on audio data, then directionality requirement becomes very high, but this makes it difficult to identify sound sources in large rooms with many participants
Solution Approach 1:
The patent combines audio data analysis with video image analysis to identify sound source locations. The sound source location calculation unit integrates phase difference information from multiple microphones with face recognition results from video images, allowing accurate sound source identification without requiring high directionality from the microphone array alone.
Solution Approach 2:
The patent introduces video image data as an intermediary to bridge the gap between audio signals and sound source location identification. By using face recognition on video frames, the system can visually identify participants and correlate their positions with audio phase differences, enabling accurate localization without high directionality requirements.
2Measurement precision
If high directionality is required for sound pickup, then sound source location can be identified accurately, but this becomes challenging in large rooms with many participants
Solution Approach 1:
The system merges audio phase difference calculation with video face recognition to identify sound sources. This combination allows the system to handle large rooms with many participants effectively, as the video component provides spatial context that complements the audio data, removing the need for high directionality that would be required in purely audio-based systems.
Solution Approach 2:
The patent adds the visual dimension by incorporating video image analysis into the sound source identification process. By using face recognition on video frames, the system creates a two-dimensional visual reference that complements the audio phase difference information, enabling effective sound source localization in complex environments without requiring high directional sensitivity.
3Device complexity
If only audio data is used for sound pickup, then system complexity is reduced, but the ability to enhance specific sound sources is limited
Solution Approach 1:
The patent combines audio data processing with video image processing to enable sound source enhancement. By integrating face recognition results with audio phase difference information, the system can identify and enhance specific participants' voices in the output, providing selective sound enhancement capability while maintaining reasonable system complexity through unified processing architecture.
Data Source
AI summary
Provided is a method of controlling an image and sound pickup device, which is includes obtaining a plurality of audio signals and a participant image, which shows a plurality of participants, and generating location information about a sound source location by using comparison information about a comparison among the plurality of audio signals and face recognition that is performed on the participant image; and generating an estimated utterer image, which displays an estimated utterer, by using the location information.


