Automatic Audio Focusing on Multiple Objects via Microphone Array
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current electronic devices with microphones for video capture can only perform audio focusing on a single object manually selected by the user, limiting the quality of audio capture in video shooting.
Innovation Solution
An electronic device equipped with a camera, microphone array, and processor that automatically identifies and focuses on objects of interest by inferring their importance through image and voice feature analysis, adjusting microphone activity to emphasize the desired audio.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If audio focusing is performed on only one object selected manually, then the device complexity is reduced and ease of operation is improved, but the audio quality for multiple objects of interest deteriorates
Solution Approach 1:
The system automatically identifies and prioritizes multiple objects of interest without requiring manual user selection. The processor analyzes video frames to detect objects and determines their importance levels autonomously, enabling the device to serve itself in selecting which objects to focus on audio-wise, thereby improving audio quality for multiple objects while maintaining ease of operation.
Solution Approach 2:
The system performs preliminary object detection and importance assessment before audio focusing is applied. By pre-identifying multiple objects of interest and their priority levels in advance, the system prepares the audio focusing parameters ahead of time, allowing seamless multi-object audio focusing without requiring real-time manual intervention.
2Reliability
If automatic audio focusing on multiple objects is implemented, then the audio quality for multiple objects of interest is improved, but the device complexity increases
Solution Approach 1:
The processor performs multiple functions using the same hardware resources: it detects objects in video frames, determines their importance levels, and controls audio focusing parameters all in one integrated process. This multi-functionality approach allows the system to achieve complex automatic multi-object audio focusing without proportionally increasing hardware complexity, as the processor handles detection, analysis, and control universally.
Solution Approach 2:
The system segments the audio focusing task by assigning different importance levels to different objects and applying weighted audio processing to each segment. By dividing the audio stream into object-specific channels with different priority weights, the complex task of multi-object audio focusing is broken down into manageable segments that can be processed independently and then combined, reducing overall system complexity.
3Loss of time
If manual selection of single object for audio focusing is used, then the processing time is reduced, but the productivity of video capturing is limited
Solution Approach 1:
The system performs object detection and importance assessment at periodic intervals during video capture, rather than continuously or manually. This periodic automatic updating of object priorities allows the system to maintain high audio quality for multiple objects while keeping processing time manageable through rhythmical, scheduled analysis cycles.
Data Source
AI summary
The present disclosure relates to a device and method of providing automatic audio focusing, the method includes: registering objects of interest; capturing a video; displaying the video on a display; recognizing at least one object included in the video; inferring at least one object of interest included in the video from the recognized at least one object; identifying distribution of the at least one object of interest in the video; and performing audio focusing on the at least one object of interest by adjusting activity of each of multiple microphones included in a microphone array on the basis of the distribution of the at least one object of interest in the video, whereby it is possible to emphasize voice of the object of interest during the video capturing of the electronic device, thereby improving the satisfaction with the video capturing result.


