Video Conference System Sound Source Tracking Sub-Image Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video conference devices face challenges in efficiently tracking sound sources and displaying corresponding close-up images with minimal computation, especially in scenarios with multiple conference members, leading to high processor resource usage.
Innovation Solution
A video conference system comprising an image detection device, a sound source detection device, and a processor that outputs a positioning signal when a sound source is detected, allowing the processor to determine if a real face image exists in a sub-image block of the conference image and display a close-up image accordingly, reducing computation by analyzing only specific sub-image blocks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the processor performs image analysis on the entire captured conference image to determine the location of a close-up face, then the accuracy of sound source tracking is improved, but the computation time and processor resource consumption increase significantly
Solution Approach 1:
The patent divides the conference image into multiple sub-images based on the positioning signal from the sound source detection device. Instead of analyzing the entire image, the processor only analyzes the specific sub-image where the sound source is located, thereby reducing computation time while maintaining tracking accuracy.
Solution Approach 2:
The patent applies local quality by focusing the image analysis only on the specific region (sub-image) corresponding to the sound source location. This allows the system to concentrate computational resources on the most relevant area rather than processing the entire image uniformly.
2Measurement precision
If the processor performs image analysis on the entire captured conference image to determine the location of a close-up face, then the accuracy of sound source tracking is improved, but the processor resource consumption increases significantly
Solution Approach 1:
The patent divides the conference image into multiple sub-images based on the positioning signal from the sound source detection device. Instead of analyzing the entire image, the processor only analyzes the specific sub-image where the sound source is located, thereby reducing computation time while maintaining tracking accuracy.
Solution Approach 2:
The patent applies local quality by focusing the image analysis only on the specific region (sub-image) corresponding to the sound source location. This allows the system to concentrate computational resources on the most relevant area rather than processing the entire image uniformly.
3Productivity
If the system automatically tracks sound sources and displays corresponding close-up images, then the video conference effectiveness is improved, but the device complexity increases
Solution Approach 1:
The patent integrates multiple functions into a unified system: the image detection device captures images, the sound source detection device identifies sound locations, and the processor coordinates both to automatically generate close-up conference images. This multi-functional integration improves video conference effectiveness while managing system complexity through coordinated operation of existing components.
Data Source
AI summary
A video conference system, a video conference apparatus and a video conference method are provided. The video conference system includes a video conference apparatus and a display apparatus. The video conference apparatus includes an image detection device, a sound source detection device, and a processor. The image detection device obtains a conference image of a conference space. When the sound source detection device detects a sound generated by a sound source in the conference space, the sound source detection device outputs a positioning signal. The processor receives the positioning signal, and determines whether a real face image exists in a sub-image block of the conference image corresponding to the sound source according to the positioning signal to output the image signal. The display apparatus displays a close-up conference image including the real face image according to the image signal.


