Video Mask Region Calculation Using Multi-Frame Detection Scores
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing systems fail to effectively utilize detection scores for calculating mask regions, particularly when face recognition does not succeed in all frames of a video, leading to incomplete or inaccurate masking.
Innovation Solution
An image processing system that calculates mask regions by associating detection boundaries and scores across multiple frames, adjusting mask regions based on detection stability and movement of the target, using techniques like deep learning for face detection and considering detection scores to determine appropriate masking.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If face detection is applied to each frame independently, then detection speed is maintained, but detection accuracy deteriorates when detection fails in certain frames
Solution Approach 1:
The system performs detection in advance on multiple frames including future frames, and stores detection results for later use. When detection fails in the current frame, the pre-detected results from other frames are utilized to maintain detection accuracy without requiring re-detection.
Solution Approach 2:
Detection scores serve as an intermediary metric to evaluate detection reliability. The system uses these scores to determine whether to trust detection results from specific frames and to selectively combine results from multiple frames, mediating between conflicting detection outcomes.
2Reliability
If mask region is calculated only from current frame detection, then processing time is reduced, but mask accuracy deteriorates when detection fails
Solution Approach 1:
Detection results and mask regions are calculated in advance for multiple frames and stored. When processing the current frame, the system retrieves pre-calculated results from other frames, avoiding redundant detection computations and reducing processing time while maintaining mask reliability.
Solution Approach 2:
The system merges detection results from multiple frames by combining detection boundaries and scores. When detection succeeds in multiple frames, the results are integrated to produce a more reliable mask region that compensates for failures in individual frames.
3Measurement precision
If detection score threshold is set high, then detection quality is improved, but detection coverage deteriorates due to more failures
Solution Approach 1:
The system combines detection results from multiple frames, allowing high thresholds to be maintained for quality while compensating for coverage losses through temporal aggregation. Frames that pass the threshold are merged with others to ensure continuous detection coverage.
Solution Approach 2:
The system uses detection scores as feedback to dynamically adjust processing. When detection scores are high, results are accepted; when scores are low or detection fails, the system retrieves alternative results from other frames, creating a feedback loop that maintains both quality and coverage.
4Reliability
If mask region is expanded to cover detection uncertainty, then masking reliability is improved, but unnecessary masking increases
Solution Approach 1:
The system applies different masking strategies to different regions based on local detection confidence. High-confidence regions receive precise masking only, while low-confidence or overlapping regions from multiple frames receive expanded masking to ensure reliability, avoiding unnecessary masking in clearly detected areas.
Solution Approach 2:
The mask region dynamically adjusts based on detection scores and temporal consistency. When detection is consistent across frames, the mask is precise; when detection varies or fails, the mask expands to cover uncertainty, creating a dynamic adaptation between precision and reliability.
Data Source
Figure 1
Figure 2A~2C
Figure 3
AI summary
To provide an image processing system that appropriately calculates a mask region. An image processing system is provided that includes a mask region calculation unit, wherein the mask region calculation unit acquires a detection result of an image detection process for an image of a frame in a video, and the mask region calculation unit calculates a mask region for an image of a first frame based on at least the detection result of an image of a second frame. By the image detection process, a detection boundary or a detection region is associated with at least the image of the second frame, and a detection score is set for the detection boundary or the detection region. The detection result includes the detection boundary or the detection region and the detection score, and the second frame is a frame at a time point different from the first frame on the time axis of the video.