Video Mask Region Calculation Using Multi-Frame Detection Scores

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing systems fail to effectively utilize detection scores for calculating mask regions, particularly when face recognition does not succeed in all frames of a video, leading to incomplete or inaccurate masking.

Innovation Solution

An image processing system that calculates mask regions by associating detection boundaries and scores across multiple frames, adjusting mask regions based on detection stability and movement of the target, using techniques like deep learning for face detection and considering detection scores to determine appropriate masking.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If face detection is applied to each frame independently, then detection speed is maintained, but detection accuracy deteriorates when detection fails in certain frames

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs detection in advance on multiple frames including future frames, and stores detection results for later use. When detection fails in the current frame, the pre-detected results from other frames are utilized to maintain detection accuracy without requiring re-detection.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Detection scores serve as an intermediary metric to evaluate detection reliability. The system uses these scores to determine whether to trust detection results from specific frames and to selectively combine results from multiple frames, mediating between conflicting detection outcomes.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If mask region is calculated only from current frame detection, then processing time is reduced, but mask accuracy deteriorates when detection fails

Engineering Contradiction:
Improvemask reliabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

Detection results and mask regions are calculated in advance for multiple frames and stored. When processing the current frame, the system retrieves pre-calculated results from other frames, avoiding redundant detection computations and reducing processing time while maintaining mask reliability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system merges detection results from multiple frames by combining detection boundaries and scores. When detection succeeds in multiple frames, the results are integrated to produce a more reliable mask region that compensates for failures in individual frames.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If detection score threshold is set high, then detection quality is improved, but detection coverage deteriorates due to more failures

Engineering Contradiction:
Improvedetection qualityVSAvoiddetection coverage
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system combines detection results from multiple frames, allowing high thresholds to be maintained for quality while compensating for coverage losses through temporal aggregation. Frames that pass the threshold are merged with others to ensure continuous detection coverage.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses detection scores as feedback to dynamically adjust processing. When detection scores are high, results are accepted; when scores are low or detection fails, the system retrieves alternative results from other frames, creating a feedback loop that maintains both quality and coverage.

Inventive Principle:
Principle #23Feedback

4Reliability

If mask region is expanded to cover detection uncertainty, then masking reliability is improved, but unnecessary masking increases

Engineering Contradiction:
Improvemasking reliabilityVSAvoidunnecessary masking
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The system applies different masking strategies to different regions based on local detection confidence. High-confidence regions receive precise masking only, while low-confidence or overlapping regions from multiple frames receive expanded masking to ensure reliability, avoiding unnecessary masking in clearly detected areas.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The mask region dynamically adjusts based on detection scores and temporal consistency. When detection is consistent across frames, the mask is precise; when detection varies or fails, the mask expands to cover uncertainty, creating a dynamic adaptation between precision and reliability.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP4730760A1Image processing system, program, and image processing method
Publication Date: 2026.04.22 EIZO CORP
  • EP4730760A1 patent drawingFigure 1
  • EP4730760A1 patent drawingFigure 2A~2C
  • EP4730760A1 patent drawingFigure 3

AI summary

To provide an image processing system that appropriately calculates a mask region. An image processing system is provided that includes a mask region calculation unit, wherein the mask region calculation unit acquires a detection result of an image detection process for an image of a frame in a video, and the mask region calculation unit calculates a mask region for an image of a first frame based on at least the detection result of an image of a second frame. By the image detection process, a detection boundary or a detection region is associated with at least the image of the second frame, and a detection score is set for the detection boundary or the detection region. The detection result includes the detection boundary or the detection region and the detection score, and the second frame is a frame at a time point different from the first frame on the time axis of the video.