Video Inpainting with Blur Mask Refinement for Complex Backgrounds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video inpainting technologies struggle to maintain accuracy and image quality when dealing with complex backgrounds and object shielding, as optical flow-based methods fail in complex movements and neural network models suffer from limited generation capabilities and blurring issues.

Innovation Solution

An image processing method that decomposes inpainting into three phases: mask processing, morphological processing on blurred regions, and inpainting based on different models to enhance image quality, using information propagation, image, and object inpainting models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If optical flow-based video padding is used, then processing speed is maintained, but image quality deteriorates in complex background movements and object shielding scenarios

Engineering Contradiction:
Improveprocessing speedVSAvoidimage quality
Core Design Contradiction:
SpeedVSManufacturing precision

Solution Approach 1:

The patent segments the video processing task into two distinct phases: optical flow-based padding for speed and neural network-based refinement for quality. The first phase uses optical flow to quickly pad masked regions, while the second phase applies a neural network model specifically to regions where image quality deteriorates, thus resolving the contradiction between speed and quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing qualities to different regions of the video frame. High-quality neural network processing is applied only to masked regions and regions adjacent to them where quality deterioration occurs, while the rest of the video maintains the faster optical flow processing quality, optimizing the balance between speed and quality.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If neural network model-based video padding is used, then image quality improves in complex movements, but generation capability is limited causing blurring in complex textures and object shielding

Engineering Contradiction:
Improveimage qualityVSAvoidgeneration capability
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent combines two different processing approaches (optical flow and neural network) into a composite processing system. The optical flow component handles motion compensation efficiently, while the neural network component handles complex texture generation and object reconstruction, creating a composite solution that leverages the strengths of both methods to overcome individual limitations.

Inventive Principle:
Principle #40Composite materials

3Device complexity

If single-model video padding is used, then processing simplicity is maintained, but image quality deteriorates when both complex movement and complex texture are present

Engineering Contradiction:
Improveprocessing simplicityVSAvoidimage quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent implements a dynamic processing system that adaptively switches between or combines different processing models based on the characteristics of the video content. The system dynamically determines whether to use optical flow, neural network, or both, depending on the complexity of movement and texture in different regions, thus maintaining simplicity when possible while ensuring quality when needed.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP4425423B1Image processing method and apparatus, device, storage medium and program product
Publication Date: 2026.01.28 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • EP4425423B1 patent drawingFigure 1~3
  • EP4425423B1 patent drawingFigure 4
  • EP4425423B1 patent drawingFigure 5~6

AI summary

Provided in the present application are an image processing method and apparatus, a device and a storage medium, the method comprising: carrying out mask processing on a first-type object contained in an acquired target video frame image to obtain an image to be processed, the first-type object being an image element to be repaired; performing repair processing on the first-type object in the image to be processed, so as to obtain a first repaired image, and generating a corresponding image initial mask template on the basis of an initial blurry region in the first repaired image; when the first number of initial blurry pixels contained in the image initial mask template reaches a first threshold value, performing morphological processing on blurry regions corresponding to the initial blurry pixels, so as to obtain an image target mask template; when the second number of intermediate blurry pixels contained in the image target mask template reaches a second threshold value, performing in the first repaired image repair processing on pixel regions corresponding to the intermediate blurry pixels, so as to obtain a second repaired image; and on the basis of the second repaired image, determining a target repaired image corresponding to the image to be processed.