Lightweight robust real-time video matting method and system

Through the lightweight and robust real-time video cutout method, the lightweight encoder and loop decoder combined with wavelet downsampling technology are used to solve the problems of high labor costs and slow speed in the existing video cutout method, and the efficient real-time video cutout effect is achieved.

CN120298447APending Publication Date: 2025-07-11NORTHEASTERN UNIV CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510415801.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing video cutout methods have problems such as high labor costs, slow speed and insufficient accuracy in real-time applications, especially in the absence of auxiliary information, which is difficult to achieve efficient video cutouts.

Method used

A lightweight and robust real-time video cutout method is adopted, and feature images are extracted using a lightweight encoder, combined with the cyclic decoder aggregation time information, edge details are preserved through wavelet downsampling technology, and a U-net network model is built during the inference process to create a U-net network model for video cutout.

Benefits of technology

It achieves a significant increase inference speed without affecting accuracy, enhances the cutout effect, is suitable for handling motion blur and rough edge scenes, is suitable for real-time video processing, and meets the needs of high precision and high speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298447A_ABST
    Figure CN120298447A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight robust real-time video matting method and system, and relates to the technical field of video image processing. The invention provides a lightweight robust real-time video matting method and system. A lightweight encoder is applied to a video matting task; features and time information are aggregated by using a cyclic decoder, so that good time consistency is ensured; more edge details are reserved by using a wavelet down-sampling technology, and the matting precision is improved; in the reasoning process, a re-parameterization technology is used to improve the reasoning speed of the model, and a real-time video matting solution with good balance between precision and speed is realized. Meanwhile, the invention provides a lightweight robust real-time video matting system, which comprises a down-sampling module, an encoder module, a feature pyramid module, a cyclic decoder module, a refining module and a high-resolution prediction module, has a high-precision matting effect for a real-time video scene, and meets the real-time application requirement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video image processing, and particularly relates to a lightweight and robust real-time video matting method and system. Background Art

[0002] With the development of 5G communication technology and edge computing, real-time session video applications have broken through the boundaries of traditional video conferencing and formed an ecosystem covering multiple dimensions such as remote work, online education, live e-commerce, social entertainment, etc. Real-time video processing technology has gradually become the core supporting technology in the field of human-computer interaction.

[0003] Video matting and background replacement technology, as a key link in real-time video processing, shows significant application-driven characteristics in its technological evolution. In the video conferencing scenario, the professional image can be effectively enhanced and environmental interference can be reduced through the dynamic background replacement function; in the field of live e-commerce, more than 20 million product display interactions can be created daily through real-time background special effects; it is worth noting that with the EU's "Artificial Intelligence Act" imposing strict control on biometric data, background replacement technology with privacy protection capabilities has become a compliance requirement for multinational enterprise video systems.

[0004] Early video matting methods usually required auxiliary tripartite graph annotation, and sampling or propagation methods were usually used as solutions. However, compared with pictures, creating tripartite graph annotation for video frames is more time-consuming and the labor cost is huge. Some subsequent methods tried to use a single tripartite graph input for video matting, but there are still limitations in real-time applications. Background-based methods require users to take an additional background picture without the subject, while mask-guided methods use predicted or annotated segmentation masks to provide guidance for the foreground area. Tripartite graph-free methods relax the requirements for tripartite graphs and save a large amount of labor costs by using background-free pictures, segmentation masks, etc. as guidance for the network, and there is also a certain improvement in accuracy and speed. Although these methods do not require users to input additional auxiliary information, they need to perform matting in two stages, and the speed is often not satisfactory. There are also some methods that explore an automatic video matting network framework that does not rely on auxiliary information. The MODNet network suppresses flickering by comparing the differences between adjacent frames to improve the consistency of video matting results. In RVM, a cyclic structure is used to utilize temporal information to further improve temporal consistency and matting quality. VideoMatt discusses and analyzes various temporal modeling methods in the matting network and discloses a new benchmark dataset. This type of method does not require any auxiliary information to directly process the input video, greatly improving the efficiency of video matting. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a lightweight and robust real-time video matting method and system in view of the above-mentioned deficiencies of the prior art, applying a lightweight encoder to the video matting task; using a cyclic decoder to aggregate features and temporal information to ensure good temporal consistency; using wavelet downsampling technology to retain more edge details and improve the matting accuracy; and using a reparameterization technology during the inference process to improve the inference speed of the model, so as to achieve a real-time video matting solution with a good balance between accuracy and speed.

[0006] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0007] On the one hand, the present invention provides a lightweight and robust real-time video matting method, including the following steps:

[0008] Step 1: Obtain the original video stream, extract consecutive image frames from the original video stream, and perform downsampling and normalization processing on the consecutive image frames to obtain low-resolution consecutive image frames;

[0009] Obtain the original video stream, and extract high-resolution consecutive image frames F i , i = 1, 2,.., N, where N is the number of consecutive image frames, and i is the serial number of the image frame F i in the original video stream. Perform preprocessing on the consecutive image frames F i to obtain N preprocessed low-resolution images I i , i = 1, 2,.., N;

[0010] The specific method of the preprocessing operation is as follows:

[0011] Adopt the bilinear interpolation method to perform downsampling on the consecutive image frames F i . Dynamically set the sampling rate according to the resolution of the input consecutive image frames, so that the resolution of the low-resolution image obtained after downsampling does not exceed the pre-set low-resolution image resolution threshold, and obtain the downsampled consecutive image frames F i ';

[0012] Perform normalization operation on the downsampled consecutive image frames F i ' to obtain the low-resolution image I i ;

[0013] Step 2: Construct a lightweight and robust real-time video matting model;

[0014] Construct a lightweight real-time video matting model based on the U-net network model, including an encoder module, a feature pyramid module, a cyclic decoder module, a refinement module, and a high-resolution prediction module; among them, the encoder module includes m encoder sub-blocks, and the cyclic decoder module includes m decoder sub-blocks;

[0015] The first to the m-th encoder sub-blocks of the encoder module extract feature images of different scales respectively; the feature image output by the m-th encoder sub-block is input into the feature pyramid module, and the feature image output by the feature pyramid module is input into the first decoder sub-block; the feature image output by the (m - 1)-th encoder sub-block is concatenated with the feature image output by the first decoder sub-block and input into the second decoder sub-block; until the feature image output by the first encoder sub-block and the feature image output by the (m - 1)-th decoder sub-block are concatenated and input into the m-th decoder sub-block; the refinement module receives the pre-processed low-resolution image and the feature image output by the m-th decoder sub-block to obtain the feature image output by the refinement module; the high-resolution prediction module receives the consecutive image frames and the feature image output by the refinement module to obtain the output foreground prediction image and Alpha prediction image;

[0016] Step 2.1: Construct an encoder module, including m encoder sub-blocks, which sequentially extract feature images of different scales;

[0017] Use m RepViT networks as the m encoder sub-blocks to construct the encoder module. The encoder module takes the low-resolution image I i as input. The m encoder sub-blocks sequentially extract feature images of different scales, and the scales of the feature maps extracted by the m encoder sub-blocks are controlled by setting the stride parameters in each RepViT network;

[0018] Step 2.2: Construct a feature pyramid module, which takes the smallest-scale feature image output by the encoder module as input, performs semantic enhancement on the feature image input into the feature pyramid module, and outputs a feature image with enhanced semantic information;

[0019] Use the LR-ASPP model to construct the feature pyramid module. The LR-ASPP model includes two branches. Among them, the first branch is sequentially a convolutional layer, a normalization layer, and a ReLU function activation layer; the second branch is sequentially an adaptive average pooling layer, a convolutional layer, and a Sigmoid function activation layer; the results obtained by processing the two branches are multiplied to obtain a feature image with enhanced semantic information;

[0020] Step 2.3: Construct a recurrent decoder module, including m decoder sub-blocks, which respectively receive feature images of different scales and output output feature images that aggregate temporal information;

[0021] Use m decoder sub-blocks to construct the recurrent decoder module. The first decoder sub-block includes a ConvGRU sub-block and an upsampling sub-block. The second to the m-th decoder sub-blocks include a Conv sub-block, a ConvGRU sub-block, and an upsampling layer; the Conv sub-block includes a convolutional layer, a normalization layer, and a ReLU function activation layer;

[0022] The m decoder sub - blocks of the recurrent decoder module respectively receive feature images of different scales and output output feature images that aggregate temporal information. Among them, the first decoder sub - block receives the feature image with enhanced semantic information output by the feature pyramid module. The second to the m - th decoder sub - blocks splice the feature images output by the first to the m - 1 - th encoder sub - blocks and the feature image output by the previous decoder sub - block, and use the spliced feature image as the input feature image of the second to the m - th decoder sub - blocks. After layer - by - layer fusion of the feature images, the feature image output by the recurrent decoder module is obtained;

[0023] In the ConvGRU sub - block, the number of channels of the feature map input at the current time step is sliced in half to the original number of channels as the decoded feature map x, and is concatenated with the hidden state variable h at the previous time step i-1 After concatenation, the resulting image is successively subjected to convolution and Sigmoid activation operations and then split by the number of channels to obtain the reset gate r and the input gate z; The decoded feature map x and the reset gate r are concatenated with the hidden state variable h at the previous time step i-1 along the channel dimension, and then successively subjected to convolution and Tanh activation operations to obtain the candidate hidden state c; The hidden state h is updated according to the input gate z and the candidate hidden state c; The hidden state variable h at the last time step is taken i as the output of the ConvGRU sub - block;

[0024] In the up - sampling layer, the feature image output by the ConvGRU sub - block is up - sampled by a factor of 2 using the bilinear interpolation method and used as the input of the next decoder sub - block;

[0025] Step 2.4: Construct a refinement module, including a wavelet down - sampling sub - block, two Conv sub - blocks, two up - sampling sub - blocks, and a mapping layer, to obtain a one - channel Alpha prediction image, a three - channel foreground prediction image, and a one - channel segmentation image;

[0026] Use a wavelet down - sampling sub - block, two Conv sub - blocks, two up - sampling sub - blocks, and a mapping layer to construct a refinement module;

[0027] The refinement module takes the feature image output by the recurrent decoder module and the low - resolution image I i as inputs, concatenates the feature image output by the recurrent encoder module and the feature image obtained by passing the low - resolution image I i through the wavelet down - sampling sub - block, successively passes through the first Conv sub - block and the first up - sampling layer, and then concatenates with the low - resolution image I i and successively passes through the second Conv sub - block and the second up - sampling layer to obtain a hidden feature image, and then passes through the mapping layer to obtain the feature image output by the refinement module: a one - channel Alpha prediction image, a three - channel foreground prediction image, and a one - channel segmentation image;

[0028] The wavelet downsampling sub-block includes two branches. The first branch extracts the low-frequency component I of the low-resolution image I i and based on the low-frequency component I L , extracts the horizontal high-frequency component H and the rough low-frequency component I L ′ of the low-resolution image I i ; The second branch extracts the high-frequency component I of the low-resolution image I L and based on the high-frequency component I i , extracts the vertical high-frequency component V and the diagonal high-frequency component D of the low-resolution image I H ; The rough low-frequency component I H ′, the horizontal high-frequency component H, the vertical high-frequency component V and the diagonal high-frequency component D are concatenated along the channel dimension; The concatenated feature image is subjected to convolution, normalization, and ReLU activation operations to obtain the feature map H output by the wavelet downsampling sub-block i ; L ′, the horizontal high-frequency component H, the vertical high-frequency component V and the diagonal high-frequency component D are concatenated along the channel dimension; The concatenated feature image is subjected to convolution, normalization, and ReLU activation operations to obtain the feature map H output by the wavelet downsampling sub-block i ;

[0029] Step 2.5: Use deep guided filtering to construct a high-resolution prediction module to output a refined high-resolution foreground prediction image and an Alpha prediction image;

[0030] The high-resolution prediction module takes the continuous image frame F i , the feature image and the hidden feature image output by the refinement module as inputs, and performs deep guided filtering on the input continuous image frame F i , a one-channel Alpha prediction image, a three-channel foreground prediction image, and the hidden feature image output by the refinement module, calculates statistics such as mean, covariance, and variance through a box filter, and combines convolution operations to extract and transform the input feature map, and outputs a refined high-resolution foreground prediction image and an Alpha prediction image;

[0031] Step 3: Use a lightweight and robust real-time video matting model to perform real-time matting on continuous image frames;

[0032] Input the continuous image frames into the lightweight and robust real-time video matting model to obtain a high-resolution foreground prediction image and an Alpha prediction image.

[0033] On the other hand, the present invention provides a lightweight and robust real-time video matting system, including a downsampling module, an encoder module, a feature pyramid module, a cyclic decoder module, a refinement module, and a high-resolution prediction module;

[0034] The downsampling module is used to receive a high-resolution input image and perform downsampling operations on it to output a low-resolution image;

[0035] The encoder module is used to receive the low-resolution image output by the downsampling module, and sequentially extract and output feature images of different scales;

[0036] The feature pyramid module is used to receive the feature image with the smallest scale output by the encoder module, perform semantic enhancement on it, and output a semantically enhanced feature image;

[0037] The cyclic encoding module is used to fuse the feature images of different scales output by the encoder module and the semantically enhanced feature image output by the feature pyramid module, and output a fused feature image;

[0038] The refinement module is used to receive the low-resolution image output by the downsampling module and the fused feature image output by the cyclic encoder, and output a low-resolution matte image;

[0039] The high-resolution prediction module is used to receive the low-resolution matte image output by the refinement module and the high-resolution input image, and output a high-resolution matte image;

[0040] The beneficial effects of adopting the above technical solutions are as follows: A lightweight and robust real-time video matte method provided by the present invention extracts feature images through a lightweight encoder module, and eliminates redundant residual structures through re-parameterization during the inference process, which can greatly improve the inference speed without affecting the matte accuracy; the semantic information of the feature image is enhanced through the feature pyramid module, enhancing the matte effect of the system; the optimized cyclic decoder module aggregates feature and temporal information to ensure good temporal consistency; the refinement module incorporating wavelet downsampling has better edge refinement ability and is more suitable for processing scenarios such as motion blur and rough edges; using a deep guidance filter, a fine high-resolution matte result can be obtained quickly, which has more application advantages for high-resolution application scenarios. At the same time, a lightweight and robust real-time video matte system provided by the present invention includes a downsampling module, an encoder module, a feature pyramid module, a cyclic decoder module, a refinement module, and a high-resolution prediction module; compared with the prior art, the present invention has a high-precision matte effect for real-time video scenarios, meets the real-time application requirements, and is a real-time video matte solution with a good balance between accuracy and speed. Description of the Drawings

[0041] Figure 1 It is a structural diagram of a lightweight real-time video matte model constructed based on the U-net network model provided by an embodiment of the present invention;

[0042] Figure 2 It is a structural diagram of the RepViT network provided by an embodiment of the present invention;

[0043] Figure 3 It is a schematic diagram of the cyclic decoder provided by an embodiment of the present invention;

[0044] Figure 4 Schematic diagram of the process for extracting components by Haar wavelet transform provided by an embodiment of the present invention. Specific implementation manners

[0045] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0046] On the one hand, the present invention provides a lightweight and robust real-time video matting method, as Figure 1 shown, including the following steps:

[0047] Step 1: Obtain an original video stream, extract consecutive image frames from the original video stream, and perform downsampling and normalization processing on the consecutive image frames to obtain low-resolution consecutive image frames;

[0048] Obtain the original video stream and extract high-resolution consecutive image frames F i , i = 1, 2,.., N, where N is the number of consecutive image frames and i is the serial number of the image frame F i in the original video stream. Preprocess the consecutive image frames F i to obtain N preprocessed low-resolution images I i , i = 1, 2,.., N;

[0049] The specific method of the preprocessing operation is as follows:

[0050] Adopt the bilinear interpolation method to perform downsampling on the consecutive image frames F i ; Dynamically set the sampling rate according to the resolution of the input consecutive image frames, so that the resolution of the low-resolution image obtained after downsampling does not exceed a pre-set low-resolution image resolution threshold, and obtain the downsampled consecutive image frames F i ';

[0051] Perform normalization operation on the downsampled consecutive image frames F i ' to obtain the low-resolution image I i ;

[0052] In this embodiment, the resolution threshold of the low-resolution image is set to 512×512 pixels. For continuous image frames of 1920×1080 pixels, the bilinear interpolation method is used, and the sampling rate is set to 0.25, that is, 4 times downsampling. For continuous image frames with a resolution lower than 512×512 pixels, the sampling rate is set to 1 to keep the original size. The continuous image frames after downsampling are normalized. The mean parameters of the RGB three channels of the continuous image frames after downsampling are set to 0.485, 0.456, and 0.406 respectively, and the standard deviation parameters of the RGB three channels of the continuous image frames after downsampling are set to 0.229, 0.224, and 0.225 respectively.

[0053] Step 2: Construct a lightweight and robust real-time video matting model;

[0054] Based on the U-net network model, a lightweight real-time video matting model is constructed, including an encoder module, a feature pyramid module, a recurrent decoder module, a refinement module, and a high-resolution prediction module; among them, the encoder module includes m encoder sub-blocks, and the recurrent decoder module includes m decoder sub-blocks;

[0055] The encoder module is used to extract feature images of different scales; the feature pyramid module is used to perform semantic enhancement on the feature images; the recurrent decoder module is used to fuse feature images of different scales; the refinement module is used to refine the fused feature images to obtain a low-resolution matting result; the high-resolution prediction module is used to predict the high-resolution matting result;

[0056] The first to the m-th encoder sub-blocks of the encoder module extract feature images of different scales respectively; the feature image output by the m-th encoder sub-block is input into the feature pyramid module, and the feature image output by the feature pyramid module is input into the first decoder sub-block; the feature image output by the m-1-th encoder sub-block is concatenated with the feature image output by the first decoder sub-block and input into the second decoder sub-block; until the feature image output by the first encoder sub-block and the feature image output by the m-1-th decoder sub-block are concatenated and input into the m-th decoder sub-block; the refinement module receives the preprocessed low-resolution image and the feature image output by the m-th decoder sub-block to obtain the feature image output by the refinement module; the high-resolution prediction module receives the continuous image frames and the feature image output by the refinement module to obtain the output foreground prediction image and Alpha prediction image;

[0057] Step 2.1: Construct an encoder module, including m encoder sub-blocks, and sequentially extract feature images of different scales;

[0058] Use m RepViT networks as m encoder sub-blocks to construct the encoder module. The encoder module takes the low-resolution image I iAs input, m encoder sub-blocks sequentially extract feature images of different scales, and the scales of the feature maps extracted by the m encoder sub-blocks are controlled by setting the stride parameters in each RepViT network;

[0059] The RepViT network structure is as Figure 2 shown. In the inference stage, the reparameterization technique is used to remove the residual structure in the RepViT block, which greatly improves the inference speed without affecting the matte extraction accuracy;

[0060] In this embodiment, 4 encoder sub-blocks are set. For the low-resolution image I of size H×W i , the 4 encoder sub-blocks respectively output feature images of scales ;

[0061] Step 2.2: Construct a feature pyramid module. Using the feature image of the smallest scale output by the encoder module as input, the feature image input to the feature pyramid module is semantically enhanced, and a feature image with enhanced semantic information is output;

[0062] The LR-ASPP model is used to construct the feature pyramid module. The LR-ASPP model includes two branches. Among them, the first branch is sequentially a convolutional layer, a normalization layer, and a ReLU function activation layer; the second branch is sequentially an adaptive average pooling layer, a convolutional layer, and a Sigmoid function activation layer; the results obtained by processing the two branches are multiplied to obtain a feature image with enhanced semantic information;

[0063] In this embodiment, the LR-ASPP model is used to construct the feature pyramid module. Using the feature image of scale output by the 4th encoder sub-block of the encoder module as input, it is semantically enhanced to obtain a semantically enhanced feature image S of scale ; i4 ;

[0064] Step 2.3: Construct a recurrent decoder module, which includes m decoder sub-blocks that respectively receive feature images of different scales and output an output feature image that aggregates temporal information;

[0065] The recurrent decoder is as Figure 3 shown. It is unidirectional in the time dimension, that is, it is only related to the input of the previous moment and the input of the current moment. By passing the hidden information of the previous frame to the next frame, more feature information is obtained, and the matte extraction effect is optimized frame by frame, which is more suitable for real-time video processing such as live streaming;

[0066] The recurrent decoder module is constructed using m decoder sub - blocks. The first decoder sub - block includes a ConvGRU sub - block and an up - sampling sub - block, and the second to the m - th decoder sub - blocks include a Conv sub - block, a ConvGRU sub - block, and an up - sampling layer; the Conv sub - block includes a convolutional layer, a normalization layer, and a ReLU function activation layer;

[0067] The m decoder sub - blocks of the recurrent decoder module respectively receive feature images of different scales and output output feature images that aggregate temporal information. Among them, the first decoder sub - block receives the feature image with enhanced semantic information output by the feature pyramid module, and the second to the m - th decoder sub - blocks splice the feature images output by the first to the m - 1 - th encoder sub - blocks and the feature image output by the previous decoder sub - block, and use the spliced feature image as the input feature image of the second to the m - th decoder sub - blocks. After layer - by - layer fusion of the feature images, the feature image output by the recurrent decoder module is obtained;

[0068] In the ConvGRU sub - block, the number of channels of the feature map input at the current time step is sliced in half to be the decoded feature map x, and it is concatenated with the hidden state variable h at the previous time step i-1 After concatenation, the resulting image is successively subjected to convolution and Sigmoid activation operations and then split by the number of channels to obtain the reset gate r and the input gate z; the decoded feature map x and the reset gate r are concatenated with the hidden state variable h at the previous time step i-1 along the channel dimension, and then successively subjected to convolution and Tanh activation operations to obtain the candidate hidden state c; the hidden state h is updated according to the input gate z and the candidate hidden state c; the hidden state variable h at the last time step is taken i as the output of the ConvGRU sub - block;

[0069] In the up - sampling layer, the feature image output by the ConvGRU sub - block is up - sampled by a factor of 2 using bilinear interpolation and used as the input for the next decoder sub - block;

[0070] In this embodiment, the recurrent decoder module includes 4 decoder sub - blocks. The first encoder sub - block includes a ConvGRU sub - block and an up - sampling layer, and the second to the fourth encoder sub - blocks include a Conv sub - block, a ConvGRU sub - block, and an up - sampling layer. The first decoder sub - block receives the feature image S with enhanced semantic information at scale The second to the fourth encoder sub - blocks receive the feature images at scale i4 output by the encoder sub - blocks. After layer - by - layer fusion of the feature images, the feature map output by the recurrent decoder module at scale is obtained;

[0071] ​Step 2.4: Construct a refinement module, including a wavelet downsampling sub-block, two Conv sub-blocks, two upsampling sub-blocks, and a mapping layer, to obtain a one-channel Alpha prediction image, a three-channel foreground prediction image, and a one-channel segmentation image;

[0072] Use a wavelet downsampling sub-block, two Conv sub-blocks, two upsampling sub-blocks, and a mapping layer to construct a refinement module;

[0073] The refinement module takes the feature image output by the recurrent decoder module and the low-resolution image I i as inputs, concatenates the feature image output by the recurrent encoder module and the low-resolution image I i with the feature image obtained through the wavelet downsampling sub-block, passes through the first Conv sub-block and the first upsampling layer in sequence, and then concatenates with the low-resolution image I i to pass through the second Conv sub-block and the second upsampling layer in sequence to obtain a hidden feature image, and then passes through the mapping layer to obtain the feature image output by the refinement module: a one-channel Alpha prediction image, a three-channel foreground prediction image, and a one-channel segmentation image;

[0074] The wavelet downsampling sub-block includes two branches, as Figure 4 shown. The first branch extracts the low-frequency component I i of the low-resolution image I L , and based on the low-frequency component I L extracts the horizontal high-frequency component H and the rough low-frequency component I i ' of the low-resolution image I L ; The second branch extracts the high-frequency component I i of the low-resolution image I H , and based on the high-frequency component I H extracts the vertical high-frequency component V and the diagonal high-frequency component D of the low-resolution image I i ; Concatenate the rough low-frequency component I L ', the horizontal high-frequency component H, the vertical high-frequency component V, and the diagonal high-frequency component D along the channel dimension; Perform convolution, normalization, and ReLU activation operations on the concatenated feature image to obtain the feature map H i output by the wavelet downsampling sub-block;

[0075] Step 2.5: Use deep guided filtering to construct a high-resolution prediction module to output a refined high-resolution foreground prediction image and Alpha prediction image;

[0076] The high-resolution prediction module takes the consecutive image frames F i , the feature image output by the refinement module, and the hidden feature image as inputs, and processes the input consecutive image frames F i, perform depth-guided filtering on a one-channel Alpha prediction image, a three-channel foreground prediction image, and the hidden feature image output by the refinement module. Calculate statistics such as mean, covariance, and variance through a box filter, and combine convolution operations to extract and transform the input feature map, outputting a refined high-resolution foreground prediction image and Alpha prediction image;

[0077] Step 3: Use a lightweight and robust real-time video matting model to perform real-time matting on consecutive image frames;

[0078] The goal of the video matting task is to accurately estimate the mapping coefficient α and the foreground image for each frame of the input video sequence. In a quantization form, given a video sequence H i ={H1, H2,..., H T}, each frame H i can be regarded as an unknown foreground image f i and a background image b i through a linear combination of the mapping coefficient α i ∈[0, 1]:

[0079] I i =α i f i +(1 - α i )b i

[0080] By obtaining an accurate α value, the extracted foreground object can be integrated into a new background to achieve the background replacement effect.

[0081] Input consecutive image frames into a lightweight and robust real-time video matting model to obtain a high-resolution foreground prediction image and an Alpha prediction image.

[0082] On the other hand, a lightweight and robust real-time video matting system of this embodiment includes a downsampling module, an encoder module, a feature pyramid module, a cyclic decoder module, a refinement module, and a high-resolution prediction module;

[0083] The downsampling module is used to receive a high-resolution input image and perform a downsampling operation on it, outputting a low-resolution image;

[0084] The encoder module is used to receive the low-resolution image output by the downsampling module, and sequentially extract and output feature images of different scales;

[0085] The feature pyramid module is used to receive the feature image with the smallest scale output by the encoder module, perform semantic enhancement on it, and output a semantically enhanced feature image;

[0086] The loop encoding module is used to fuse the feature images of different scales output by the encoder module and the semantic enhancement feature image output by the feature pyramid module, and output a fused feature image;

[0087] The refinement module is used to receive the low-resolution image output by the downsampling module and the fused feature image output by the loop encoder, and output a low-resolution matte image;

[0088] The high-resolution prediction module is used to receive the low-resolution matte image output by the refinement module and the high-resolution input image, and output a high-resolution matte image;

[0089] In this embodiment, the encoder module uses the RepViT model as the backbone network, including 4 encoder sub-blocks of different scales. The scales of the feature images output by the 4 encoder sub-blocks of different scales are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image in sequence; the LR-ASPP model is used as the feature pyramid module. The feature pyramid module receives the feature image with a scale of 1 / 32 of the original image output by the encoder module, performs semantic enhancement on it, and outputs a semantic enhancement feature image with a scale of 1 / 32 of the original image; the ConvGRU model and the upsampling sub-block are used as the loop decoder module; the refinement module receives the low-resolution image and the feature image output by the loop decoder, and outputs a one-channel Alpha prediction image, a three-channel foreground prediction image, and a one-channel segmentation image of the low-resolution image; the high-resolution prediction module receives the input high-resolution input image and the one-channel Alpha prediction image, three-channel foreground prediction image, and one-channel segmentation image of the low-resolution image output by the refinement module, and outputs a refined high-resolution foreground prediction image and Alpha prediction image.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.

Claims

1. A lightweight and robust real-time video matting method, characterized in that: It includes the following steps: Step 1: Obtain the original video stream, extract consecutive image frames from the original video stream, perform downsampling and normalization processing on the consecutive image frames to obtain low-resolution consecutive image frames; Step 2: Construct a lightweight and robust real-time video matting model; Construct a lightweight real-time video matting model based on the U-net network model, including an encoder module, a feature pyramid module, a recurrent decoder module, a refinement module, and a high-resolution prediction module; among them, the encoder module includes m encoder sub-blocks, and the recurrent decoder module includes m decoder sub-blocks; The first to the m-th encoder sub-blocks of the encoder module respectively extract feature images of different scales; input the feature image output by the m-th encoder sub-block into the feature pyramid module, and input the feature image output by the feature pyramid module into the first decoder sub-block; splice the feature image output by the (m - 1)-th encoder sub-block and the feature image output by the first decoder sub-block, and input it into the second decoder sub-block; until the feature image output by the first encoder sub-block and the feature image output by the (m - 1)-th decoder sub-block are spliced and input into the m-th decoder sub-block; the refinement module receives the preprocessed low-resolution image and the feature image output by the m-th decoder sub-block to obtain the feature image output by the refinement module; The high-resolution prediction module receives the consecutive image frames and the feature image output by the refinement module to obtain the output foreground prediction image and Alpha prediction image; Step 3: Use the lightweight and robust real-time video matting model to perform real-time matting on the consecutive image frames.

2. A lightweight and robust real-time video matting method according to claim 1, characterized in that: The said Step 1 includes: Obtain the original video stream and extract high-resolution continuous image frames F i , i = 1, 2,.., N, where N is the number of continuous image frames and i is the serial number of the image frame F i in the original video stream. Preprocess the continuous image frames F i to obtain N preprocessed low-resolution images I i , i = 1, 2,.., N; Perform a normalization operation on the downsampled consecutive image frames F i ′ to obtain the low-resolution image I i .

3. A lightweight and robust real-time video matting method according to claim 2, characterized in that: The specific method of the preprocessing operation in Step 1 is: Perform downsampling on the continuous image frame F using the bilinear interpolation method i ; Dynamically set the sampling rate according to the resolution of the input continuous image frame, so that the resolution of the low-resolution image obtained after downsampling does not exceed a pre-set low-resolution image resolution threshold, and obtain the downsampled continuous image frame F i '.

4. A lightweight and robust real-time video matting method according to claim 3, characterized in that: The said Step 2 includes: Step 2.1: Construct an encoder module, including m encoder sub-blocks, and sequentially extract feature images of different scales; Step 2.2: Construct a feature pyramid module, use the smallest-scale feature image output by the encoder module as the input, perform semantic enhancement on the feature image input into the feature pyramid module, and output the feature image with enhanced semantic information; Step 2.3: Construct a recurrent decoder module, including m decoder sub-blocks, respectively receive feature images of different scales, and output the output feature image that aggregates temporal information; Step 2.4: Construct a refinement module, including a wavelet downsampling sub-block, two Conv sub-blocks, two upsampling sub-blocks, and a mapping layer, to obtain a one-channel Alpha prediction image, a three-channel foreground prediction image, and a one-channel segmentation image; Step 2.5: Use depth-guided filtering to construct a high-resolution prediction module, and output the refined high-resolution foreground prediction image and Alpha prediction image.

5. A lightweight and robust real-time video matting method according to claim 4, characterized in that: The specific method of the said Step 2.1 includes: An encoder module is constructed using m RepViT networks as m encoder sub - blocks. The encoder module takes a low - resolution image I i as input. The m encoder sub - blocks sequentially extract feature images of different scales, and the scale of the feature maps extracted by the m encoder sub - blocks is controlled by setting the stride parameters in each RepViT network.

6. A lightweight and robust real-time video matting method according to claim 5, characterized in that: The specific method of the said Step 2.2 is: Use the LR-ASPP model to construct the feature pyramid module. The LR-ASPP model includes two branches. Among them, the first branch is sequentially a convolutional layer, a normalization layer, and a ReLU function activation layer; the second branch is sequentially an adaptive average pooling layer, a convolutional layer, and a Sigmoid function activation layer; the results obtained by processing the two branches are multiplied to obtain the feature image with enhanced semantic information.

7. A lightweight and robust real-time video matting method according to claim 6, characterized in that: The specific method of the said Step 2.3 is: Construct a recurrent decoder module using m decoder sub - blocks. The first decoder sub - block includes a ConvGRU sub - block and an upsampling sub - block. The second to the m - th decoder sub - blocks include a Conv sub - block, a ConvGRU sub - block, and an upsampling layer. The Conv sub - block includes a convolutional layer, a normalization layer, and a ReLU activation layer. The m decoder sub - blocks of the recurrent decoder module respectively receive feature images of different scales and output output feature images that aggregate temporal information. Among them, the first decoder sub - block receives the feature image with enhanced semantic information output by the feature pyramid module. The second to the m - th decoder sub - blocks concatenate the feature images output by the first to the m - 1 - th encoder sub - blocks and the feature image output by the previous decoder sub - block, and use the concatenated feature image as the input feature image of the second to the m - th decoder sub - blocks. After layer - by - layer fusion of the feature images, the feature image output by the recurrent decoder module is obtained. In the ConvGRU sub-block, the number of channels of the feature map input at the current moment is sliced in half to the original number of channels as the decoded feature map x, which is concatenated with the hidden state variable h at the previous moment. i-1 After concatenation, the resulting map is successively subjected to convolution and Sigmoid activation operations and then split by the number of channels to obtain the reset gate r and the input gate z. The decoded feature map x and the reset gate r are concatenated with the hidden state variable h at the previous moment i-1 along the channel dimension, and then successively subjected to convolution and Tanh activation operations to obtain the candidate hidden state c. The hidden state h is updated according to the input gate z and the candidate hidden state c. The hidden state variable h at the last time step is taken i as the output of the ConvGRU sub-block. In the upsampling layer, the feature image output by the ConvGRU sub - block is upsampled by a factor of 2 using bilinear interpolation and used as the input of the next decoder sub - block.

8. A lightweight and robust real-time video matting method according to claim 7, characterized in that: The specific method of step 2.4 is as follows: Construct a refinement module using a wavelet downsampling sub - block, two Conv sub - blocks, two upsampling sub - blocks, and a mapping layer. The refinement module takes the feature image output by the cyclic decoder module and the low-resolution image I i As inputs, the feature image output by the cyclic encoder module and the low-resolution image I i The feature images obtained through the wavelet downsampling sub-blocks are concatenated, and then pass through the first Conv sub-block and the first upsampling layer in sequence, and then concatenated with the low-resolution image I i Are concatenated, pass through the second Conv sub-block and the second upsampling layer in sequence to obtain a hidden feature image, and then pass through a mapping layer to obtain the feature image output by the refinement module: a one-channel Alpha prediction image, a three-channel foreground prediction image, and a one-channel segmentation image; The wavelet downsampling sub-block includes two branches. The first branch extracts the low-frequency component I i of the low-resolution image I L , and based on the low-frequency component I L , extracts the horizontal high-frequency component H and the rough low-frequency component I i ' of the low-resolution image I L ; The second branch extracts the high-frequency component I i of the low-resolution image I H , and based on the high-frequency component I H , extracts the vertical high-frequency component V and the diagonal high-frequency component D of the low-resolution image I i ; Concatenate the rough low-frequency component I L ', the horizontal high-frequency component H, the vertical high-frequency component V, and the diagonal high-frequency component D along the channel dimension; Perform convolution, normalization, and ReLU activation operations on the concatenated feature image to obtain the feature map H i output by the wavelet downsampling sub-block.

9. A semantic segmentation method based on HDR dynamic neural radiance fields according to claim 8, characterized in that: The specific method of step 2.5 is as follows: The high-resolution prediction module takes the consecutive image frames F i , the feature image and the hidden feature image output by the refinement module as inputs, and performs depth-guided filtering on the input consecutive image frames F i , a one-channel Alpha prediction image, a three-channel foreground prediction image, and the hidden feature image output by the refinement module. The module calculates statistics such as mean, covariance, and variance through a box filter, and combines convolution operations to extract and transform features, and outputs the refined high-resolution foreground prediction image and Alpha prediction image.

10. A lightweight and robust real-time video matting system that implements real-time video matting based on the method described in claim 1, characterized in that: It includes a downsampling module, an encoder module, a feature pyramid module, a recurrent decoder module, a refinement module, and a high - resolution prediction module. The downsampling module is used to receive a high - resolution input image, perform a downsampling operation on it, and output a low - resolution image. The encoder module is used to receive the low - resolution image output by the downsampling module, sequentially extract and output feature images of different scales. The feature pyramid module is used to receive the feature image with the smallest scale output by the encoder module, perform semantic enhancement on it, and output a semantically enhanced feature image. The recurrent encoding module is used to fuse the feature images of different scales output by the encoder module and the semantically enhanced feature image output by the feature pyramid module, and output a fused feature image. The refinement module is used to receive the low - resolution image output by the downsampling module and the fused feature image output by the recurrent encoder, and output a low - resolution matte image. The high - resolution prediction module is used to receive the low - resolution matte image output by the refinement module and the high - resolution input image, and output a high - resolution matte image.

Citation Information

Cited By

  • Double-branch video matting method

    CN121482565A