An Interactive Video Matting System Based on Mask Propagation Network
Through user interaction, the mask generation and the mask propagation network and the temporal feature fusion module are used to realize the automation and efficiency of video cutouts, solving the problems of high cost and difficulty in space-time consistency in the existing technology.
Patent Information
- Application Number
- CN202210193688.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-01
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-01
AI Technical Summary
The existing video cutout method requires providing three-point pictures or sketches for each frame of the image, which leads to high time and labor costs, and is difficult to maintain consistency in the space-time dimension, which easily leads to artifacts and flickering.
An interactive video cutout system based on mask propagation network is adopted, and the mask is automatically inferred and optimized to ensure space-time consistency by users clicking or graffitiing on any frame of the video by users, and a mask time domain propagation module and a subdivision module based on spatiotemporal feature fusion.
It greatly reduces the workload of video clipping, achieves high-performance clipping effect, and effectively solves the problem of space-time consistency between frames, reducing artifacts and flickering.
Smart Images

Figure CN114549574B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a video matting system based on a mask propagation network and feature fusion. Background Art
[0002] Image Matting is a technology focused on foreground object extraction. Its core idea is to mathematically model an image, considering the image as a convex combination of the foreground and background parts with a certain weight (transparency mask), and separating the foreground and background parts through a determined transparency mask (Alpha Matte). The solution formula of its mathematical model is as follows:
[0003] I z = α z F z + (1 - α z )B z (1)
[0004] Where z represents a certain pixel point with coordinates (x, y) in the image, and I z represents the RGB color value of the pixel point z, F z represents the color value of the foreground pixel point, B z represents the color value of the background pixel point, and α z represents the transparency mask value of z, with a value range of [0, 1]. To solve this formula, additional supplementary constraints need to be introduced. Common supplementary inputs include trimap, scribbles, background image, foreground coordinates, etc.
[0005] Video matting is a task of extracting moving foreground objects from a given video based on image matting. Compared with image matting, video matting methods bring two challenges.
[0006] Firstly, video matting methods perform matting on each frame of the video. Therefore, supplementary inputs such as trimap or scribbles need to be provided for each frame. As the data volume increases, if manual annotation is used for the video, the time and labor costs will be huge. For this reason, researchers often strive to weaken the input of supplementary information and let the network learn the movement of the foreground object to predict the transparency mask of the object in the video; in addition, there are also methods that choose to use an input background video or background image, compare it with the original video to obtain foreground prior information, so as to avoid a large amount of manual annotation. However, such methods have higher requirements for the shooting environment and scene and have certain limitations.
[0007] Secondly, the transparency mask (alpha matte) obtained by video matting needs to be consistent in both the temporal and spatial dimensions. If the image matting algorithm is directly applied to each frame of the video and the results are then stitched together to form a video, artifacts and flickering will inevitably occur in moving objects or details. The traditional solution is to calculate the motion of the foreground by finding the local or non-local affinities between pixel colors, but the results are often unsatisfactory, especially when dealing with complex scenes, such as complex backgrounds and fast-moving foregrounds. More recent methods use optical flow estimation to predict the motion of the foreground, but optical flow estimation often cannot handle large semi-transparent areas well. Summary of the Invention
[0008] The present invention proposes an interactive video matting system based on a mask propagation network and feature fusion. The user only needs to provide a small number of clicks or scribbles on any frame of the video to indicate whether the position is the foreground or the background, and then the matting of all video frames can be completed without providing a trimap for each frame image, greatly reducing the workload of video matting. At the same time, it has performance comparable to that of advanced matting algorithms. In addition, it solves the spatio-temporal consistency problem of foreground objects between frames.
[0009] An interactive video matting system based on a mask propagation network includes a cache module, an interactive image rough segmentation module, a mask temporal propagation module, and a fine segmentation module based on spatio-temporal feature fusion.
[0010] The cache module is used to cache the video in the form of video frames to obtain the original input image of each frame; at the same time, it is used to cache the memory frames marked by the mask temporal propagation module.
[0011] The interactive object rough segmentation module is used to interact with the input image. The interaction includes two interaction methods: clicking and scribbling. The user can choose any interaction method according to the actual situation. By single-clicking or scribbling, the foreground object information (indicator map) of the original input image is obtained, and it is combined with the original input image and input into the image segmentation network to obtain a preliminary mask (Mask).
[0012] The user can optimize the mask by repeating clicks or scribbles until a sufficiently accurate mask is obtained and then send it to the mask temporal propagation module.
[0013] The mask temporal propagation module includes a spatio-temporal memory frame reader based on an attention mechanism (memoryframe reader);
[0014] The described spatio-temporal memory frame reader based on the attention mechanism includes a memory encoder, a query encoder, and a query decoder.
[0015] After obtaining the mask corresponding to a single-frame original image, the mask time-domain propagation module performs mask propagation in both forward and backward time-domain directions. The principle is to predict the mask of the query frame based on the memory frames already existing in the current cache module, then mark the query frame with the predicted mask as a memory frame and store it in the cache module, and take the next frame of the video as a new query frame, repeating the above operations until the next frame is a memory frame or the last frame of the video, which means that the masks of all frames have been obtained.
[0016] The specific propagation method is to use the current interactive frame as the memory frame and the adjacent frame as the query frame, match through the key feature maps of the memory frame and the query frame, then multiply the value feature map of the memory frame by the weight generated by the key feature matching, and finally connect the value feature of the query frame. Figure 1 And send it to the query decoder for decoding to finally predict the mask of the query frame.
[0017] The described fine segmentation module based on spatio-temporal feature fusion includes a fine segmentation encoder, a fine segmentation decoder, an ASPP atrous convolution pooling pyramid, a spatio-temporal feature fusion module, and a progressive refinement module.
[0018] The fine segmentation module based on spatio-temporal feature fusion predicts an accurate alpha matte based on all the video frame masks output by the mask time-domain propagation module and the original video frames, and uses the spatio-temporal information between frames to eliminate possible artifacts and flickering phenomena in video matting.
[0019] The described fine segmentation module based on spatio-temporal feature fusion performs the following operations on each original image F in the video i Execute the following: Take F i and the two adjacent original images F i-1 、F i+1 and the corresponding masks M i M i-1 、M i+1Separate into three groups of four-channel input data and input them into the fine segmentation encoder for multi-level feature extraction. The encoded features at the bottom layer of the fine segmentation encoder are input into the ASPP atrous convolution pooling pyramid for multi-scale feature extraction and fusion, and then the features are output to the bottom layer of the fine segmentation decoder for decoding layer by layer upwards. At the same time, each layer in the fine segmentation encoder will output the extracted feature maps. The feature maps at each level are output to the corresponding level of the spatio-temporal feature fusion module through skip connections for feature alignment and fusion. The spatio-temporal feature fusion module outputs the aligned and fused feature maps to the corresponding level of the fine segmentation decoder through skip connections, and adds them to the feature maps decoded at the previous level of the fine segmentation decoder for decoding at the current level. The features decoded at the previous level of the fine segmentation decoder refer to the features obtained by outputting from the ASPP atrous convolution pooling pyramid to the bottom layer of the fine segmentation decoder and then decoding layer by layer upwards. In addition, progressive refinement modules are respectively connected to the output parts of the second, third, and fifth layers of the fine segmentation decoder, so that the matte results will be progressively refined during the upward decoding process of the fine segmentation decoder, and finally an accurate transparency matte (alpha matte) is obtained.
[0020] A method of using an interactive video matte extraction system based on a mask propagation network is as follows:
[0021] Step (1): Cache the video to be processed in the form of video frames through the cache module, so as to obtain the original image of each frame;
[0022] Step (2): Coarsely segment the foreground object in the original input image through the interactive object coarse segmentation module, so as to extract the foreground object mask (Mask);
[0023] The user selects the original image of any frame from the cache module as the original input image, and obtains the corresponding indication map of the original input image by clicking or scribbling on the foreground object. The indication map refers to a single-channel map in which the values of the pixels clicked or scribbled by the user are set to 1, and is connected to the three channels of the original input image to jointly form a four-channel input, which is input into the image segmentation network; the image segmentation network coarsely segments the foreground object in the original input image according to the semantic information provided by the indication map, so as to extract the foreground object mask (Mask).
[0024] The user can optimize the mask by repeating clicking or scribbling until a sufficiently accurate mask is obtained, and then send it to the mask temporal propagation module.
[0025] Step (3): Obtain the masks of all frames in the video through the mask temporal propagation module.
[0026] After obtaining the mask (Mask) corresponding to a single-frame original image, the mask time-domain propagation module will perform mask propagation in both forward and backward time domains. Its principle is to predict the mask of the query frame based on the existing memory frames in the current cache module, then mark the query frame with the predicted mask as a memory frame and store it in the cache module, and take the next frame of the video as the new query frame, repeating the above operations until the next frame is a memory frame or the last frame of the video frame, which means that the masks of all frames have been obtained.
[0027] Step (4): The user determines whether they are satisfied with the masks of all the obtained frames.
[0028] If the user is not satisfied with the masks of all the obtained frames, select the original image corresponding to the unsatisfied mask as the original input image, and obtain the masks of all frames again through Step (2) and Step (3) until the user obtains the masks of all frames that satisfy them.
[0029] Step (5): When the user obtains the masks of all frames that satisfy them, predict the precise transparency mask through the fine segmentation module based on spatio-temporal feature fusion.
[0030] The fine segmentation module based on spatio-temporal feature fusion predicts the precise transparency mask according to all the video frame masks output by the mask time-domain propagation module and the original input image of each frame of the video stored in the cache module.
[0031] For each original image F in the video through the fine segmentation module based on spatio-temporal feature fusion i Perform the following operations: Let F i and the two adjacent original images F i-1 、F i+1 and the corresponding masks M i M i-1 、M i+1They are respectively grouped into three sets of four-channel input data and fed into the fine segmentation encoder for multi-level feature extraction. The encoded features at the bottom layer of the fine segmentation encoder are input into the ASPP atrous convolution pooling pyramid for multi-scale feature extraction and fusion, and then the features are output to the bottom layer of the fine segmentation decoder for decoding layer by layer upward. At the same time, each layer in the fine segmentation encoder will output the extracted feature maps, and the feature maps at each level are output to the spatio-temporal feature fusion module at the corresponding level through skip connections for feature alignment and fusion. The spatio-temporal feature fusion module outputs the aligned and fused feature maps to the corresponding level of the fine segmentation decoder through skip connections, and adds them to the feature maps decoded at the previous level of the fine segmentation decoder for decoding at the current level. The features decoded at the previous level of the fine segmentation decoder refer to the features obtained by outputting from the ASPP atrous convolution pooling pyramid to the bottom layer of the fine segmentation decoder and then decoding layer by layer upward. The output parts of the second, third, and fifth layers of the fine segmentation decoder are respectively connected to the progressive refinement module, and the progressive refinement matte results are finally used to obtain an accurate transparency matte (alpha matte).
[0032] The beneficial effects of the present invention are as follows:
[0033] Compared with the existing video matte extraction methods, the method of the present invention only needs to click or scribble on the foreground object of any frame of the video to achieve matte extraction of the foreground object of the entire video, without providing a trimap for each frame, greatly reducing the workload of users, and achieving the effect of advanced matte extraction algorithms. In addition, through the spatio-temporal feature fusion module, the spatio-temporal consistency problem between video frames is effectively solved, and the artifacts and flickering phenomena that may occur in the details of moving objects are effectively suppressed. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flowchart of the entire interactive video matte extraction method and system;
[0035] Figure 2 It is a flowchart of the spatio-temporal memory frame reader;
[0036] Figure 3 It is a flowchart of the mask temporal propagation module;
[0037] Figure 4 It is a structural diagram of the feature fusion network module;
[0038] Figure 5 It is a structural diagram of the fine segmentation network based on the spatio-temporal feature aggregation module;
[0039] Figure 6 It is a matte extraction effect diagram of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The embodiments of the present invention will be clearly and detailedly described below in conjunction with the accompanying drawings, so that the advantages, features and technical solutions of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.
[0041] As Figure 1 shown, an interactive video matting system based on a mask propagation network includes a cache module, an interactive image rough segmentation module, a mask temporal propagation module, and a fine segmentation module based on spatio-temporal feature fusion:
[0042] I. Cache module:
[0043] The cache module is used to cache the video in the form of video frames, so as to obtain the original input image of each frame; at the same time, it is used to cache the memory frames marked by the mask temporal propagation module.
[0044] II. Interactive target rough segmentation module:
[0045] As Figure 1 shown in the upper half, in this module, the user selects the original input image of any frame from the cache module, and obtains the corresponding indication map of the original input image by clicking or scribbling on the foreground target. The indication map is a single-channel map in which the values of the pixels clicked or scribbled by the user are set to 1, and the three channels of the original input image are connected to form a four-channel input, which is input to the image segmentation network; the image segmentation network performs rough segmentation on the foreground target in the original input image according to the semantic information provided by the indication map, so as to extract the foreground target mask (Mask).
[0046] In the embodiments of the present invention, a user operation interface GUI is built. The user interacts with the input image through the user operation interface GUI. The user can select two ways of clicking and scribbling for interaction to generate a single-channel indication map. In addition, the user can click and scribble multiple times to adjust the generated foreground target mask.
[0047] In the embodiments of the present invention, the image segmentation network uses the DeeplabV3+ network as the backbone. This network accepts six-channel input, where three channels are RGB images, one channel is a mask, and two channels are positive and negative scribbling maps. There are two situations for the mask. At the initial interaction, the mask is empty. When adjusting the foreground target mask that has been generated, the mask is a single-channel map containing the error area.
[0048] In the embodiments of the present invention, the DeeplabV3+ network is selected to be trained on the publicly available dataset PASCAL VOC 2012 Segmentation Competition. In order for the DeeplabV3+ network to learn the interaction method of user scribbles, it is necessary to collect training data of user scribble interactions, but this will bring a huge workload. Therefore, we set the random experience of whether the mask is empty to 0.5. When the mask is not empty, erosion and dilation are performed on the groundtruth alphamatte to obtain the training mask. The groundtruth alphamatte is provided by the publicly available dataset. Then, for the error area of the mask, a thinning or random Bezier curve strategy is used to generate the corresponding input scribbles to simulate user scribbles.
[0049] III. Mask temporal propagation module:
[0050] The mask temporal propagation module includes a spatio-temporal memory frame reader based on an attention mechanism;
[0051] The spatio-temporal memory frame reader based on the attention mechanism includes a memory encoder, a query encoder, and a mask decoder.
[0052] The original input image F selected by the interactive object rough segmentation module i and the generated mask M i predict the corresponding masks of all the remaining video frames.
[0053] In the spatio-temporal memory frame reader, when processing a video, generally, each frame of the picture is processed sequentially starting from the second frame. We regard the previous video frames with object masks as memory frames and the current video frame without a mask as a query frame.
[0054] The spatio-temporal memory frame reader based on the attention mechanism includes a memory encoder, a query encoder, and a mask decoder. For both the memory encoder and the query encoder, the ResNet50 is used as the backbone network for both of these two encoder networks, and the feature map of stage-4 (res4) of ResNet50 is used as the basic feature map for calculating the key-value feature map.
[0055] For the input part, the memory encoder adds additional input channels in the first convolutional layer. Its input is the image and the mask, while the input of the query encoder is only the image.
[0056] As Figure 2 shown, two convolutional layers are added to the end of both the memory encoder and the query encoder, generating two feature maps, namely the Key Map and the Value Map, respectively, for calculating the similarity of the key features between the query frame and the memory frame. The Key Map and the Value Map are represented by and respectively, where HW represents the size of the original image, and C k and C v are set to 128 and 512 respectively.
[0057] As can be seen from Figure 2 , for each memory frame T, the spatio-temporal memory frame reader calculates its key-value feature map through convolution operations and concatenates the outputs into the memory key map K M and the memory value map V M . The query key map K Q and the memory key map K M are matched through dot product, and the formula is as follows:
[0058] F = (K M ) T K Q (2)
[0059] where the entity F ∈ R THW*HW represents the affinity between the query point and the memory point.
[0060] For the spatio-temporal memory reading operation, first measure the similarity of all pixels between the query key map and the memory key map to calculate the weight of V M . Multiply V M by the weight and then add it to V Q and input them together into the mask decoder.
[0061] After the mask decoder obtains the output of the spatio-temporal memory reading operation, it reconstructs the target mask of the query frame. Using the mask refinement network proposed by Facebook as a building block, a convolutional layer and a residual block are used to compress the output of the spatio-temporal memory reading operation to 256 channels, and then the compressed read operation output is gradually enlarged through three mask refinement modules, doubling each time, and the mask refinement module at each stage is connected to the query encoder through a skip connection to obtain the output and feature map of the previous stage. The output of the last mask refinement module is fed into a convolutional layer to reconstruct the object mask. Each convolutional layer of the decoder uses a 3×3 convolutional filter to produce an output of 256 channels, and the last convolutional layer outputs a predicted mask at 1 / 4 the scale of the original image.
[0062] The steps of the embodiment of the present invention mainly include:
[0063] 1) Use the mask extracted by interactive segmentation and the single-frame original image as memory frames, and use the adjacent frames to be predicted as query frames.
[0064] 2) Perform a convolutional operation on the query frame to obtain the key feature map K Q and value feature map V Q of the query frame.
[0065] 3) Calculate the affinity of the pixels of the key feature map K Q of the query frame and the key feature map K M of the memory frame, and then multiply it by the value feature map V M of the memory frame to obtain an aligned value feature map V Q of the query frame.
[0066] 4) Add the aligned feature map and the value feature map of the query frame, and then decode it by the decoder to obtain the mask (Mask) of the query frame.
[0067] 5) Put the query frame into the cache (Cache) module as a memory frame. Continue to predict the next frame.
[0068] 6) Repeat the above operations until the next frame is a memory frame or the last frame of the video frame, and then stop the propagation.
[0069] The above spatio-temporal memory frame reader realizes the entire derivation process of obtaining the mask prediction of the query frame through the memory frame. In order to use a single-frame picture and mask to derive the masks of the entire video frame, it is also necessary to specify the corresponding mask propagation strategy.
[0070] As Figure 3 shown, we use the original input image F i and the mask M iAs a benchmark, it propagates in both forward and backward directions in the time domain dimension to other frames. In each direction, there are the following strategies:
[0071] Each time it propagates from the current frame to the next frame in this direction, marks the query frame with the predicted mask as a memory frame and stores it in the Cache module, and continues this propagation process until the next frame is a memory frame or the last frame of the video frame.
[0072] Through the mask time-domain propagation module, the user can obtain the masks corresponding to all video frames.
[0073] IV. Fine segmentation module based on spatio-temporal feature fusion:
[0074] The fine segmentation module based on spatio-temporal feature fusion predicts an accurate alpha matte according to all video frame masks output by the mask time-domain propagation module and the original video frame images, and uses the spatio-temporal information between frames to eliminate possible artifacts and flickering phenomena in video matting.
[0075] As can be seen from Figure 5 the fine segmentation module based on spatio-temporal feature fusion includes a fine segmentation encoder, a fine segmentation decoder, an ASPP atrous convolution pooling pyramid, a spatio-temporal feature fusion module, and a progressive refinement module.
[0076] Connect the three consecutive frame images F r-1 、F r 、F r+1 in the cache module and the corresponding masks M r-1 、M r 、M r+1 as three groups of 4-channel inputs, and respectively input them into the fine segmentation encoder to extract three groups of depth feature maps Fea r-1 、Fea r 、Fea r+1 of different scales. The encoding features at the bottom layer of the fine segmentation encoder are input into the ASPP atrous convolution pooling pyramid for multi-scale feature extraction and fusion, and then the features are output to the bottom layer of the fine segmentation decoder for layer-by-layer upward decoding. The three groups of feature maps Fea r-1 、Fea r 、Fea r+1, it is output to the spatio-temporal feature fusion module through skip connections for feature alignment and feature fusion, and then the aligned and fused features of different scales are respectively input into the corresponding levels of the fine segmentation decoder, and added to the feature map decoded by the previous level of the fine segmentation decoder for decoding at the current level, eliminating incorrect predictions and providing guiding information. The features decoded by the previous level of the fine segmentation decoder refer to the features obtained by outputting the ASPP atrous convolution pooling pyramid to the bottom layer of the fine segmentation decoder and then decoding layer by layer upward. In the second, third, and fourth layers of the fine segmentation decoder, progressive refinement modules are respectively introduced to gradually refine the transparency mask prediction results during the upward decoding process.
[0077] The ASPP atrous convolution pooling pyramid module in the embodiments of the present invention mainly draws on the practice method adopted in DeeplabV3+. Its main purpose is to capture semantic information of different scales according to atrous convolutions with different sampling rates and fuse them. The specific principle will not be elaborated here.
[0078] Next, the fine segmentation encoder, the fine segmentation decoder, the spatio-temporal feature fusion module, and the progressive refinement module will be introduced.
[0079] (1) Fine segmentation encoder and fine segmentation decoder:
[0080] As Figure 5 shown, a custom U-Net structure is used for the fine segmentation encoder and decoder networks. At the input part of the fine segmentation encoder, an RGB image plus a guidance map form a four-channel feature input S0 ∈ R 4 *512*512 , with the number of channels being 4 and the size set to 512 * 512 according to the input size, because the input image is usually cropped during network data loading. The input features pass through two layers of convolution to obtain a two-fold downsampled feature map S1 ∈ R 32 *256*256 , and after each layer of convolution, spectral normalization operation and batch normalization are performed. The purpose of this is to add Lipschitz constant constraints to the network to make the training more stable. Then, it passes through the second layer of convolution and the first residual block Res1 in sequence to obtain the feature S2 ∈ R 64*128*128 , and then passes through the second residual block Res2 of the third layer to obtain the feature S3 ∈ R 128*64*64 , and then passes through the third residual block Res3 of the fourth layer and the fourth residual block Res4 of the fifth layer to obtain a 16-fold downsampled feature map S4 ∈ R 256 *32*32 and a 32-fold downsampled map S5 ∈ R 512*16*16 .
[0081] In the fine segmentation decoder part, as Figure 5As shown on the right, the feature maps decoded at each layer are combined with the features output by the spatio-temporal feature fusion module at the corresponding layer, and then upsampled and decoded. In addition, at the second, third, and fifth layers, transparency masks of different scales are predicted through convolution. These predicted transparency masks and the predictions of the next layer are used as inputs to the progressive refinement module to infer the transparency mask of the next level.
[0082] (2) Spatio-temporal feature fusion module
[0083] As Figure 4 shown, the spatio-temporal feature fusion module includes a spatio-temporal feature alignment module and a spatio-temporal feature aggregation module. The spatio-temporal feature alignment module consists of two 3×3 ordinary convolutions and a 3×3 deformable convolution network, and the spatio-temporal feature aggregation module is composed of a cascaded channel attention network, a spatial attention network, and a global convolution network.
[0084] In this embodiment, the features Fea r-1 and Fea r of two adjacent frames are input into the spatio-temporal feature alignment module to obtain the feature Fea r-1 aligned with Fea r . After performing the same operation on Fea r and Fea r+1 , Fea r+1 is aligned with Fea r . Feature alignment is performed with a group of three adjacent frames of features to obtain two features aligned with Fea r .
[0085] After concatenating the two aligned features together, they are input into the spatio-temporal feature aggregation module. The input features are weighted by the channel attention network for their channels, and then the spatial information is weighted by the spatial attention network. Finally, the features are output after passing through a global convolution.
[0086] The above feature alignment relies on the principle characteristics of deformable convolution. Deformable convolution can apply weights to the offsets of the input features to align the features. For the pixel point P in the image frame I t at time t, the offset of P is represented by Δp, and the aligned feature F * is expressed as:
[0087]
[0088] Among them, k is the convolution sum position of the deformable convolution, wk is the weight at that position. Δpk is the linear offset of the feature at times t and t+Δt. Aligning the features and learning such offsets allows the model to automatically map the same or similar regions and pixels, and also encodes temporal information into the aligned features.
[0089] The spatio-temporal feature aggregation module processes the aligned features through an attention mechanism, guiding the model to utilize the importance of different channels and the regions of interest on a certain channel from both the channel dimension and the spatial dimension. The channel attention network in it uses a global average pooling layer to pool the input features, followed by a fully connected layer to calculate the channel attention weight map, and multiplies this map with the aligned features. The spatial attention network obtains two 1×H×W features from the input features through the global average pooling layer and the global max pooling layer, concatenates them together, reduces the number of channels through a convolution and a sigmoid activation layer, outputs a 1×H×W spatial attention weight map, then multiplies the original input features with the spatial attention weight map to obtain the feature map with spatial weights applied, and finally uses a 1×1 convolutional layer to reduce the number of channels and expands the receptive field through a global convolutional network.
[0090] (3) Progressive refinement module
[0091] The purpose of this module is to use high-level features to identify the internal regions of the object and use low-level features to depict the foreground information near the boundary, so as to improve the final matte extraction effect.
[0092] The principle is as follows:
[0093] Assume that the current decoder layer is l. Then the progressive refinement module first upsamples the transparency mask α output by the decoder at layer l-1 l-1 to the scale of the current layer α l Then, the self-guiding map g l (x,y) is obtained using the following formula:
[0094]
[0095] As can be seen from formula (4), the self-guiding map represents the unknown regions of the predicted mask in the previous layer. The self-guiding map is a single-channel map with pixel values consisting of 0 and 1, where 1 represents the unknown regions and 0 represents the determined regions. Through the self-guiding map, the determined regions of the mask α l at the current level are replaced by the determined regions of the mask α l-1 in the previous layer, so that both high-level and low-level information are utilized. The formula is as follows:
[0096] α l =α l g l +αl-1 (1 - g l ) (5)
[0097] After applying the progressive refinement module to the second layer, the third layer, and the fifth layer of the decoder respectively, as the features are decoded upward, the unknown area represented by the self-guiding map also shrinks, and the predicted transparency mask is gradually refined.
[0098] The above process is for predicting the transparency mask of a single video frame image. The fine segmentation module of spatio-temporal feature fusion is applied to each original image and the corresponding mask to obtain a fine matting result for the video.
[0099] Figure 6 This is the matting effect diagram of the embodiment of the present invention.
[0100] The fine segmentation module based on spatio-temporal feature fusion in the embodiment of the present invention is trained on the public video dataset provided by Deep Video Matting. This dataset contains 408 foreground pictures, 87 foreground videos, and a total of 6,659 background videos. I selected 48 video foregrounds and 231 picture foregrounds, and randomly selected 15 from the 6,659 background video datasets as the background. After synthesis, 4,185 training data are obtained as the training set. In addition, 248 video samples are selected from the validation set as the validation set.
Claims
1. An interactive video matting system based on a mask propagation network, characterized in that, it includes a cache module, an interactive image rough segmentation module, a mask temporal propagation module, and a fine segmentation module based on spatio-temporal feature fusion; The cache module is used to cache the video in the form of video frames to obtain the original input image of each frame; and is also used to cache the memory frames marked by the mask temporal propagation module; The interactive object rough segmentation module is used to interact with the input image. The interaction includes two interaction methods: clicking and scribbling. The user can choose any interaction method according to the actual situation. By single clicking or scribbling, the foreground object information of the original input image, that is, the indication map, is obtained. Combining it with the original input image and inputting it into the image segmentation network to obtain a preliminary mask; The user can optimize the mask by repeating clicking or scribbling until a sufficiently accurate mask is obtained and then send it to the mask temporal propagation module; The mask temporal propagation module includes a spatio-temporal memory frame reader based on an attention mechanism; the spatio-temporal memory frame reader based on an attention mechanism includes a memory encoder, a query encoder, and a mask decoder; After the mask temporal propagation module obtains the mask corresponding to a single-frame original image, it will perform mask propagation in both forward and backward temporal directions; The principle is to predict the mask of the query frame according to the existing memory frames in the current cache module, then mark the query frame with the predicted mask as a memory frame and store it in the cache module, and take the next frame of the video as a new query frame, repeating the above operations until the next frame is a memory frame or the last frame of the video frame, at which point the propagation stops, indicating that the masks of all frames have been obtained; The fine segmentation module based on spatio-temporal feature fusion includes a fine segmentation encoder, a fine segmentation decoder, an ASPP atrous convolution pooling pyramid, a spatio-temporal feature fusion module, and a progressive refinement module; The fine segmentation module based on spatio-temporal feature fusion predicts an accurate transparency mask according to all the video frame masks output by the mask temporal propagation module and the original video frame images, and uses the spatio-temporal information between frames to eliminate possible artifacts and flickering phenomena in video matting.
2. An interactive video matting system based on a mask propagation network according to claim 1, characterized in that, The specific propagation method is to use the current interactive frame as the memory frame and the adjacent frame as the query frame. Match through the key feature maps of the memory frame and the query frame, then multiply the value feature map of the memory frame by the weight generated by the key feature matching, and finally connect the value feature map of the query frame and send it to the mask decoder for decoding, and finally predict the mask of the query frame.
3. An interactive video matting system based on a mask propagation network according to claim 2, characterized in that, The described fine segmentation module based on spatio-temporal feature fusion processes each frame of the original image F in the video i as follows: Combine F i with two adjacent frames of the original image F i-1 and F i+1 , as well as the corresponding masks M i M i-1 and M i+1 respectively to form three groups of four-channel input data, which are fed into the fine segmentation encoder for multi-level feature extraction. The encoded features at the bottom layer of the fine segmentation encoder are input into the ASPP atrous convolution pooling pyramid for multi-scale feature extraction and fusion, and then the features are output to the bottom layer of the fine segmentation decoder for layer-by-layer upward decoding; at the same time, each layer in the fine segmentation encoder will output the extracted feature maps, and the feature maps at each level are output to the spatio-temporal feature fusion module at the corresponding level through skip connections for feature alignment and fusion. The spatio-temporal feature fusion module outputs the aligned and fused feature maps to the corresponding level of the fine segmentation decoder through skip connections, and adds them to the feature maps decoded at the previous level of the fine segmentation decoder for decoding at the current level; the features decoded at the previous level of the fine segmentation decoder refer to the features obtained by outputting from the ASPP atrous convolution pooling pyramid to the bottom layer of the fine segmentation decoder and then decoding layer by layer upward; in addition, progressive refinement modules are respectively connected to the output parts of the second, third, and fifth layers of the fine segmentation decoder, so that the matte results will be progressively refined during the upward decoding process of the fine segmentation decoder, and finally an accurate transparency mask is obtained.
4. An interactive video matting system based on a mask propagation network according to claim 1 or 2 or 3, characterized in that, The image segmentation network of the described interactive target rough segmentation module uses the DeeplabV3+ network as the backbone. This network accepts six-channel inputs, where three channels are RGB images, one channel is a mask, and two channels are positive and negative scribble maps. There are two cases for the mask. When initially interacting, the mask is empty. When adjusting the foreground target mask that has already been generated, the mask is a single-channel map containing the error region.
5. An interactive video matting system based on a mask propagation network according to claim 4, characterized in that, for the memory encoder and the query encoder, both of these encoder networks use ResNet50 as the backbone network, and use the feature map of stage-4 of ResNet50 as the basic feature map for calculating the key-value feature map; for the input part, the memory encoder adds additional input channels in the first convolutional layer, and its input is the image and the mask, while the input of the query encoder is only the image; Two convolutional layers are added to the ends of the memory encoder and the query encoder respectively to generate a key map and a value map for calculating the similarity of the key features between the query frame and the memory frame. The key map and the value map are represented by and respectively, where HW represents the original image size, and C k and C v are set to 128 and 512 respectively; For each memory frame T, the spatio-temporal memory frame reader calculates its key-value feature maps through convolution operations and concatenates the outputs into a memory key map K M and a memory value map V M , while the query key map K Q and the memory key map K M are matched through dot product, and the formula is as follows: F=(K M ) T K Q (2) where the entity F ∈ R THW*HW represents the affinity between the query point and the memory point; Perform a spatio-temporal memory reading operation. First, calculate the weights of V by measuring the similarity of all pixels between the query key map and the memory key map. M Multiply V M by the weights and then add the result to V Q and input the sum into the mask decoder. After the mask decoder obtains the output of the spatio-temporal memory reading operation, it reconstructs the target mask of the query frame. Using the mask refinement network proposed by Facebook as the building block, use a convolutional layer and a residual block to compress the output of the spatio-temporal memory reading operation to 256 channels, and then gradually amplify the compressed read operation output through three mask refinement modules, doubling the magnification each time, and each stage of the mask refinement module is connected to the query encoder through a skip connection to obtain the output and feature map of the previous stage; the output of the last mask refinement module is passed into a convolutional layer to reconstruct the object mask, and each convolutional layer of the decoder uses a 3×3 convolutional filter to produce 256-channel outputs, and the last convolutional layer outputs a predicted mask with a scale of 1 / 4 of the original image.
6. An interactive video matting system based on a mask propagation network according to claim 5, characterized in that, The described fine segmentation encoder and decoder network uses a custom U-Net structure. In the input part of the fine segmentation encoder, an RGB image plus a guidance map form a four-channel feature input \(S_0\in\mathbb{R}\) 4*512*512 , with 4 channels and a size of \(512\times512\) set according to the input size; the input features pass through two convolutional layers to obtain a two-fold downsampled feature map \(S_1\in\mathbb{R}\) 32*256*256 , and after each convolutional layer, spectral normalization and batch normalization are performed. The purpose of this is to add a Lipschitz constant constraint to the network to make the training more stable; then it passes through the second convolutional layer and the first residual block Res1 in sequence to obtain the feature \(S_2\in\mathbb{R}\) 64*128*128 , then passes through the second residual block Res2 in the third layer to obtain the feature \(S_3\in\mathbb{R}\) 128*64*64 , and then passes through the third residual block Res3 in the fourth layer and the fourth residual block Res4 in the fifth layer to obtain a 16-fold downsampled feature map \(S_4\in\mathbb{R}\) 256*32*32 and a 32-fold downsampled map \(S_5\in\mathbb{R}\) 512*16*16 ; In the fine segmentation decoder part, the feature map decoded by each layer is combined with the feature output by the spatio-temporal feature fusion module of the corresponding layer, and then upsampled and decoded; in addition, in the second layer, the third layer, and the fifth layer, transparency masks of different scales are predicted through convolution; these predicted transparency masks and the predictions of the next layer are used as the input of the progressive refinement module to deduce the transparency mask of the next level.
7. An interactive video matting system based on a mask propagation network according to claim 6, characterized in that, The spatio-temporal feature fusion module includes a spatio-temporal feature alignment module and a spatio-temporal feature aggregation module. The spatio-temporal feature alignment module consists of two 3×3 ordinary convolutions and a 3×3 deformable convolution, and the spatio-temporal feature aggregation module is composed of a series-connected channel attention network, a spatial attention network, and a global convolutional network; Input the features Fea of two adjacent frames r-1 and Fea r into the spatio-temporal feature alignment module to obtain Fea r-1 aligned with Fea r features. After performing the same operation on Fea r and Fea r+1 , Fea r+1 is aligned with Fea r features. Group the features of three adjacent frames for feature alignment to obtain two features aligned with Fea r ; After concatenating the two aligned features together, input them into the spatio-temporal feature aggregation module. The input features are weighted by the channel attention network for their channels, then weighted by the spatial attention network for their spatial information, and finally output features after passing through a global convolution. The above feature alignment relies on the principle characteristics of deformable convolution. Deformable convolution can apply weights to the offsets of the input features to align the features; For the image frame I at time t t For the pixel point P in * it, using Δp to represent the offset of P, the aligned feature F is expressed as: where k is the convolution sum position of the deformable convolution, and w k is the weight at this position; Δpk is the linear offset of the features at times t and t + Δt; The spatio-temporal feature aggregation module processes the aligned features through an attention mechanism, guiding the model to utilize the importance of different channels and the regions of interest on a certain channel from the channel dimension and the spatial dimension. Among them, the channel attention network uses a global average pooling layer to perform pooling operations on the input features, followed by a fully connected layer to calculate the channel attention weight map, and multiplies this map with the aligned features; the spatial attention network is to obtain two 1×H×W features from the input features through the global average pooling layer and the global max pooling layer, concatenate them together, reduce the channels through a convolution and a sigmoid activation layer, output a 1×H×W spatial attention weight map, then multiply the original input features with the spatial attention weight map to obtain the feature map with spatial weights applied, and finally use a 1×1 convolution layer to reduce the number of channels and expand the receptive field through a global convolutional network.
8. An interactive video matting system based on a mask propagation network according to claim 7, characterized in that, The step-by-step refinement module uses high-level features to identify the internal regions of the object and uses low-level features to depict the foreground information near the boundary, so as to improve the final matting effect; The principle is as follows: Assume that the current decoder layer is \(l\). Then, the progressive refinement module first upsamples the transparency mask \(\alpha\) output by the decoder of layer \(l - 1\) to the scale of the current layer \(\alpha\). Then, the self-guidance map \(g\) is obtained using the following formula: l-1 upsampled to the scale of the current layer \(\alpha\); l and then the self-guidance map \(g\) is obtained using the following formula l (x, y): As can be seen from formula (4), the self-guiding map represents the unknown region of the predicted mask in the previous layer. The self-guiding map is a single-channel map with pixel values composed of 0 and 1, where 1 represents the unknown region and 0 represents the determined region; Mask the current level α through the self-guiding map l Replace the determined area of l-1 with the determined area of the upper-level mask α, so that both high-level and low-level information can be utilized; the formula is as follows: α l = α l g l + α l-1 (1 - g l ) (5) After applying the step-by-step refinement module to the second layer, the third layer, and the fifth layer of the decoder respectively, as the features are decoded upward, the unknown region represented by the self-guiding map will also shrink, and the predicted transparency mask is gradually refined.
9. A method for using an interactive video matting system based on a mask propagation network, characterized in that, The steps are as follows: Step (1), cache the video to be processed in the form of video frames through the cache module, so as to obtain the original input image of each frame; Step (2), perform rough segmentation on the foreground object in the original input image through the interactive object rough segmentation module, so as to extract the foreground object mask; The user selects an original input image of any frame from the cache module, and obtains an indication map corresponding to the original input image by clicking or scribbling on the foreground object. The indication map is a single-channel map in which the pixel values clicked or scribbled by the user are set to 1, and the three channels of the original input image are connected together to form a four-channel input, which is input into the image segmentation network; The image segmentation network performs rough segmentation on the foreground object in the original input image according to the semantic information provided by the indication map, so as to extract the foreground object mask; The user can optimize the mask by repeating clicking or scribbling until a sufficiently accurate mask is obtained, and then send it to the mask temporal propagation module; Step (3), obtain the masks of all frames in the video through the mask temporal propagation module; After the mask temporal propagation module obtains the mask corresponding to a single-frame original image, it will perform mask propagation in both forward and backward temporal directions; The principle is to predict the mask of the query frame based on the memory frames already existing in the current cache module, then mark the query frame with the predicted mask as a memory frame and store it in the cache module, and take the next frame of the video as the new query frame, repeating the above operations until the next frame is a memory frame or the last frame of the video frames, which means that the masks of all frames have been obtained; Step (4): The user determines whether they are satisfied with the masks of all the obtained frames; If the user is not satisfied with the masks of all the obtained frames, select the original image corresponding to the unsatisfied mask as the original input image, and obtain the masks of all frames again through Step (2) and Step (3) until the user obtains satisfactory masks of all frames; Step (5): When the user obtains satisfactory masks of all frames, predict an accurate transparency mask through the fine segmentation module based on spatio-temporal feature fusion; The fine segmentation module based on spatio-temporal feature fusion predicts an accurate transparency mask according to all the video frame masks output by the mask temporal propagation module and the original input image of each frame of the video stored in the cache module; For each frame of the original image F in the video through the fine segmentation module based on spatio-temporal feature fusion i perform the following operations: Take F i and the two adjacent frames of the original image F i-1 , F i+1 and the corresponding masks M i M i-1 , M i+1 respectively to form three groups of four-channel input data, and input them into the fine segmentation encoder for multi-level feature extraction. The encoding features at the bottom layer of the fine segmentation encoder are input into the ASPP dilated convolutional pooling pyramid for multi-scale feature extraction and fusion, and then the features are output to the bottom layer of the fine segmentation decoder for decoding layer by layer from bottom to top; at the same time, each layer in the fine segmentation encoder will output the extracted feature maps, and the feature maps at each level are output to the spatio-temporal feature fusion module at the corresponding level through skip connections for feature alignment and fusion. The spatio-temporal feature fusion module outputs the aligned and fused feature maps to the corresponding level of the fine segmentation decoder through skip connections, and adds them to the feature maps decoded at the previous level of the fine segmentation decoder for decoding at the current level; The features decoded at the upper level of the said fine segmentation decoder refer to the features obtained by outputting from the ASPP atrous convolution pooling pyramid to the bottom layer of the fine segmentation decoder and then decoding layer by layer upwards. The output parts of the second, third, and fifth layers of the fine segmentation decoder are respectively connected to the progressive refinement module, and the matte results are progressively refined to finally obtain an accurate transparency mask.
Citation Information
Patent Citations
Video object detection and segmentation method based on space-time double-branch network
CN110097568A
Video target segmentation method of space-time component graph
CN111652899A