Video object segmentation method based on mask feature aggregation and object enhancement
By designing a multi-scale mask feature aggregation unit and a target-enhanced attention mechanism in the video target segmentation method, the problem of failing to fully utilize the reference target mask features in the prior art is solved, and efficient target segmentation in complex scenarios is achieved.
Patent Information
- Application Number
- CN202210569043.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-05-24
AI Technical Summary
When using the reference target mask, the existing video target segmentation method fails to fully explore its features and fails to effectively integrate the target mask features and reference frame image features, resulting in poor segmentation effect in complex scenarios.
A video target segmentation method based on mask feature aggregation and target enhancement is designed. Through an optimized multi-scale mask feature aggregation unit and target enhancement attention mechanism, effective fusion and feature matching of target mask features and reference frame image features are achieved.
This method can accurately and quickly segment the target in complex environments, improve the contour accuracy and robustness to the target's rapid movement, deformation and occlusion, and achieve a high segmentation effect.
Smart Images

Figure CN115035437B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and in particular to a video target segmentation method based on mask feature aggregation and target enhancement. Background Art
[0002] In recent years, video object segmentation (VOS) has attracted widespread attention due to its wide applications in video manipulation and editing, among which the semi-supervised video object segmentation task has attracted particular attention. The semi-supervised video object segmentation task greatly simplifies video manipulation and editing applications. This task segments objects in a video sequence given an initial object mask. It allows users to provide an object mask only in the first frame, and then specific objects in the remaining frames will be automatically segmented without the user having to laboriously process the entire video.
[0003] The reference frame and the corresponding reference target mask are two crucial reference information in video target segmentation, which are mainly used to remember historical target information and match current target features. The reference frame refers to the original RGB image of the historical frame that has been segmented, which contains the complete information of the target and the background environment. The reference target mask refers to the target mask corresponding to the reference frame (the first frame is the true value, and the remaining frames are predicted values). It contains the edge and contour features of the target and clearly expresses the area and boundary of the target in the background environment.
[0004] Although the reference target mask helps the algorithm to accurately segment the target, how to properly utilize the reference target mask and effectively fuse it with the reference frame to better remember and match the target remains an open problem. Most previous methods only perform simple auxiliary processing on the reference target mask, and do not further explore the features in the target mask, nor do they further explore how to effectively fuse the target mask features with the reference frame image features. In addition, they also ignore the impact of the target mask on the feature matcher. For example, MaskTrack and RGMP simply concatenate the reference frame and the target mask in the channel dimension as the input of the network. FEELVOS directly uses the target mask to distinguish foreground and background pixels. Until recently, the use of reference target masks in video target segmentation has attracted attention. For example, SwiftNet generates target mask features through convolution and sub-pixel modules, and fuses them with reference frame image features to achieve efficient reference feature encoding. However, in addition to using target masks, these methods usually have other specific designs and different experimental settings to improve their performance, such as network structure, training and inference configuration, hyperparameters, other special modules, etc. It is difficult to find out the most effective way and whether there is a better way to use the reference object mask. In addition, previous research on the use of object masks mainly focused on the feature encoding part, ignoring the application of object masks in the feature matching process, which is also a very critical link in video object segmentation. Summary of the invention
[0005] In view of the above problems, the present invention proposes a video target segmentation method based on mask feature aggregation and target enhancement, which can achieve accurate and fast target segmentation in many difficult practical scenarios.
[0006] In order to achieve the above object, the present invention provides a video target segmentation method based on mask feature aggregation and target enhancement, comprising the following steps:
[0007] S1. Design and obtain an optimized multi-scale mask feature aggregation unit;
[0008] S2, using the target-enhanced attention mechanism to obtain a target-enhanced feature matching unit;
[0009] S3. Use the server to train the network model, optimize the network parameters by reducing the network loss function until the network converges, and obtain a video target segmentation method based on multi-scale mask feature aggregation and target enhancement;
[0010] S4. Segment a given target in a new video sequence using the video target segmentation method based on multi-scale mask feature aggregation and target enhancement.
[0011] Preferably, the step S1 specifically includes the following steps:
[0012] S11. Design a low-level fusion mask feature aggregation unit I1, using the same query encoding unit and reference encoding unit of the backbone network to extract query image key-value encoding pairs (k Q ,v Q ) and the reference key-value coding pair (k R ,v R ), the superscript Q refers to the query image, the superscript R refers to the reference set, the query encoding unit is an image feature encoder with an input channel of 3, and the end of the image feature encoder with an input channel of 3 is attached with two convolutions in parallel for generating a query image key-value encoding pair (k Q ,v Q ), the reference encoding unit is an image feature encoder with 4 input channels, and the end of the image feature encoder with 4 input channels is attached with two convolutions in parallel to generate a reference key-value encoding pair (k R ,v R ), the reference encoding unit first concatenates the reference frame and the reference target mask in the channel dimension and then sends them together to the image feature encoder with an input channel of 4. The mathematical expression is:
[0013] R = {Concate(I i ,M i )} N
[0014] Among them, I i ,M i They represent the i-th reference frame RGB image and reference target mask in the reference set R respectively; N is the reference set size; Concate represents the concatenation operation along the channel dimension;
[0015] S12. Design three high-level fusion mask feature aggregation units I2, I3, and I4 to aggregate the target mask or target mask feature with the image feature after the feature extraction stage for foreground discovery. The I2 is composed of a reference encoding unit and a query encoding unit. The query encoding unit is an image feature encoder. The end of the image feature encoder is attached with two parallel convolutions for generating a query image key-value coding pair (k Q ,v Q ), the reference encoding unit consists of an image feature encoder that shares features with the query encoder and a mask feature aggregation module to generate reference key-value encoding pairs (k R ,v R), the reference coding unit directly samples the original target mask, and uses a feature aggregation module to fuse it with the reference frame features output by the image feature encoder; an independent mask feature encoder is used in the reference coding unit of I3 to extract features from the target mask, and then the reference frame features output by the shared image feature encoder are fused using the feature aggregation module; I4 further shares the mask feature encoder and the image feature encoder based on I3;
[0016] S13, design four multi-scale fusion mask feature aggregation units I5, I6, I7, I8, the four feature extraction units are composed of a reference encoding unit and a query encoding unit, and output the query image key-value encoding pair (k Q ,v Q ) and the reference key-value coding pair (k R ,v R ) and the target mask feature F M , the I5 adopts the SwiftNet architecture, the image feature encoders in the reference coding unit and the query coding unit share features, and the reference coding unit fuses the downsampled target mask information into the reference frame features extracted by the image feature encoder after the first and fourth stages of the backbone network; the reference coding unit of I6 uses a separate mask feature encoder to extract the reference target mask feature F M , rather than simply downsampling, and then using the AFC module to fuse it with the reference frame features extracted by the image feature encoder in the first four stages of the backbone network, the image feature encoder is the main branch, and the image feature encoders in the query encoding unit and the reference encoding unit are not shared; the structure of I7 is basically the same as that of I6, with the mask encoder as the main branch, and the parameters of the image feature encoders in the query encoding unit and the reference encoding unit are shared; the difference between I8 and I6 is that only the parameters of the image feature encoders in the query encoding unit and the reference encoding unit are shared;
[0017] S14, using the default feature matching unit and decoding unit, after step S3 and step S4, compare the effects of various feature extraction units to obtain the optimal multi-scale mask feature aggregation unit I8.
[0018] Preferably, the feature aggregation module in step S12 is composed of two parallel convolution branches, one of which is composed of a 1×7 convolution and a 7×1 convolution connected in series; the other branch is composed of a 7×1 convolution and a 1×7 convolution connected in series;
[0019] Preferably, the step S2 specifically includes the following steps:
[0020] S21. Use the reference mask feature F generated by the mask encoderM To generate the target attention map w R ; Then the target attention map w R and reference value encoding feature v R Multiply to get the reference value encoding feature of the target enhancement
[0021] S22: According to the similarity between the query frame and the previous frame, the target attention map corresponding to the previous frame is Transform to the query frame and obtain the target attention map w corresponding to the query frame Q ; The target attention map w Q and query value encoding feature v Q Multiply to get the target enhanced query value encoding feature
[0022] S23, using the query image key-value encoding pair after target enhancement Retrieve a reference key-value coding pair The information in the query value is encoded with the feature The final matching features are obtained after splicing.
[0023] Preferably, the information retrieval process in step S23 is specifically as follows: firstly, the similarity between the query key feature and the reference key feature is calculated, and after the reference frame dimension is normalized, the reference frame value feature is weighted and summed as the weight, and then the reference frame value feature is concatenated with the query value encoding feature, that is:
[0024]
[0025] Where p and q represent pixels in the query key encoded feature and the reference key encoded feature respectively, [] represents concatenation, σ represents the Softmax function, and y is the output of the feature matching unit.
[0026] Preferably, the step S3 specifically includes the following steps:
[0027] S31, using the server to execute a training video segment generating unit to generate a training video segment of length T, where T≥2;
[0028] S32, using the server to execute the feature encoding unit, to perform the query image key-value encoding (k Q ,v Q ), reference key-value coding pair (k R ,v R ) and the reference target mask feature F M Extraction of
[0029] S33, using the server to execute the target enhanced feature matching unit described in step S2, according to the query image key-value coding pair (k Q ,vQ ) and the reference target mask feature F M To retrieve the reference key-value coding pair (k R ,v R ) to obtain the final matching features;
[0030] S34, using the server to execute the decoding unit and output the final segmentation result of the query frame;
[0031] S35, using the server to perform network training in an end-to-end manner; the mathematical expression of the segmentation loss function L is:
[0032] L(Y,M)=L ce (Y,M)+α·L IoU (Y,M)
[0033] in, represents the cross entropy loss; represents the mask intersection-over-union loss; Y represents the true value of the target mask; M represents the target mask prediction result; Ω represents the set of all pixels in the target mask; T represents the length of the training video segment; α is a hyperparameter;
[0034] S36. Utilize the server to optimize the objective function and obtain the local optimal network parameters.
[0035] Preferably, the step S31 specifically includes the following steps:
[0036] S311, randomly extracting T images at intervals from any video in the multiple video data sets;
[0037] S312, performing T different affine transformations on the T images respectively to form training video clips, where the affine transformations include translation, scaling, flipping, rotation and shearing.
[0038] Preferably, the step S32 is specifically as follows: for the reference set, using a shared image feature encoder and a mask feature encoder to extract features from the input reference frame image and the reference frame target mask prediction result respectively; then adding the features of each stage of the mask feature encoder and the features of the corresponding stage of the image feature encoder respectively through a squeeze excitation fusion module; then injecting the added features into the image feature encoder; finally, the image feature encoder outputs a reference key-value coding pair (k R ,v R ), the mask feature encoder outputs the mask feature F M ; Reference key-value coding pair (k R ,v R ) is directly stored in the memory; for the query frame, the image feature encoder is directly used to encode the features to obtain the query image encoding feature pair (k Q ,vQ ).
[0039] Preferably, the step S34 is specifically as follows: using multiple residual blocks as decoders, converting the matching features in the step S33 and the query image encoding feature pairs (k Q ,v Q ) is used as input, upsampling is performed by a factor of 2 at each stage, and the final segmentation result is output.
[0040] Preferably, the step S4 specifically includes the following steps:
[0041] S41, initializing the segmentation target, the mask of the target to be segmented is given in the first frame of the new video sequence, and the reference set is initialized using the first frame and its target mask; the segmentation starts from the second frame of the video sequence;
[0042] S42, the current frame image and the reference set are extracted by a feature extraction unit to obtain a query image encoding feature pair (k Q ,v Q ), refer to the key-value coding pair (k R ,v R ), the mask feature encoder outputs the mask feature F M ;
[0043] S43, executing the target enhanced matching unit to obtain matching features;
[0044] S44, pairing the matching feature in step S43 with the query image encoding feature in step S42 (k Q ,v Q ) is input to the decoding unit to obtain the target mask prediction result of the current frame;
[0045] S45. Put the current frame and its target mask prediction results into the reference set every 5 frames.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] The video target segmentation method based on mask feature aggregation and target enhancement provided by the present invention provides an optimal feature encoder configuration by designing and comparing eight different feature encoder designs: a multi-scale mask feature aggregation encoder, wherein the feature encoder makes full use of edge contour information in a target mask through multi-scale mask feature aggregation, strengthens the learning of target appearance representation, and makes the segmentation result have better contour accuracy; through target enhancement type attention, more attention is paid to a given target in the first frame, the interference of targets with similar appearance features and colors in the background is reduced, and the robustness of the method to challenges such as rapid movement, deformation and occlusion of the target is enhanced, so that the system can accurately segment targets in complex environments, and can accurately and quickly segment targets in many difficult actual scenarios, with a J&F value of 91.1% on the DAVIS2016 verification set, 85.5% on the DAVIS2017 verification set, and an overall score of 81.9% on the YouTube-VOS 2018 verification set, which has very good results. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Eight feature extraction unit diagrams designed for the present invention;
[0049] Figure 2 The target enhanced feature matching unit diagram designed for the present invention;
[0050] Figure 3 This is an algorithm framework diagram of a video target segmentation method based on multi-scale mask feature aggregation and target enhancement according to the present invention. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] In view of the problems and shortcomings existing in the prior art, the present invention proposes a video target segmentation method based on mask feature aggregation and target enhancement, which mainly includes the steps of multi-scale target mask feature aggregation encoder design, target enhancement feature matching unit design, model training and model inference.
[0053] The present invention proposes a video target segmentation method based on mask feature aggregation and target enhancement, comprising the following steps:
[0054] S1. Design and obtain an optimized multi-scale mask feature aggregation unit;
[0055] S2, using the target-enhanced attention mechanism to obtain a target-enhanced feature matching unit;
[0056] S3. Use the server to train the network model, optimize the network parameters by reducing the network loss function until the network converges, and obtain a video target segmentation method based on multi-scale mask feature aggregation and target enhancement;
[0057] S4. Segment a given target in a new video sequence using the video target segmentation method based on multi-scale mask feature aggregation and target enhancement.
[0058] The following is a detailed description of each step.
[0059] Step S1: Design and obtain an optimized multi-scale mask feature aggregation unit. Eight feature extraction units are designed through summarization, such as Figure 1 The figure shows eight kinds of feature extraction units designed in the present invention, and summarizes a feature extraction unit configuration with the best effect, namely, a multi-scale target mask feature aggregation unit.
[0060] The specific implementation process is as follows:
[0061] S11. Design a low-level fusion mask feature aggregation unit I1, which uses the same query encoding unit and reference encoding unit of the backbone network to extract query image key-value encoding pairs (k Q ,v Q ) and the reference key-value coding pair (k R ,v R ), the superscript Q refers to the query image, and the superscript R refers to the reference set; the query encoding unit consists of an image feature encoder with 3 input channels (with two convolutions in parallel at the end to generate query feature encoding pairs (k Q ,v Q ), while the reference encoding unit consists of an image feature encoder with 4 input channels (with two convolutions in parallel at the end to generate query feature encoding pairs (k R ,v R )) is composed of: the reference encoding unit first concatenates the reference frame and the reference target mask in the channel dimension and then sends them together to the image feature encoder with an input channel of 4. The mathematical expression is:
[0062] R = {Concate(I i ,M i )} N
[0063] Among them, I i ,M iThey represent the i-th reference frame RGB image and reference target mask in the reference set R respectively; N is the reference set size; Concate represents the concatenation operation along the channel dimension;
[0064] S12, design three high-level fusion mask feature aggregation units I2, I3, I4, aggregate the target mask or target mask feature with the image feature after the feature extraction stage for foreground detection; I2 consists of a reference encoding unit and a query encoding unit. The query encoding unit consists of an image feature encoder (with two convolutions connected in parallel at the end to generate a query feature encoding pair (k Q ,v Q ), the reference encoding unit consists of an image feature encoder that shares features with the query encoder and a mask feature aggregation module, which outputs a reference feature encoding pair (k R ,v R ); the reference coding unit directly samples the original target mask and uses the feature aggregation module to fuse it with the reference frame features output by the image feature encoder; unlike I2, an independent mask feature encoder is used in the reference coding unit of I3 to extract features from the target mask, and then the feature aggregation module is used to fuse the reference frame features output by the shared image feature encoder; I4 further shares the mask feature encoder and the image feature encoder on the basis of I3; specifically, the feature aggregation module consists of two parallel convolution branches, one of which is composed of a 1×7 convolution and a 7×1 convolution in series; the other branch is composed of a 7×1 convolution and a 1×7 convolution in series;
[0065] S13, design four multi-scale fusion mask feature aggregation units I5, I6, I7, I8, the four feature extraction units are composed of reference encoding units and query encoding units, and output reference feature encoding pairs (k R ,v R ) and query feature encoding pair (k Q ,v Q ) and the target mask feature F M I5 basically adopts the architecture of SwiftNet. The image feature encoders in the reference coding unit and the query coding unit share features. After the first and fourth stages of the backbone network, the reference coding unit fuses the downsampled target mask information into the reference frame features extracted by the image feature encoder. The reference coding unit of I6 uses a separate mask feature encoder to extract the reference target mask feature F. M, instead of simple downsampling, the AFC module is then used to fuse it with the reference frame features extracted by the image feature encoder in the first four stages of the backbone network (the image feature encoder is the main branch), and the image feature encoders in the query coding unit and the reference coding unit are not shared; the structures of I7 and I6 are basically the same, with the mask encoder as the main branch, and the parameters of the image feature encoders in the query coding unit and the reference coding unit are shared; compared with I6, I8 only shares the parameters of the image feature encoders in the query coding unit and the reference coding unit.
[0066] S14, using the default feature matching unit and decoding unit, after step S3 and step S4, compare the effects of various feature extraction units to obtain the optimal multi-scale mask feature aggregation unit I8.
[0067] The present invention designs eight different feature extraction units to find an effective way to use the target mask in the feature extraction unit. In order to test their effectiveness, the present invention provides a unified benchmark, which keeps the same architecture (feature matching unit and decoder), the same hyperparameters and training / inference configuration except the feature extraction unit. The present invention empirically summarizes two key findings from the comparison: (i) It is necessary to use a separate encoder to independently extract the target mask features, which is more beneficial than using the original mask or simple downsampling; (ii) Multi-scale aggregation of target mask features and reference frame image features can improve performance, indicating that both low-level and high-level mask features are useful. The present invention also finally selects the multi-scale target mask feature aggregation unit ( Figure 1 , instance I8) as our final feature extraction unit.
[0068] Step S2: Obtain a target-enhanced feature matching unit using a target-enhanced attention mechanism. Figure 2 The figure shows the target enhanced feature matching unit designed by the present invention, and the specific steps are as follows:
[0069] S21. Use the reference mask feature F generated by the mask encoder M To generate the target attention map w R ; Then the target attention map w R and reference value encoding feature v R Multiply to get the reference value encoding feature of the target enhancement Specifically:
[0070] w R =Conv(F M )
[0071]
[0072] Among them, Conv represents 1×1 convolution, represents the Hartmann product;
[0073] S22: According to the similarity between the query frame and the previous frame, the target attention map corresponding to the previous frame is Transform to the query frame and obtain the target attention map w corresponding to the query frame Q ; The target attention map w Q and query value encoding feature v Q Multiply to get the target enhanced query value encoding feature Specifically:
[0074]
[0075]
[0076] Among them, Conv represents 1×1 convolution, represents the Hadamard product, × represents matrix multiplication, σ represents the Softmax function, and the ss subscript -1 represents the last element of the reference set, that is, the previous frame of the current frame;
[0077] S23. Use the target enhanced query key-value coding pair Retrieve a reference key-value coding pair The information in the query value is encoded with the feature After splicing, the final matching features are obtained, specifically:
[0078]
[0079] Where p and q represent pixels in the query key encoded feature and the reference key encoded feature respectively, [] represents concatenation, σ represents the Softmax function, and y is the output of the feature matching unit.
[0080] The present invention explores the use of target masks in feature matching units. This is often overlooked in previous methods, but the present invention finds that it helps eliminate background interference. Commonly used feature matching units in video target segmentation use non-local attention. However, the attention between the query frame (current frame) and the reference set (reference frame and reference target mask) in this feature matching unit involves a large number of unnecessary feature pairs (such as the relationship between the background), and therefore contains too much background noise and interference. To address this problem, the present invention improves the above problem simply and effectively by explicitly introducing the reference target mask information into the feature matching unit. Unlike previous feature matching units, the present invention proposes a new target enhanced feature matching unit, which uses target enhanced attention to first generate a mask attention map using target mask features, and then use the mask attention map to enhance the target area and suppress the background (such as Figure 3 ).
[0081] Step S3: Use the server to train the network model, optimize the network parameters by reducing the network loss function, until the network converges, and obtain a video target segmentation method based on multi-scale mask feature aggregation and target enhancement. Figure 3 FIG. 1 is an algorithm framework diagram of a video target segmentation method based on multi-scale mask feature aggregation and target enhancement of the present invention, and the specific steps are as follows:
[0082] S31, using the server to execute the training video segment generation unit to generate a training video segment of length T, where T ≥ 2; specifically, randomly extracting T images at intervals from any video in the multiple video data sets, and performing T different affine transformations (combinations of translation, scaling, flipping, rotation, and shearing) on the T images to form a training video segment; or, randomly extracting an image from the image data set, and performing T different affine transformations to form a training video segment;
[0083] S32, using the server to execute the feature encoding unit to extract the query image key-value encoding pair, the reference key-value encoding pair, and the reference target mask feature, where the query image key-value encoding pair is (k Q ,v Q ), the reference key-value coding pair is (k R ,v R ), the reference target mask feature is F M , the superscript Q refers to the query image, the superscript R refers to the reference set, and the superscript M refers to the reference target mask. Specifically, for the reference set, the shared image feature encoder and mask feature encoder are used to extract features from the input reference frame image and the reference frame target mask prediction result respectively; then the features of each stage of the mask feature encoder and the features of the corresponding stage of the image feature encoder are added after passing through a squeeze-excitation fusion module (AFC module); then the added features are injected into the image feature encoder; finally, the image feature encoder outputs the reference key-value coding pair (k R ,v R ), the mask feature encoder outputs the mask feature F M ; Key-value coding pair (k R ,v R ) is directly stored in the memory; for the query frame, the image feature encoder is directly used to encode the features to obtain the query image encoding feature pair (k R ,v R );
[0084] S33, using the server to execute the target enhanced feature matching unit described in step S2; according to the query image key-value coding pair (k Q ,v Q ) and the reference target mask feature F MTo retrieve the reference key-value coding pair (k R ,v R ) to obtain the final matching features;
[0085] S34, using the server to execute the decoding unit and output the final segmentation result of the query frame; specifically, using multiple residual blocks as decoders, taking the matching features in step S33 and the query encoding features in step S32 introduced through the jump connection as input, performing 2 times upsampling in each stage, and finally outputting the final segmentation result.
[0086] S35, using the server to perform network training in an end-to-end manner; the mathematical expression of the segmentation loss function L is:
[0087] L(Y,M)=L ce (Y,M)+α·L IoU (Y,M)
[0088] in, represents the cross entropy loss;
[0089] represents the mask intersection-over-union loss; Y represents the true value of the target mask; M represents the target mask prediction result; Ω represents the set of all pixels in the target mask; T represents the length of the training video segment; α is a hyperparameter;
[0090] S36. Use the server to optimize the objective function and obtain the local optimal network parameters. Specifically, use the loss function L in step S35 as the objective function, use the AdamW optimizer to iteratively update the network parameters, and reduce the target loss function until it converges to the local optimum. The training is now completed, and the trained network weights for video target segmentation based on multi-scale mask feature aggregation and target enhanced attention are obtained.
[0091] Step S4: segment the given target in the new video sequence using the video target segmentation method based on multi-scale mask feature aggregation and target enhancement. The specific steps are as follows:
[0092] S41, initializing the segmentation target, the mask of the target to be segmented is given in the first frame of the new video sequence, and the reference set is initialized using the first frame and its target mask; the segmentation starts from the second frame of the video sequence;
[0093] S42, the current frame (query frame) image and the reference set are extracted by a feature extraction unit to obtain a query frame key-value coding pair (k Q ,v Q ), refer to the key-value coding pair (k R ,v R ), the mask feature encoder outputs the mask feature F M ;
[0094] S43, executing the target enhanced matching unit to obtain matching features;
[0095] S44, inputting the matching features in step S43 and the query encoding features in step S42 into a decoding unit to obtain a target mask prediction result of the current frame;
[0096] S45. Put the current frame and its target mask prediction results into the reference set every 5 frames.
[0097] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. It should therefore be understood that many modifications may be made to the exemplary embodiments and that other arrangements may be devised without departing from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in a manner different from that described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be used in other described embodiments.
Claims
1. A video target segmentation method based on mask feature aggregation and target enhancement, characterized in that: The following steps are involved: S1. Design and obtain an optimized multi-scale mask feature aggregation unit; S2, using the target-enhanced attention mechanism to obtain a target-enhanced feature matching unit; S3. Use the server to train the network model, optimize the network parameters by reducing the network loss function until the network converges, and obtain a video target segmentation method based on multi-scale mask feature aggregation and target enhancement; S4, segmenting a given target in a new video sequence using the video target segmentation method based on multi-scale mask feature aggregation and target enhancement; The step S1 specifically includes the following steps: S11, design a low-level fusion mask feature aggregation unit I1; S12, design three high-level fusion mask feature aggregation units I2, I3, I4; S13, design four multi-scale fusion mask feature aggregation units I5, I6, I7, I8, the four feature extraction units are composed of a reference encoding unit and a query encoding unit, and output the query image key-value encoding pair (k Q ,v Q ) and reference key-value coding pairs (k R ,v R ) and the target mask feature F M , the I5 adopts the SwiftNet architecture, the image feature encoders in the reference coding unit and the query coding unit share features, and the reference coding unit fuses the downsampled target mask information into the reference frame features extracted by the image feature encoder after the first and fourth stages of the backbone network; the reference coding unit of I6 uses a separate mask feature encoder to extract the reference target mask feature F M , rather than simply downsampling, and then using the AFC module to fuse it with the reference frame features extracted by the image feature encoder in the first four stages of the backbone network, the image feature encoder is the main branch, and the image feature encoders in the query encoding unit and the reference encoding unit are not shared; the structure of I7 is basically the same as that of I6, with the mask encoder as the main branch, and the parameters of the image feature encoders in the query encoding unit and the reference encoding unit are shared; the difference between I8 and I6 is that only the parameters of the image feature encoders in the query encoding unit and the reference encoding unit are shared; S14, using the default feature matching unit and decoding unit, after step S3 and step S4, compare the effects of various feature extraction units to obtain the optimal multi-scale mask feature aggregation unit I8; The step S2 specifically includes the following steps: S21. Use the reference mask feature F generated by the mask encoder M To generate the target attention map w R ; Then the target attention map w R and reference value encoding feature v R Multiply to get the reference value encoding feature of the target enhancement S22: According to the similarity between the query frame and the previous frame, the target attention map corresponding to the previous frame is Transform to the query frame and obtain the target attention map w corresponding to the query frame Q ; The target attention map w Q and query value encoding feature v Q Multiply to get the target enhanced query value encoding feature S23, using the query image key-value encoding pair after target enhancement Retrieve a reference key-value coding pair The information in the query value is encoded with the feature The final matching features are obtained after splicing.
2. The video target segmentation method based on mask feature aggregation and target enhancement according to claim 1, characterized in that: Step S11 specifically includes: using the same query encoding unit and reference encoding unit of the backbone network to extract the query image key-value encoding pair (k Q ,v Q ) and reference key-value coding pairs (k R ,v R ), the superscript Q refers to the query image, the superscript R refers to the reference set, the query encoding unit is an image feature encoder with an input channel of 3, and the end of the image feature encoder with an input channel of 3 is attached with two convolutions in parallel for generating a query image key-value encoding pair (k Q ,v Q ), the reference encoding unit is an image feature encoder with 4 input channels, and the end of the image feature encoder with 4 input channels is attached with two convolutions in parallel to generate a reference key-value encoding pair (k R ,v R ), the reference encoding unit first concatenates the reference frame and the reference target mask in the channel dimension and then sends them together to the image feature encoder with an input channel of 4. The mathematical expression is: R={Concate(I i ,M i )} N Among them, I i ,M i They represent the i-th reference frame RGB image and reference target mask in the reference set R respectively; N is the reference set size; Concate represents the concatenation operation along the channel dimension; Step S12 is specifically as follows: aggregating the target mask or target mask features with the image features after the feature extraction stage for foreground discovery, wherein I2 is composed of a reference encoding unit and a query encoding unit, wherein the query encoding unit is an image feature encoder, and the end of the image feature encoder is attached with two convolutions connected in parallel for generating a query image key-value encoding pair (k Q ,v Q ), the reference encoding unit consists of an image feature encoder that shares features with the query encoder and a mask feature aggregation module to generate reference key-value encoding pairs (k R ,v R ), the reference coding unit directly samples the original target mask, and uses a feature aggregation module to fuse it with the reference frame features output by the image feature encoder; an independent mask feature encoder is used in the reference coding unit of I3 to extract features from the target mask, and then a feature aggregation module is used to fuse the reference frame features output by the shared image feature encoder; I4 further shares the mask feature encoder and the image feature encoder based on I3.
3. The video target segmentation method based on mask feature aggregation and target enhancement according to claim 2, characterized in that: The feature aggregation module in step S12 is composed of two parallel convolution branches, one of which is composed of a 1×7 convolution and a 7×1 convolution connected in series; the other branch is composed of a 7×1 convolution and a 1×7 convolution connected in series.
4. The video target segmentation method based on mask feature aggregation and target enhancement according to claim 1, characterized in that: The information retrieval process in step S23 is specifically as follows: firstly, the similarity between the query key feature and the reference key feature is calculated, and then the reference frame value feature is weighted and summed as the weight after normalization of the reference frame dimension, and then concatenated with the query value encoding feature, that is: Where p and q represent pixels in the query key encoded feature and the reference key encoded feature respectively, [] represents concatenation, σ represents the Softmax function, and y is the output of the feature matching unit.
5. The video target segmentation method based on mask feature aggregation and target enhancement according to claim 1, characterized in that: The step S3 specifically comprises the following steps: S31, using the server to execute the training video segment generation unit to generate a training video segment of length T, wherein T≥2; S32, using the server to execute the feature encoding unit, to perform the query image key-value encoding (k Q ,v Q ), reference key-value coding pair (k R ,v R ) and the reference target mask feature F M Extraction of S33, using the server to execute the target enhanced feature matching unit described in step S2, according to the query image key-value coding pair (k Q ,v Q ) and the reference target mask feature F M To retrieve the reference key-value coding pair (k R ,v R ) to obtain the final matching features; S34, using the server to execute the decoding unit and output the final segmentation result of the query frame; S35, using the server to perform network training in an end-to-end manner; the mathematical expression of the segmentation loss function L is: L(Y,M)=L ce (Y,M)+α·L IoU (Y,M) in, represents the cross entropy loss; represents the mask intersection-over-union loss; Y represents the true value of the target mask; M represents the target mask prediction result; Ω represents the set of all pixels in the target mask; T represents the length of the training video segment; α is a hyperparameter; S36. Utilize the server to optimize the objective function and obtain the local optimal network parameters.
6. The video target segmentation method based on mask feature aggregation and target enhancement according to claim 5, characterized in that: The step S31 specifically includes the following steps: S311, randomly extracting T images at intervals from any video in the multiple video data sets; S312, performing T different affine transformations on the T images respectively to form training video clips, where the affine transformations include translation, scaling, flipping, rotation and shearing.
7. The video target segmentation method based on mask feature aggregation and target enhancement according to claim 5, characterized in that: The step S32 specifically includes: for the reference set, using the shared image feature encoder and the mask feature encoder to extract features from the input reference frame image and the reference frame target mask prediction result respectively; then adding the features of each stage of the mask feature encoder and the features of the corresponding stage of the image feature encoder through a squeeze excitation fusion module respectively; Then the added features are injected into the image feature encoder; finally, the image feature encoder outputs the reference key-value coding pair (k R ,v R ), the mask feature encoder outputs the mask feature F M ; Reference key-value coding pair (k R ,v R ) is directly stored in the memory; for the query frame, the image feature encoder is directly used to encode the features to obtain the query image encoding feature pair (k Q ,v Q ).
8. The video target segmentation method based on mask feature aggregation and target enhancement according to claim 5, characterized in that: The step S34 is specifically as follows: using multiple residual blocks as decoders, converting the matching features in the step S33 and the query image encoding feature pairs (k Q ,v Q ) is used as input, upsampling is performed by a factor of 2 at each stage, and the final segmentation result is output.
9. The video target segmentation method based on mask feature aggregation and target enhancement according to claim 1, characterized in that: The step S4 specifically comprises the following steps: S41, initializing the segmentation target, the mask of the target to be segmented is given in the first frame of the new video sequence, and the reference set is initialized using the first frame and its target mask; the segmentation starts from the second frame of the video sequence; S42, the current frame image and the reference set are extracted by a feature extraction unit to obtain a query image encoding feature pair (k Q ,v Q ), refer to the key-value coding pair (k R ,v R ), the mask feature encoder outputs the mask feature F M ; S43, executing the target enhanced matching unit to obtain matching features; S44, pairing the matching feature in step S43 with the query image encoding feature in step S42 (k Q ,v Q ) is input to the decoding unit to obtain the target mask prediction result of the current frame; S45. Put the current frame and its target mask prediction results into the reference set every 5 frames.
Citation Information
Patent Citations
Video target segmentation method of space-time component graph
CN111652899A
Quick real-time video target segmentation method based on bimodal interaction and state feedback
CN113807322A