A Semi-Supervised Video Object Segmentation Method Based on Spatiotemporal Decoupling and Region Enhancement
By using the technology of space-time decoupling and region enhancement in the semi-supervised video target segmentation method, structured long-term attention and regional bar attention are calculated, and the shortcomings of the existing methods in the processing of space-time information and target importance are solved, and a more accurate and robust target segmentation effect is achieved.
Patent Information
- Application Number
- CN202410735421.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-06-07
AI Technical Summary
The existing semi-supervised video target segmentation method has shortcomings in processing spatiotemporal and target importance information, and cannot effectively extract the distinctive characteristics of the target, and is easily disturbed by other similar targets in the prospect.
The semi-supervised video target segmentation method based on spatiotemporal decoupling and region enhancement is adopted to calculate the similarity between the current frame features and the target area by calculating the structured long-term attention and regional bar attention of the current frame features, separate the spatiotemporal and target importance parts, and enhance the model's attention to the target.
It effectively solves the occlusion and background interference problems between similar objects, extracts the spatial and temporal relationships of pixels and significant information on the target, and improves the accuracy and robustness of target segmentation.
Smart Images

Figure CN118447436B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video image processing, and particularly relates to a semi-supervised video object segmentation method based on spatio-temporal decoupling and region enhancement. Background Art
[0002] Video object segmentation is an important computer vision task and has wide applications in fields such as video surveillance, autonomous driving, and human-computer interaction. With the emergence of a large number of videos, traditional supervised video object segmentation methods are becoming increasingly unsuitable for today's environment. Therefore, the present invention particularly focuses on semi-supervised video object segmentation, in which real annotations are only used in the first frame, aiming to track and segment objects in all subsequent frames.
[0003] Currently, most semi-supervised video object segmentation methods use matching-based methods. Although effective, they do not fully utilize the historical information in the model. To overcome this limitation, some methods use a memory mechanism to encode more frames and corresponding masks into feature embeddings and store them in a memory bank for auxiliary segmentation. Although this method is effective, there are still two problems: only focusing on spatio-temporal information and not considering target importance information, unable to extract the significant features of the target and unable to exclude the interference of other similar targets in the foreground to the target; ignoring the position information of the object in the mask and lacking the prior information of the object position in the memory frame, which will cause the interference of similar objects or noise in the background to the model matching process. Summary of the Invention
[0004] Aiming at the above deficiencies of the prior art, the present invention provides a semi-supervised video object segmentation method based on spatio-temporal decoupling and region enhancement to solve the above technical problems.
[0005] The present invention provides a semi-supervised video object segmentation method based on spatio-temporal decoupling and region enhancement. The method aims at the current frame features extracted by the Transformer network and the long-term memory frame masks in the semi-supervised video object segmentation model, and includes the following steps:
[0006] Calculate the structured long-term attention of the current frame features, and combine the region strip attention of the current frame features and the long-term memory frame masks to calculate the similarity between the current frame features and the target region;
[0007] Divide the similarity into a spatio-temporal part and a target importance part, and output the feature result after element-wise addition of the spatio-temporal part and the target importance part, where the spatio-temporal part is used to capture the spatio-temporal relationship between the current frame and the memory frame, and the target importance part is used for the model to be able to mine the main features of the target.
[0008] Further, it includes: calculating the structured long-term attention of the current frame feature, and combining the regional strip attention of the current frame feature and the long-term memory frame mask to calculate the similarity between the current frame feature and the target region, including:
[0009] The query feature Q is combined with the key feature K and the result of its global average pooling, and the feature sum is obtained by expanding to the input feature size;
[0010] The key feature K is refined through the regional strip attention module RSA, and the structured long-term similarity between the query feature Q and the key feature K is calculated by combining them, and the formula is:
[0011]
[0012] The sum of the value feature V and the recognition embedding E of the long-term memory mask is refined through the regional strip attention module RSA, and the structured long-term attention of the current frame feature is calculated by combining the structured long-term similarity SLTSim(Q, K) between the query feature Q and the key feature K, and the formula is:
[0013]
[0014] Further, the regional strip attention module RSA includes:
[0015] The input feature embedding is cropped according to the corresponding target region;
[0016] Horizontal and vertical strip pooling operations are performed on the cropped feature embedding to generate horizontal and vertical strip feature embeddings respectively;
[0017] The ReLU activation function is selected to obtain the importance degree of each feature point of the embedding and ensure that all values are non-negative;
[0018] The strip feature embedding is restored to the same size as the input feature embedding by means of zero padding, and 1×1 convolution is used to enhance the feature information of the feature embedding;
[0019] A Sigmoid function is selected to convert the padded feature embedding into a probability map, and the probability value that all feature embeddings belong to the target region is set;
[0020] The Hadamard product of the input feature embedding and the corresponding probability value is calculated as the output of the regional strip attention module RSA.
[0021] The beneficial effects of the present invention are as follows: The present invention proposes a semi-supervised video object segmentation structured Transformer with regional strip attention, which combines the Transformer architecture and the regional strip attention module, enabling the object segmentation model to solve two challenging scenarios: occlusion between similar objects and background interference, while extracting the spatio-temporal relationships of pixels and the significant information of the object. This structure makes full use of the valuable information in the long-term memory frames and the current frame, enabling the model to obtain not only the spatio-temporal information of the object but also the importance information of the object. To make full use of the target position information of the mask, this patent proposes a regional strip attention module to calculate the exact region of the target and calculate the strip attention within the target region in all long-term memory frames, thereby enhancing the model's attention to the target. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is an architecture diagram in the semi-supervised video object segmentation model of an embodiment of the present invention.
[0024] Figure 2 It is an architecture diagram of the structured long-term attention of an embodiment of the present invention.
[0025] Figure 3 It is an architecture diagram of the regional strip attention of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0027] The following explains the key terms that appear in the present invention.
[0028] Transformer: An encoder-decoder model based on the attention mechanism.
[0029] Semi-supervised: The supervision information is incomplete. Specifically, only the first frame of the video has annotation information, and all subsequent frames have no annotation information.
[0030] Sigmoid: An "S"-shaped function.
[0031] Norm: A normalization operation to reduce differences.
[0032] The embodiments of the present invention provide a semi-supervised video object segmentation method based on spatio-temporal decoupling and region enhancement. The method is directed to the current frame features and long-term memory frame masks extracted by the Transformer network in the semi-supervised video object segmentation model as shown in Figure 1 In the figure, ID represents the recognition mechanism in AOT; LN represents layer normalization. The encoder is implemented using the existing ResNet, and the decoder is implemented using the existing AOT. E = ID(Y|D) is the recognition embedding of the long-term memory mask, Y is the long-term memory mask, and D is the recognition library. For semi-supervised video object segmentation, this patent only uses annotation information in the first frame. Figure 1
[0033] Figure 2 As shown in the embodiments of the present invention include the following steps:
[0034] Calculate the structured long-term attention of the current frame features, and combine the region bar attention of the current frame features and the long-term memory frame masks to calculate the similarity between the current frame features and the target region;
[0035] Divide the similarity into a spatio-temporal part and a target importance part, and output the feature result after element-wise addition of the spatio-temporal part and the target importance part. Among them, the spatio-temporal part is used to capture the spatio-temporal relationship between the current frame and the memory frame, and the target importance part is used for the model to be able to mine the main features of the target.
[0036] To reduce the interference of the same type of targets and highlight the differences between different targets, this patent adopts the idea of image de-averaging. This patent proposes a structured long-term attention that decomposes the feature similarity between the query feature Q embedding and the key feature K embedding into two elements: a spatio-temporal part and a target importance part.
[0037] The query feature Q in the current frame features and the key feature K j are respectively processed by global average pooling GAP to obtain,;
[0038] Calculate the dot product of the query feature Q i and the key feature K j , and the formula is:
[0039]
[0040] where i and j respectively represent any spatial positions of Q and K, and T represents matrix transpose;
[0041]
[0041] Calculate the query feature Q i and the key feature K j for similarity. The formula is:
[0042]
[0043] Therefore, the corresponding matrices of [terms] only relate to [terms]. Therefore, decouple the above similarity formula into two terms to ensure that they do not affect each other, expressed as:
[0044]
[0045] Optionally, as an embodiment of the present invention, as Figure 2 shown, calculate the structured long-term attention of the current frame feature, and combine the regional strip attention of the current frame feature and the long-term memory frame mask to calculate the similarity between the current frame feature and the target area, including:
[0046] Extend the query feature Q, the result of the key feature K and its global average pooling to obtain feature [terms];
[0047] Refine [terms], the key feature K through the regional strip attention module RSA, and combine [terms] to calculate the structured long-term similarity between the query feature Q and the key feature K. The formula is:
[0048]
[0049] In this embodiment, in order to reduce the interference of the same type of targets and highlight the differences between different targets, this patent adopts the idea of image de-averaging. This patent proposes a structured long-term attention that decomposes the feature similarity between the query feature Q embedding and the key feature K embedding into two elements: a spatio-temporal part and a target importance part. The first term in the formula represents the spatio-temporal part, which explores the spatio-temporal relationship between the unique information of the current frame and the unique information of the long-term memory frame. The latter term is the target importance part, which utilizes the target importance information of the current frame and the long-term memory frame.
[0050] SLTSim(Q, K) represents the structured long-term similarity between Q and K. The first term in the above formula represents the spatio-temporal part, which explores the spatio-temporal relationship between the unique information of the current frame and the unique information of the long-term memory frame. The latter term is the target importance part, which utilizes the representative information of the current frame and the target importance information of the long-term memory frame.
[0051] Refine the sum of the value feature V and the recognition embedding E of the long-term memory mask through the regional strip attention module RSA, and combine the structured long-term similarity SLTSim(Q, K) between the query feature Q and the key feature K to calculate the structured long-term attention of the current frame feature. The formula is:
[0052] 。
[0053] Optionally, as an embodiment of the present invention, as Figure 3 shown, the regional bar attention module RSA includes:
[0054] Crop the input feature embedding according to the corresponding target region to obtain a feature embedding with the size of H ’ ×W ’ ×C;
[0055] Perform horizontal and vertical bar pooling operations on the cropped feature embedding to generate horizontal and vertical bar feature embeddings respectively, and apply convolutional layers with the same size as the two embeddings; in order to increase the receptive field of each feature point, 3×1 and 1×3 convolutions Conv are applied to the two embeddings respectively; Figure 3 where HSP is horizontal bar pooling, VSP is vertical bar pooling, and Exp is the expansion operation.
[0056] Select the ReLU activation function to obtain the importance degree of each feature point of the embedding, and ensure that all values are non-negative;
[0057] Restore the bar feature embedding to the same size as the input feature embedding by zero-padding, and use 1×1 convolution to enhance the feature information of the feature embedding;
[0058] Select a Sigmoid function to convert the padded feature embedding into a probability map, and set the probability values of all feature embeddings belonging to the target region;
[0059] Calculate the Hadamard product of the input feature embedding and the corresponding probability value as the output of the regional bar attention module RSA.
[0060] In this embodiment, in order to restore the embedding to the same size as the input, this study uses zero-padding to fill it, mainly considering that the ReLU activation function processes all values as non-negative, and the minimum value is zero. This method can maintain the continuity of the attention values between the target region and the non-target region, and regard the two regions as a unified whole.
[0061] In this embodiment, in the probability map, the probability values belonging to the target region are within the interval [0,1], and the probability values of the non-target region are 0.5. The higher the probability value, the more attention the corresponding point in the feature embedding receives.
[0062] Although the present invention has been described in detail by reference to the accompanying drawings and in conjunction with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope covered by the present invention. Or any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A semi-supervised video object segmentation method based on spatiotemporal decoupling and region enhancement, characterized in that: The method targets the current frame features and long-term memory frame masks extracted by the Transformer network in the semi-supervised video object segmentation model, and includes the following steps: Step 1: Calculate the structured long-term attention of the current frame feature, and combine the regional strip attention of the current frame feature and the long-term memory frame mask to calculate the similarity between the current frame feature and the target area, including: The query feature Q is combined with the key feature K and their global average pooling result and Expand to input feature size Get Features and ; Will , the key feature K is refined by the regional strip attention module RSA, and combined with and Calculate the structured long-term similarity between the query feature Q and the key feature K. The formula is: ; The sum of the value feature V and the recognition embedding E of the long-term memory mask is refined by the regional strip attention module RSA, and combined with the structured long-term similarity SLTSim(Q,K) between the query feature Q and the key feature K, the structural long-term attention of the current frame feature is calculated, and the formula is: ; Step 2: Divide the similarity into the spatiotemporal part and the target importance part, and output the feature result after element-by-element addition of the spatiotemporal part and the target importance part. The spatiotemporal part is used to capture the spatiotemporal relationship between the current frame and the memory frame, and the target importance part is used for the model to mine the main features of the target.
2. The method according to claim 1, characterized in that The regional strip attention module RSA comprises: The input feature embedding is cropped according to the corresponding target area; Perform horizontal and vertical strip pooling operations on the cropped feature embeddings to generate horizontal and vertical strip feature embeddings respectively; Select the ReLU activation function to get the importance of each feature point in the embedding and ensure that all values are non-negative; The strip feature embedding is restored to the same size as the input feature embedding by zero padding, and 1×1 convolution is used to enhance the feature information of the feature embedding; Select a Sigmoid function to convert the padded feature embedding into a probability map and set the probability value of all feature embeddings belonging to the target region; The Hadamard product of the input feature embedding and the corresponding probability value is calculated as the output of the regional strip attention module RSA.
Citation Information
Patent Citations
Video target segmentation method based on space-time decoupling attention mechanism
CN116416553A