A Multi-Task Learning Video Instance Segmentation Method Based on Spatiotemporal Information Enhancement

The multi-task learning framework with a progressive tracker, refinement compensator, and spatial interaction module addresses inefficiencies in spatiotemporal feature modeling and robustness issues, improving video instance segmentation precision and robustness.

CN120071223BActive Publication Date: 2025-07-15ZHONGKE (SHENZHEN) WIRELESS SEMICON CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510542202.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-15
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

When the existing video instance segmentation method handles dynamic video scenes, the spatiotemporal information modeling is complex and the calculation is large, and the robustness is poor. The task weight is unbalanced in multi-task learning, resulting in instance loss or tracking failure.

Method used

A multi-task learning video instance segmentation method based on space-time information enhancement is adopted, and a progressive tracker, a refined compensator and a spatial interaction module are used to combine multi-head attention mechanisms and self-attention mechanisms to perform video instance segmentation. Through interference information filtering and instance interaction information fusion, segmentation accuracy and robustness are improved.

Benefits of technology

It realizes efficient instance segmentation in dynamic video scenarios, improves the accuracy and robustness of video instance segmentation, can handle complex scenes and long-term tracking, and reduces instance loss and tracking failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071223B_ABST
    Figure CN120071223B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-task learning video instance segmentation method based on spatio-temporal information enhancement, which relates to the field of computer vision technology. Video instance segmentation is performed based on a multi-task learning video instance segmentation model including a progressive tracker, a refinement compensator, and a spatial interaction module. The progressive tracker is used to filter interference information from the segmentation instance representation of the current segment by using the denoised instance query of the previous segment based on the multi-head attention mechanism. The refinement compensator is used to dynamically adjust the attention area and mine relevant information from the global context. The spatial interaction module is used to calculate the correlation between the denoised instance representations. The present invention realizes online and semi-online instance segmentation of target objects in a video sequence. By combining the anti-interference instance representation of past frames and the temporal-based multi-head attention mechanism, the video temporal information is fully exploited. Through dynamically adjusting the attention area and the interactive information representation of instances between frames, the discriminative information of the tracker is enriched.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology. Specifically, it relates to a multi-task learning video instance segmentation method based on enhanced spatio-temporal information, which is applicable to online and accurate tracking of video instances. Background Art

[0002] Video instance segmentation is an important task in computer vision, aiming to perform instance-level segmentation on each frame of a video, that is, to identify different object instances and track their changes in the time dimension. It has wide applications in fields such as autonomous driving, surveillance analysis, artificial intelligence drones, or embodied robot vision. Traditional video instance segmentation methods often rely on static image segmentation techniques. However, in dynamic video scenarios, factors such as the movement of objects, changes in the external environment, and occlusion make the task more complex. Therefore, how to perform effective video instance segmentation under the synergistic effect of spatio-temporal information has become a hot issue in research.

[0003] However, existing video instance segmentation methods still face some challenges:

[0004] (1) The modeling of spatio-temporal information is often complex and computationally intensive. How to efficiently extract and fuse spatio-temporal features and avoid redundant calculations is still an urgent problem to be solved.

[0005] (2) Factors such as the rapid movement of objects in the video, long-term occlusion, and scale changes make the existing methods less robust in dealing with these situations. Especially when facing complex scenarios or long-term tracking, instance loss or tracking failure may occur.

[0006] (3) Multi-task learning may lead to conflicts between tasks in some cases. How to design a reasonable task weight and sharing mechanism to ensure the balance of each task is also a key issue. Summary of the Invention

[0007] The present invention aims to provide a multi-task learning video instance segmentation method based on enhanced spatio-temporal information, which can obtain more temporal context information while retaining real-time performance, thereby improving the accuracy and robustness of video instance segmentation.

[0008] To solve the above problems, the technical solutions adopted by the present invention are as follows:

[0009] The present invention provides a multi-task learning video instance segmentation method based on enhanced spatio-temporal information, which performs video instance segmentation based on a multi-task learning video instance segmentation model including a progressive tracker, a refinement compensator, and a spatial interaction module;

[0010] The progressive tracker is used to utilize the denoised instance query of the previous segment based on the multi-head attention mechanism Segmentation instance representation of the current segment Perform interference information filtering to obtain the denoised instance representation after filtering the interference information ;

[0011] The refinement compensator is used to query the denoised instance of the previous segment and the denoised instance representation after filtering the interference information as inputs, dynamically adjust the region of interest, mine relevant information from the global context, and output the refined compensation information ;

[0012] The spatial interaction module is used to query the denoised instance of the previous segment and the denoised instance representation after filtering the interference information as inputs, calculate the correlation between the denoised instance representations, and output the interaction information of the inter-frame instances ;

[0013] Fuse the refined compensation information and the interaction information of the inter-frame instances to obtain the denoised instance query of the current segment .

[0014] As a further description of the above technical solution, for a complete video sequence segmented into segments, and each segment contains frames, when , the multi-task learning video instance segmentation model runs in an online manner, and when , the multi-task learning video instance segmentation model runs in a semi-online manner.

[0015] As a further description of the above technical solution, use the pre-trained parameter frozen segmenter Mask2Former to perform an initial match on the video to obtain the segmentation instance representation of the current segment .

[0016] As a further description of the above technical solution, the progressive tracker includes interference filters, each interference filter includes a short-term temporal convolutional block, a multi-head attention mechanism, a self-attention mechanism, and a feed-forward neural network, where the instance identity of the current segment obtained by introducing the Hungarian match is used, and based on the similarity relationship between the denoised instance query of the previous segment, the linear change of the current segment with respect to the instance representation and the linear transformation , the denoised instance representation after filtering the interference information is obtained; the short-term temporal convolutional block is used to focus on the segmentation instance representation of the current segment and the denoised instance query of the previous segment Similarity, capture short-term temporal correlations in time series data, extract patterns and features within a short time window, and model the dynamic context of instances.

[0017] As a further description of the above technical solution, for a certain segment, the progressive tracker uses a loss function to supervise the network, performs binary matching on each element in the predicted value and the label one by one, and the optimal matching result is obtained based on minimizing the loss function, that is:

[0018]

[0019] Among them, is the matching cost between instances adopted by the pre-trained parameter frozen segmenter Mask2Former, is the index of the target instance, represents all possible matching schemes, is a constant used to adjust the matching cost, is the set of indices of the denoising instance queries that have been occupied, represents the index of the predicted instance for matching, represents the accurate annotation information matched with the th predicted object, represents the total number of accurate annotation information for the current segment, represents the indicator function, which takes the value of 1 when is in the set and 0 otherwise.

[0020] As a further description of the above technical solution, use the prediction of the pre-trained parameter frozen segmenter Mask2Former to replace the prediction of the progressive tracker for Hungarian matching to obtain the optimal matching result , and during the model update process, update the index set according to the index of the newly matched query to reflect the new matching situation.

[0021] As a further description of the above technical solution, the refinement compensator includes an enhanced local attention mechanism, a self-attention module, and a feed-forward neural network; the enhanced local attention mechanism only focuses on a small number of sampled points near the query point, and each query item is assigned fixed key information.

[0022] As a further description of the above technical solution, the interactive information between inter-frame instances

[0023]

[0024]

[0025] Among them, represents a feedforward neural network, represents a self-attention mechanism, represents normalization, represents an interactive attention matrix, indicating the asymmetric interaction information between different trajectories, represents a cross-attention mechanism, represents a mask matrix, is the sigmoid function, is an interactive bounding constant, represents the attention matrix obtained after N asymmetric convolutions.

[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0027] (1) For the task of video instance segmentation, a multi-task learning video instance segmentation framework including a progressive tracker, a refinement compensator, and a spatial interaction module is proposed, realizing online and semi-online instance segmentation of target objects in a video sequence.

[0028] (2) The progressive tracker combines the anti-interference instance representation of past frames and the temporal-based multi-head attention mechanism, fully mining the video temporal information, enhancing the deep semantic representation of the instance, and suppressing the problem of insufficient feature expressiveness caused by interference information in some frames, partial occlusion, etc.

[0029] (3) The refinement compensator and the spatial interaction module enrich the discriminative information of the tracker by dynamically adjusting the attention region and the interactive information representation between instances in frames, enabling it to better adapt to extreme occlusion and unconventional deformations caused by entangled targets.

[0030] To make the above objects, features, and advantages of the present invention more obvious and understandable, specific embodiments of the present invention are hereinafter given, and in conjunction with the accompanying drawings, the details are described as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0032] Figure 1 is the system model diagram of the method of the present invention.

[0033] Figure 2 is the schematic diagram of the progressive tracker of the method of the present invention.

[0034] Figure 3 Schematic diagram of the refinement compensator and spatial compensation module of the method of the present invention.

[0035] Figure 4 Experimental comparison result graph of the method of the present invention. Detailed implementation manners

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention.

[0037] The embodiments of the present invention propose a multi - task learning video instance segmentation method based on spatio - temporal information enhancement, and build a multi - task learning video instance segmentation model (Anti - distanbance Video Instance Segmentation Framework, ADVISOR) including a progressive tracker, a refinement compensator, and a spatial interaction module for video instance segmentation.

[0038] As Figure 1 shown, the present invention uses a pre - trained parameter - frozen segmenter (Mask2Former) for preliminary instance segmentation. The progressive tracker filters and associates the interference information of the current clip instance by combining the instance information of the previous clip provided by the refinement compensator and the spatial interaction module.

[0039] Most video instance segmentation algorithms use Hungarian matching for association, whether it is a query - based tracker or a clip - based tracker. For a complete video sequence, the present invention divides it into clips (each clip contains frames), that is,

[0040] clips .

[0041] When and , the model ADVISOR of the present invention runs in an online and semi - online manner respectively. ADVISOR reads the clips in sequence and outputs the segmentation instance representation. The segmentation instance representation is filtered by the progressive tracker to obtain the denoised instance representation. The denoised instance representation obtains the category and mask information outputs through the classification head and the mask head respectively. In addition, the segmentation instance representation of the current clip The instance representation obtained through the RCP (Refinement Compensator) and the Space Interaction Module (SIM) is saved in the memory. These two modules, RCP and SIM, can improve the noise reduction efficiency of the progressive tracker. Based on conventional matching, the progressive tracker can also perform inter-frame association on the segmented instance representations with IDs. While retaining the temporal information, the progressive tracker filters out interference information from the instance representation of the current frame, enabling online and semi-online video instance segmentation. In addition, the refinement compensator and the space interaction module respectively add refined instance information and interactive information between instances to counteract unconventional deformations caused by low frame rate, occlusion, etc., providing richer instance information for the progressive tracker and achieving a more accurate alignment effect.

[0042] The specific description of the video instance segmentation method according to the embodiments of the present invention is as follows:

[0043] 1. Progressive Tracker

[0044] The starting point for designing the progressive tracker is to make full use of temporal information while overcoming the lack of attention to interference information and achieving accurate positioning of spatial information. As Figure 2 shown, the progressive tracker is used to filter out interference information from the segmented instance representation of the current segment based on the multi-head attention mechanism and by using the noise-reduced instance query of the previous segment to obtain the noise-reduced instance representation after filtering out interference information , specifically as follows:

[0045] The interference filter is a key module of the progressive tracker, which uses the segmented instance representation of the current segment (i.e., the initially matched instance representation) and the noise-reduced instance query of the previous segment ( represents the index of the current segment) to effectively associate multiple instances and avoid the problem of rough alignment caused by the interaction of similar instances.

[0046] In the present invention, the Hungarian matching is used to perform an initial match on the instance representation output by the segmenter to obtain the segmented instance representation of the current segment . The segmented instance representation of the current segment contains interference information, and the progressive tracker filters out the interference information based on this rough tracking result. During the filtering process, the progressive tracker will refer to the noise-reduced instance query of the previous segment , and then use the RCP and SIM modules to achieve the fusion of refined compensation information and interactive information between instances.

[0047] The progressive tracker consists of interference filters. In the embodiments of the present invention,​ , each interference filter includes a short - term temporal convolution block, a multi - head attention mechanism, a self - attention mechanism, and a feed - forward neural network. The short - term temporal convolution block is the core component of the filter, which can reduce the noise or instability in the data and make the model more robust. The present invention uses the short - term temporal convolution block to focus on the current segment and the similarity with the previous segment , capture the short - term temporal correlation in the time - series data, extract the patterns and features within a short - time window, and model the dynamic context of the instance. In the online mode ( ), the present invention needs to face highly similar adjacent - frame instance representations. In this case, mutations caused by occlusion and low frame rate will introduce more noise, resulting in an unsatisfactory filtering effect for interference information. While the semi - online mode ( ) can effectively alleviate this problem: on the one hand, the progressive tracker avoids the large - scale accumulation and propagation of interference information, and at the same time, compared with the online mode, it can avoid tracking interruption caused by occlusion; on the other hand, the multi - frame aggregated denoising instance query provided by the segment can generate more accurate trajectory associations than frame - by - frame matching. The interference filter uses the multi - frame aggregated denoising instance query of the previous segment , the linear change of the current segment regarding the instance representation and the linear transformation to obtain the denoised instance representation after filtering out the interference information . As Figure 2 shown, each interference filter can be expressed as:

[0048] (1)

[0049] where DF represents the interference filter, represents the short - term temporal convolution, represents the multi - head attention mechanism, represents the instance identity of the current segment obtained by the Hungarian matching, and FFN represents the feed - forward neural network. The present invention introduces , that is, reduces the difficulty of the interference information filtering task, and avoids introducing too much noise when using the stored denoised instance query of the previous segment as the denoised instance query of the current segment

[0050] For a certain segment, the progressive tracker uses a loss function to supervise the network. The present invention performs binary matching on each element in the predicted value and the label . contains class and mask predictions of the denoised instance representation after query - based filtering of interference information; contains the A ground truth (accurate annotation information). Taking as the th predicted object that matches the accurate annotation information ( represents the index of the predicted instance used for matching), the optimal matching result can be obtained by minimizing the loss function:

[0051] (2)

[0052] where is the matching cost between instances adopted by Mask2former. is the index of the target instance. represents all possible matching schemes. is a constant used to adjust the matching cost. is the set of indices of the denoising instance queries that have been occupied. represents the total number of accurate annotation information in the current segment. represents the indicator function, which takes the value of 1 when the matching prediction is in the set and 0 otherwise. The present invention aims to prevent these already matched instance prototypes from matching with newly added ground truth again. To achieve this, the present invention modifies the matching cost by simply adding to the matching cost. The purpose of this part is to add a large cost to the already matched instance prototypes, making it less likely for them to match with new ground truth again. In addition, the present invention also sets an additional iterative constraint to judge the credibility before and after progressive tracking, that is, in the initial stage of the experiment, the present invention uses a segmenter (mask2former) with pre-trained parameters frozen to predict to replace the prediction of the progressive tracker for Hungarian matching to obtain the optimal matching result :

[0053] (3)

[0054] Finally, during the update process, the present invention will update the index set according to the indices of the newly matched queries to reflect the new matching situation:

[0055] (4)

[0056] is defined as and the union of the indices of a set of newly matched queries. The indices of the newly matched queries are composed of where is the number of newly matched queries. For predicted values that do not match the ground truth, the present invention determines them as a special label, that is, false targets or unrecognized targets.

[0057] 2. Refinement compensator

[0058] When querying cross-frame matching targets through denoising instances, the long-term entanglement between multiple targets will seriously affect the association process. The consequences caused by entanglement include drastic deformations caused by partial occlusion, unconventional deformations, and the accumulation of long-term interference information, etc. Neither the algorithm based on the attention mechanism nor the frame-by-frame association algorithm can show sufficient robustness, especially for long-term low-frame-rate video sequences. The essential reason is that traditional decoders cannot directly parse object context and unconventional deformations from dense spatio-temporal features, resulting in the propagation and accumulation of errors that damage the distinguishability of target features. Different from previous modules, the refinement compensator extracts refinement compensation information based on segments and stores it in the memory module, which has two advantages: on the one hand, it provides sufficient compensation information for the initial denoising instance query of the next segment; on the other hand, it provides compensation information as a reference for interference information filtering.

[0059] In the present invention, as Figure 3 shown on the left, the refinement compensator is used for the denoising instance query of the previous segment and the denoising instance representation after filtering interference information as inputs, dynamically adjusts the attention area, mines relevant information from the global context, and outputs the refinement compensation information , specifically as follows:

[0060] The refinement compensator includes an Enhanced Local Self-attention (ELSA), a self-attention module, and a feed-forward neural network FFN.

[0061] (5)

[0062] The ELSA in the refinement compensator adapts to irregular targets and can embed more powerful local details. Considering the influence of the number of object queries and the label embedding dimension, traditional self-attention modules will query all positions on the feature map, so when the feature map resolution is high, it will bring a high computational complexity. While ELSA only focuses on a small part of the sampling points near the query point and solves the problem of high computational complexity by assigning a fixed and small number of Keys to each Query. It realizes the adaptive adjustment of the attention area to better capture the local details of the target and realizes the extraction of the target shape information.

[0063] The introduction of the self-attention mechanism enables the model to perform non-linear weight allocation for instances in each frame according to different unconventional deformation situations, so that conventional and unconventional deformations in the denoising instance query have different attentions, distinguish the interference information in the instance representation, and then provide constraints on unconventional deformations for the recursive tracker.

[0064] 3. Space Interaction Module (SIM)

[0065] To mine the interaction information between instances and filter the interference information between instances based on the interaction information in the progressive tracker, the present invention designs a space interaction module.

[0066] As Figure 3 shown on the right, the space interaction module uses the denoising instance query of the previous segment and the denoising instance representation after filtering the interference information as inputs, calculates the correlation between the denoising instance representations, and outputs the interaction information of the inter-frame instances , specifically as follows:

[0067] In the current segment, asymmetric interactions between different instances can be extracted through cross-attention and interaction attention modules, that is, the correlations of different instance representations are different, which can be used as judgment information for dividing the interaction information space. To reduce the number of parameters and computational resources, asymmetric convolution is used to focus on more critical interactions, and then information screening is performed through matrix multiplication. Finally, the obtained instance interaction information is stored in Memory for the progressive tracker of the next segment.

[0068] Cross-attention not only focuses on the spatial information between instances, but also, thanks to the introduction of, it is also sensitive to the temporal relationship. Cross-attention captures the asymmetric interactions between different instances by establishing the spatio-temporal relationship between the previous segment and the current segment:

[0069] (6)

[0070] In the formula, is the cross-attention mechanism, and are respectively the transformed key matrix and value matrix, , and respectively represent from to , to , to linear transformation parameter matrix.

[0071] Inspired by self-attention, the present invention designs an interactive attention matrix , which uses the current representation for interactive self-attention modeling:

[0072] (7)

[0073] (8)

[0074] (9)

[0075] (10)

[0076] where represents linear transformation, is the instance representation after linear transformation of spatio-temporal interaction information. The present invention uses the interactive attention matrix to represent the asymmetric interaction information between different trajectories. The -th element in represents the influence of instance on instance

[0077] On this basis, in order to further consider the overall instance interaction information in the current segment, such as the interaction behavior between multiple instances, the present invention models higher-level interactions through the cascading of asymmetric convolutions:

[0078] (11)

[0079] represents the attention matrix obtained after asymmetric convolutions, , . Where and are asymmetric convolution kernels with dimensions [1, k] and [k, 1] respectively, represents PReLU. Asymmetric convolution can dynamically modify parameters during backpropagation, reducing redundant information between each feature set.

[0080] The present invention converts the obtained attention matrix into a mask matrix through an activation function and a sign function, which is used to cover the irrelevant information contained in the matrix:

[0081] (12)

[0082] where is the sigmoid function, is an interaction boundary constant used to partition the current interaction information.

[0083] To capture the significant interactions between instances, the present invention passes to retain high attention values for the instance interaction regions, is element-wise multiplication. The present invention normalizes all non-zero elements in and then obtains the final instance-to-instance interaction information through residual and self-attention:

[0084] (13)

[0085] where, represents the instance-to-instance interaction information. represents the denoised instance query of the previous segment. represents the denoised instance representation after filtering out interference information. represents a feed-forward neural network. represents a self-attention mechanism. represents normalization. represents a cross-attention mechanism. is an interaction attention matrix, representing the asymmetric interaction information between different trajectories. represents a mask matrix.

[0086] 4. Fuse the refined compensation information and the interaction information of the inter-frame instances to obtain the denoised instance query of the current segment.

[0087] The method ADVISOR of the present invention conducts a large number of quantitative and qualitative experiments with existing methods (MinVIS, VITA, GenVIS) on the benchmark dataset Youtube-VIS 2022, including the average precision (AP, AP50, AP75) and average recall rate (AR1, AR10) under multiple thresholds. The experiments show that a multi-task learning video instance segmentation method based on spatio-temporal information enhancement proposed by the present invention achieves state-of-the-art performance. The long-term video instance segmentation still has the best performance, as Figure 4 shown, demonstrating the excellence and necessity of the present invention.

[0088] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-task learning video instance segmentation method based on spatio-temporal information enhancement, characterized in that, Video instance segmentation is performed based on a multi-task learning video instance segmentation model including a progressive tracker, a refinement compensator, and a spatial interaction module; The progressive tracker is used to query the denoised instances of the previous segment based on the multi-head attention mechanism for the segmentation instance representation of the current segment to filter out interference information and obtain the denoised instance representation after filtering out interference information The refinement compensator is used to query the noise reduction instance of the previous segment and the noise reduction instance representation after filtering interference information As input, dynamically adjust the region of interest, mine relevant information from the global context, and output the refined compensation information The spatial interaction module is used to query the noise reduction instances of the previous segment and the noise reduction instance representation after filtering out interference information As inputs, calculate the correlation between the noise reduction instance representations, and output the interaction information of the inter-frame instances Refined compensation information and interaction information of the inter-frame instance are fused to obtain the noise reduction instance query of the current segment Among them, the progressive tracker includes n interference filters, each interference filter includes a short-term temporal convolutional block, a multi-head attention mechanism, a self-attention mechanism, and a feed-forward neural network, where the current segment instance identity ID obtained by introducing Hungarian matching is used, and based on the similarity relationship between the denoised instance query of the previous segment, the linear transformation K and the linear transformation V of the current segment with respect to the instance representation, the denoised instance representation after filtering out interference information is obtained; the short-term temporal convolutional block is used to focus on the segmentation instance representation of the current segment with the denoised instance query of the previous segment similarity, capture the short-term temporal correlation in the time series data, extract the patterns and features within a short time window, and model the dynamic context of the instance; For a certain segment, the progressive tracker uses a loss function to supervise the network, performs binary matching on each element in the predicted value and the label one by one, and the optimal matching result is obtained based on minimizing the loss function, that is: Among them, is the matching cost between instances adopted by the pre-trained parameter frozen segmenter Mask2Former, g is the index of the target instance, σ represents all possible matching schemes, Γ is a constant used to adjust the matching cost, and Φ i-1 is the set of indices of the denoising instance queries that have been occupied, and σ(g) represents the index of the predicted instance used for matching. represents the g-th predicted object that matches the accurate annotation information represents the total number of accurate annotation information of the current segment. represents the indicator function, which takes the value of 1 when σ(g) is in the set Φ i-1 and 0 otherwise.

2. The multi-task learning video instance segmentation method based on spatio-temporal information enhancement according to claim 1, wherein, For a complete video sequence divided into N v segments, and each segment contains N f frames, when N f = 1, the multi-task learning video instance segmentation model runs in an online manner. When N f = 3, the multi-task learning video instance segmentation model runs in a semi-online manner.

3. The multi-task learning video instance segmentation method based on spatio-temporal information enhancement according to claim 1, characterized in that Use the Mask2Former segmenter with frozen pre-trained parameters to perform an initial matching on the video, obtaining the segmentation instance representation of the current clip 4. The method for multi-task learning video instance segmentation based on spatio-temporal information enhancement according to claim 1, wherein Predictions using the Mask2Former segmenter with frozen pre-trained parameters Replace the predictions of the progressive tracker Perform Hungarian matching to obtain the optimal matching result During the model update process, update the index set Φ according to the indices of the newly matched queries to reflect the new matching situation.

5. The multi-task learning video instance segmentation method based on spatio-temporal information enhancement according to claim 1, wherein The refinement compensator includes an enhanced local attention mechanism, a self-attention module, and a feed-forward neural network; the enhanced local attention mechanism only focuses on a small part of the sampling points near the query point, and each query item is assigned fixed key information.

6. The multi-task learning video instance segmentation method based on spatio-temporal information enhancement according to claim 1, wherein Interaction information of inter-frame instances The calculation formula is as follows: Among them, FFN represents a feed-forward neural network, SAT represents a self-attention mechanism, Norm represents normalization, and A I represents an interactive attention matrix, representing the asymmetric interaction information between different trajectories, and A C represents a cross-attention mechanism, and A mask represents a mask matrix, is the sigmoid function, ξ ∈ [0, 1] is an interaction bounding constant, and A N represents the attention matrix obtained after N asymmetric convolutions.

Citation Information

Patent Citations

  • Video multi-target tracking method based on multi-scale channel feature aggregation

    CN117173217A