A lightweight semi-supervised video object segmentation method based on multi-scale pyramid
Through the multi-scale pyramid structure and adaptive fusion module, the memory problem of the semi-supervised video object segmentation method in long video scenarios is solved, and the model is lightweight and efficient segmentation is achieved.
Patent Information
- Application Number
- CN202510582225.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing semi-supervised video object segmentation method faces the problem of high memory usage in long video scenarios, resulting in an exponential increase in GPU memory demand, affecting training and deployment efficiency.
The multi-scale pyramid structure and scale adaptive fusion module are adopted, combined with the pyramid structure long and short-term memory module, and the memory usage is optimized through the cross-attention mechanism, the propagation module is simplified, the redundant data storage is reduced, and the model is lightweighted.
While maintaining the segmentation accuracy unchanged, the model volume is significantly reduced, the inference speed is improved, and efficient processing and stable segmentation of long video sequences are achieved.
Smart Images

Figure CN120388321B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a lightweight semi-supervised video object segmentation method based on multi-scale pyramid, and belongs to the field of computer vision. Background Art
[0002] The task of semi-supervised video object segmentation is to segment the target object in subsequent frames frame by frame, given a manually annotated segmentation mask from one or more frames. Unlike unsupervised video object segmentation, which typically segments salient objects within a video frame, semi-supervised methods obtain the annotation information of specific target objects before inference. As deep learning evolves towards efficiency and intelligence, matching methods based on memory networks have become the current mainstream paradigm due to their end-to-end learning capabilities and dynamic memory storage mechanisms. Typical methods such as STM and STCN have become widely compared baseline models in semi-supervised video object segmentation methods, with their simple construction and high scalability.
[0003] However, methods based on memory networks still face the problem of insufficient GPU memory requirements. As the length of the video increases, the data that the memory storage needs to store also grows exponentially, which leads to an increase in the model's demand for GPU memory. The training and deployment processes are facing severe challenges. The AOT series of methods introduce object association mechanisms and Transformer architectures to propose a hierarchical propagation model, which significantly enhances the segmentation stability in complex scenes through the object association mechanism. The Cutie network proposed by Cheng et al. adopts a target-level memory mechanism and object-level Transformer to enhance the segmentation capability of complex scenes. The LiVOS method proposed by Liu et al. reduces the computational complexity and improves the processing efficiency of long video sequences by introducing a linear attention mechanism. MobileVOS proposed by Miles et al. uses knowledge distillation and contrastive learning techniques to further improve the lightweight and inference speed of the Transformer model.
[0004] Despite this, the aforementioned methods still face the problem of high memory usage in long video scenarios, requiring continuous memory storage maintenance. This leads to an exponential increase in GPU memory requirements as the video length increases. Therefore, a semi-supervised video object segmentation algorithm is needed to strike a balance between performance and computational overhead, significantly reducing the model size while maintaining relatively unchanged segmentation accuracy and improving inference speed. Summary of the Invention
[0005] Based on this, the present invention provides a lightweight semi-supervised video object segmentation method based on multi-scale pyramid.
[0006] The present invention is implemented by the following scheme:
[0007] Step 1: Build a semi-supervised video object segmentation dataset;
[0008] Step 2: Preprocess the semi-supervised video object segmentation dataset and perform model pre-training;
[0009] Step 3: Establish a lightweight semi-supervised video object segmentation model based on multi-scale pyramid:
[0010] A multi-scale pyramid structure is established to aggregate target object features and ID information at different scales and restore occluded targets.
[0011] A scale-adaptive fusion module is established to integrate the target object features of the ID branch and the video frame branch, achieving fusion at different scales to avoid the dominance of information from one branch.
[0012] A pyramid-structured long-short-term memory module is established to reduce memory usage, optimize the long-term propagation module, simplify the short-term propagation module, and use a cross-attention mechanism to filter and update the information of historical frames to avoid storing redundant data.
[0013] Step 4: Construct the loss function, update the model parameters, set the training parameters, train the model, and obtain the optimal weights;
[0014] Step 5: Detect the test set image based on the optimal weight and obtain the final segmentation result.
[0015] Furthermore, in step 1, a semi-supervised video object segmentation dataset is constructed: DAVIS2017, YouTube-VOS2019, and LVOS datasets for long videos are obtained respectively, and the dataset distribution is adjusted for model training and reasoning.
[0016] Furthermore, in step 2, the DAVIS2017, YouTube-VOS2019 and LVOS datasets are randomly cropped for data augmentation, and a memory library is pre-established, and the memory storage ratio is set to update the memory library every 2 frames during training and every 10 frames during testing.
[0017] Furthermore, the step 3 establishes a lightweight semi-supervised video object segmentation model based on a multi-scale pyramid: a multi-scale pyramid structure is established to aggregate target object features and ID information of different scales to restore occluded targets; a scale-adaptive fusion module is established to integrate target object features of the ID branch and the video frame branch, and fusion is achieved at different scales to avoid the information of a certain branch from dominating; a pyramid structure long-short-term memory module is established to reduce memory usage, optimize the long-term propagation module, simplify the short-term propagation module, and use a cross-attention mechanism to filter and update the information of historical frames to avoid storing redundant data.
[0018] Furthermore, the multi-scale pyramid structure includes a memory frame processing unit, a query frame processing unit and a multi-scale ID library; the memory frame processing unit is: for a memory frame containing a mask, the model retrieves the corresponding mask and the ID library of different scales S from the memory library i , and then pass the memory frame mask through S at different scales i Perform ID vector embedding to obtain The memory frame image is encoded by the memory frame encoder The query frame processing unit is: for the current frame without mask, the query frame encoder encodes The multi-scale ID library is initialized as follows: It contains M vectors of dimension C, which are masks for embedding different targets, and each target is randomly assigned a unique identification feature vector.
[0019] Furthermore, the scale-adaptive fusion module is composed of two autonomously learning gating units and two double-layer perceptrons; the gating unit autonomously learns the specific weight matrix of each branch, thereby realizing adaptive weighting of pixel values, and adaptively adjusts the amount of information of each channel through the gating mechanism, so that the key features are enhanced and the noise information is suppressed, thereby improving the discrimination ability of the target features; the scale-adaptive fusion module processes the ID branch and video frame branch data through the gating unit, and then processes them through the double-layer perceptron respectively, and obtains the final fusion feature through element-by-element multiplication and element-by-element addition.
[0020] Furthermore, the pyramid structure long short-term memory module is composed of three core components: short-term propagation, long-term propagation and self-propagation. Before the propagation process, feature preprocessing is required through ID embedding, and after the propagation process, feature expression needs to be enhanced with the help of residual connection. The feature preprocessing is: using the scale adaptive fusion module to obtain and multi-scale pyramid structure ID embedding is performed, and the embedded features are then injected into the short-term and long-term propagation modules respectively; the short-term propagation module strengthens the expression of repeated features by aggregating target information between adjacent frames; the long-term propagation module uses a cross-attention operation to efficiently aggregate target information from long-term memory frames to generate a reference frame and a dynamic frame; the enhanced feature expression process is as follows: the features of the short-term propagation module and the long-term propagation module are added, and then through a linear normalization layer and a residual connection, the features required for generating decoding are obtained with the help of the self-propagation module.
[0021] Furthermore, the memory representation of the reference frame in the long-term memory is After δ frames, the new memory frame is expressed as Add it to the long-term memory bank to form M r ,M t The cross-attention operation propagates the information in the past frame to the dynamic frame M by processing local and global features at different granularity levels. t+δ .
[0022] Furthermore, the step 4 constructs a loss function, updates the model parameters, sets the training parameters, performs model training, and obtains the optimal weight: the training process selects ResNet-50 and MoblieNetV2 as encoders, wherein ResNet-50 uses the pre-training parameters of MaskR-CNN; the training parameters: the pyramid structure long short-term memory module is set to four scales of 16×, 16×, 8×, and 4×, and the number of layers corresponding to each scale is 2, 1, 1, and 0 layers, respectively, wherein the feature map of the 4× scale is only used for target matching, and the feature maps of the other scales are used for generating segmentation masks.
[0023] Beneficial effects of the present invention:
[0024] The present invention proposes a lightweight semi-supervised video object segmentation method based on a multi-scale pyramid; the present invention establishes a multi-scale pyramid structure for aggregating target object features and ID information at different scales to restore occluded targets; the present invention establishes a scale-adaptive fusion module for integrating target object features of ID branches and video frame branches, realizing fusion at different scales to avoid the information of a certain branch from dominating; the present invention establishes a pyramid-structured long-short-term memory module to reduce memory usage, optimizes the long-term propagation module, simplifies the short-term propagation module, and adopts a cross-attention mechanism to filter and update the information of historical frames to avoid storing redundant data. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0026] Figure 1 This is a flowchart of the overall implementation of an embodiment of the present invention;
[0027] Figure 2 A framework diagram of the overall network model in an embodiment of the present invention;
[0028] Figure 3 This is a schematic diagram of a scale-adaptive fusion module in an embodiment of the present invention;
[0029] Figure 4 Schematic diagram of a pyramid-structured long short-term memory module according to an embodiment of the present invention;
[0030] Figure 5 This is a diagram showing the segmentation effect of an embodiment of the present invention on the test dataset DAVIS2017;
[0031] Figure 6 This is a diagram of the segmentation effect of an embodiment of the present invention on the test dataset YouTube-VOS2019. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present invention will be further described below with reference to the accompanying drawings in the embodiments of the present invention, but are not intended to limit the present invention.
[0033] like Figure 1 As shown, the present invention provides a lightweight semi-supervised video object segmentation method based on a multi-scale pyramid, the steps are as follows:
[0034] Step 1: Build a semi-supervised video object segmentation dataset;
[0035] Step 2: Preprocess the semi-supervised video object segmentation dataset and perform model pre-training;
[0036] Step 3: Establish a lightweight semi-supervised video object segmentation model based on multi-scale pyramid:
[0037] A multi-scale pyramid structure is established to aggregate target object features and ID information at different scales and restore occluded targets.
[0038] A scale-adaptive fusion module is established to integrate the target object features of the ID branch and the video frame branch, achieving fusion at different scales to avoid the dominance of information from one branch.
[0039] A pyramid-structured long-short-term memory module is established to reduce memory usage, optimize the long-term propagation module, simplify the short-term propagation module, and use a cross-attention mechanism to filter and update the information of historical frames to avoid storing redundant data.
[0040] Step 4: Construct the loss function, update the model parameters, set the training parameters, train the model, and obtain the optimal weights;
[0041] Step 5: Detect the test set image based on the optimal weight and obtain the final segmentation result.
[0042] Furthermore, in step 1, a semi-supervised video object segmentation dataset is constructed: DAVIS2017, YouTube-VOS2019, and LVOS datasets for long videos are obtained respectively, and the dataset distribution is adjusted for model training and reasoning.
[0043] Furthermore, in step 2, the DAVIS2017, YouTube-VOS2019 and LVOS datasets are randomly cropped for data augmentation, and a memory library is pre-established, and the memory storage ratio is set to update the memory library every 2 frames during training and every 10 frames during testing.
[0044] Furthermore, the step 3 establishes a lightweight semi-supervised video object segmentation model based on a multi-scale pyramid: a multi-scale pyramid structure is established to aggregate target object features and ID information of different scales to restore occluded targets; a scale-adaptive fusion module is established to integrate target object features of the ID branch and the video frame branch, and fusion is achieved at different scales to avoid the information of a certain branch from dominating; a pyramid structure long-short-term memory module is established to reduce memory usage, optimize the long-term propagation module, simplify the short-term propagation module, and use a cross-attention mechanism to filter and update the information of historical frames to avoid storing redundant data.
[0045] like Figure 2 As shown in the figure, the lightweight semi-supervised video object segmentation model based on multi-scale pyramid is as follows: first, the memory frame and query frame are processed separately, and a multi-scale ID library is established to embed the ID vector. The multi-scale ID is the multi-scale target identity; secondly, for the memory frame containing the mask, the model retrieves the corresponding mask and the ID library of different scales S from the memory library. i , and then pass the memory frame mask through S at different scales i Perform ID vector embedding to obtain The memory frame image is encoded by the memory frame encoder For the current frame without mask, obtain the encoding by querying the frame encoder Again, the multi-scale ID library is initialized as It contains M vectors of dimension C. In order to embed the masks of different targets, each target is randomly assigned a unique identification feature vector; from this, the memory frame is embedded And obtained by embedding the ID vector With the help of scale adaptive fusion module fusion, the enhanced features are obtained Finally, the enhanced features Embedded with query frame After three propagation, upsampling, and decoding processes, the final segmentation result is obtained. The process is defined as follows, where Represents the pyramid structure long short-term memory module, which plays a core role in the propagation and matching mechanism. + represents the upsampling function, R (i) is the image decoder, Represents the decoding features of the previous layer, and the whole process is performed recursively.
[0046]
[0047] like Figure 3 As shown, the scale adaptive fusion module divides the ID branch and video frame branch After being processed by the gate unit, the final fusion feature is obtained through a two-layer perceptron The gating unit autonomously learns the weight matrix specific to each branch, thereby achieving adaptive weighting of pixel values. The gating mechanism adaptively adjusts the amount of information in each channel, enhancing key features while suppressing noise information, thereby improving the discriminative ability of target features. The scale-adaptive fusion module reweights the original target features of the two branches, effectively reducing the influence of interfering features. This process enhances the stability of the fused features and achieves more effective key and feature value representation in the query frame. The process is defined as follows, where ⊙ represents an element-by-element multiplication operation, and GM1 and GM2 are two learnable gating units.
[0048]
[0049] like Figure 4 As shown in the figure, the pyramid structure long short-term memory module consists of three core components: short-term propagation, long-term propagation and self-propagation. Before the propagation process, feature preprocessing needs to be performed through ID embedding, and after the propagation process, feature expression needs to be enhanced with the help of residual connection. First, the scale-adaptive fusion module is used to obtain and multi-scale pyramid structure ID embedding is performed, and the embedded features are injected into the short-term and long-term propagation modules respectively; secondly, the short-term propagation module strengthens the expression of repeated features by aggregating target information between adjacent frames, and the long-term propagation module uses cross-attention operation to efficiently aggregate target information from long-term memory frames to generate reference frames and a unique dynamic frame; thirdly, the memory representation of the reference frame in the long-term memory is After δ frames, the new memory frame is expressed as Add it to the long-term memory bank to form M r ,M t ; From this, the cross attention operation propagates the information in the past frame to the dynamic frame M by processing local and global features at different granularity levels t+δ Finally, by adding the features of the short-term propagation module and the long-term propagation module, and then passing through the linear normalization layer and residual connection, the features required for generative decoding are obtained with the help of the self-propagation module.
[0050] The long-term propagation module will Set as query, Set as key and value respectively. First, the similarity between query and key is calculated using multi-scale Softmax; second, for the context information of the value, multi-scale context information is obtained through the gated aggregation unit; finally, the multi-scale context information is integrated with the similarity between query and key through element-by-element multiplication to obtain the output. The specific process is defined as follows, where TA represents the traditional attention operator and f q 、f k and f v are the projections of queries, keys, and values, respectively. is the scaling factor.
[0051]
[0052] The improved attention operation mechanism is expressed as follows, where SM represents the similarity measurement operator, SC represents the multi-scale context information, and GA represents the gated aggregation unit.
[0053]
[0054] In the process of multi-scale contextual information processing, memory context First, it is projected through a linear layer, and then a series of depthwise convolution operations are applied to the projected feature map to generate The process is expressed as follows, where Represented as feature maps at different levels, ReLU is the universal activation function, and DWConv is the depthwise convolution operator.
[0055]
[0056] The long-term propagation output is expressed as follows, where SA represents the scalable cross-attention operator, f q and f k are the query and key projections used to calculate the attention matrix, respectively.
[0057]
[0058] Furthermore, in step 4, the loss function is constructed, the model parameters are updated, the training parameters are set, and the model is trained to obtain the optimal weights: the lightweight semi-supervised video object segmentation model based on multi-scale pyramid is implemented using PyTorch. The AdamW optimizer is used during the training process, with a total of 100,000 iterations, a batch size of 8, and a sequence length of 4. The initial learning rate is set to 2×10 on DAVIS2017 and YouTube-VOS2019. -4 , set to 1×10 on the LVOS dataset -5 , the memory bank is updated every δ frames, which is set to 2 frames and 10 frames during training and testing respectively.
[0059] Furthermore, in the training process of step 4: ResNet-50 and MoblieNetV2 are selected as encoders, wherein ResNet-50 uses the pre-trained parameters of Mask R-CNN.
[0060] Furthermore, the training parameters of step 4 are as follows: the pyramid structure long short-term memory module is set to four scales of 16×, 16×, 8×, and 4×, and the number of layers corresponding to each scale is 2, 1, 1, and 0 respectively. The feature map of the 4× scale is only used for target matching.
[0061] Furthermore, the loss function in step 4: When training the model, the loss function is composed of BCE loss and boundary IOU loss in equal proportions. The total loss function is expressed as follows, where L BCE represents the binary cross entropy loss, L B-IoU represents the boundary IOU loss, and α is the balance coefficient.
[0062] L=α·L BCE +(1-α)·L B-IoU
[0063] Furthermore, the step 5 detects the test set image based on the optimal weight to obtain the final segmentation result.
[0064] like Figure 5 、 6The figure shows a comparison of the segmentation results of our method. In the DAVIS2017 dataset, despite the relatively low segmentation difficulty, both Comparison Method 1 and Comparison Method 2 failed to correctly segment the vehicle's windshield wipers in the first video sequence. This is due to the persistent and recurring occlusion between the wipers and the rearview mirror, which poses a significant challenge to methods lacking long-term modeling capabilities. In contrast, our method, leveraging the long-term modeling capabilities of its multi-scale pyramid structure, accurately segmented the wipers in nearly all frames. In the second video sequence, the skier disappears in frame 25 and does not reappear until frame 50. Only our method successfully captures this change and achieves accurate segmentation, further demonstrating the advantages of its long-term modeling capabilities. In the YouTube-VOS dataset, the video content is diverse and the segmentation difficulty is relatively high. In the first video sequence, all methods easily segmented similar individuals with discontinuous spatial positions. However, when objects suddenly appear and overlap, other methods struggle to effectively handle these changes. Our method, leveraging its multi-scale modeling capabilities and feature fusion mechanism, fully leverages information from memorized frames to successfully segment similar individuals. In the second video sequence, the proposed method accurately segments the car and its components despite challenging lighting variations. This is due to the stability provided by the recursive design of the pyramid-structured long short-term memory module and the decoding module. Visual analysis results demonstrate that the proposed method performs well in complex scenes and outperforms other advanced video object segmentation methods in both segmentation accuracy and stability.
[0065] Tables 1 and 2 show the comparative experimental results of different methods on the YouTube-VOS and DAVIS17 datasets. The values in bold represent the optimal results for that metric. As can be seen from the table, our method improves segmentation accuracy while also increasing inference speed and enabling efficient modeling of long-term temporal relationships.
[0066] Table 1 Comparison results of the present invention with other advanced methods on YouTube-VOS
[0067]
[0068] Table 2 Comparison results of the present invention with other advanced methods on DAVIS17
[0069]
[0070] The above are specific embodiments of the present invention. It should be noted that the present invention is not limited to the above specific embodiments. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A lightweight semi-supervised video object segmentation method based on multi-scale pyramid, characterized by: The method comprises the following steps: Step 1: Build a semi-supervised video object segmentation dataset; Step 2: Preprocess the semi-supervised video object segmentation dataset and perform model pre-training; Step 3: Establish a lightweight semi-supervised video object segmentation model based on multi-scale pyramid: A multi-scale pyramid structure is established to aggregate target object features and ID information at different scales and restore occluded targets. The multi-scale pyramid structure includes a memory frame processing unit, a query frame processing unit and a multi-scale ID library; The memory frame processing unit is: for a memory frame containing a mask, the model retrieves the corresponding mask and the ID library of different scales from the memory library , and then pass the memory frame mask through different scales Embed the ID vector to get , the memory frame image is encoded by the memory frame encoder ; The query frame processing unit is: for the current frame without mask, the query frame encoder encodes ; The multi-scale ID library is initialized as follows: , which contains M vectors of dimension C, which are masks for embedding different targets. Each target is randomly assigned a unique identification feature vector; A scale-adaptive fusion module is established to integrate the target object features of the ID branch and the video frame branch, achieving fusion at different scales to avoid the dominance of information from one branch. The scale-adaptive fusion module consists of two autonomously learning gate units and two double-layer sensors; The gating unit autonomously learns the specific weight matrix of each branch to achieve adaptive weighting of pixel values. The gating mechanism adaptively adjusts the amount of information in each channel, so that key features are enhanced and noise information is suppressed, thereby improving the ability to discriminate target features. The scale-adaptive fusion module processes the ID branch and video frame branch data through the gating unit, and then processes them through the two-layer perceptron respectively, and obtains the final fusion feature through element-by-element multiplication and element-by-element addition. ; A pyramid-structured long-short-term memory module is established to reduce memory usage, optimize the long-term propagation module, simplify the short-term propagation module, and use a cross-attention mechanism to filter and update the information of historical frames to avoid storing redundant data. The pyramid-structured long short-term memory module consists of three core components: short-term propagation, long-term propagation, and self-propagation. Before the propagation process, feature preprocessing is required through ID embedding, and after the propagation process, feature expression is enhanced with the help of residual connections. The feature preprocessing is: using the scale adaptive fusion module to obtain and multi-scale pyramid structure Perform ID embedding and then inject the embedded features into the short-time and long-time propagation modules respectively; The short-term propagation module enhances the expression of repeated features by aggregating target information between adjacent frames; The long-term propagation module uses a cross-attention operation to efficiently aggregate target information from the long-term memory frame to generate a reference frame and a dynamic frame; The enhanced feature expression process is as follows: by adding the features of the short-term propagation module and the long-term propagation module, and then passing through a linear normalization layer and a residual connection, the features required for generating decoding are obtained with the help of the self-propagation module; Step 4: Construct the loss function, update the model parameters, set the training parameters, train the model, and obtain the optimal weights; Step 5: Detect the test set image based on the optimal weight and obtain the final segmentation result.
2. The lightweight semi-supervised video object segmentation method based on multi-scale pyramid according to claim 1, characterized in that: The memory representation of the reference frame in the long-term memory is , after The new memory frame after the frame is represented as , adding it to the long-term memory bank to form ; The crisscross attention operation propagates information from past frames to dynamic frames by processing local and global features at different granularity levels. .
3. The lightweight semi-supervised video object segmentation method based on multi-scale pyramid according to claim 1, characterized in that: The training process: ResNet-50 and MoblieNetV2 are selected as encoders, where ResNet-50 uses the pre-trained parameters of Mask R-CNN.
4. The lightweight semi-supervised video object segmentation method based on multi-scale pyramid according to claim 1, characterized in that: The training parameters are as follows: the pyramid structure long short-term memory module is set to four scales of 16×, 16×, 8×, and 4×, and the number of layers corresponding to each scale is 2, 1, 1, and 0 layers, respectively. Among them, the feature map of the 4× scale is only used for target matching, and the feature maps of the other scales are used for segmentation mask generation.
Citation Information
Patent Citations
Video semantic segmentation method based on ConvLSTM convolutional neural network
CN111860386A
Transform-based semi-supervised video target segmentation method
CN114429607A