Weak supervision video sequence segmentation method based on state space model

Through a weakly supervised video sequence segmentation method based on state space model, deep convolutional neural network and selective scanning visual model are used for feature extraction and comparison learning, which solves the problem of efficient computing and rich spatiotemporal information capture in the existing technology, and achieves efficient video segmentation effect.

CN120564098AActive Publication Date: 2025-08-29HARBIN INST OF TECH AT WEIHAI +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510641944.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-29
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

Existing video segmentation methods are difficult to capture rich spatiotemporal information while efficiently computed. In particular, traditional convolutional neural networks can only capture limited spatial features, while networks based on self-attention mechanisms have high computational complexity, resulting in high computational costs.

Method used

Using a weakly supervised video sequence segmentation method based on the state space model, spatial features are extracted through deep convolutional neural networks, combined with non-overlapping feature block division and selective scanning vision state space model encoder, pixel-level and block-level comparison learning is performed, multi-layer features are fused, and lightweight decoder predicted segmentation mask is generated.

Benefits of technology

While ensuring efficient calculations, it significantly enhances the ability to capture feature details and global semantic information, and generates high-discrimination and consistency semantic segmentation maps, which are suitable for scenes with difficult labeling such as medical image analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564098A_ABST
    Figure CN120564098A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning semantic segmentation, in particular to a weak supervision video sequence segmentation method based on a state space model, which realizes segmentation of an image sequence through weak supervision learning and a visual state space structure model, does not depend on expensive pixel-level annotation data, and improves the segmentation efficiency. The time sequence and space relation between the images in the image sequence is fully extracted, the calculation complexity and time consumption in image sequence segmentation are greatly reduced, the problems of poor image segmentation effect and poor model generalization ability are solved, and the segmentation effect and generalization ability of an image sequence segmentation model in a segmentation task are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning semantic segmentation, and in particular to a weakly supervised video sequence segmentation method based on a state-space model. Background Art

[0002] With the rapid development of computer vision and deep learning technologies, semantic segmentation, a core task in image analysis, has found widespread application in fields such as medical imaging, autonomous driving, and video surveillance. Currently, fully supervised semantic segmentation models based on pixel-level labels have achieved excellent segmentation accuracy. However, such models rely heavily on precise and costly annotated data, which typically needs to be provided by experienced pathologists or domain experts. This is not only expensive but also susceptible to subjective errors.

[0003] At the same time, the rapid growth of video data, along with the development of multimedia technology, has brought new challenges. Traditional image segmentation methods often only process single-frame static images, ignoring the rich temporal and spatial correlation information in video sequences, making it difficult to meet the requirements of high-precision semantic segmentation. Existing video segmentation methods aim to exploit the temporal correlation between consecutive frames, but traditional convolutional neural networks can only capture spatial features within a limited range and are unable to effectively handle long-range temporal dependencies. Furthermore, although networks based on self-attention mechanisms can capture long-range dependencies, their computational complexity increases significantly with sequence length, resulting in high computational costs.

[0004] Therefore, there is an urgent need for a segmentation method that can capture rich spatiotemporal information while ensuring efficient computation. Summary of the Invention

[0005] The purpose of the present invention is to provide a weakly supervised video sequence segmentation method based on a state-space model, aiming to solve the technical problem that the segmentation methods in the prior art are difficult to capture rich spatiotemporal information while ensuring efficient calculation.

[0006] To achieve the above object, the present invention adopts a weakly supervised video sequence segmentation method based on a state space model, comprising the following steps:

[0007] The image sequence is input into the deep convolutional neural network for spatial feature extraction to obtain the spatial feature matrix;

[0008] The extracted spatial feature matrix is ​​input into the decoder for decoding to obtain a pixel-level feature map;

[0009] Input the image sequence into the non-overlapping feature block partitioning layer, divide the input image sequence into image feature block sequences, and establish context and semantic relationships;

[0010] The image feature block sequence is input into the state space model encoder based on selective scanning vision to extract the temporal feature sequence;

[0011] The extracted image feature block sequence and pixel-level feature map are input into the contrastive learning state space model to perform pixel-level and block-level contrastive learning to obtain a contrastive learning feature matrix;

[0012] The spatial feature matrix, contrastive learning feature matrix and temporal feature sequence are input into the feature fusion module, and the pixel-level and block-level spatial features and temporal features are fused to obtain a fused feature sequence;

[0013] The fused feature sequence is input into a lightweight decoding head based on a convolutional neural network to predict the segmentation mask, and visual visualization is performed to generate a semantic segmentation map.

[0014] Among them, when the image sequence is input into the deep convolutional neural network for spatial feature extraction and the spatial feature matrix is ​​obtained:

[0015] The image sequence is input into a deep convolutional neural network to extract the spatial feature matrix. Specifically, the first three convolution stages of ResNet-50 are used as the backbone network to capture rich pixel-level features and provide basic support for subsequent feature encoding.

[0016] Among them, when the extracted spatial feature matrix is ​​input into the decoder for decoding and the pixel-level feature map is obtained:

[0017] 1×1 convolution is used to compress the number of channels in the spatial feature matrix, and bilinear upsampling is used to restore the original image size to generate a pixel-level feature map. This processing method preserves spatial information while reducing the amount of computation.

[0018] Among them, the image sequence is input into the non-overlapping feature block division layer, the input image sequence is divided into image feature block sequences, and the context and semantic relationship are established:

[0019] The image sequence is divided into 4×4 blocks through overlapping feature block partitioning layers to better capture local and global semantic information, which helps the model understand the image content and improve segmentation performance.

[0020] Among them, when the image feature block sequence is input into the state space model encoder based on selective scanning vision to extract the temporal feature sequence:

[0021] The image feature block sequence is input into the state space model encoder based on selective scanning vision to extract temporal features;

[0022] The selective scanning module adopts a multi-path cross-scanning strategy to expand the image block along four different paths and uses the state space structure to process the data of each path to construct a global receptive field;

[0023] The multi-level and multi-scale temporal feature extraction process enables the model to capture complex spatiotemporal relationships.

[0024] Among them, the extracted image feature block sequence and pixel-level feature map are input into the contrastive learning state space model to perform pixel-level and block-level contrastive learning to obtain the contrastive learning feature matrix:

[0025] The input image is considered as a bag, and each pixel in the image is considered as a separate instance. In this way, contrastive learning is performed between pixel-level and block-level features to achieve multi-level feature alignment and consistency. By constructing a mapping relationship between pixel-level and block-level features, it is ensured that during contrastive learning, pixel-level features can capture the context and semantic information contained in block-level features, while block-level features can further refine the detailed description of pixel-level features. The contrastive learning loss function is used to optimize the model so that similar features have greater similarity in the high-dimensional feature space, while different features have more obvious distinction. Through this joint contrastive learning of pixel and block levels, the ability to capture feature details and global semantic information can be significantly enhanced, and ultimately a contrastive learning feature matrix F with higher discrimination and good consistency is generated. c .

[0026] Among them, when the spatial feature matrix, contrast learning feature matrix and temporal feature sequence are input into the feature fusion module, the pixel-level and block-level spatial features and temporal features are fused to obtain the fused feature sequence:

[0027] The multi-layer feature input is linearly transformed into the fully connected layer, and the feature sequence after linear transformation is spliced ​​in the spatial dimension to form a comprehensive feature matrix, which is input into the multi-layer perceptron for multi-level feature fusion. The obtained feature is represented as:

[0028]

[0029] Each feature sequence F i Contains spatial information and semantic information at different scales. Linear represents a linear transformation layer. Concat represents concatenation on the feature dimension. MLP is a multi-layer perceptron containing multiple fully connected layers and activation functions.

[0030] Finally, the feature sequence fused by the multi-layer perceptron is used as the output of the decoding head.

[0031] Among them, when the fused feature sequence is input into the lightweight decoding head based on the convolutional neural network to predict the segmentation mask, and visual visualization is performed to generate a semantic segmentation map:

[0032] The feature sequence fused by the multi-layer perceptron is input into a lightweight decoding head based on a convolutional neural network. The convolutional neural network is used to process the fused feature sequence layer by layer to extract and learn features of different categories in the image. The output of the last convolution layer is activated by the softmax function to generate the probability distribution of each pixel belonging to each category, forming the final semantic segmentation mask.

[0033] The segmentation mask is used to generate a visually interpretable semantic segmentation map based on the prediction results, the semantic category label of each area in the image is displayed, and a visual image stream is generated and output; at the same time, the last convolution layer is visually interpreted, and the forward and backward propagation gradient weights of the last convolution layer are recorded using the two hook functions register_forward_hook and register_backward_hook. The Grad-CAM is generated using the gradient weights and normalized to obtain an interpretable weight map. The threshold Seg_threshold is used to select pixels, and pixels exceeding the threshold are set to white, otherwise they are set to black to obtain a segmentation map.

[0034] The present invention discloses a weakly supervised video sequence segmentation method based on a state-space model. First, an image sequence is input into a deep convolutional neural network for spatial feature extraction to obtain a spatial feature matrix; then the extracted spatial feature matrix is ​​input into a decoder for decoding to obtain a pixel-level feature map; then the image sequence is input into a non-overlapping feature block partitioning layer to divide the input image sequence into an image feature block sequence, and context and semantic relationships are established; then the image feature block sequence is input into a state-space model encoder based on selective scanning vision to extract a temporal feature sequence; and the extracted image feature block sequence and the pixel-level feature map are input into a contrastive learning state-space model to perform pixel-level and block-level contrastive learning to obtain a contrastive learning feature matrix; then the spatial feature matrix, the contrastive learning feature matrix and the temporal feature sequence are input into a feature fusion module to fuse the pixel-level and block-level spatial features and temporal features to obtain a fused feature sequence; finally, the fused feature sequence is input into a lightweight decoding head based on a convolutional neural network to predict a segmentation mask, and visual visualization is performed to generate a semantic segmentation map, thereby solving the technical problem that the segmentation method in the prior art is difficult to capture rich spatiotemporal information while ensuring efficient calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0036] Figure 1 is a flow chart of the weakly supervised video sequence segmentation method based on the state space model of the present invention;

[0037] Figure 2 1 is a network structure diagram of the weakly supervised video sequence segmentation method based on the state space model of the present invention;

[0038] Figure 3 Schematic diagram of a contrastive learning module of the weakly supervised video sequence segmentation method based on a state-space model of the present invention;

[0039] Figure 4 Schematic diagram of a contrastive learning module of the weakly supervised video sequence segmentation method based on a state-space model of the present invention;

[0040] Figure 5 Schematic diagram of a two-dimensional selective scanning state space model module of the weakly supervised video sequence segmentation method based on the state space model of the present invention;

[0041] Figure 6 It is a schematic diagram of a selective scanning module of the weakly supervised video sequence segmentation method based on a state space model of the present invention. DETAILED DESCRIPTION

[0042] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, but should not be understood as limiting the present invention.

[0043] See also Figures 1 to 6 ,in Figure 1 is a flow chart of the weakly supervised video sequence segmentation method based on the state space model of the present invention; Figure 2 1 is a network structure diagram of the weakly supervised video sequence segmentation method based on the state space model of the present invention; Figure 3 Schematic diagram of a contrastive learning module of the weakly supervised video sequence segmentation method based on a state-space model of the present invention; Figure 4 Schematic diagram of a contrastive learning module of the weakly supervised video sequence segmentation method based on a state-space model of the present invention; Figure 5 Schematic diagram of a two-dimensional selective scanning state space model module of the weakly supervised video sequence segmentation method based on the state space model of the present invention;

[0044] Figure 6 It is a schematic diagram of a selective scanning module of the weakly supervised video sequence segmentation method based on a state space model of the present invention.

[0045] The present invention provides a weakly supervised video sequence segmentation method based on a state space model, comprising the following steps:

[0046] S1. Input the image sequence into the deep convolutional neural network for spatial feature extraction to obtain the spatial feature matrix;

[0047] For this specific implementation, the input image is sent to ResNet50 for feature extraction.

[0048] S2. Input the extracted spatial feature matrix into the decoder for decoding to obtain a pixel-level feature map;

[0049] For this specific implementation, a 1×1 convolution is used to reduce the spatial feature matrix output channel size to 1, and a pixel-level feature map is created by restoring the image to its original size after a bilinear upsampling process.

[0050] S3, input the image sequence into the non-overlapping feature block division layer, divide the input image sequence into image feature block sequences, and establish context and semantic relationships;

[0051] For this specific embodiment, the image sequence is divided into blocks of size 4×4 by overlapping feature block partitioning layers.

[0052] S4, inputting the image feature block sequence into the state space model encoder based on selective scanning vision to extract the temporal feature sequence;

[0053] For this specific embodiment, the block sequence is input into the layered encoder, and the selective scanning module first expands the input block into a sequence along four different traversal paths (i.e., cross scanning), and uses a separate structured state space module with a selection mechanism and scanning calculation to process each block sequence in parallel, and then reshapes and merges the obtained sequences to form an output map (i.e., cross merging). By adopting a complementary one-dimensional traversal path, the selective scanning module enables each pixel in the image to effectively integrate information from all other pixels in different directions, thereby facilitating the establishment of a global receptive field in the 2D space. The temporal feature F is obtained. t .

[0054] S5. Input the extracted image feature block sequence and pixel-level feature map into a contrastive learning state space model to perform pixel-level and block-level contrastive learning to obtain a contrastive learning feature matrix;

[0055] For this specific embodiment, the input image is regarded as a package, and each pixel in the image is regarded as a separate instance. In this way, contrastive learning is performed between pixel-level and block-level features to achieve multi-level feature alignment and consistency. By constructing a mapping relationship between pixel-level and block-level features, it is ensured that when performing contrastive learning, pixel-level features can capture the context and semantic information contained in block-level features, while block-level features can further refine the detailed description of pixel-level features. The contrastive learning loss function is used to optimize the model so that features of the same type have greater similarity in the high-dimensional feature space, while features of different types have more obvious distinguishability. Through this joint contrastive learning of pixel and block levels, the ability to capture feature details and global semantic information can be significantly enhanced, and ultimately a contrastive learning feature matrix F with higher discrimination and good consistency is generated. c .

[0056] S6. Input the spatial feature matrix, the contrast learning feature matrix and the temporal feature sequence into the feature fusion module, and fuse the pixel-level and block-level spatial features and temporal features to obtain a fused feature sequence;

[0057] For this specific implementation, multi-layer features are input to the fully connected layer for linear transformation, and the feature sequences after linear transformation are spliced ​​in the spatial dimension to form a comprehensive feature matrix, which is input to the multi-layer perceptron for multi-level feature fusion. The obtained feature is represented as follows:

[0058] F fused =(MLP(Concat(Linear(F s ),Linear(F t ),Linear(F c ))))

[0059] Each feature sequence F i It contains spatial and semantic information at different scales. Linear represents a linear transformation layer, Concat represents concatenation on the feature dimension, and MLP is a multi-layer perceptron consisting of multiple fully connected layers and activation functions. Ultimately, the feature sequence fused by the multi-layer perceptron serves as the output of the decoding head.

[0060] S7. Input the fused feature sequence obtained by fusion into the lightweight decoding head based on convolutional neural network to predict the segmentation mask, and perform visual visualization to generate a semantic segmentation map.

[0061] For this specific implementation, the feature sequence fused by the multi-layer perceptron is input into a lightweight decoding head based on a convolutional neural network. The convolutional neural network is used to process the fused feature sequence layer by layer to extract and learn features of different categories in the image. The output of the last layer of convolution passes through the softmax activation function to generate the probability distribution of each pixel belonging to each category, forming the final semantic segmentation mask.

[0062] The segmentation mask is used to generate a visually interpretable semantic segmentation map based on the prediction results, displaying the semantic category labels for each region in the image, generating a visual image stream, and outputting it. At the same time, the last convolutional layer can be visually interpreted. The two hook functions register_forward_hook and register_backward_hook are used to record the forward and backward propagation gradient weights of the last convolution layer. The gradient weights are used to generate a Grad-CAM, which is then normalized to obtain an interpretable weight map. The threshold Seg_threshold is used to select pixels, setting them to white if they exceed the threshold, and black otherwise, to obtain the segmentation map.

[0063] See also Figure 2 The input image is respectively subjected to a deep convolutional network, a contrastive learning state space model, and a two-dimensional selective scanning state space model to extract features of different scales. The extracted features of different scales are then fused through a multi-layer perceptron. Finally, the fused feature sequence is input into a lightweight decoding head based on a convolutional neural network to predict the semantic segmentation mask.

[0064] Further, see Figure 2 To illustrate this embodiment, in step S6 described in this embodiment, the specific process of using the weakly supervised learning framework to train the network under the supervision of the true label and the predicted label is as follows: through the extracted comprehensive features, the model will calculate the probability of each pixel belonging to each category. These predictions are not directly compared with the pixel-level labels (because there are no pixel-level labels), but are aggregated into image-level predictions. The network is weakly supervised using the image-level labels. The loss of the weakly supervised network is defined as:

[0065]

[0066] Where m is the total number of categories, u represents the possible positive class labels in the predicted image; v represents the possible negative class labels in the predicted image; Y n represents image-level labels; represents the image-level prediction probability, and I is the indicator function.

[0067] See also Figure 3, a block partitioning module is used to divide the input image into non-overlapping blocks and perform the mapping process. The flattened sequence is normalized through layer normalization before being fed into the contrastive learning state-space model and deep convolutional layers. The contrastive learning state-space model module creates contrast between lesion images at the pixel and block levels, while the deep convolution function aims to preserve intricate details.

[0068] See also Figure 4 The contrastive learning state-space model module combines information from different granularities of pathology images for comparative modeling, explicitly exploring the relationship between patch-level and pixel-level features. Each pathology image is first expanded into a sequence along four different directions via a scan expansion operation. These sequences are then processed for feature extraction by the S6 module, ensuring that information from all directions is thoroughly scanned to capture distinct features. A contrast correlation operation is then designed to obtain a contrast map.

[0069] See also Figure 5 After the input data enters the feature extraction module, it is first standardized by layer normalization. The standardized data undergoes a series of linear transformation, normalization, selective scanning, activation and depth convolution processing. The processed data is added to the original input through residual connection, and then processed by LayerNorm again. The result enters the feedforward neural network for linear transformation and activation processing. The final output is added to the output of the previous layer through residual connection to obtain the final output of the module.

[0070] See also Figure 6 The selective scanning module first expands the input block into a sequence along four different traversal paths (i.e., cross scanning), processes each block sequence in parallel using a separate structured state space module with a selection mechanism and scan computation, and then reshapes and merges the resulting sequences to form the output map (i.e., cross merging). By adopting complementary one-dimensional traversal paths, the selective scanning module enables each pixel in the image to effectively integrate information from all other pixels in different directions, thereby facilitating the establishment of a global receptive field in 2D space.

[0071] The present invention uses a weakly supervised video sequence segmentation method based on a state space model. When it is used, the image sequence is first input into a deep convolutional neural network for spatial feature extraction to obtain a spatial feature matrix; the extracted spatial feature matrix is ​​then input into a decoder for decoding to obtain a pixel-level feature map; the image sequence is then input into a non-overlapping feature block partitioning layer to divide the input image sequence into an image feature block sequence, and context and semantic relationships are established; the image feature block sequence is then input into a state space model encoder based on selective scanning vision to extract a temporal feature sequence; the extracted image feature block sequence and the pixel-level feature map are input into a contrastive learning state space model to perform pixel-level and block-level contrastive learning to obtain a contrastive learning feature matrix; the spatial feature matrix, the contrastive learning feature matrix and the temporal feature sequence are then input into a feature fusion module to fuse the pixel-level and block-level spatial features and temporal features to obtain a fused feature sequence; finally, the fused feature sequence is input into a lightweight decoding head based on a convolutional neural network to predict a segmentation mask, and visual visualization is performed to generate a semantic segmentation map. In this way, the technical problem that the segmentation method in the prior art is difficult to capture rich spatiotemporal information while ensuring efficient calculation is solved.

[0072] Based on the above method, spatial features are extracted from the image at the convolutional block layer to generate pixel-level semantic features. The image sequence is partitioned at the feature block partitioning layer to extract features and establish contextual and semantic relationships. Contrastive learning is used to enhance the consistent expression of block-level and pixel-level features. A state-space model encoder based on selective scan vision is used to extract temporal feature sequences. A feature fusion module integrates multi-scale features to generate fused prediction features. Finally, a lightweight decoder head based on a convolutional neural network is used to predict the segmentation mask and visualize the segmentation results. By extracting spatial features using a deep convolutional neural network and capturing the temporal dependencies in the image sequence through a state-space model, the computational complexity of traditional methods is effectively reduced. Through block-level and pixel-level contrastive learning, this method can better learn the consistency and discriminability of image features, thereby enhancing the segmentation effect of the model. By integrating spatial and temporal information, the feature fusion module integrates multi-level features, improving the overall performance of the model, especially maintaining good segmentation accuracy in weakly supervised environments. This method does not rely on expensive pixel-level annotation data and can achieve accurate semantic segmentation using only coarse-grained annotations, making it suitable for scenarios where annotation is difficult, such as medical image analysis.

[0073] The above disclosure is only a preferred embodiment of the present invention, and certainly cannot be used to limit the scope of the rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.

Claims

1. A weakly supervised video sequence segmentation method based on a state-space model, characterized in that: The steps include: The image sequence is input into the deep convolutional neural network for spatial feature extraction to obtain the spatial feature matrix; The extracted spatial feature matrix is ​​input into the decoder for decoding to obtain a pixel-level feature map; Input the image sequence into the non-overlapping feature block partitioning layer, divide the input image sequence into image feature block sequences, and establish context and semantic relationships; The image feature block sequence is input into the state space model encoder based on selective scanning vision to extract the temporal feature sequence; The extracted image feature block sequence and pixel-level feature map are input into the contrastive learning state space model to perform pixel-level and block-level contrastive learning to obtain a contrastive learning feature matrix; The spatial feature matrix, contrastive learning feature matrix and temporal feature sequence are input into the feature fusion module, and the pixel-level and block-level spatial features and temporal features are fused to obtain a fused feature sequence; The fused feature sequence is input into a lightweight decoding head based on a convolutional neural network to predict the segmentation mask, and visual visualization is performed to generate a semantic segmentation map.

2. The weakly supervised video sequence segmentation method based on the state space model according to claim 1, characterized in that When the image sequence is input into the deep convolutional neural network for spatial feature extraction and the spatial feature matrix is ​​obtained: The image sequence is input into a deep convolutional neural network to extract the spatial feature matrix. Specifically, the first three convolution stages of ResNet-50 are used as the backbone network to capture rich pixel-level features and provide basic support for subsequent feature encoding.

3. The weakly supervised video sequence segmentation method based on the state space model according to claim 2, characterized in that The extracted spatial feature matrix is ​​input into the decoder for decoding to obtain the pixel-level feature map: A 1×1 convolution is used to compress the number of channels of the spatial feature matrix, and the original image size is restored through bilinear upsampling to generate a pixel-level feature map.

4. The weakly supervised video sequence segmentation method based on the state space model according to claim 3, characterized in that The image sequence is input into the non-overlapping feature block partitioning layer, which divides the input image sequence into a sequence of image feature blocks. When establishing context and semantic relationships: The image sequence is divided into 4×4 blocks through overlapping feature block partitioning layers to better capture local and global semantic information.

5. The weakly supervised video sequence segmentation method based on the state space model according to claim 4, characterized in that When the image feature block sequence is input into the state space model encoder based on selective scanning vision to extract the temporal feature sequence: The image feature block sequence is input into the state space model encoder based on selective scanning vision to extract temporal features; The selective scanning module adopts a multi-path cross-scanning strategy to expand the image block along four different paths and uses the state space structure to process the data of each path to construct a global receptive field; The multi-level and multi-scale temporal feature extraction process enables the model to capture complex spatiotemporal relationships.

6. The weakly supervised video sequence segmentation method based on the state space model according to claim 5, characterized in that The extracted image feature block sequence and pixel-level feature map are input into the contrastive learning state space model to perform pixel-level and block-level contrastive learning. When the contrastive learning feature matrix is ​​obtained: The input image is regarded as a bag, and each pixel in the image is regarded as a separate instance. Multi-level feature alignment is achieved through contrastive learning between pixel-level and block-level features. A mapping relationship between the two is constructed to complement each other's information. The model is optimized using the contrastive learning loss function to enhance the ability to capture feature details and global semantics, and generate a contrastive learning feature matrix F with high discrimination and good consistency. c .

7. The weakly supervised video sequence segmentation method based on the state space model according to claim 6, characterized in that When the spatial feature matrix, contrast learning feature matrix and temporal feature sequence are input into the feature fusion module, the pixel-level and block-level spatial features and temporal features are fused to obtain a fused feature sequence: The multi-layer feature input is linearly transformed into the fully connected layer, and the feature sequence after linear transformation is spliced ​​in the spatial dimension to form a comprehensive feature matrix, which is input into the multi-layer perceptron for multi-level feature fusion. The obtained feature is expressed as: F fused =(MLP(Concat(Linear(F s ),Linear(F t ),Linear(F c )))) Each feature sequence F i Contains spatial information and semantic information at different scales. Linear represents a linear transformation layer. Concat represents concatenation on the feature dimension. MLP is a multi-layer perceptron containing multiple fully connected layers and activation functions. Finally, the feature sequence fused by the multi-layer perceptron is used as the output of the decoding head.

8. The weakly supervised video sequence segmentation method based on the state space model according to claim 7, characterized in that: When the fused feature sequence is input into the lightweight decoding head based on the convolutional neural network to predict the segmentation mask and perform visual visualization to generate the semantic segmentation map: The feature sequence fused by the multi-layer perceptron is input into a lightweight decoding head based on a convolutional neural network. The convolutional neural network is used to process the fused feature sequence layer by layer to extract and learn features of different categories in the image. The output of the last convolution layer is activated by the softmax function to generate the probability distribution of each pixel belonging to each category, forming the final semantic segmentation mask. The segmentation mask is used to generate a visually interpretable semantic segmentation map based on the prediction results, the semantic category label of each area in the image is displayed, and a visual image stream is generated and output; at the same time, the last convolution layer is visually interpreted, and the forward and backward propagation gradient weights of the last convolution layer are recorded using the two hook functions register_forward_hook and register_backward_hook. The Grad-CAM is generated using the gradient weights and normalized to obtain an interpretable weight map. The threshold Seg_threshold is used to select pixels, and pixels exceeding the threshold are set to white, otherwise they are set to black to obtain a segmentation map.

Citation Information

Patent Citations

  • Weak supervision image detection method and system based on reinforcement learning of visual attention mechanism

    CN110084245A

  • Weak supervision time sequence action positioning method based on comparative learning

    CN114494941A

  • Weak supervision semantic segmentation method based on background prior

    CN118015282A

  • Intestinal ultrasound image segmentation method, system and equipment based on diffusion network and depth metric learning, and medium

    CN118762040A

  • Weak supervision semantic segmentation method based on large model auxiliary supervision

    CN119152209A