Vision Mama-based implicit contrast learning method
By introducing Vision Mamba encoder and guided gradient stop method in the contrast learning model, the shortcomings of traditional contrast learning methods in capturing long-range dependencies in image and achieving negative samples away effects are solved, and more efficient feature extraction and model performance improvement are achieved.
Patent Information
- Application Number
- CN202510000038.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-01
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-01
AI Technical Summary
The existing contrast learning methods are insufficient in capturing long-range dependencies in images, especially the lack of negative sample pairs in the traditional SimSiam model, which makes it difficult to achieve the effect of negative samples away from each other in the feature space.
Vision Mamba is used as a visual encoder, combined with the method of guiding gradient stop, the feature representation of the positive sample pair is controlled to be in one direction in the feature space, thereby achieving the effect of negative samples being distant, and implementing an implicit contrast learning method.
Capturing the context information of the image through the Vision Mamba encoder, obtaining more efficient and performing feature vectors, and guiding gradient stops in a self-supervised learning network framework significantly improves the performance of the model and enables it to show higher accuracy in image classification tasks.
Smart Images

Figure CN119992130A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image self-supervised learning, and in particular to an implicit contrastive learning method based on Vision Mamba. Background Art
[0002] Contrastive learning is a self-supervised learning method in the field of machine learning and deep learning. Its core goal is to drive the model to learn good features or representations by comparing the similarities or differences between data samples. Specifically, contrastive learning ensures that similar data samples are close to each other as positive sample pairs in the feature space, while dissimilar data samples are far away from each other as negative sample pairs, so as to better learn the characteristics of the target. Contrastive learning has many applications in fields such as computer vision and natural language processing. In the field of computer vision, it is often used for downstream tasks such as image classification, target detection, image segmentation, and prediction. Specifically, contrastive learning can be applied to scenarios such as medical image analysis and diagnosis, face recognition and expression analysis, and video action segmentation and prediction. Contrastive learning is still an active research field. Its theory and application are still developing and improving, and it has broad research prospects and research value.
[0003] In computer vision, contrastive learning methods extract feature representations of images through encoders. Convolutional neural networks (CNNs) are commonly used encoders in contrastive learning models. Due to their local receptive field and weight sharing characteristics, they can capture local features well, and have relatively few parameters and high computational efficiency. They have proven their powerful feature extraction capabilities in large-scale image recognition tasks, but their ability to capture long-range dependencies is weak. Vision Transformer is another visual encoder that performs well in processing sequence data and can capture long-range dependencies. Due to its self-attention mechanism, Vision Transformer can process all input elements in parallel, which can provide more global context information when processing images, but usually requires more computing resources and memory. Vision Mamba is the latest visual encoder that can capture and understand global and local information in images, encode positions, and thus improve the model's ability to understand visual data. It has become a new infrastructure in the field of images. The present invention attempts to apply Vision Mamba to contrastive learning models, replace the original visual encoder in the model, and enhance the model's ability to extract features to improve the performance of the model.
[0004] One of the core contents of contrastive learning is to construct positive and negative sample pairs and learn feature representations through symmetric networks. Strictly speaking, self-supervised learning methods with positive and negative sample pairs are called contrastive learning methods. However, in models such as SimSiam and BYOL, there are only positive sample pairs but no negative sample pairs, so they can only be regarded as self-supervised learning methods with symmetric structures similar to contrastive learning. Due to the existence of a unilateral gradient stop structure in the SimSiam model network, in previous studies, it is necessary to exchange the positions of the source encoder and the target encoder when calculating the loss function, and take the average of the two loss function calculations. Therefore, the previous conclusion is that the feature representations of the positive sample pairs in the SimSiam model are close to each other in the feature space. We found that when the loss function is calculated once through the unilateral gradient stop structure in the network, the feature representation of the positive sample pair in the feature space is that the features on one side are close to the features on the other side in a unidirectional direction. Therefore, based on this discovery, by guiding the gradient descent to stably control the unidirectional approach between the feature representations, the SimSiam model with only positive sample pairs can also achieve the effect of negative samples being far away, realizing an implicit contrastive learning method. Summary of the invention
[0005] The purpose of the present invention is to propose an implicit contrastive learning method based on Vision Mamba. On the basis of the traditional SimSiam model, the Vision Mamba encoder is used to extract the feature representation of the image, and the feature representation of the positive sample pair is controlled to approach in one direction in the feature space by guiding the gradient stop method, so as to achieve the effect of moving the negative samples away, thereby realizing an implicit contrastive learning method.
[0006] The technical solution for implementing the present invention is an implicit learning method based on Vision Mamba, comprising the following steps:
[0007] Step 1: Given an image x 1 、x 2 , construct the positive sample pair of the image through data enhancement, that is, view x 11 、x 12 and x 21 、x 22 ;
[0008] Step 2: Split each view into image blocks, obtain the embedding vector, add the position embedding and category label to form the initial input sequence T 0 ;
[0009] Step 3: Process the input sequence T of the Vision Mamba encoder layer l l-1 , divided into the principal line vector T a and branch vector T b , the main line vector T aAfter the one-dimensional convolution layer and the S6 layer of the Mamba model, the forward state space vector y is obtained. forward and the backward state space vector y backward ;
[0010] Step 4: Use branch vector T b The gated state space vector y forward and backward , perform the first residual connection, and then add it to the input sequence T l-1 Perform the second residual connection to obtain the output sequence T of the l-th layer encoder l , and then get the final output sequence T of the Vision Mamba encoder as a whole L ;
[0011] Step 5: Process the final output sequence T L Get the feature vector z, image x 1 、x 2 The four views correspond to the eigenvectors z 11 、z 12 、z 21 、z 22 , determine the closest feature vector z in the feature space 11 、z 21 , by guiding the gradient stop, guiding the positive sample to z 11 、z 12 With z 21 、z 22 Approach in one direction to achieve negative sample z 11 、z 21 away from;
[0012] Step 6: Calculate z 11 With z 21 The cosine similarity between them is used as the loss function, and the parameters of the self-supervised learning network framework are optimized according to the loss function to achieve implicit contrastive learning.
[0013] Preferably, in step 1, a given image x is constructed by data augmentation 1 、x 2 Positive sample pair, i.e. view x 11 、x 12 and x 21 、x 22 , and construct the SimSiam self-supervised learning network architecture for the positive sample pairs.
[0014] Preferably, in step 2, each view is divided into image blocks, an embedding vector is obtained through a linear projection layer, and then the position embedding and category label are added to form the initial input sequence T 0 , and input into the VisionMamba encoder.
[0015] Preferably, step 3 converts the input sequence T of the first layer of the Vision Mamba encoder into l-1 The main line vector T a and branch vector T b , the main line vector T a Then process to get the forward vector T af and the backward vector T ab , respectively processed by the one-dimensional convolution layer and the S6 layer of the state space model Mamba, and the hidden layer parameters in S6 are updated according to the input to obtain the state space vector y forward and backward .
[0016] Preferably, step 4 uses the branch vector T obtained by step 3 b , for the state space vector y forward and backward Gating is performed, and the two state space vectors are connected for the first time with residuals, and their sum is then added to the input sequence T l-1 Perform the second residual connection to obtain the output sequence T of the l-th layer encoder l After L layers of encoders, the final output sequence T of the VisionMamba encoder is obtained. L .
[0017] Preferably, step 5 converts the final output sequence T outputted in step 4 into L After further processing, the feature vector z is obtained through the normalization layer and the linear projection layer. 1 Two views of x 11 、x 12 The corresponding eigenvector z 11 、z 12 , and another image x 2 The two views correspond to the eigenvector z 21 、z 22 , determine the closest feature vector z in the feature space based on the Euclidean distance 11 、z 21 . By guiding the gradient to stop, control z 11 One-way approach to z 12 , z 21 One-way approach to z 22 , thus achieving z 11 With z 21 , and then realize the eigenvector z 1 With z 2 That is, by guiding the positive sample pairs to approach, the negative samples are moved away from each other.
[0018] Preferably, step 6 calculates z according to the result of guiding the gradient to stop in step 511 With z 21 The cosine similarity between them is used as the loss function of the model, and the gradient is returned according to the loss function to optimize the parameters of the self-supervised learning network framework. Since the loss function of the self-supervised learning framework only reflects the unidirectional approach of positive sample pairs, but actually achieves the effect of moving negative samples away, an implicit contrastive learning method is implemented.
[0019] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the program.
[0020] A computer-readable storage medium stores a computer program, which implements the steps of the above method when executed by a processor.
[0021] A computer program product comprises a computer program, which implements the steps of the above method when executed by a processor.
[0022] Compared with the prior art, the present invention has the following significant advantages: (1) the present invention processes images through the Vision Mamba encoder, effectively encodes spatial information through position embedding, and uses a bidirectional S6 layer to capture the contextual information of the image from both the forward and backward directions, thereby achieving a richer understanding of the image context and obtaining feature vectors with higher efficiency and performance; (2) based on the traditional self-supervised learning network framework Simsiam, the present invention implements guided gradient stopping, controls the feature vectors of the positive sample pairs to approach in a single direction in the feature space, and achieves the actual effect of moving the negative samples away, thereby realizing an implicit contrastive learning method and improving the model performance.
[0023] The present invention is described in further detail below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic diagram of the process of the present invention.
[0025] Figure 2 This is the overall network framework diagram of the present invention.
[0026] Figure 3 It is a framework diagram of the Vision Mamba module in the present invention.
[0027] Figure 4 It is a framework diagram of the Vision Mamba encoder module in the present invention.
[0028] Figure 5 It is a framework diagram of the S6 layer in the present invention.
[0029] Figure 6Schematic diagram of the guided gradient stop module in the present invention. DETAILED DESCRIPTION
[0030] like Figure 1 , Figure 2 As shown in Figure 1, an implicit contrastive learning method based on Vision Mamba and guided gradient stopping is implemented on the self-supervised learning network architecture by guiding gradient stopping. Given an image x 1 、x 2 , construct the positive sample pair of the image through data enhancement, that is, view x 11 、x 12 and x 21 、x 22 ; Split each view into image blocks, obtain the embedding vector, add the position embedding and category label to form the initial input sequence T 0 ; Process the input sequence T of the Vision Mamba encoder layer l l-1 , divided into the principal line vector T a and branch vector T b , the main line vector T a After the one-dimensional convolution layer and Mamba's S6 layer, the forward state space vector y is obtained. forward and the backward state space vector y backward ; Through the branch vector T b The gated state space vector y forward and backward , perform the first residual connection, and then add it to the input sequence T l-1 Perform the second residual connection to obtain the output sequence T of the l-th layer encoder l , and then get the final output sequence T of the Vision Mamba encoder as a whole L ; Process the final output sequence T L Get the feature vector z, image x 1 、x 2 The four views correspond to the eigenvectors z 11 、z 12 、z 21 、z 22 , determine the closest feature vector z in the feature space 11 、z 21 , by guiding the gradient stop, guiding the positive sample to z 11 、z 12 With z 21 、z 22 Approach in one direction to achieve negative sample z 11 、z 21 ; calculate z 11 With z 21The cosine similarity between them is used as the loss function, and the parameters of the self-supervised learning network framework are optimized according to the loss function to achieve implicit contrastive learning. The specific steps are as follows:
[0031] Step 1: Figure 2 As shown, for a given image x 1 、x 2 , construct positive sample pairs through data enhancement, that is, view x 11 、x 12 and x 21 、x 22 And construct the SimSiam self-supervised learning network architecture for the positive sample pairs. Data enhancement mainly uses random cropping and random horizontal flipping in geometric enhancement, color jitter and random grayscale in color enhancement, Gaussian blur in blur enhancement, and one-time standardization.
[0032] Step 2: Figure 3 As shown, Figure 3 The network architecture of the VisionMamba module is shown. Figure 2 Each view is divided into J identical image blocks of size p, where each block is Represented by x, the overall image block sequence p Represents. Input the image block into the linear projection layer to obtain the embedding vector t p , add position embedding E pos and the class label t cls , forming the initial input sequence T 0 , input to the Vision Mamba encoder. Position embedding enables the use of spatial information in sequence data, and the category label provides a vector that can aggregate the entire image information for classification or other downstream tasks in the last layer of the model.
[0033] Step 3: Figure 4 As shown, Figure 4 The network architecture of the VisionMamba encoder module is shown. Figure 3 The VisionMamba encoder module has L = 24 layers. For any layer l, the input sequence T l-1 The main line vector T a and branch vector T b , the main line vector T a After the linear projection layer, we get the forward vector T af and the backward vector T in the opposite direction ab . T af and T abAfter passing through the forward one-dimensional convolution layer and the backward one-dimensional convolution layer respectively, the local features and patterns in the sequence data are captured to obtain the forward convolution vector T af0 and the backward convolution vector T ab0 . T af0 and T ab0 After the activation function layer introduces nonlinearity, the forward activation vector x is obtained forward and the backward activation vector x backward .x forward and x backward After selective scanning by Mamba’s S6 layer, important context information is retained and irrelevant content is filtered out to obtain the state space vector y forward and backward The network architecture of the S6 layer is as follows: Figure 5 As shown, it is Figure 4 The expansion of the two modules with the same name in S6 is a state space model, where each hidden state h k All based on the current input x k and the previous hidden state h k-1 Calculated, output y k Based on the hidden state h k The state equation and output equation are calculated as follows:
[0034]
[0035] Among them C s is the parameter matrix, A s , B s , Δ s The same is the parameter matrix, Is to use A s , B s , Δ s And the discrete parameter matrix obtained according to the zero-order hold technique is expressed as follows:
[0036]
[0037] Where I is the identity matrix.
[0038] Step 4: Use the branch vector T obtained in step 3 b , after the activation function layer, the state space vector y forward and backward The gate is used to perform the first residual connection between the two state space vectors. After the sum is passed through the linear projection layer, it is combined with the input sequence T l-1 Perform the second residual connection to obtain the output sequence T of the l-th layer encoder l The initial input sequence T 0After L layers of Vision Mamba encoder modules, the final output sequence T is obtained L .
[0039] Step 5: Figure 2 and Figure 3 As shown, the final output sequence T output in step 4 L After the normalization layer and the linear projection layer, the feature vector z is obtained. Image x 1 Two views of x 11 、x 12 The corresponding eigenvector z 11 、z 12 , and another image x 2 The two views correspond to the eigenvector z 21 、z 22 The position relationship of the above eigenvectors in the feature space is as follows: Figure 6 As shown in (a) in , by calculating the Euclidean distance between feature vectors, we can find the feature vectors from different views that are closest to each other in the feature space, such as Figure 6 The calculation formula of Euclidean distance d is as follows:
[0040]
[0041] Where i, j, k, l are the numbers of the feature vectors, z ij 、z kl represents vectors from different images, ||·|| 2 is the two-norm, d ij,kl represents the distance between two vectors from different images, d m represents the minimum Euclidean distance. Assume that d m =d 11,21 , then the closest eigenvector in the feature space is z 11 、z 21 .like Figure 2 As shown, select z 11 、z 21 The Vision Mamba encoder on the network side is used as the source encoder, z 12 、z 22 The VisionMamba encoder on the network side is used as the target encoder; the feature vector enters the guided gradient stop module, where z 11 、z 21 Input the linear prediction layer to get the prediction vector p 11 、p 21 ; Calculate p 11 With z 12 、p 21 With z 22 The loss function L 11,12and L 21,22 , the calculation formula is as follows:
[0042]
[0043] Where sg(·) represents the gradient stop operator and D(·) represents the cosine similarity. The source encoder parameters are updated by gradient backpropagation, while the target encoder gradient stops and the parameters are not updated by gradient backpropagation. After unilateral gradient backpropagation and parameter update, z 11 、z 21 Will approach z in a single direction in the feature space 12 、z 22 ,like Figure 6 As shown in (c) in the figure. By calculating the Euclidean distance, the closest feature vectors are ensured to be far away from each other, the gradient on the corresponding side is updated, and the gradient on the other side is stopped, so as to achieve stable guided gradient stopping. By guiding the gradient to stop, z is stably achieved 11 With z 21 The distance in feature space, such as Figure 6 As shown in (d) in .
[0044] Step 6: Calculate z based on the result of guiding the gradient to stop in step 5 11 With z 21 The cosine similarity between them is used as the loss function of the entire self-supervised learning network framework. The network framework parameters are optimized according to the loss function. The calculation formula is as follows:
[0045]
[0046] Since the loss function of the self-supervised learning network framework only reflects the proximity of positive samples (z 11 Close to 12 , z 21 Close to 22 ), but actually achieves negative samples away from (z 11 Stay away from z 21 ), that is, negative samples are not reflected in the loss function of self-supervised learning, but the same negative sample removal effect as contrastive learning is achieved. Therefore, this method realizes implicit contrastive learning.
[0047] Table 1 Comparison of linear evaluation results of image classification by different methods
[0048]
[0049] Table 1 is a comparison of the effects of the implicit contrastive learning method in the present invention and other contrastive learning methods or self-supervised learning methods with symmetrical structures on image classification tasks, where the method of the present invention is denoted as Vim+GSG. The dataset used is ImageNet, the training round epoch=100, and the evaluation index is the linear evaluation accuracy, which is a common index in image classification tasks. The symbol "↑" indicates that the higher the value, the higher the classification accuracy, and the better the accuracy of the method. It can be found that the method of the present invention has achieved the highest ranking in this index, which fully proves that the present invention has achieved better results in image classification problems compared with the existing contrastive learning methods.
[0050] The present invention uses the Vision Mamba encoder to capture the contextual information of the image, obtains a more efficient feature vector, and guides the gradient to stop in the self-supervised learning network framework to achieve implicit contrastive learning, thereby improving the model performance.
[0051] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An implicit contrastive learning method based on Vision Mamba, characterized in that: The steps include: Step 1: Given images x1 and x2, construct a positive pair of images through data augmentation, i.e., view x 11 、x 12 and x 21 、x 22 ; Step 2: Divide each view into image blocks, obtain the embedding vector, add the position embedding and category label to form the initial input sequence T0; Step 3: Process the input sequence T of the Vision Mamba encoder layer l l-1 , divided into the principal line vector T a and branch vector T b , the main line vector T a After the one-dimensional convolution layer and the S6 layer of the Mamba model, the forward state space vector y is obtained. forward and the backward state space vector y backward ; Step 4: Use branch vector T b The gated state space vector y forward and backward , perform the first residual connection, and then add it to the input sequence T l-1 Perform the second residual connection to obtain the output sequence T of the l-th layer encoder l , and then get the final output sequence T of the Vision Mamba encoder as a whole L ; Step 5: Process the final output sequence T L Get the feature vector z, the four views of images x1 and x2 correspond to the feature vector z 11 、z 12 、z 21 、z 22 , determine the closest feature vector z in the feature space 11 、z 21 , by guiding the gradient stop, guiding the positive sample to z 11 、z 12 With z 21 、z 22 Approach in one direction to achieve negative sample z 11 、z 21 away from; Step 6: Calculate z 11 With z 21 The cosine similarity between them is used as the loss function, and the parameters of the self-supervised learning network framework are optimized according to the loss function to achieve implicit contrastive learning.
2. The implicit contrastive learning method based on Vision Mamba according to claim 1, characterized in that: Step 2 divides each view into image blocks, obtains the embedding vector, adds the position embedding and category label to form the initial input sequence T0, and inputs it into the Vision Mamba encoder.
3. The implicit contrastive learning method based on Vision Mamba according to claim 1, characterized in that: Step 3: Transform the input sequence T of the Vision Mamba encoder layer l into l-1 The main line vector T a and branch vector T b , the main line vector T a Then process to get the forward vector T af and the backward vector T ab , processed by the one-dimensional convolution layer and the Mamba S6 layer, the state space vector y is obtained forward and backward .
4. The implicit contrastive learning method based on Vision Mamba according to claim 1, characterized in that: Step 4: Use the branch vector T obtained in step 3 b , for the state space vector y forward and backward Gating is performed, and the two state space vectors are connected for the first time with residuals, and their sum is then added to the input sequence T l-1 Perform the second residual connection to obtain the output sequence T of the l-th layer encoder l , and then get the final output sequence T of the Vision Mamba encoder as a whole L .
5. The implicit contrastive learning method based on Vision Mamba according to claim 1, characterized in that: Step 5: The final output sequence T output by step 4 L Processing is performed to obtain the feature vector z and the two views x of image x1 11 、x 12 The corresponding eigenvector z 11 、z 12 , and the two views of another image x2 correspond to feature vectors z 21 、z 22 , determine the closest feature vector z in the feature space based on the Euclidean distance 11 、z 21 ; By guiding the gradient to stop, z 11 One-way approach to z 12 , z 21 One-way approach to z 22 , thus achieving z 11 With z 21 That is, the distance between negative samples is achieved by guiding the positive sample pairs to approach each other.
6. The implicit contrastive learning method based on Vision Mamba according to claim 1, characterized in that: Step 6: Calculate z based on the result of step 5 to guide the gradient to stop 11 With z 21 The cosine similarity between them is used as the loss function; the parameters of the self-supervised learning network framework are optimized according to the loss function; since the loss function of the self-supervised learning network framework only reflects the proximity of positive sample pairs, but actually achieves the effect of negative samples moving away, implicit contrastive learning is achieved.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the implicit contrastive learning method based on Vision Mamba as described in any one of claims 1 to 6 is implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the implicit contrastive learning method based on Vision Mamba as described in any one of claims 1 to 6 is implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the implicit contrastive learning method based on Vision Mamba described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Visual representation method and device based on bidirectional state space model
CN117876845A
Ankylosing spondylitis rating method based on visual state space model
CN118570202A