Self-supervised contrast learning method and system based on dynamic mask
By introducing a dynamic masking mechanism and a dual network architecture in self-supervised comparison learning, the problem of insufficient feature mining of mask processing is solved, and effective processing of multi-scale features is achieved, which significantly improves the robustness of the model and the diversity of feature representations.
Patent Information
- Application Number
- CN202510268643.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The existing mask image modeling cannot be targeted for masking, resulting in insufficient feature mining and the inability to dynamically constrain multi-scale features, affecting the model output results.
A self-supervised comparison learning method based on dynamic masks is adopted, and the mask ratio between the guide mask and the random mask is dynamically adjusted, and the dual network architecture of the online network and the target network is combined with the comparison loss calculation method, so as to achieve targeted processing of different semantic levels and multi-scale features.
It significantly improves the integrity and richness of image feature representation, improves the robustness of the model and the diversity of feature representations, reduces the dependence on complex data augmentation strategies, and improves the efficiency and accuracy of feature extraction.
Smart Images

Figure CN120198776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of self-supervised learning visual representation, and in particular to a self-supervised contrastive learning method and system based on dynamic mask. Background Art
[0002] Deep learning is a branch of machine learning, but its model training usually relies on a large amount of labeled data, which requires high labor costs. Self-supervised learning, on the other hand, does not rely on manual labeling. It makes full use of unlabeled data to automatically generate pseudo-labels to guide the model to learn. That is, it defines preset tasks through internal patterns in the data, and uses the model to solve these tasks to learn features that are helpful for downstream tasks. It usually uses instance discrimination as a preset task, and uses two enhanced views generated from the same image as positive sample pairs, and views generated from other images as negative samples. It completes the learning goal by shortening the distance between positive sample pairs and pushing the distance between negative sample pairs in the feature space.
[0003] Although contrastive learning has made significant progress in the field of visual representation, it still has some inherent limitations: traditional contrastive methods are mainly built around instance-level representations, which make it difficult to capture fine-grained semantic structures and content relationships within images; at the same time, model performance is heavily dependent on preset data enhancement strategies, and the choice of enhancement method directly determines feature quality and adaptability to downstream tasks. In addition, in order to achieve good performance, existing methods usually require the design of complex enhancement processes, which not only increases computational costs, but also introduces artificial design preferences, limiting the adaptability of the model in different scenarios.
[0004] In order to solve these problems, many researchers have tried to improve the contrastive learning framework, such as combining global and local features, designing explicit or implicit guidance models to capture finer image details; optimizing the sample selection strategy of contrastive learning through attention mechanisms, dynamic weight allocation of samples, and other methods. At the same time, mask image modeling, as an emerging self-supervised learning method, masks specific areas of the input image to guide the model to understand the internal structure and contextual information of the image, and shows unique advantages in fine-grained feature modeling; however, this type of method lacks pertinence and dynamism in capturing semantic associations and instance discrimination between images. The existing training model is relatively random in the division or segmentation of images, and is not very pertinent. It cannot ensure that the mask operation is performed on semantically important areas, which reduces the semantic association between images; and the existing mask processing method is relatively simple and static, and cannot dynamically make targeted constraints on linear or nonlinear features. It is relatively random and affects the final model output results. Summary of the invention
[0005] To solve the technical problems that existing masked image modeling cannot perform targeted masking, resulting in insufficient feature mining and inability to dynamically constrain multi-scale features, affecting the model output results, the purpose of the present invention is to provide a self-supervised contrast learning method based on dynamic masking, and the specific technical solutions are as follows:
[0006] Collect images and divide them into a training set and a test set, and perform data augmentation on each image in the training set;
[0007] Generate a target region map corresponding to each image based on the augmented image, divide the target region map into several image patches, calculate the average pixel value of each image patch, and respectively determine a first screening set and a second screening set according to the average pixel value;
[0008] Dynamically adjust the masking ratio of the guiding mask and the random mask during each round of training for masking, where the guiding mask and the random mask respectively cover the image patches in the first screening set and the second screening set;
[0009] Obtain an online network and a target network, input the masked image into the online network, and input the augmented and unmasked image into the target network to extract feature representations once in sequence; input the augmented and unmasked image into the online network, and input the masked image into the target network to extract feature representations twice in sequence;
[0010] Calculate the bidirectional contrast loss according to the first-order feature representation and the second-order feature representation, and combine them to form a total loss function;
[0011] Update the online network through backpropagation based on the total loss function, and update the target network in an exponentially weighted moving average manner through the parameters of the online network to complete the training of the online network and the target network;
[0012] Use the trained online network to perform image classification on the training set and the test set, calculate the Top-1, Top-5, and K-nearest neighbor classification accuracies respectively and evaluate them to achieve self-supervised contrast learning.
[0013] Preferably, generating a target region map corresponding to each image based on the augmented image, dividing the target region map into several image patches, calculating the average pixel value of each image patch, and respectively determining a first screening set and a second screening set according to the average pixel value includes:
[0014] Obtain a pre-trained model, input the augmented image into the pre-trained model, and perform optimization calculation on each image through an optimization objective to generate a target region map, where the optimization objective includes a sparsity constraint term and a smoothing regularization term;
[0015] Divide the target region map into several image patches, and calculate the average pixel value of each image patch;
[0016] Arrange the average pixel values in descending order, and determine the first screening set and the second screening set respectively. The first screening set is the image blocks with higher rankings, and the second screening set is the remaining image blocks.
[0017] Preferably, optimize each image through an optimization objective, and the corresponding calculation formula is:
[0018]
[0019] Among them, x represents the input enhanced image; f θ represents the pre-trained model parameters; m represents the target region map; φ represents the introduced noise; ⊙ represents the dot product symbol; SR(·) represents the smoothing regularization term; α represents the hyperparameter of the sparsity constraint term; β represents the hyperparameter of the smoothing regularization term.
[0020] Preferably, dynamically adjust the mask ratio of the guiding mask and the random mask during each round of training for masking, and the corresponding calculation formula is:
[0021]
[0022] Among them, λ t represents the guiding mask ratio corresponding to the current training cycle t; T represents the total number of training cycles; t represents the current training cycle; λ0 and λ T both represent hyperparameters.
[0023] Preferably, the online network includes an online encoder with a Vision Transformer architecture a projector p with three fully connected layers o and a predictor q with two fully connected layers; the target network includes a target encoder with a Vision Transformer architecture and a projector p with three fully connected layers t .
[0024] Preferably, input the masked image into the online network, and input the enhanced and unmasked image into the target network to extract feature representations once in sequence; input the enhanced and unmasked image into the online network, and input the masked image into the target network to extract feature representations twice in sequence, including:
[0025] Define the masked image as x m and the enhanced and unmasked image as x a respectively;
[0026] Input the masked image x m into the online encoder to obtain the encoded features Through projector p o Perform low-dimensional mapping on the encoded features To generate a feature representation Input the feature representation z m Into predictor q to generate a query feature representation q m = q(z m )); Input the enhanced and unmasked image x a Into the target encoder To obtain the encoded features Input the encoded encoded features Into projector p t To project the high-dimensional feature representation into a low-dimensional contrast space and generate a key feature representation Wherein, the query feature representation and the key feature representation are jointly defined as a primary feature representation;
[0027] Similarly, input the enhanced and unmasked image into the online network, and input the masked image into the target network to sequentially extract the secondary feature representations.
[0028] Preferably, calculate the bidirectional contrast loss according to the primary feature representation and the secondary feature representation, and combine them to form a total loss function, including:
[0029] Divide the loss calculation into a first part and a second part based on the primary feature representation and the secondary feature representation, and calculate the loss of the first part according to the primary feature representation. The corresponding calculation formula is:
[0030]
[0031] Wherein, Represents the loss of the first part; q m Represents the primary feature representation of the online network; z a Represents the primary feature representation of the target network; Represents the negative sample set; z' represents A negative sample in; τ represents the temperature coefficient;
[0032] Calculate the loss of the second part according to the secondary feature representation. The corresponding calculation formula is:
[0033]
[0034] Wherein, Represents the loss of the second part; q a Represents the secondary feature representation of the online network; z m Represents the secondary feature representation of the target network;
[0035] Combine the losses of the first part and the second part to form a total loss function. The corresponding calculation formula is:
[0036]
[0037] Among them, represents the total loss function.
[0038] Preferably, the online network is updated by backpropagation based on the total loss function, and the target network is updated in an exponentially weighted moving average manner through the parameters of the online network, completing the training of the online network and the target network, including:
[0039] Based on the total loss function, the parameters of the online network are updated by backpropagating the gradient by minimizing the total loss value, completing the training of the online network;
[0040] The target network is updated in an exponentially weighted moving average manner through the parameters of the online network, completing the training of the target network, and the corresponding calculation formula is:
[0041] θ target = mθ target +(1 - m)θ online
[0042] Among them, θ target represents the parameters of the target network; θ online represents the parameters of the online network; m represents the momentum coefficient.
[0043] To solve the above technical problems, the present invention provides another technical solution as follows: A self-supervised contrastive learning system based on dynamic masking, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of a self-supervised contrastive learning method according to any one of the foregoing are implemented.
[0044] The present invention has the following beneficial effects:
[0045] 1. By combining the dynamic masking mechanism, the model can simultaneously focus on image content at different semantic levels, including high-level semantic structures and fine-grained visual details, significantly enhancing the integrity and richness of image feature representation. It can make targeted processing of features at different scales, improving accuracy. That is, the ingenious combination of dynamic masking and random masking enhances the robustness of the model and the diversity of feature representation. Adaptive masking processing is performed according to the semantic importance of the image. Through the alternating processing strategy of masked images and only enhanced images input into the online network and the target network, the deep integration of multi-level information is achieved, solving the problem of insufficient feature mining in masking processing. The dual-network architecture of the online network and the target network, combined with the innovative contrast loss calculation method, enables the model to understand image content from different perspectives, capture high-level semantic structures and fine-grained texture information simultaneously, optimize the model, significantly enhancing the model's ability to understand multi-scale features and the perception ability of internal structures, reducing the dependence on complex data augmentation strategies, and improving the efficiency and robustness of feature extraction. Through the obtained total loss function, the parameters of the online network and the target network are effectively adjusted to complete the training of the network, providing a higher-quality feature supervision signal for the self-supervised learning process, greatly improving the model training efficiency and feature extraction performance, and thus achieving better performance in various downstream visual tasks.
[0046] 2. The present invention also provides a self-supervised contrast learning system based on dynamic masking for implementing the self-supervised contrast learning method based on dynamic masking provided above. This system has the same beneficial effects as the above-mentioned self-supervised contrast learning method based on dynamic masking and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 It is a flowchart of the implementation of a self-supervised contrast learning method based on dynamic masking provided by an embodiment of the present invention;
[0049] Figure 2 It is a schematic diagram of the model framework of a self-supervised contrast learning method based on dynamic masking provided by an embodiment of the present invention;
[0050] Figure 3A distribution map of t-SNE visualization performed by a self-supervised contrastive learning method based on dynamic masking and MoCo v3 on the CIFAR10 dataset provided by an embodiment of the present invention. Detailed implementation manners
[0051] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following describes in detail the specific implementation manners, structures, features and effects of a self-supervised contrastive learning method and system based on dynamic masking proposed according to the present invention with reference to the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0053] The following specifically describes the specific solutions of a self-supervised contrastive learning method and system based on dynamic masking provided by the present invention with reference to the accompanying drawings.
[0054] Existing methods partition images rather randomly, lacking pertinence and dynamics in capturing semantic associations and instance discrimination between images, unable to ensure that the masking operation is performed on semantically important regions, reducing the semantic relevance between images, unable to dynamically make targeted constraints on linear or non-linear features, and affecting the output results of the model; the first embodiment of the present invention provides a self-supervised contrastive learning method based on dynamic masking. This method combines dynamic masking and random masking skillfully through a dual-network architecture of an online network and a target network, innovates the contrast loss calculation method, effectively updates the parameters of the online network and the target network, enables the model to simultaneously focus on image content at different semantic levels, improves the integrity and richness of image feature representation, can perform targeted processing on features of different scales, and improves accuracy; to implement a self-supervised contrastive learning method based on dynamic masking, a self-supervised contrastive learning system based on dynamic masking is provided. This system is essentially a software system composed of units that implement corresponding functions. The specific steps in this method are introduced in detail below.
[0055] Please refer to Figure 1 and Figure 2 , which respectively show the implementation flowchart and the schematic diagram of the model framework of a self-supervised contrastive learning method based on dynamic masking provided by an embodiment of the present invention, and illustrate that after the input original image is specifically processed by a self-supervised contrastive learning method based on dynamic masking provided by the present application, denoted as RFMaCo, the self-supervised contrastive learning of the original image is realized; the method includes:
[0056] Step S1: collect images and divide them into training set and test set, and perform data enhancement on each image in the training set;
[0057] Step S2: generating a target region map corresponding to each image based on the enhanced image, dividing the target region map into equal parts to obtain a number of image blocks, calculating the average pixel value of each image block, and sorting them according to the average pixel values to determine the first screening set and the second screening set respectively;
[0058] Step S3: dynamically adjusting the mask ratio of the guide mask to the random mask during each round of training to perform masking, wherein the guide mask and the random mask respectively mask the image blocks in the first screening set and the second screening set;
[0059] Step S4: obtaining an online network and a target network, inputting the masked image into the online network and the enhanced and unmasked image into the target network to extract a primary feature representation in sequence; inputting the enhanced and unmasked image into the online network and the masked image into the target network to extract a secondary feature representation in sequence;
[0060] Step S5: Calculate the bidirectional contrast loss based on the primary feature representation and the secondary feature representation, and combine them to form a total loss function;
[0061] Step S6: Based on the total loss function, the online network is updated by back propagation, and the target network is updated by exponential average movement of the parameters of the online network to complete the training of the online network and the target network;
[0062] Step S7: Use the trained online network to classify the images of the training set and the test set, calculate and evaluate the Top-1, Top-5 and K nearest neighbor classification accuracies, and realize self-supervised comparative learning.
[0063] To better explain, self-supervised learning refers to defining preset tasks through internal patterns of data, and using models to solve these tasks to learn features that are helpful for downstream tasks; for example, contrastive learning usually uses instance discrimination as a preset task, and uses two enhanced views generated from the same image as positive sample pairs, while views generated from other images are used as negative samples. The learning goal is achieved by shortening the distance between positive sample pairs and pushing the distance between negative sample pairs in the feature space. In this process, the model can effectively learn discriminative feature representations; masking refers to the use of binary codes to perform bitwise AND operations on the target field to mask the current input bit.
[0064] Understandably, there are already methods in the prior art that combine contrastive learning with masked image modeling. For example, the ACoMIM (Attention Guided Contrastive Masked Image Modeling for Transformer-Based Self-Supervised Learning) method guides the model to further optimize the learning of global and local features by applying masks to some semantically key regions in the framework of contrastive learning; the CMAE (Contrastive Masked Autoencoder) model solves the problem of insufficient global feature mining in masked image modeling by adding a contrastive learning objective in the masked autoencoder framework. However, although existing research has made certain progress in this regard, there are still limitations. In the ACoMIM method, the quality of the attention map may vary due to different image contents, training stages, or model structures, which may lead to unstable performance; although the pixel offset method proposed by the CMAE model is better than random cropping, it may still be insufficient to capture more complex image transformations, limiting the model's adaptability to various perspectives and scene changes. Therefore, a self-supervised contrastive learning method based on dynamic masks is proposed in this application.
[0065] Specifically, in step S1, the images are divided into a training set and a test set to complete the training of the model through the image data in the training set; and the same series of data augmentations are performed on each image in the training set, that is, any series of the same data augmentation operations including but not limited to rotation, scaling, random cropping, color adjustment, etc. are performed on each image to ensure the diversity of the data set and the generalization ability of the model, and improve the adaptability and accuracy of the model when facing new data.
[0066] Optionally, tests are performed based on the collected images, which include the name of the image data set, the number of categories, the image size, and the number of samples in the training set and the test set; that is, the situation of the image data set involved in this application is described in combination with Table 1.
[0067] Description of the image data set involved in Table 1
[0068] Dataset Number of classes Image size Number of training sets Number of test sets CIFAR10 10 32x32 50000 10000 CIFAR100 100 32x32 50000 10000 TinyImageNet 200 64x64 100000 10000 ImageNet100 100 Not unique 70000 40000
[0069] Furthermore, in step S2, it includes:
[0070] Step S21: Obtain a pre-trained model, input the augmented images into the pre-trained model, and perform optimization calculations on each image through the optimization objective to generate a target region map. The optimization objective includes a sparsity constraint term and a smoothing regularization term.
[0071] As an alternative implementation, in this embodiment, the pre-trained model is ResNet-18; that is, a deep residual network containing 18 weighted layers, consisting of an input layer, a 7×7 initial convolutional layer, a max pooling layer, 8 residual units (each containing 2 3×3 convolutional layers and skip connections), an average pooling layer, and a final fully connected output layer; it uses ReLU as the activation function, and the residual connection structure effectively solves the problems of gradient disappearance and gradient explosion in deep networks.
[0072] It should be noted that the target region map refers to the saliency region map, that is, the prominent part in the image, which is used to simulate the human visual system's perception ability of significant or prominent regions in the image; it is generated by optimizing the calculation of a specific region through the pre-trained model and is used to guide the division of regions in the dynamic mask.
[0073] It can be explained that adding a sparsity constraint term to the optimization objective aims to make the target region map more concise, highlight the key regions in the image, reduce redundant information, and enable the pre-trained model to focus more on the salient regions in the image; the smoothing regularization term is used to ensure the spatial continuity and smoothness of the target region map, avoid over-segmentation of the salient regions in the image or the generation of abrupt edges, and improve the accuracy and robustness of the target region map; that is, by combining the sparsity constraint term and the smoothing regularization term, a more reliable basis is provided for subsequent data analysis.
[0074] Furthermore, in step S21, each image is optimized and calculated through the optimization objective, and the corresponding calculation formula is:
[0075]
[0076] Among them, x represents the enhanced input image; f θ represents the pre-trained model parameters; m represents the target region map; φ represents the introduced noise; ⊙ represents the dot product symbol; SR(·) represents the smoothing regularization term; α represents the hyperparameter of the sparsity constraint term; β represents the hyperparameter of the smoothing regularization term.
[0077] It should be noted that x represents the enhanced input image, which is usually a three-dimensional array with a dimension of w×h×c. Among them, w represents the width, h represents the height; c represents the number of channels. For example, when processing a color image, each pixel has three color channels: red, green, and blue, so c = 3. At this time, each element in the three-dimensional array represents the intensity value of a certain position in the image on a specific color channel.
[0078] Step S22: Divide the target region map into several image patches equally, and calculate the average pixel value of each image patch.
[0079] Specifically, the target region map is divided into several non-overlapping image patches of equal size according to the equal division principle, that is, for the target region map According to the fixed image patch size p, for example, p = 16, it is divided into L = hw / p 2 non-overlapping image patches; then calculate the average pixel value of each image patch, denoted as {μ1, μ2, …, μ L}.
[0080] Step S23: Sort the average pixel values in descending order, and determine the first screening set and the second screening set respectively. The first screening set is the image patches with higher rankings, and the second screening set is the remaining image patches.
[0081] Specifically, based on the average pixel values obtained above, sort them from high to low to obtain a sorting index sequence, denoted as I sorted = argsort({μ1, μ2, …, μ L}, descending = True), and divide the image patches into high-significance regions and low-significance regions according to the sorting index sequence; it can be explained that the total number of image patches is L, the total masking rate is γ, and the guiding masking rate is λ t , then the size of the first screening set is λ t ·γ·L image patches, that is, select the first λ t ·γ·L image patches from the image patches sorted in descending order of average pixel value as the first screening set for guiding masking; and the remaining (1 - λ t )·γ·L image patches to be masked are used as the second screening set for random masking; that is, the first screening set includes the image patches with higher rankings, referring to the high-significance regions; the second screening set is the remaining image patches, representing the low-significance regions; this provides a basis for the subsequent selection of guiding masking, enabling the algorithm to focus on the image regions that the model considers the most important first, avoiding random partitioning or segmentation of the image, and reducing the model's ability to understand image data.
[0082] It can be understood that by combining guiding masking and random masking, different masked enhanced views are dynamically generated; that is, for the target region map, a "guiding masking region" that needs to be masked in the target region map is dynamically divided, which refers to the most concerned region, representing the image patches in the first screening set; then randomly select some non-significant regions as the "random masking region", representing the image patches in the second screening set, ensuring that the coverage range of the image mask is diverse and more targeted, and ensuring that the subsequent masking operation is for semantically important regions; and combining these two modes of guiding masking and random masking can provide complementary learning constraints for semantic key regions and non-key regions, enhancing the multi-dimensional understanding ability of feature representation; that is, providing diverse training samples and enhancing the robustness of the model.
[0083] Further, in step S3, the masking is performed by dynamically adjusting the masking ratio of the guiding mask to the random mask in each round of training process, and the corresponding calculation formula is:
[0084]
[0085] where λ t represents the guiding mask ratio corresponding to the current training cycle t; T represents the total number of training cycles; t represents the current training cycle; λ0 and λ T both represent hyperparameters.
[0086] It should be noted that λ0 and λ T represent hyperparameters, and their value ranges are [0, 1]; that is, an adaptive masking strategy is adopted, and the masking ratio is automatically adjusted according to a linear function as the training progresses, enabling the model to gradually learn more complex feature representations; in practical applications, the guiding mask based on the salient region can guide the model to focus on the semantically important regions in the image, while the random mask helps the model explore diverse feature distributions, avoiding the singularity in the feature learning process, significantly improving the adaptability and generalization ability of the model, and achieving more efficient self-supervised feature learning.
[0087] Further, the online network includes an online encoder with a Vision Transformer architecture a projector p with three fully connected layers o and a predictor q with two fully connected layers; the target network includes a target encoder with a Vision Transformer architecture and a projector p with three fully connected layers t .
[0088] It can be explained that the online network is different from the target network, and a dual-network architecture is used to achieve multi-scale feature extraction of different input images, that is, through cross-view contrast learning to cooperate with each other to achieve feature transfer; among them, the encoder with a Vision Transformer architecture refers to the part of the network responsible for processing and encoding image data to capture the features and relationships in the image to complete the tasks arranged in the network; the projector refers to a network model containing three fully connected layers, which is used to project high-dimensional data into a low-dimensional space to achieve the dimensionality reduction task; the predictor with two fully connected layers refers to a small network used to process features in the self-supervised learning framework, usually composed of two linear layers, with a non-linear activation function (such as ReLU or GELU) in the middle, predicting the feature representation of another view from the feature representation of one view, helping the model learn view-invariant features and better adapt to downstream tasks.
[0089] Further, in step S4, it includes:
[0090] Define the masked image as x respectivelym The enhanced and unmasked image is x a ;
[0091] The masked image x m is input into the online encoder to obtain the encoded features Through the projector p o for the encoded features perform low-dimensional mapping to generate the feature representation The feature representation z m is input into the predictor q to generate the query feature representation q m = q(z m ); The enhanced and unmasked image x a is input into the target encoder to obtain the encoded features The encoded encoded features are input into the projector p t to project the high-dimensional feature representation into the low-dimensional contrast space to generate the key representation wherein, the query feature representation and the key feature representation are jointly defined as a primary feature representation;
[0092] Similarly, the enhanced and unmasked image is input into the online network, and the masked image is input into the target network to sequentially extract the secondary feature representation.
[0093] It should be noted that in this embodiment, the primary feature representation refers to the set of feature representations extracted when the masked image is input into the online network and the enhanced and unmasked image is input into the target network; the secondary feature representation is the set of feature representations extracted after the input is exchanged. Through this cross-input method, features can be extracted from different perspectives, promoting the diversity of feature learning and improving the comprehensiveness of feature representation; and when the input changes, it can improve the adaptability of the network to input changes, enhance the robustness of the model, avoid the domain gap between upstream and downstream tasks, and better generalize to other downstream tasks.
[0094] It can be understood that for the masked image and the enhanced and unmasked image, the input to the online network and the target network is swapped to ensure that the two different networks can extract feature representations from different perspective views respectively, so as to make full use of the advantages of masked contrastive learning.
[0095] Preferably, in this embodiment, the definition of the contrast loss function for the online network and the target network is:
[0096]
[0097] Among them, q represents the query vector, that is, the first - order feature representation output by the online network; k represents the positive sample vector, that is, positive sample pairs are formed through feature representations for calculating the contrastive loss; κ represents the negative sample set, that is, the set of data points that do not meet the conditions or do not have feature representations; τ represents the temperature coefficient, which is used to adjust the sharpness of the contrastive loss.
[0098] It is explained that a two - way contrastive loss is made based on the online network and the target network. That is, through the way of contrastive learning, the network can better capture the internal structure and features of the input image, improving the generalization ability of the model; and through the comparison between the online network and the target network, the over - fitting phenomenon in the training process can be reduced, accelerating the convergence speed of model training. By comparing the outputs of the two networks, the model can more quickly identify important feature differences, improving the learning efficiency, and thus reducing the consumption of computing resources.
[0099] Further, in step S5, it includes:
[0100] Based on the first - order feature representation and the second - order feature representation, the loss calculation is divided into a first part and a second part. The loss of the first part is calculated according to the first - order feature representation, and the corresponding calculation formula is:
[0101]
[0102] Among them, represents the loss of the first part; q m represents the first - order feature representation of the online network; z a represents the first - order feature representation of the target network; represents the negative sample set; z' represents a negative sample in κ; τ represents the temperature coefficient;
[0103] The loss of the second part is calculated according to the second - order feature representation, and the corresponding calculation formula is:
[0104]
[0105] Among them, represents the loss of the second part; q a represents the second - order feature representation of the online network; z m represents the second - order feature representation of the target network;
[0106] The losses of the first part and the second part are combined to form the total loss function, and the corresponding calculation formula is:
[0107]
[0108] Among them, represents the total loss function.
[0109] Further, in step S6, it includes:
[0110] Step S61: Based on the total loss function, update the parameters of the online network by backpropagating the gradient to minimize the total loss value, and complete the training of the online network.
[0111] Explanation is made that backpropagating the gradient of the parameters of the online network, that is, performing backpropagation to realize the update of the online network. Based on minimizing the total loss value, call the corresponding function to clear the previous gradient information to avoid the cumulative effect; then calculate the gradient of the loss with respect to the parameters of the online network to realize the update of the parameters of the online network. Optionally, an optimizer is included in the network, and the optimizer can store the gradient values of the parameters of the online network to enable the network to make adaptive adjustments.
[0112] Step S62: Update the target network in an exponentially weighted moving average manner with the parameters of the online network, and complete the training of the target network. The corresponding calculation formula is:
[0113] θ target = mθ target +(1 - m)θ online
[0114] where θ target represents the parameters of the target network; λ online represents the parameters of the online network; m represents the momentum coefficient.
[0115] Explanation is made that during the training process, the smoothing coefficient is combined with the parameters of the online network to update the target network, making the parameters of the target network smoother and capable of tracking the long-term average behavior of the online network, which helps to improve the stability of the target network and enhance the reliability.
[0116] It can be explained that in step S7, it includes:
[0117] Use the trained online network to perform image classification on the training set and the test set, calculate the Top-1, Top-5 and K-nearest neighbor classification accuracies respectively and evaluate them to realize self-supervised contrastive learning.
[0118] Specifically, a self-supervised contrastive learning method based on dynamic masking proposed in the first embodiment of the present application is compared with other classical self-supervised contrastive learning methods in terms of the Top-1 and Top-5 classification accuracies of the four original image datasets of CIFAR10, CIFAR100, TinyImageNet, and ImageNet100 proposed in this embodiment. That is, an explanation is made in combination with Table 2. Among them, Top-1 means taking the largest one in the last probability vector as the prediction result. If the predicted classification result of the largest one is correct, it means that the prediction result value is correct; Top-5 means taking the first five largest ones in the last probability vector as the prediction result. Among these five, if one prediction is correct, the predicted classification result is correct.
[0119] Table 2 Top-1 and Top-5 accuracies of different self-supervised contrastive learning methods on four original image datasets
[0120]
[0121] Preferably, in this embodiment, the Batchsize is set to 256, the temperature coefficient is set to 0.2, and the values of λ0 and λ T are 0 and 0.5 respectively, and the total masking rate is 25%.
[0122] Furthermore, for the four original image datasets of CIFAR10, CIFAR100, Tiny ImageNet, and ImageNet100 proposed in this embodiment, the k-nearest neighbor classification accuracy is calculated. Optionally, a comparative analysis is made with K being 20 or 100 respectively. That is, an explanation is made in combination with Table 3. Among them, the k-nearest neighbor (k-Nearest Neighbors, abbreviated as KNN) algorithm is used for classification and regression problems. That is, given a training dataset, for a new input instance, the K instances closest to this instance are found in the training dataset.
[0123] Table 3 k-nearest neighbor classification accuracies of different self-supervised contrastive learning methods on four original image datasets
[0124]
[0125] It is noted that in this embodiment, a self-supervised contrastive learning method based on dynamic masking is denoted as RFMaCo. In terms of data comparison, RFMaCo performs excellently in the original image datasets, and the effectiveness of the technical solution of the present application is further verified by combining the k-nearest neighbor classification accuracy. It can be seen that RFMaCo still remains leading under different settings of the K value.
[0126] Understandably, by combining with the dynamic masking mechanism, the model can simultaneously focus on image content at different semantic levels, including high-level semantic structures and fine-grained visual details, significantly enhancing the integrity and richness of image feature representation, being able to make targeted processing of features at different scales and improve accuracy. That is, the ingenious combination of dynamic masking and random masking enhances the robustness of the model and the diversity of feature representation. Through adaptive masking processing according to the semantic importance of the image, and by means of the alternating processing strategy of the masked image and the image input only with enhancement processing in the online network and the target network, the deep integration of multi-level information is achieved, and the problem of insufficient feature mining in masking processing is solved. The dual-network architecture of the online network and the target network, combined with the innovative contrast loss calculation method, enables the model to understand image content from different perspectives, capture high-level semantic structures and fine-grained texture information at the same time, optimize the model, significantly enhancing the model's understanding ability of multi-scale features and the perception ability of internal structures, reducing the dependence on complex data augmentation strategies, and improving the efficiency and robustness of feature extraction. Through the obtained total loss function, the parameters of the online network and the target network are effectively adjusted to complete the training of the network, providing a higher-quality feature supervision signal for the self-supervised learning process, greatly improving the model training efficiency and feature extraction performance, and thus obtaining better performance in various downstream visual tasks.
[0127] Please refer to Figure 3 , which shows the distribution diagram of t-SNE (t-distributed Stochastic Neighbor Embedding) visualization performed on the CIFAR10 dataset by a self-supervised contrast learning method based on dynamic masking provided by an embodiment of the present invention and MoCo v3 (Momentum Contrast for Unsupervised Visual Representation Learning); preferably, to better reflect the technical advantages brought by RFMaCo, a comparison is made with the prior art solution, that is, the feature distribution comparison after t-SNE dimensionality reduction visualization based on RFMaCo and MoCo v3; among them, Figure (a) is the feature distribution diagram obtained by the MoCo v3 method, and Figure (b) is the feature distribution diagram obtained by the RFMaCo method.
[0128] It can be illustrated that there are obvious defects in the feature distribution of MoCo v3. From Figure 3It can be seen that there is a large range of overlap among image data points of multiple categories, especially in the middle region of the image. The category boundaries are blurred, resulting in a decrease in the distinguishability between categories. Moreover, the data points of some categories are scattered, lacking good internal aggregation. In contrast, the RFMaCo proposed in this embodiment shows significant advantages. The data points of each category form a more compact clustering structure, and the distance between the data points within the category is closer, indicating a higher consistency in feature representation. At the same time, the boundaries between different categories are clearer, and the category intervals are more distinct, reducing classification ambiguity. Especially in the complex category region, RFMaCo can still maintain good distinguishability. That is, by comparing and analyzing different self-supervised contrast learning methods, it can be known that RFMaCo is effective in feature representation learning. Through the dynamic mask guidance strategy, it combines the guidance mask and the random mask, and conducts two-way contrast learning based on the online network and the target network, successfully constructing a feature space with richer semantics and more reasonable structure, providing a better feature representation basis for subsequent classification tasks, and thus enabling more accurate category recognition and more robust model performance in practical applications.
[0129] An embodiment of the present invention also proposes a self-supervised contrast learning system based on dynamic masks, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a self-supervised contrast learning method based on dynamic masks as described in any of the foregoing embodiments.
[0130] It should be noted that a self-supervised contrast learning system based on dynamic masks has the same beneficial effects as the self-supervised contrast learning method based on dynamic masks provided above, and will not be elaborated here.
[0131] It should be noted that the above order of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0132] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.
Claims
1. A self-supervised contrastive learning method based on dynamic mask, characterized in that: The method comprises: Collect images and divide them into training sets and test sets, and perform data augmentation on each image in the training set; Generate a target area map corresponding to each image based on the enhanced image, divide the target area map equally into a number of image blocks, calculate the average pixel value of each image block, and sort them according to the average pixel value to determine a first screening set and a second screening set respectively; Dynamically adjust the mask ratio of the guide mask to the random mask during each round of training to perform masking, wherein the guide mask and the random mask respectively mask the image blocks in the first screening set and the second screening set; An online network and a target network are obtained, and the masked image is input into the online network and the enhanced and unmasked image is input into the target network to extract a primary feature representation in sequence; the enhanced and unmasked image is input into the online network and the masked image is input into the target network to extract a secondary feature representation in sequence; Calculate the bidirectional contrast loss based on the primary feature representation and the secondary feature representation, and combine them to form the total loss function; Based on the total loss function, the online network is updated through back propagation, and the target network is updated through the parameters of the online network in an exponential average moving manner to complete the training of the online network and the target network; The trained online network is used to classify images in the training set and test set, and the Top-1, Top-5 and K nearest neighbor classification accuracies are calculated and evaluated respectively to achieve self-supervised comparative learning.
2. A self-supervised contrastive learning method based on dynamic mask according to claim 1, characterized in that: Generate a target area map corresponding to each image based on the enhanced image, divide the target area map equally to obtain a number of image blocks, calculate the average pixel value of each image block, and sort according to the average pixel value to determine the first filter set and the second filter set, including: Obtain a pre-trained model, input the enhanced image into the pre-trained model, perform optimization calculation on each image by optimizing the target, and generate a target region map, wherein the optimization target includes a sparsity constraint term and a smooth regularization term; The target area map is equally divided into several image blocks, and the average pixel value of each image block is calculated; The average pixel values are arranged in descending order to determine a first filter set and a second filter set, respectively. The first filter set is the image blocks with the highest order, and the second filter set is the remaining image blocks.
3. A self-supervised contrastive learning method based on dynamic mask according to claim 2, characterized in that: Each image is optimized by optimizing the target, and the corresponding calculation formula is: Where x represents the enhanced image of the input; f θ represents the pre-training model parameters; m represents the target area map; φ represents the introduced noise; ⊙ represents the dot product symbol; SR(·) represents the smoothing regularization term; α represents the hyperparameter of the sparsity constraint term; β represents the hyperparameter of the smoothing regularization term.
4. The self-supervised contrastive learning method based on dynamic mask according to claim 1, characterized in that: The mask ratio of the guided mask to the random mask is dynamically adjusted during each round of training. The corresponding calculation formula is: Among them, λ t represents the guide mask ratio corresponding to the current training cycle τ; T represents the total number of training cycles; t represents the current training cycle; λ0 and λ T All represent hyperparameters.
5. The self-supervised contrastive learning method based on dynamic mask according to claim 1, characterized in that: The online network includes an online encoder of the Vision Transformer architecture The projector p of three fully connected layers o and two layers of fully connected layers of predictors q; the target network includes a target encoder of the Vision Transformer architecture and three fully connected layers of projectors p t .
6. A self-supervised contrastive learning method based on dynamic mask according to claim 5, characterized in that: The masked image is input into the online network, and the enhanced and unmasked image is input into the target network to extract feature representations one by one; The enhanced and unmasked image is input into the online network, and the masked image is input into the target network to extract secondary feature representations in sequence, including: Define the masked image as x m , the enhanced and unmasked image is x a ; The masked image x m Input online encoder Get the encoding features Through the projector o Encoding features Perform low-dimensional mapping to generate feature representation Denote the feature z m Input predictor q to generate query feature representation q m =q(z m ); The enhanced and unmasked image x a Input target encoder The encoded features are obtained from The encoded features Input projector p t In the process, the high-dimensional feature representation is projected into the low-dimensional contrast space to generate the key feature representation Among them, the query feature representation and the key feature representation are defined together as a primary feature representation; Similarly, the enhanced and unmasked image is input into the online network, and the masked image is input into the target network to extract secondary feature representations in turn.
7. A self-supervised contrastive learning method based on dynamic mask according to claim 6, characterized in that: The bidirectional contrast loss is calculated based on the primary feature representation and the secondary feature representation, and combined to form the total loss function, including: Based on the primary feature representation and the secondary feature representation, the loss calculation is divided into the first part and the second part. The loss of the first part is calculated according to the primary feature representation. The corresponding calculation formula is: in, represents the loss of the first part; q m represents a feature representation of the online network; z a Represents a feature representation of the target network; represents the negative sample set; z' represents A negative sample in the equation; τ represents the temperature coefficient; The loss of the second part is calculated based on the secondary feature representation. The corresponding calculation formula is: in, represents the loss of the second part; q a represents the secondary feature representation of the online network; z m Represents the secondary feature representation of the target network; The losses of the first and second parts are combined to form the total loss function, and the corresponding calculation formula is: in, Represents the total loss function.
8. The self-supervised contrastive learning method based on dynamic mask according to claim 1, characterized in that: Based on the total loss function, the online network is updated through back propagation, and the target network is updated through the parameters of the online network in an exponential average moving manner to complete the training of the online network and the target network, including: Based on the total loss function, the parameters of the online network are updated by gradient backpropagation by minimizing the total loss value to complete the training of the online network. The target network is updated by exponentially moving the parameters of the online network to complete the training of the target network. The corresponding calculation formula is: i target =mθ target +(1-m)θ online Among them, θ target represents the parameters of the target network; θ online represents the parameters of the online network; m represents the momentum coefficient.
9. A self-supervised contrastive learning system based on dynamic masking, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of a self-supervised contrastive learning method based on dynamic masking as described in any one of claims 1 to 8 are implemented.