Model robustness improvement method based on comparative learning
Through two-stage training and explanation of heat map similarity constraints, the problem of model concerning regional dispersion and semantic inconsistency in comparison learning is solved, which improves the feature extraction ability and robustness of the model, and improves the performance of downstream tasks.
Patent Information
- Application Number
- CN202510337338.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
The existing comparison learning technology has models that focus on regional dispersion, contrast is not based on semantic consistency, and lack of stability constraints, resulting in feature learning instability and discriminant ability.
Using a two-stage training method, the neural network model is initially trained by data-enhanced image dataset, and then the secondary enhancement is performed by calculating image similarity and interpreting heat map similarity loss. The total loss of the model is calculated by combining comparative learning loss and interpreting heat map similarity loss, and the model parameters are adjusted to improve feature extraction capabilities.
The feature extraction capability and robustness of the model in image classification and object detection tasks are improved, and the processing performance of downstream tasks is enhanced.
Smart Images

Figure CN120259697A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence explainable technology, and specifically relates to a method for improving model robustness based on contrastive learning. Background Art
[0002] In recent years, contrastive learning, as a self-supervised learning method, has made significant progress in the field of computer vision. Contrastive learning constructs positive sample pairs (semantically similar image pairs) and negative sample pairs (semantically dissimilar image pairs) to enable the model to shorten the distance between positive sample pairs in the feature space and push the distance between negative sample pairs away, thereby learning more robust feature representations. Typical contrastive learning methods include MoCo and SimCLR, which optimize the model through contrastive loss functions, so that it can extract differentiated features based on different instances in an unsupervised environment, and show excellent performance in downstream tasks (such as image classification, object detection, etc.).
[0003] However, existing contrastive learning techniques still have the following problems:
[0004] 1. The model's focus area is scattered and irrelevant features may be learned: During the training process, the contrastive learning model may focus on non-discriminative features in the image. For example, in some data sets, the model will compare the background area of two images too much and reduce the contrast of the foreground area;
[0005] 2. The model's comparison of positive or negative pairs is not necessarily based on semantic consistency: Existing contrastive learning methods mainly rely on data augmentation or instance discrimination, but in the comparison process, the model may use low-level features (such as color, texture) for distinction rather than high-level semantic features. For example, for a picture of a white bird and a picture of any category with a white background, the network considers the two pictures to be similar only because of their similar colors;
[0006] 3. Lack of clear contrastive learning constraint mechanism makes the features learned by the model unstable: Due to the lack of control over the feature contrast mechanism, the model may learn inconsistent features for the same category of images under different data augmentation methods, resulting in unstable training. For example, although some contrastive learning methods can optimize the distribution of feature space, they fail to ensure the consistency of contrast features between samples of the same category, which affects the model's discrimination ability. Summary of the invention
[0007] In view of the shortcomings of the prior art, the present invention proposes a method for improving model robustness based on contrastive learning, which includes:
[0008] S1: Obtain an image dataset and perform data augmentation. Use the augmented image dataset to conduct the first-stage contrastive learning training on the neural network model to obtain a preliminarily trained neural network model, enabling the model to have a certain feature extraction ability.
[0009] S2: Perform secondary augmentation on the image dataset to obtain a first-augmented image and a second-augmented image;
[0010] S3: Input the first-augmented image and the second-augmented image of each image into the preliminarily trained neural network model for processing, and calculate the similarity and contrastive learning loss between the two augmented images of the same image according to the processing results;
[0011] S4: Calculate the similarity interpretation heatmap of the positive pair images according to the similarity between the two images;
[0012] S5: Conduct the second-stage contrastive learning training on the preliminarily trained neural network model according to the similarity interpretation heatmap of the positive pair images and calculate the interpretation heatmap similarity loss;
[0013] S6: Calculate the total model loss according to the contrastive learning loss and the interpretation heatmap similarity loss; Adjust the model parameters according to the total model loss to obtain a trained neural network model.
[0014] Preferably, the process of performing secondary augmentation on the augmented image dataset includes:
[0015] Perform non-spatial transformation data augmentation on the image to obtain a first-augmented image;
[0016] Perform spatial transformation data augmentation on the image to obtain an intermediate image; Perform non-spatial transformation data augmentation on the intermediate image to obtain a second-augmented image.
[0017] Preferably, the process of calculating the similarity between two images includes:
[0018] The image is processed by the neural network model, and the cosine similarity between the feature encodings extracted by the neural network model for the two images is calculated.
[0019] Preferably, the formula for calculating the similarity interpretation heatmap of the positive pair images is:
[0020]
[0021] where, e i represents the similarity interpretation heatmap of the first-augmented image x i relative to the second-augmented image x j and e j represents the similarity interpretation heatmap of the second-augmented image x j relative to the first-augmented image x iThe similarity interpretation heat map, where Z represents the total number of elements in the feature map, f(·) represents the neural network model, and sim(·) represents the similarity calculation. represents the first enhanced image x i the element at the m-th row and n-th column in the k-th feature map of the last layer of the network during the forward propagation process represents the second enhanced image x j the element at the m-th row and n-th column in the k-th feature map of the last layer of the network during the forward propagation process.
[0022] Preferably, the process of contrastive learning training of the neural network model based on the similarity interpretation heat map of the positive pair images includes:
[0023] Randomly crop, scale, and flip the similarity interpretation heat map of one image in the positive pair images to obtain an enhanced interpretation heat map;
[0024] Use the enhanced interpretation heat map and the similarity interpretation heat map of the other image in the positive pair images as positive samples; use the enhanced interpretation heat map and the similarity interpretation heat maps generated from other images in the same batch as negative samples to perform contrastive learning training on the neural network model.
[0025] Preferably, the formula for calculating the similarity loss of the interpretation heat map is:
[0026]
[0027] where represents the similarity loss of the interpretation heat maps between image i and image j, u(·) represents the spatial transformation data augmentation operation, N represents the batch size, e i represents the first enhanced image x i relative to the second enhanced image x j of the similarity interpretation heat map, e j represents the second enhanced image x j relative to the first enhanced image x i of the similarity interpretation heat map, e k represents the k-th similarity interpretation heat map in the same batch, and τ represents the temperature hyperparameter.
[0028] Preferably, the formula for calculating the total loss of the model is:
[0029]
[0030] where represents the total loss of the model, represents the contrastive learning loss between image i and image j, represents the similarity loss of the interpretation heat maps between image i and image j, and λ represents the weight of the similarity loss of the interpretation result.
[0031] The beneficial effects of the present invention are as follows: The present invention is a method for contrastive learning training. When performing image classification or object detection tasks, the contrastive learning model serves as a feature extractor. After fine-tuning, it will extract more discriminative features between categories, thereby indirectly improving the accuracy of the final result of the entire downstream task. The present invention uses the similarity of the interpretation results to constrain the network to compare the key features between pictures, thereby enhancing the feature extraction ability of the network, improving the processing performance of the downstream task, improving the robustness of the model, and having good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flowchart of the method for improving the robustness of the model based on contrastive learning in the present invention;
[0033] Figure 2 It is an interpretation heat map of the present invention and the comparative method in model decision-making under different training strategies in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] The present invention proposes a method for improving the robustness of the model based on contrastive learning, which is implemented based on the CLCE (contrastive learning framework constrained by interpretive results) designed by the present invention, as Figure 1 shown. The method includes the following contents:
[0036] S1: Obtain an image data set and perform data augmentation, and use the augmented image data set to perform the first-stage contrastive learning training on the neural network model to obtain a preliminarily trained neural network model.
[0037] The present invention adopts two-stage training. The first stage:
[0038] Obtain an image data set and perform data augmentation, including random cropping, random flipping, and random color transformation; use the image contrastive learning method to train the neural network model f(·) (such as MoCo, SimCLR, etc.). Train for p epochs, and save the trained model weights for the second model training. The loss function used for the contrastive learning method training is the contrastive learning loss, and a preliminarily trained neural network model is obtained; the first-stage training enables the model to have a certain feature extraction ability.
[0039] S2: Perform secondary augmentation on the image data set to obtain a first augmented image and a second augmented image.
[0040] In the second stage, an additional constraint on the similarity of the explanatory results is added on the basis of conventional contrastive learning. In the second stage:
[0041] Given an image, after performing non-spatial transformation enhancement on it (such as random color transformation, Gaussian blur, etc., but without cropping and flipping), the first enhanced image x is obtained i ; perform spatial transformation data augmentation on the image, such as random cropping, scaling, and flipping, etc., and record this operation as u(·) to obtain an intermediate image; perform non-spatial transformation data augmentation on the intermediate image to obtain the second enhanced image x j ; use x i and x j as positive sample pairs, and use x i and all other images in a batch except x j as negative sample pairs.
[0042] S3: Input the first enhanced image and the second enhanced image of each image into the preliminarily trained neural network model for processing, and calculate the similarity and contrastive learning loss of the two enhanced images of the same image according to the processing results.
[0043] The process of calculating the similarity of the two enhanced images includes:
[0044] The image is processed by the neural network model, and the cosine similarity between the feature encodings extracted by the neural network model for the two images is calculated.
[0045] Calculate the contrastive learning loss:
[0046] In some preferred embodiments, the neural network is trained using the SimCLR method. The two views of the first enhanced image and the second enhanced image are mapped to feature representations h i and h j through an encoder network (such as ResNet-50), and further obtain representation vectors z i and z j through a non-linear mapping head. SimCLR aims to minimize the distance between the positive sample pair z i and z j in the representation space, while maximizing the difference from other samples in the same batch. Its target loss function, that is, the contrastive learning loss, is defined as:
[0047]
[0048] where 1 [k≠i] ∈ {0, 1} is a mapping function with a value of 1 when k ≠ i, τ is the temperature hyperparameter, and N is the batch size.
[0049] S4: Calculate the similarity interpretation heatmap of the positive pair images based on the similarity of the two images.
[0050] After calculating the similarity, according to the Grad-CAM method, the similarity interpretation heatmap of the positive pair images is calculated by taking the derivative and weighting the cosine similarity, that is, e i and e j . Among them, e i represents the similarity feature interpretation of the input sample x i relative to x j , and e j is the same. The formula for calculating the similarity heatmap using Grad-CAM is as follows:
[0051]
[0052] Among them, e i represents the similarity interpretation heatmap of the first augmented image x i relative to the second augmented image x j , e j represents the similarity interpretation heatmap of the second augmented image x j relative to the first augmented image x i , Z represents the total number of elements in the feature map, f(·) represents the neural network model, sim(·) represents the similarity calculation, represents the element at the m-th row and n-th column of the k-th feature map in the last layer of the network during the forward propagation of the first augmented image x i , represents the element at the m-th row and n-th column of the k-th feature map in the last layer of the network during the forward propagation of the second augmented image x j .
[0053] S5: Perform the second-stage contrastive learning training on the preliminarily trained neural network model according to the similarity interpretation heatmap of the positive pair images and calculate the similarity loss of the interpretation heatmap.
[0054] Since x i in the positive pair retains the complete image content information of x, and x j is generated by randomly augmenting the cropped x, so x i contains x j . After generating the interpretation result, the interpretation heatmap e j should also be a part of e i . After performing the u(·) operation on e i , we get and e j are in a similar relationship. Therefore, and e j can be used as a positive pair, The generated explanation heatmaps for all other positive pairs in a batch are used as negative pairs, and contrastive learning training is performed again; thus, the aim is to make the explanations corresponding to different data augmentation versions of the same image similar, and the explanations of different images in the same batch dissimilar. The formula for calculating the similarity loss of the explanation heatmaps is as follows:
[0055]
[0056] Among them, represents the similarity loss of the explanation heatmaps between image i and image j, u(·) represents the spatial transformation data augmentation operation, N represents the batch size. Since there are N images in a batch, after data augmentation, there will be 2N image samples, and only the positive pairs are used for similarity explanations, so there will be 2N explanation heatmaps; e i represents the similarity explanation heatmap of the first augmented image x i relative to the second augmented image x j , e j represents the similarity explanation heatmap of the second augmented image x j relative to the first augmented image x i , e k represents the k-th similarity explanation heatmap in the same batch, and τ represents the temperature hyperparameter.
[0057] S6: Calculate the total model loss according to the contrastive learning loss and the similarity loss of the explanation heatmaps; adjust the model parameters according to the total model loss to obtain a trained neural network model.
[0058] The present invention takes the weighted sum of the contrastive learning loss and the similarity loss of the explanation heatmaps as the total model loss, which is expressed as:
[0059]
[0060] Among them, represents the total model loss, represents the contrastive learning loss between image i and image j, represents the similarity loss of the explanation heatmaps between image i and image j, and λ represents the weight of the similarity loss of the explanation results.
[0061] Update the network by backpropagation according to the total model loss function, and train for q epochs according to the steps of S2 - S6. The trained model can be used for feature extraction work in downstream tasks (such as image classification, object detection).
[0062] Evaluate the present invention:
[0063] Compare the present invention with the comparative method, such as Figure 2As shown, for the positive pair of images (the first two groups), when the original model compares the dissimilar features of the two images, its attention will be scattered to the foreground objects with the same original category, while the model trained under the interpretability constraint will focus its attention on the background area. In the negative pair of samples (the bottom example), when the original model compares the similarity of the two images, due to the scattered attention mechanism, there are actually inconsistent semantic feature comparisons, such as comparing the light gray floor with the light gray fur of the white fox in another image; while the model trained with the interpretability constraint can maintain semantic consistency and accurately achieve local feature alignment, such as the matching of the foot features of the dog in the explanatory image with the feet in the comparison image, proving that the method of the present invention enhances the feature extraction ability of the network and improves the processing performance of downstream tasks.
[0064] In summary, the present invention uses the similarity of the interpretation results of two images to constrain the model to compare the discriminative features between images, thereby improving the feature extraction ability and robustness of the model, and further enhancing the performance of the model when performing downstream tasks.
[0065] The above-mentioned embodiments further elaborate on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for improving the robustness of a model based on contrastive learning, characterized in that, Including: S1: Obtain an image dataset and perform data augmentation. Use the augmented image dataset to conduct the first-stage contrastive learning training on a neural network model to obtain a preliminarily trained neural network model. S2: Perform secondary augmentation on the image dataset to obtain a first augmented image and a second augmented image. S3: Input the first augmented image and the second augmented image of each image into the preliminarily trained neural network model for processing, and calculate the similarity and contrastive learning loss between the two augmented images of the same image according to the processing results. S4: Calculate the similarity interpretation heatmap of the positive pair images according to the similarity between the two images. S5: Conduct the second-stage contrastive learning training on the preliminarily trained neural network model according to the similarity interpretation heatmap of the positive pair images and calculate the interpretation heatmap similarity loss. S6: Calculate the total model loss according to the contrastive learning loss and the interpretation heatmap similarity loss; adjust the model parameters according to the total model loss to obtain a trained neural network model.
2. The method for improving the robustness of a model based on contrastive learning according to claim 1, wherein The process of performing secondary augmentation on the augmented image dataset includes: Perform non-spatial transformation data augmentation on the image to obtain a first augmented image. Perform spatial transformation data augmentation on the image to obtain an intermediate image; perform non-spatial transformation data augmentation on the intermediate image to obtain a second augmented image.
3. A method for improving the robustness of a model based on contrastive learning according to claim 1, characterized in that, The process of calculating the similarity between two images includes: The image is processed by the neural network model, and the cosine similarity between the feature encodings extracted by the neural network model for the two images is calculated.
4. A method for improving the robustness of a model based on contrastive learning according to claim 1, characterized in that The formula for calculating the similarity interpretation heatmap of the positive pair images is: Among them, e i represents the similarity interpretation heat map of the first enhanced image x i relative to the second enhanced image x j , and e j represents the similarity interpretation heat map of the second enhanced image x j relative to the first enhanced image x i . Z represents the total number of elements of the feature map, f(·) represents the neural network model, and sim(·) represents the similarity calculation represents the element at the m-th row and n-th column of the k-th feature map in the last layer of the network during the forward propagation process of the first enhanced image x i , and represents the element at the m-th row and n-th column of the k-th feature map in the last layer of the network during the forward propagation process of the second enhanced image x j .
5. A method for improving the robustness of a model based on contrastive learning according to claim 1, characterized in that The process of conducting contrastive learning training on the neural network model according to the similarity interpretation heatmap of the positive pair images includes: Randomly crop, scale, and flip the similarity interpretation heatmap of one image in the positive pair images to obtain an augmented interpretation heatmap. Use the augmented interpretation heatmap and the similarity interpretation heatmap of the other image in the positive pair images as positive samples; use the augmented interpretation heatmap and the similarity interpretation heatmaps generated from other images in the same batch as negative samples to conduct contrastive learning training on the neural network model.
6. A method for improving the robustness of a model based on contrastive learning according to claim 1, characterized in that, The formula for calculating the interpretation heatmap similarity loss is: Among them, represents the interpretability heatmap similarity loss between image i and image j, u(·) represents the spatial transformation data augmentation operation, N represents the batch size, and e i represents the first augmented image x i with respect to the second augmented image x j of the similarity interpretability heatmap, and e j represents the second augmented image x j with respect to the first augmented image x i of the similarity interpretability heatmap, and e k represents the k-th similarity interpretability heatmap in the same batch, and τ represents the temperature hyperparameter.
7. A method for improving the robustness of a model based on contrastive learning according to claim 1, characterized in that The formula for calculating the total model loss is: Among them, l represents the total loss of the model, and l i,j represents the contrastive learning loss between image i and image j, represents the similarity loss of the interpretation heatmaps between image i and image j, and λ represents the weight of the similarity loss of the interpretation result.