Adversarial Training Method for Object Detection Models Based on Contrastive Learning

By introducing a contrast learning module into the object detection model and combining an adversarial training method, the problem of performance balance between adversarial samples and clean samples in the prior art is solved, and higher robustness and accuracy are achieved.

CN117197577BActive Publication Date: 2025-06-10YUNNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311218767.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-20
Publication Date
2025-06-10
Estimated Expiration
2043-09-20

AI Technical Summary

Technical Problem

Existing target detection models are not robust enough in the face of adversarial attacks, resulting in a significant decrease in accuracy on adversarial samples but a difficult balance between the two.

Method used

The object detection model adversarial training method based on contrast learning is adopted, and the object detection model is divided into two parts: backbone network and detection head, and a comparison learning module is added after the backbone network. The features are further extracted through the contrast learning module, and the training is combined with clean samples and adversarial samples, and the loss function is optimized to balance the performance between the two.

Benefits of technology

While improving the robustness of the model to adversarial samples, maintain or improve the accuracy on clean samples, achieving a balance between adversarial samples and clean samples performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117197577B_ABST
    Figure CN117197577B_ABST
Patent Text Reader

Abstract

The present invention discloses an adversarial training method for an object detection model based on contrastive learning. The object detection model is divided into a backbone network and a detection head. A contrastive learning module is added after the backbone network to form a network model. A number of training sample pairs are collected according to the needs of the object detection model and the network model is trained. The contrastive learning module is removed from the trained network model to restore the object detection model for actual object detection applications. By combining the characteristics of contrastive learning during the adversarial training process, the present invention enables the object detection model to learn more robust feature representations of adversarial samples and clean samples during training, so that the object detection model can achieve better accuracy on clean samples while obtaining higher robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of object detection. More specifically, it relates to an adversarial training method for an object detection model based on contrastive learning. Background Art

[0002] Object detection is an important computer vision task, which aims to recognize the surrounding environment by predicting the categories and locations of objects of interest in images. Object detection has extensive requirements and values in many practical applications, such as visual surveillance, autonomous driving, face recognition, etc. Object detection usually uses sensors such as lidar and cameras to obtain the depth and semantic information of the scene, and uses deep learning models for feature extraction and prediction.

[0003] However, the robustness issue of deep learning models has always been an obstacle affecting their performance and security. Research shows that deep learning models are vulnerable to adversarial attacks, that is, adding some tiny perturbations to the original image, which makes the model predict incorrect results. These adversarial perturbations are often imperceptible to the human eye, but can cause the model to completely fail. This is unacceptable for applications such as autonomous driving where safety is crucial. To improve the robustness of object detection models, it is necessary to conduct quantitative analysis and evaluation on them and design effective defense methods. Currently, many research works have explored and solved the robustness issue of object detection models from different perspectives.

[0004] Currently, the research on model robustness mainly falls into three categories. The first is to modify the model, that is, to achieve defense by modifying the model structure or parameter information learned from the data, such as methods like network distillation and gradient regularization. The second category is to add auxiliary models to enhance robustness by introducing additional network models, such as adversarial detection and ensemble defense. The third category is to modify the data, that is, to achieve defense by modifying the data and features during the training or testing phase, such as feature compression and adversarial training. Among them, adversarial training has been proven to be a very effective method for improving the adversarial robustness of deep learning models. By incorporating adversarial examples into the training data, the deep model becomes more resilient to attacks and performs well in real-world scenarios. However, the main problem with current adversarial training is that while it can increase the model's robustness to adversarial samples to a certain extent, it also significantly reduces the accuracy of the model on clean samples. Therefore, how to balance clean samples and adversarial samples is the biggest obstacle to adversarial training. Zhang et al. first introduced adversarial training to enhance the adversarial robustness of object detection models by selecting the most adversarial adversarial samples in different task domains as the enhanced data for adversarial training, thus making the model have good robustness. Although their method improves the robustness of the object detection model to adversarial samples, since the model learns the decision boundary of the most adversarial features, it results in too much loss of accuracy on clean samples. Chen et al. proposed CWAT, which improved the generation method of adversarial examples, but still did not solve the above problem. Xie et al. used knowledge extraction and feature alignment to reduce the loss on clean samples, but with limited effect. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide an adversarial training method for object detection models based on contrastive learning. By combining the characteristics of contrastive learning during the adversarial training process, the object detection model can learn more robust feature representations of adversarial samples and clean samples during training, enabling the object detection model to achieve better accuracy on clean samples while obtaining higher robustness.

[0006] To achieve the above invention purpose, the adversarial training method for object detection models based on contrastive learning of the present invention includes the following steps:

[0007] S1: Divide the object detection model M into a backbone network and a detection head. The backbone network extracts features from the input image to obtain a feature image f, and the detection head completes object detection based on the feature image f. Add a contrastive learning module after the backbone network to form a network model M c , and the contrastive learning module is used to receive the feature image f output by the backbone network for further feature extraction to obtain the feature g for contrastive learning;

[0008] S2: Collect a number of training sample pairs according to the requirements of the target detection model. Each training sample pair includes a clean sample image and the corresponding adversarial sample image, and mark the target area in each sample image. Use the annotation information of the target area as the label of the sample image;

[0009] S3: Use the clean sample images and adversarial sample images in the training sample pairs collected in step S2 to train the network model M c to obtain the trained network model M c , and the calculation formula of the loss function L for each training batch in the training process is as follows:

[0010] L = αL det (x) + βL det (x adv ) + L cl

[0011] where L det (x) represents the target detection loss of the clean sample image x in the current training batch, and L det (x adv ) represents the target detection loss of the adversarial sample image x adv in the current training batch, α and β represent preset weight coefficients, and L cl represents the contrastive learning loss between the clean sample image and the adversarial sample image in the current training batch;

[0012] S4: Remove the contrastive learning module from the trained network model M c to restore the target detection model M for actual target detection applications.

[0013] The adversarial training method for the target detection model based on contrastive learning in the present invention divides the target detection model into a backbone network and a detection head. A contrastive learning module is added after the backbone network to form a network model. A number of training sample pairs are collected according to the requirements of the target detection model and the network model is trained. The contrastive learning module is removed from the trained network model to restore the target detection model for actual target detection applications.

[0014] The present invention proposes to combine the characteristics of contrastive learning in the training stage of the target detection model, and use the principles of homogeneous clustering and heterogeneous exclusion of contrastive learning to help the target detection model obtain more robust adversarial features and clean feature representations. Experiments show that the present invention can achieve good results on both adversarial samples and clean samples. Brief Description of the Drawings

[0015] Figure 1 is a flowchart of the specific implementation of the adversarial training method for the target detection model based on contrastive learning in the present invention;

[0016] Figure 2 It is a schematic diagram of adding a contrastive learning module to the target detection model in this embodiment;

[0017] Figure 3 It is the network model structure diagram obtained by adding a contrastive learning module to the SSD model in this embodiment;

[0018] Figure 4 It is the network model structure diagram obtained by adding a contrastive learning module to the YOLO model in this embodiment. Specific implementation manners

[0019] The following describes the specific implementation manners of the present invention with reference to the accompanying drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.

[0020] Embodiment

[0021] Figure 1 It is the flowchart of the specific implementation manner of the adversarial training method for the target detection model based on contrastive learning of the present invention. As Figure 1 shown, the specific steps of the adversarial training method for the target detection model based on contrastive learning of the present invention include:

[0022] S101: Add a contrastive learning module:

[0023] There are usually two components in the target detection model, namely Backbone and Head. Among them, Backbone is the backbone network, mainly referring to the convolutional neural network for feature extraction (for example: ResNet-50, Darknet53, etc.). The backbone network converts the input image into a representation in a high-dimensional feature space and has usually been pre-trained on large datasets (such as ImageNet|COCO, etc.) and has pre-trained weights; Head, that is, the detection head, is mainly used to convert the abstract high-dimensional features into data representations familiar to humans, that is, to predict the types and positions (bounding boxes) of targets.

[0024] In the present invention, the target detection model M is divided into two parts: the backbone network and the detection head. The backbone network extracts features from the input image to obtain the feature image f, and the detection head completes target detection according to the feature image f. In practical applications, the detection head can be determined first, and then the part before the detection head is used as the backbone network. Then a contrastive learning module is added after the backbone network to form the network model M c , and the contrastive learning module is used to receive the feature image f output by the backbone network for further feature extraction to obtain the feature g for contrastive learning.

[0025] Figure 2 It is a schematic diagram of adding a contrastive learning module to the target detection model in this embodiment. As Figure 2 shown, the contrastive learning module in this embodiment includes a global average pooling layer, a first fully connected layer, a ReLU activation function layer, and a second fully connected layer, where:

[0026] The global average pooling layer is used to perform global average pooling on the feature image f and send the obtained features to the first fully connected layer.

[0027] The first fully connected layer is used to integrate the received features and send the obtained features to the ReLU activation function layer.

[0028] The ReLU activation function layer is used to process the received features using the ReLU activation function and send the obtained features to the second fully connected layer.

[0029] The second fully connected layer is used to integrate the received features and use the obtained features as the features g for contrastive learning.

[0030] S102: Obtain training samples:

[0031] According to the requirements of the target detection model, collect a number of training sample pairs. Each training sample pair includes a clean sample image and a corresponding adversarial sample image, and mark the target area in each sample image. Use the annotation information of the target area as the label of the sample image.

[0032] In this embodiment, the adversarial sample image is generated based on the network model M c , and the specific method is as follows:

[0033] Record the number of collected clean sample images as D. Randomly arrange the D clean sample images twice, and denote the obtained sequences as X = {x 1 , x 2 , …, x D}, Y = {y 1 , y 2 , …, y D}, where x d and y d respectively represent the d-th clean sample image in sequences X and Y. Input the clean sample images in sequences X and Y into the network model M c in sequence, obtain the features g(x d ) and g(y d ) output by the contrastive learning module, and then calculate the perturbation amount ε d of the d-th clean sample image x d according to the following formula:

[0034]

[0035] Among them, η represents a preset perturbation coefficient, and sign() represents the sign function. represents the gradient calculated based on the loss loss(x d , y d ). The calculation formula of the loss loss(x d , y d ) is as follows:

[0036]

[0037] Among them, exp() represents the exponential function with the natural constant e as the base, τ represents a preset temperature coefficient, and g(y d′ ) represents the feature output by the contrastive learning module after the d'-th clean sample image in the sequence Y is input into the network model M c . 1 [d′≠d] represents the indicator function, which is 1 when d'≠d [d′≠d] = 1, otherwise 1 [d′≠d] = 0, and sim() represents calculating the similarity of features. In this embodiment, the feature output by the contrastive learning module is a feature vector, so the cosine similarity can be used as the similarity of the features f(x d ), f(y d ).

[0038] Add the perturbation amount ε n to the clean sample image x n to obtain the corresponding adversarial sample image x adv,n .

[0039] S103: Compare and train the network model:

[0040] Use the clean sample images and adversarial sample images in the training sample pairs collected in step S2 to train the network model M c to obtain the trained network model M c . During the training process, the calculation formula of the loss function L for each training batch is as follows:

[0041] L = αL det (x) + βL det (x adv ) + L cl

[0042] Among them, L det (x) represents the object detection loss of the clean sample image x in the current training batch, and L det (x adv ) represents the object detection loss of the adversarial sample image x adv in the current training batch. α and β represent preset weight coefficients, and L clRepresents the contrastive learning loss of the current training batch.

[0043] Since the object detection task is generally considered a multi-task learning task, the loss function for the original object detection task is mainly determined by the localization task and the classification task. Therefore, in this embodiment, the object detection loss L det (z) is calculated as follows:

[0044] L det (z) = L cls (z) + L loc (z)

[0045] where z ∈ {x, x adv},L cls (z) represents the classification loss of the sample image z, usually using the Cross-Entropy loss, and L loc (z) represents the localization loss of the sample image z, usually using the IOU (Intersection over Union) loss.

[0046] The contrastive learning loss L cl is obtained based on the clean sample image and the adversarial sample image, and the calculation method is as follows:

[0047] Denote the number of training sample pairs in the current training batch as N. Denote the serial number of the clean sample image in the nth training sample pair as 2n - 1, and the serial number of the adversarial sample image as 2n. Then, the contrastive learning loss L of the current training batch is calculated as follows cl :

[0048]

[0049] where l(2n - 1, 2n) represents the contrast loss of the clean sample image in the nth pair of training samples, and l(2n, 2n - 1) represents the contrast loss of the adversarial sample image in the nth pair of training samples. The calculation formula is as follows:

[0050]

[0051] where g i , g j , g k respectively represent the features output by the contrastive learning module after the sample images with serial numbers i, j, and k are input into the network model M c , and 1 [k≠i] represents the indicator function. When k′ ≠ i, 1 [k≠i] = 1, otherwise 1 [k≠i] = 0.

[0052] S104: Separate the object detection model:

[0053] Remove the contrastive learning module from the trained network model M c to obtain the object detection model M by restoration for actual object detection applications.

[0054] To achieve the above invention objective, this embodiment uses specific examples to conduct experimental verification on the present invention.

[0055] In this embodiment, two object detection models are selected, namely the SSD (Single Shot Detector) model and the YOLO (You Only Look Once) model. First, add a contrastive learning module to the object detection model. Figure 3 is the network model structure diagram obtained by adding a contrastive learning module to the SSD model in this embodiment. As Figure 3 shown, for the SSD model, select the VGG part in the network as the backbone network because it is found that the main feature representation in the SSD model is in this layer, and the subsequent extra part plays a role in aggregation and more advanced semantic extraction. Therefore, add the contrastive learning module after the main network structure of VGG. Figure 4 is the network model structure diagram obtained by adding a contrastive learning module to the YOLO model in this embodiment. As Figure 4 shown, the YOLO model has a more complex structure and features of multi-scale fusion, and three output structures of the DARKNET network need to be considered simultaneously. Therefore, for the YOLO model, jointly input the outputs of the three scales of the backbone network into the contrastive learning module, and simultaneously perform contrastive feature learning of clean samples and adversarial samples in three different feature spaces.

[0056] This embodiment uses two datasets for experiments, namely the PASCAL-VOC dataset and the MS-COCO dataset. During the experiment, resize the sizes in the dataset for different object detection modules. In this embodiment, the input image size of the SSD model is 300x300, and the input image size of the YOLO model is 416x416.

[0057] In this experiment, 11 existing methods are selected as comparison methods to conduct object detection tests on adversarial samples obtained by different attack methods, and the accuracy of object detection is statistically analyzed. The comparison methods used are:

[0058] AT-CLS, AT-LOC, AT-CON, MTD, MTD-fast, see the literature "Zhang H, Wang J. Towards adversarially robust object detection[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019: 421-430.";

[0059] TOAT-6, OWAT, CWAT, see the literature "Chen P C, Kung B H, Chen J C. Class-aware robust adversarial training for object detection[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021: 10420-10429.";

[0060] KDFA, SSFA, FA, see the literature "Xu W, Huang H, Pan S. Using feature alignment can improve clean average precision and adversarial robustness in object detection[C] / / 2021 IEEE International Conference on Image Processing(ICIP). IEEE, 2021: 2184-2188.".

[0061] The attack methods include attacks based on the classification task domain (A cls ), attacks based on the classification localization domain (A cls ), dense adversarial generation (DAG), class-wise attack (CWA), and projected gradient descent attack PGD.

[0062] First, the PASCAL-VOC dataset is used for experimental verification. VOC2012 trainval + VOC2007 train is used as the training dataset, and VOC2007 test is used as the validation set. Table 1 is the statistical table of the object detection accuracy of the present invention and the comparative methods for different samples of the PASCAL-VOC dataset when using the SSD model in this embodiment.

[0063]

[0064]

[0065] Table 1

[0066] Table 2 is a statistical table of the object detection accuracy of the present invention and the comparative method for different samples of the PASCAL-VOC dataset when the YOLO model is adopted in this embodiment.

[0067]

[0068] Table 2

[0069] As shown in Table 1 and Table 2, on the PASCAL-VOC dataset, the accuracy and mean average precision of the present invention on clean samples are both higher than those of the comparative method, which also reflects the advantage of the present invention in the robustness of object detection.

[0070] Then, the MS-COCO dataset is used for experimental verification, with COCO2017train as the training dataset and COCO2017val as the validation dataset. Table 3 is a statistical table of the object detection accuracy of the present invention and the comparative method for different samples of the MS-COCO dataset when the SSD model is adopted in this embodiment.

[0071]

[0072] Table 3

[0073] Table 4 is a statistical table of the object detection accuracy of the present invention and the comparative method for different samples of the MS-COCO dataset when the YOLO model is adopted in this embodiment.

[0074]

[0075]

[0076] Table 4

[0077] As shown in Table 3 and Table 4, on the MS-COCO dataset, the accuracy and mean average precision of the present invention on clean samples are also higher than those of the comparative method, and the performance in terms of robustness is also better.

[0078] Although the above describes the illustrative specific embodiments of the present invention for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

Claims

1. An adversarial training method for an object detection model based on contrastive learning, characterized in that, it includes the following steps: S1: Divide the target detection model M into two parts: a backbone network and a detection head. The backbone network extracts features from the input image to obtain a feature image f, and the detection head completes target detection based on the feature image f; add a contrastive learning module after the backbone network to form a network model M c , the contrastive learning module is used to receive the feature image f output by the backbone network for further feature extraction to obtain a feature g for contrastive learning; the contrastive learning module includes a global average pooling layer, a first fully connected layer, a ReLU activation function layer, and a second fully connected layer, where: The global average pooling layer is used to perform global average pooling on the feature image f, and the obtained features are sent to the first fully connected layer; The first fully connected layer is used to integrate the received features and send the obtained features to the ReLU activation function layer; The ReLU activation function layer is used to process the received features using the ReLU activation function and send the obtained features to the second fully connected layer; The second fully connected layer is used to integrate the received features and use the obtained features as the features g for contrastive learning; S2: Collect a number of training sample pairs according to the needs of the object detection model. Each training sample pair includes a clean sample image and a corresponding adversarial sample image, and the target area is marked in each sample image. The annotation information of the target area is used as the label of the sample image; S3: Use the clean sample images and adversarial sample images in the training sample pairs collected in step S2 to train the network model M c to obtain a trained network model M c , and the calculation formula of the loss function L for each training batch during the training process is as follows: L = αL det (x) + βL det (x adv ) + L cl Among them, L det (x) represents the object detection loss of the clean sample image x in the current training batch, L det (x adv ) represents the object detection loss of the adversarial sample image x adv in the current training batch, α and β represent preset weight coefficients, and L cl represents the contrastive learning loss between the clean sample image and the adversarial sample image in the current training batch; S4: Remove the contrastive learning module from the trained network model M c to restore the object detection model M for actual object detection applications.

2. The adversarial training method for an object detection model according to claim 1, characterized in that, the generation method of the adversarial sample image in the step S2 is: Let the number of collected clean sample images be D. The D clean sample images are randomly permuted twice, and the obtained sequences are denoted as X = {x 1 , x 2 , …, x D}, Y = {y 1 , y 2 , …, y D}, where x d and y d represent the d-th clean sample images in sequences X and Y respectively. The clean sample images in sequences X and Y are input into the network model M c in order, and the features g(x d ) and g(y d ) output by the contrastive learning module are obtained. Then, the perturbation amount ε d of the d-th clean sample image x d is calculated according to the following formula: Among them, η represents a preset perturbation coefficient, and sign() represents the sign function. represents the gradient calculated based on the loss loss(x d , y d ), and the calculation formula of the loss loss(x d , y d ) is as follows: Among them, exp() represents the exponential function with the natural constant e as the base, τ represents the preset temperature coefficient, and g(y d′ ) represents the feature output by the contrastive learning module after the d'-th clean sample image in the sequence Y is input into the network model M c and 1 [d′≠d] represents the indicator function, which is 1 when d'≠d [d′≠d] and 1 [d′≠d] otherwise 0, and sim() represents calculating the similarity of features; Add the perturbation amount ε n to the clean sample image x n to obtain the corresponding adversarial sample image x adv,n .

3. The adversarial training method for an object detection model according to claim 1, characterized in that, The target detection loss L det (z) in step S3 is calculated as follows: L det ψ(z) = L cls φ(z)+L loc χ(z) where \(z\in\{x, x\) adv}\), \(L\) cls (z) represents the classification loss of the sample image \(z\), and \(L\) loc (z) represents the localization loss of the sample image \(z\).

4. The adversarial training method for an object detection model according to claim 1, characterized in that, The calculation method of the contrastive learning loss L cl in step S3 is as follows: Let the number of training sample pairs in the current training batch be N. Denote the serial number of the clean sample image in the nth training sample pair as 2n - 1, and the serial number of the adversarial sample image as 2n. Then, calculate the contrastive learning loss L of the current training batch as follows cl : where l(2n - 1, 2n) represents the contrastive loss of the clean sample image in the nth pair of training samples, and l(2n, 2n - 1) represents the contrastive loss of the adversarial sample image in the nth pair of training samples. The calculation formula is as follows: Among them, g i , g j , g k respectively represent the features output by the contrastive learning module after the sample images numbered i, j, and k are input into the network model M c , 1 [k≠i] represents the indicator function, and 1 [k≠i] = 1 when k′≠i, otherwise 1 [k≠i] = 0.