Adversarial sample anomaly detection method for traffic sign board, terminal and medium
Through the anti-sample abnormal detection method for traffic signs, unsupervised training and semantic contradiction detection ideas are used to solve the problem of uncertain anti-sample detection performance in the field of autonomous driving, efficient and accurate anti-sample detection is achieved, and detection performance and adaptability are significantly improved.
Patent Information
- Application Number
- CN202411986792.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
In the field of autonomous driving, existing adversarial sample detection methods are difficult to effectively deal with diverse attack methods, resulting in uncertain detection efficiency and high cost of generating adversarial samples, affecting the robustness of the model.
A method of detecting anomaly of antagonistic samples for traffic signs is proposed. By utilizing the trained target classification model, automatic encoder model and abnormality detection model, combined with unsupervised training, the adversarial samples are identified and detected. This method does not require the generation of adversarial samples, which reduces time and calculation costs, and improves detection accuracy through semantic contradiction detection ideas.
It realizes more efficient and accurate detection of traffic sign images that have been maliciously tampered with, improves the performance and adaptability of anti-sample detection, and the detection recall rate reaches 90%-100%, significantly reducing the risk of anti-sample fraud target classification model.
Smart Images

Figure CN119942494A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence security, and specifically relates to an adversarial sample anomaly detection method, terminal and medium for traffic signs. Background Art
[0002] In the cutting-edge and critical application field of autonomous driving, the rapid maturity of deep learning technology has brought it unprecedented development opportunities. However, while deep neural networks (DNNs) show excellent performance, they also have potential security vulnerabilities and risks that cannot be ignored. One of the most difficult problems is adversarial sample attacks. Adversarial samples cleverly embed tiny perturbations that are almost imperceptible to human vision into the original driving scene data, changing the classification decisions of the autonomous driving system and causing serious safety issues.
[0003] This carefully designed adversarial attack can cause the image recognition module of the autonomous driving system to misjudge key road signs, thereby making dangerous driving maneuvers. What is more serious is that once the model is deceived by such an attack, the confidence of its wrong prediction is often extremely high, and the same small perturbation can span multiple different deep learning models to achieve a wide range of deception effects. In the field of autonomous driving cars, which is related to public safety and personal life and property, such attacks undoubtedly pose a huge threat, which may lead to frequent traffic accidents and seriously threaten the safety of road users.
[0004] Since the concept of adversarial attacks was proposed, researchers have explored a variety of attack methods. Their diversity and complexity make it extremely difficult to build a defense system that can fully defend against all known and potential attack methods. Existing defense strategies have their own advantages, but they are also accompanied by their own limitations. For example, gradient masking technology aims to build models that are difficult for attackers to exploit, but it is easily bypassed by black-box attacks; although adversarial training can improve the robustness of the model to adversarial samples, it often comes at the cost of reducing the recognition accuracy of normal inputs; and although detection-based methods can circumvent the above defects, they have to bear additional model training costs, and their detection performance is also uncertain when facing new or more covert attacks.
[0005] Traditional adversarial sample detection methods, such as adversarial training, require pre-collection of adversarial samples for a specific model and mixed training with normal sample data to generate a detector that can distinguish adversarial samples from normal images. However, this step itself faces huge challenges: the diversity of attack methods not only increases the cost of generating adversarial samples that fully cover all attack types, but also makes the detection model likely to fail when facing new types of attacks. Therefore, how to effectively deal with the threat of adversarial samples in the field of autonomous driving and ensure driving safety has become a key issue that needs to be solved urgently. Summary of the invention
[0006] The purpose of the present invention is to overcome the deficiencies in the prior art and to provide an adversarial sample anomaly detection method, terminal and medium for traffic signs. The present invention can more efficiently and accurately detect traffic sign images that have been maliciously tampered with, further improving the performance and adaptability of adversarial sample detection.
[0007] To achieve the above object, the present invention is implemented by adopting the following technical solutions:
[0008] In a first aspect, the present invention provides an adversarial sample anomaly detection method for traffic signs, comprising:
[0009] Inputting a normal sample into a trained target classification model, wherein the normal sample is a real traffic sign image that has not been tampered with, and the target classification model is used to identify the category of the traffic sign image and output a corresponding category label;
[0010] Filter correctly classified normal samples from the output of the target classification model and apply adversarial perturbations to them to generate new samples;
[0011] Normal samples and generated new samples are mixed to form a test set, and each sample in the test set is input into the autoencoder model trained with normal samples. The autoencoder model generates the corresponding reconstructed image based on the category label of the input sample, and calculates the difference map between the input sample and the reconstructed image.
[0012] The difference map is input into the anomaly detection model trained with normal samples, the energy value of the input sample is calculated, and the energy value is compared with the energy threshold of the normal sample. If the energy value is greater than the energy threshold, the sample is judged to be an adversarial sample.
[0013] Furthermore, the energy threshold of the normal sample is obtained by the following steps:
[0014] The normal sample is input into the trained anomaly detection model to obtain the energy value of the normal sample, and the energy threshold is determined according to the energy value.
[0015] Further, the autoencoder model and the anomaly detection model are trained by the following steps:
[0016] The normal samples are used to train the autoencoder model to obtain a difference map between the input sample and the reconstructed image; the difference map is input into the anomaly detection model, and the anomaly detection model is trained by an unsupervised learning method.
[0017] Furthermore, the autoencoder model is an improved conditional variational autoencoder generative adversarial network cVAE-GAN model, and the cVAE-GAN model adds a self-attention module at the output ends of the encoder and decoder.
[0018] Furthermore, the anomaly detection model is a deep auto-encoding Gaussian mixture model DAGMM, and the fully connected layer in the compression network of the DAGMM is replaced by a convolutional layer.
[0019] Furthermore, correctly classified normal samples are selected from the output of the target classification model, and adversarial perturbations are applied to them to generate new samples, including:
[0020] The adversarial perturbation applies adversarial perturbations of different strengths to normal samples correctly classified by the target classification model, and adopts multiple attack strategies to generate new samples.
[0021] Furthermore, the strength of the anti-disturbance is determined by Norm control; the attack strategies include FGSM attack, BIM attack, PGD attack, DeepFool attack and C&W attack.
[0022] Furthermore, the target classification model is constructed based on ResNet18, Vgg19, DenseNet169 or MobileNet network structure respectively.
[0023] In a second aspect, the present invention provides an electronic terminal, comprising a processor and a memory connected to the processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the above-mentioned adversarial sample anomaly detection method for traffic signs are performed.
[0024] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the program is executed by a processor, the steps of the above-mentioned adversarial sample anomaly detection method for traffic signs are implemented.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] The adversarial sample anomaly detection method for traffic signs proposed in the present invention does not need to generate adversarial samples to train the autoencoder model and the anomaly detection model, thereby greatly reducing the time and computing cost, while avoiding the uncertainty and efficiency problems that may be caused when generating adversarial samples. In addition, the method does not directly change the structure and parameters of the target classification model, so it will not affect the classification ability of the target classification model for normal images, and maintains its classification accuracy and stability.
[0027] The present invention adopts an unsupervised training method, and only uses normal samples to train the model, without the need for adversarial samples to participate in the training. By reconstructing normal samples and calculating the difference between them and the original samples, this method can train the anomaly detection module to effectively distinguish normal samples from adversarial samples. Unsupervised training not only saves the time of generating adversarial samples, but also significantly shortens the training cycle of the detection module. At the same time, it can still maintain a high detection accuracy when facing a variety of different types of attacks. The detection recall rate mostly reaches 90%-100%, ensuring that malicious adversarial samples can be accurately detected, significantly reducing the risk of adversarial samples deceiving the target classification model.
[0028] The present invention proposes a detection method based on "semantic contradiction", which uses an autoencoder model to generate a reconstructed image based on the category label, and determines whether the input sample is an adversarial sample by analyzing whether there is a significant semantic difference between the reconstructed image and the original image. The combination of the autoencoder model and the anomaly detection module enables the present invention to discover adversarial samples more efficiently and accurately, further improving the performance and adaptability of adversarial sample detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a flow chart of an adversarial sample anomaly detection method for traffic signs provided in Embodiment 1 of the present invention;
[0030] Figure 2 is a schematic diagram of the training phase of the autoencoder model provided in the first embodiment of the present invention;
[0031] Figure 3 It is a schematic diagram of the training phase and the final testing phase of the autoencoder model and anomaly detection model provided in the first embodiment of the present invention. DETAILED DESCRIPTION
[0032] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. In the absence of conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.
[0033] The terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of this disclosure / application, unless otherwise specified, "plurality" means two or more.
[0034] Embodiment 1:
[0035] Figure 1 This is a flow chart of the adversarial sample anomaly detection method for traffic signs in the first embodiment of the present invention. This flow chart only shows the logical sequence of the method described in this embodiment. In other possible embodiments of the present invention, different methods may be used without conflict. Figure 1 The steps shown or described are performed in the order shown. Figure 1 The method of this implementation specifically includes the following steps:
[0036] Inputting a normal sample into a trained target classification model, wherein the normal sample is a real traffic sign image that has not been tampered with, and the target classification model is used to identify the category of the traffic sign image and output a corresponding category label;
[0037] Filter correctly classified normal samples from the output of the target classification model and apply adversarial perturbations to them to generate new samples;
[0038] Normal samples and generated new samples are mixed to form a test set, and each sample in the test set is input into the autoencoder model trained with normal samples. The autoencoder model generates the corresponding reconstructed image based on the category label of the input sample, and calculates the difference map between the input sample and the reconstructed image.
[0039] The difference map is input into the anomaly detection model trained with normal samples, the energy value of the input sample is calculated, and the energy value is compared with the energy threshold of the normal sample. If the energy value is greater than the energy threshold, the sample is determined to be an adversarial sample. The energy threshold of the normal sample is obtained by the following steps: the normal sample is input into the trained anomaly detection model to obtain the energy value of the normal sample, and the energy threshold is determined according to the energy value.
[0040] The autoencoder model in this embodiment uses a model that combines an improved conditional variational autoencoder cVAE and a generative adversarial network GAN, referred to as cVAE-GAN. Figure 2 As shown in the figure, this model consists of three main parts: encoder, decoder, and discriminator. The encoder is used to extract the classification category features and common texture features of the input sample and compress them into a category vector and a low-dimensional vector respectively. The decoder uses the two vectors to reconstruct and restore the image. The discriminator is used to distinguish between the original image and the conditional reconstructed image. It is combined with the decoder for adversarial training to improve the conditional reconstruction performance of the model.
[0041] The steps for training the autoencoder model for conditional reconstruction performance are as follows:
[0042] (1) Image and the corresponding true label Input cVAE to get the reconstructed image generated by the decoder and the predicted labels generated by the encoder , calculate the input sample With the reconstructed image The reconstruction loss and KL divergence loss ; Calculate the true label of the input sample The predicted labels generated by the encoder The classification loss The calculation formulas for each loss are as follows:
[0043] ;
[0044] Where: represents the binary cross entropy loss; Represents the input sample With the reconstructed image The structural similarity of
[0045] ;
[0046] ;
[0047] ;
[0048] Where: and Represents the input samples With the reconstructed image The average value of , which measures the brightness information; and Represents the input samples With the reconstructed image The variance of , which measures the contrast information; For input sample With the reconstructed image covariance, which measures the input sample With the reconstructed image structural similarity; and A constant used to maintain stability and to avoid the denominator being 0; and is a small constant, ; Represents the dynamic range of pixel values.
[0049] ;
[0050] Where: j represents the dimension index of the low-dimensional representation z, and d represents the dimension size of the latent space; represents the mean of the low-dimensional representation z in the jth dimension, which is used to control the distribution center position of the low-dimensional representation; is the logarithmic variance of the low-dimensional representation z in the jth dimension, which is used to measure the uncertainty of the distribution; is the actual variance after logarithmic variance reduction, which is used to indicate the extended range of the potential distribution; and They are used to constrain the mean and variance of the low-dimensional representation z distribution to be close to the standard normal distribution; It means that the differences in all dimensions are summarized and normalized.
[0051] ;
[0052] (2) Next, randomly sample low-dimensional representations in the latent space and combine them with labels Input decoder to generate image samples corresponding to the conditions , and the reconstructed image samples generated by the decoder Together with the real input samples, they are input into the discriminator for authenticity discrimination. The discriminator is responsible for classifying the real input samples, the reconstructed image samples generated by the decoder, and the randomly generated conditional samples as true or false, and outputs the corresponding discrimination labels.
[0053] In the design of the loss function, four key points are defined: The true label of the input sample in the training data is used to supervise the discriminator's ability to distinguish real samples; is the discriminant label of the real training data, indicating its authenticity judgment result of the real input sample; The discriminant label of the reconstructed image generated by the decoder is used to evaluate whether the image generated by the decoder is close to the real input sample; Use the image ground truth labels for the random low-dimensional representations for the discriminator Generated conditional samples The discriminant label of is used to measure the performance of the decoder in the conditional generation task.
[0054] Generate loss The reconstruction ability and random generation ability of the decoder are comprehensively considered, and the samples generated by the decoder are optimized to be as close to the real samples as possible, thereby deceiving the discriminator; the identification loss It is used to improve the discriminative ability of the discriminator, so that it can effectively distinguish between real samples, reconstructed samples and randomly generated samples. With identification loss Through adversarial optimization, the decoder and discriminator continuously improve each other's performance. The quality of samples generated by the decoder is significantly improved, while the classification ability of the discriminator is more accurate, thus achieving the effect of adversarial training. and identification loss The calculation formula is as follows:
[0055] ;
[0056] ;
[0057] Where: and They are all 1 and all 0 labels respectively, indicating that the image used to calculate the loss is a real input traffic sign image or a reconstructed traffic sign image.
[0058] (3) Finally passed , , , The weighted combination of the GAN-combined cVAE network gives the overall loss function , as shown below:
[0059] ;
[0060] Where: , , is a hyperparameter used to balance the contribution of different losses;
[0061] Through the overall loss function and identification loss Update the parameters of the cVAE network and the discriminator network separately.
[0062] like Figure 2 As shown, a self-attention module is added to the encoder and decoder of the autoencoder model, and the self-attention module is used to optimize the image reconstruction process. The autoencoder model stabilizes the training process through spectral normalization technology to improve the quality of the reconstructed image. When an adversarial sample is input, the autoencoder model reconstructs based on the category label of the input sample. Since there is a contradiction between the adversarial sample label and the semantic information of the input sample, the difference between the reconstructed image and the input sample is significantly increased, and the difference is used to generate a difference map for anomaly detection.
[0063] The self-attention module allows the autoencoder model to focus on more important local or overall areas in the image, so as to better capture the details and contextual information of the image and improve the reconstruction quality. Spectral normalization technology is used to optimize the GAN training process to prevent the autoencoder model from overfitting or diverging, thereby further improving the image reconstruction effect. In the adversarial sample anomaly detection method, the autoencoder model reconstructs the input sample according to its category label. The label represents the category of the image, such as "speed limit 60". When the input is an adversarial sample, the real information of the input sample and the semantic information of the label will conflict because the adversarial sample has been tampered with (perturbation is added). Specifically, the autoencoder model tries to reconstruct the image based on the label "speed limit 60", but because the information in the adversarial sample has been tampered with, there is a large difference between the reconstructed image and the input sample. By calculating the difference map between the input sample and the reconstructed image, it can be determined whether the input sample is abnormal. If the difference is large, it means that the input sample is abnormal (i.e., adversarial sample); if the difference is small, it means that the input sample is normal.
[0064] like Figure 3 As shown, the anomaly detection model in this embodiment is a deep automatic encoding Gaussian mixture model DAGMM, and the DAGMM includes a compression network and an estimation network;
[0065] The compression network is used to reduce the dimension of the features of the input samples through a deep autoencoder to extract a low-dimensional representation, and the compression network uses a convolutional layer instead of a fully connected layer;
[0066] The estimation network is used to receive the low-dimensional representation, and estimate the probability distribution of the data under the framework of the Gaussian mixture model GMM, and predict the membership of each sample, where the membership is used to indicate the possibility that the sample belongs to a normal sample or an adversarial sample.
[0067] The specific steps of training the anomaly detection model through unsupervised learning methods are as follows: First, the difference map generated by the autoencoder model, that is, the difference between the input sample and the reconstructed image, is input into the DAGMM model; then, the compression network is used to reduce the dimensionality of the features of the difference map through the deep autoencoder to extract a simplified low-dimensional representation. Subsequently, the estimation network inputs the extracted low-dimensional features into the Gaussian mixture model GMM, and predicts the "membership" of each sample according to the GMM framework, that is, the possibility that the sample belongs to a normal sample. Based on the output of the estimation network, the energy value of each sample is calculated. The lower the energy value, the more likely the sample is to be a normal sample. Finally, the calculated sample energy value is compared with the preset energy threshold. If the energy value is higher than the threshold, the sample is determined to be an adversarial sample.
[0068] The following is an explanation of the steps for determining the sample energy value in the unsupervised anomaly detection module using a specific formula:
[0069] The low-dimensional representation provided by the compressed network It consists of two parts:
[0070] (1) Simplified low-dimensional representation learned by deep autoencoders ;
[0071] (2) Characteristics of reconstruction error and ;
[0072] The characteristics of the reconstruction error and They are calculated by Euclidean distance and cosine similarity respectively. , and The calculation formula is as follows:
[0073] ;
[0074] ;
[0075] ;
[0076] ;
[0077] Where: Indicates the quantity Input sample; h represents the encoding function; Represents the parameters of the encoder; represents the reconstruction sample; represents the decoding function; Represents the parameters of the decoder.
[0078] The low-dimensional representation obtained by the deep autoencoder and the characteristics of the reconstruction error and Splicing to obtain a comprehensive low-dimensional representation Z=[ Z enc , Z cos , Z euc ] . Low dimensional representation It is input to the estimation network, which maps it to the parameter space of the Gaussian mixture model GMM for modeling sample distribution. The estimation network predicts membership as follows:
[0079] ;
[0080] ;
[0081] Where: The representation estimation network is a multi-layer fully connected neural network used to obtain the low-dimensional representation from the comprehensive Extract features and map them to GMM components; yes The output of the parameterized multi-layer network represents the score of the sample on each GMM component; the softmax function is used to convert the unnormalized scores Convert to probability distribution ; represents the membership of the sample in each GMM component; K is the number of GMM components.
[0082] Combine the sample's membership in each GMM component , further calculate the parameters in GMM, as shown in the following formula:
[0083] ;
[0084] ;
[0085]
[0086] Where: N represents the number of samples; Represents the probability that the sample belongs to the i-th component; , , Respectively represent The probability, mean and covariance matrix of a distribution in GMM.
[0087] Then bring it into the Gaussian mixture distribution formula to calculate the sample energy, as shown in the following formula:
[0088] ;
[0089] The calculated sample energy is compared with the preset energy threshold, and samples with energy values higher than the threshold are judged as adversarial samples.
[0090] Next, multiple target classification models are trained using the traffic sign dataset. The target classification models are used to identify the categories of traffic sign images and output corresponding category labels. The classification accuracy of the classification models is verified through the test dataset to ensure that the classification accuracy reaches more than 90%. The category labels of the traffic sign dataset include "speed limit 100km / h", "no vehicles allowed", "straight ahead", etc. The target classification models are constructed based on ResNet18, Vgg19, DenseNet169 or MobileNet network structures. The four target classification models are trained using the same batch of training data, and the classification performance is verified through the test dataset to ensure that the classification accuracy of each model reaches more than 90%.
[0091] The normal samples correctly classified by the above target classification model are selected, and adversarial perturbations of different strengths are applied to the normal samples correctly classified in each model, and new samples are generated using multiple attack strategies. The strength of the adversarial perturbation is determined by Norm control, the disturbance intensity includes 0.05, 0.1 and 0.15; the attack strategies include FGSM attack, BIM attack, PGD attack, DeepFool attack and C&W attack.
[0092] The FGSM (Fast Gradient Sign Method) attack calculates the most effective perturbation through gradient information; the BIM (Basic Iterative Method) attack generates stronger adversarial samples based on the iterative optimization of FGSM; the PGD (Projected Gradient Descent) attack is an efficient and powerful attack method that generates the most effective adversarial samples through multiple iterations; the DeepFool attack fine-tunes the image until it crosses the classification boundary, resulting in classification errors; the C&W (Carlini & Wagner) attack is a more optimized adversarial attack method that aims to generate adversarial samples by minimizing the trade-off between perturbation and classification errors.
[0093] The strength of the adversarial perturbation refers to the degree of modification to the image. The larger the value, the more severely the image is modified. This embodiment uses different perturbation strengths (0.05, 0.1, 0.15), which represent the magnitude of the perturbation. Different perturbations generate adversarial samples of different strengths, and the difficulty and effect of the attack will also be different. At the same time, in the process of generating adversarial samples, this embodiment uses The norm limits the magnitude of the perturbation and emphasizes that the calculation method of the adversarial perturbation is unified when generating adversarial samples.
[0094] For the normal samples correctly classified in the above four target classification models, the aforementioned attack strategies and perturbation strengths are applied to generate new samples. In the process of generating new samples, the attack strategies and perturbation strengths can be arbitrarily combined according to different models, so as to generate diversified adversarial samples for each target classification model and comprehensively evaluate the robustness of each target classification model under different adversarial conditions.
[0095] The generated new samples are mixed with normal samples in a 1:1 ratio to form a test set, and the samples in the test set are input into the autoencoder model and the anomaly detection model in turn. First, the autoencoder model generates a reconstructed image and calculates the difference map between the input sample and the reconstructed image; then, the difference map is input into the anomaly detection model to calculate the energy value of the sample. Finally, the energy value is compared with the preset energy threshold. If the energy value exceeds the threshold, the sample is determined to be an adversarial sample.
[0096] Embodiment 2:
[0097] An embodiment of the present invention also provides an electronic terminal, characterized in that it includes a processor and a memory connected to the processor, a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the adversarial sample anomaly detection method for traffic signs described in the above-mentioned embodiment 1 are executed.
[0098] Embodiment three:
[0099] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the adversarial sample anomaly detection method for traffic signs described in the above-mentioned embodiment 1 are first implemented.
[0100] The computer-readable storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0101] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0102] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0103] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0105] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for detecting anomalies of adversarial samples for traffic signs, characterized in that: include: Inputting a normal sample into a trained target classification model, wherein the normal sample is a real traffic sign image that has not been tampered with, and the target classification model is used to identify the category of the traffic sign image and output a corresponding category label; Filter correctly classified normal samples from the output of the target classification model and apply adversarial perturbations to them to generate new samples; Normal samples and generated new samples are mixed to form a test set, and each sample in the test set is input into the autoencoder model trained with normal samples. The autoencoder model generates the corresponding reconstructed image based on the category label of the input sample, and calculates the difference map between the input sample and the reconstructed image. The difference map is input into the anomaly detection model trained with normal samples, the energy value of the input sample is calculated, and the energy value is compared with the energy threshold of the normal sample. If the energy value is greater than the energy threshold, the sample is judged to be an adversarial sample.
2. The adversarial sample anomaly detection method for traffic signs according to claim 1 is characterized in that: The energy threshold of the normal sample is obtained by the following steps: The normal sample is input into the trained anomaly detection model to obtain the energy value of the normal sample, and the energy threshold is determined according to the energy value.
3. The adversarial sample anomaly detection method for traffic signs according to claim 1, characterized in that: The autoencoder model and the anomaly detection model are trained by the following steps: The normal samples are used to train the autoencoder model to obtain a difference map between the input sample and the reconstructed image; the difference map is input into the anomaly detection model, and the anomaly detection model is trained by an unsupervised learning method.
4. The adversarial sample anomaly detection method for traffic signs according to claim 1, characterized in that: The autoencoder model is an improved conditional variational autoencoder generative adversarial network cVAE-GAN model, and the cVAE-GAN model adds a self-attention module at the output end of the encoder and decoder.
5. The adversarial sample anomaly detection method for traffic signs according to claim 1, characterized in that: The anomaly detection model is a deep auto-encoding Gaussian mixture model DAGMM, and the fully connected layer in the compression network of the DAGMM is replaced by a convolutional layer.
6. The adversarial sample anomaly detection method for traffic signs according to claim 1, characterized in that: Filter correctly classified normal samples from the output of the target classification model and apply adversarial perturbations to them to generate new samples, including: The adversarial perturbation applies adversarial perturbations of different strengths to normal samples correctly classified by the target classification model, and adopts multiple attack strategies to generate new samples.
7. The adversarial sample anomaly detection method for traffic signs according to claim 6, wherein the strength of the adversarial disturbance is Norm control; the attack strategies include FGSM attack, BIM attack, PGD attack, DeepFool attack and C&W attack.
8. The adversarial sample anomaly detection method for traffic signs according to claim 1, characterized in that: The target classification models are constructed based on ResNet18, Vgg19, DenseNet169 or MobileNet network structures respectively.
9. An electronic terminal, characterized in that: It includes a processor and a memory connected to the processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the adversarial sample anomaly detection method for traffic signs as described in any one of claims 1 to 8 are executed.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the adversarial sample anomaly detection method for traffic signs described in any one of claims 1 to 8 are implemented.