Image data attack defense method and system based on trusted distribution modeling

By constructing the distribution difference boundary of image data through an improved variational autoencoder, adversarial example labels are detected and corrected, solving the problems of high training overhead and poor generalization in existing technologies, and achieving efficient and accurate adversarial attack defense.

CN120912951APending Publication Date: 2025-11-07GUANGDONG POWER GRID CO LTD INFORMATION CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510981560.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing image adversarial attack defense methods suffer from problems such as high training overhead, poor generalization, and low efficiency, making it difficult to effectively detect and defend against various adversarial attacks.

Method used

An improved variational autoencoder (σ-VAE) is used to learn the distribution of real image data, construct the distribution difference boundary for each category, detect adversarial examples by calculating the distance between the input sample and the boundary, and correct the label using maximum likelihood estimation.

Benefits of technology

It improves the accuracy and robustness of adversarial attack defense, reduces data requirements and computational overhead, is applicable to a variety of adversarial attacks, and has good versatility and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912951A_ABST
    Figure CN120912951A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information, and relates to an image data attack defense method and system based on trusted distribution modeling. The method comprises the following steps: learning distribution of real image data by using an improved variational auto-encoder to obtain distribution representation of the image data; based on the distribution representation of the image data, constructing a distribution difference boundary of each category of images; calculating the distance between the input sample and the constructed distribution difference boundary, and detecting an adversarial sample according to the distance; and correcting the labels of the adversarial samples by using maximum likelihood estimation so as to defend adversarial attacks. According to the method, the abnormal input caused by the adversarial attack can be effectively detected, and the method has obvious advantages in the aspects of detection accuracy, data utilization efficiency, operation real-time performance and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of information technology, and relates to the technical field of defending against adversarial attacks through image data, and is also applicable to defending against adversarial attacks of image data in videos, and in particular to an image data attack defense method and system based on modeling of a trusted distribution. BACKGROUND

[0002] An adversarial attack is a technique designed to mislead machine learning models through subtle perturbations carefully designed in input data. In the field of image processing, such an attack is called an image adversarial attack, and is usually manifested as subtle modifications to images, which are often difficult to be detected by the human eye, but can cause misjudgment of image recognition systems. In the field of image adversarial attacks, an adversarial sample usually represents a picture to which perturbations have been added by an adversarial attack algorithm.

[0003] Among existing defense methods against image adversarial attacks, the following schemes are generally used:

[0004] 1) Adversarial training. The adversarial training method usually uses adversarial samples to train the deep learning model to be defended. Specifically, pictures with added perturbations are used as training samples to train the deep learning model during the training process of the model.

[0005] 2) Denoising defense. This method regards adversarial attacks as noise added to pictures, and uses adversarial samples to train a denoising model to predict the noise perturbations on the input samples, and then removes them to restore the original pictures.

[0006] 3) Defense based on generative models. This method uses a pre-trained image generative model to modify the input adversarial samples to generate samples closer to the original pictures.

[0007] The problems of the above existing technologies are as follows:

[0008] 1) The adversarial training method usually needs a large number of adversarial samples to participate in the training process of the neural network. This will result in greater training overhead, such as longer training time, more training data collection, etc. In addition, the explicit use of adversarial samples for training will damage the generalization of the defense, and may cause the model to overfit to a certain specific attack method.

[0009] 2) Using the denoising method for adversarial attack defense has the same problem, and also needs specific adversarial samples as a training set to train the denoising model. This will cause the model to overfit to certain specific noise, damaging the generalization of the defense.

[0010] 3) When using the purification defense method based on generative models, the generative models are generally slow to run, resulting in high complexity and low efficiency of the defense algorithm. SUMMARY

[0011] The present application aims at the above problems, and provides an image data attack defense method and system based on trusted distribution modeling, which solves the problems of detection blind area and insufficient robustness in existing image adversarial attack defense technology.

[0012] The technical scheme adopted by the present application is as follows:

[0013] An image data attack defense method based on trusted distribution modeling comprises the following steps:

[0014] An improved variational autoencoder is used to learn the distribution of real image data, and the distribution representation of the image data is obtained;

[0015] Based on the distribution representation of the image data, the distribution difference boundary of the image of each category is constructed;

[0016] The distance between the input sample and the constructed distribution difference boundary is calculated, and the adversarial sample is detected according to the distance;

[0017] The label of the adversarial sample is corrected by using maximum likelihood estimation, so as to defend against adversarial attacks.

[0018] Further, the improved variational autoencoder comprises an encoder and a decoder; the encoder is implemented by using a multi-layer convolutional network with residual connection; after a plurality of convolutional operations, the input picture is mapped into a multi-channel feature map; two fully connected layers are used to map the feature map into two low-dimensional vectors respectively, representing the mean and variance in the low-dimensional hidden space; the mean vector and the variance vector are combined together to form a multi-dimensional Gaussian distribution of the input picture in the hidden space; the decoder maps the latent vector sampled from the hidden space to a higher dimension, and gradually reconstructs an image with the same size as the input picture through a series of transpose convolution.

[0019] Further, based on the distribution representation of the image data, the distribution difference boundary of the image of each category is constructed, comprising:

[0020] Based on the distribution of the training sample in the hidden space, the overall Gaussian distribution of the image of each category is calculated by weighted average;

[0021] The KL divergence between each sample and the overall Gaussian distribution of the category to which it belongs is used to construct the distribution difference boundary of the image of each category in each dimension by statistically calculating the divergence values of all samples in each dimension of the hidden space and using the confidence interval.

[0022] Further, the maximum and minimum values of all KL divergences constitute a threshold interval, and the boundary of the threshold interval is the distribution difference boundary.

[0023] Further, the distance between the input sample and the constructed distribution difference boundary is calculated, and the adversarial sample is detected according to the distance, including:

[0024] For an input image, a prediction label of the input image is given by using a pre-trained classifier, an encoder is used to obtain a hidden space representation of the input image, and the distance between the hidden space representation and the overall distribution of the predicted category is calculated in each dimension;

[0025] The distance is compared with the constructed distribution difference boundary, if the distance exceeds the distribution difference boundary, the input image is determined as an adversarial sample, and an abnormal alarm is given, otherwise the input image is regarded as a normal sample.

[0026] Further, the distance is a KL divergence distance, a Wasserstein distance, or a JS divergence distance.

[0027] Further, the label of the adversarial sample is corrected by using maximum likelihood estimation, including: for the input image determined as the adversarial sample, the distance between the hidden space distribution of the input image and the overall distribution of all other categories is calculated, and the category with the nearest distance is selected as the potential true label of the input image.

[0028] An image data attack defense system based on a trusted distribution modeling includes:

[0029] A data distribution learning module is configured to learn the distribution of real image data by using an improved variational autoencoder to obtain a distribution representation of the image data.

[0030] A distribution difference boundary construction module is configured to construct a distribution difference boundary of images of each category based on the distribution representation of the image data.

[0031] An adversarial sample detection module is configured to calculate the distance between an input sample and the constructed distribution difference boundary, and detect an adversarial sample according to the distance.

[0032] A label correction module is configured to correct the label of the adversarial sample by using maximum likelihood estimation, so as to defend against adversarial attacks.

[0033] The beneficial effects of the present application are as follows:

[0034] 1) The present application adopts an improved sigma-VAE, which extends the fixed standard Gaussian assumption in the traditional VAE to a learnable variance parameter, so that the encoder can more accurately capture the statistical characteristics and uncertainty of the real data when mapping the input image to the hidden space. This method breaks the problem of information loss and limited generation ability in the traditional VAE, so that the hidden space can retain more discriminative features.

[0035] 2) The present application utilizes the distribution information of each dimension in the latent space of the training sample, and calculates the overall distribution of each category by weighted average. This method makes the statistical characteristics of each category in the latent space truly reflect the actual data distribution, and improves the accuracy of distribution fitting.

[0036] 3) The present application calculates the Kullback-Leibler divergence between each sample and the overall distribution of the category to which it belongs, and constructs a distribution difference boundary by using the confidence interval method. The boundary abstract method can accurately capture the normal range of the latent space manifold, effectively distinguish abnormal samples (adversarial samples), and stably set the detection threshold in different dimensions, so that the accuracy and robustness of the adversarial sample detection are superior to those of the traditional method, especially when facing multiple adversarial attacks (such as FGSM, PGD, FAB, etc.).

[0037] 4) In the reasoning stage of the present application, the preliminary label is obtained by the pre-trained classifier, and then the latent space representation is obtained by using the encoder, the KL divergence of each dimension is calculated, and if it exceeds the distribution boundary, it is determined as an adversarial sample, and the maximum likelihood estimation is used to correct the potential true label. This mechanism can realize the detection and classification label correction of abnormal input through latent space analysis without retraining the defense model, so as to restore the original classification result damaged by adversarial attack;

[0038] 5) The overall framework of the present application is independent of the internal structure of the defended model and the specific adversarial attack method, and only relies on the deep fitting of the real data distribution and the comparison of the latent space features. The independence of the framework greatly improves the universality and expansibility of the system, and can be widely used in various scenes. The present application is significantly superior to the existing black box defense method in accuracy, precision and real-time running time, and provides a significant competitive advantage for actual deployment.

[0039] 6) Through the above key technical measures, the present application can not only effectively detect abnormal input caused by adversarial attack, but also has obvious advantages in detection accuracy, data utilization efficiency and running real-time. These technical effects prove that the present application can solve the problems of inaccurate adversarial sample detection, large data demand and high computational overhead in the prior art, thereby realizing the comprehensive protection of the image classification system. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 is the overall implementation framework diagram of the method of the present application.

[0041] Figure 2 is the structure diagram of the variational autoencoder of the present application.

[0042] Figure 3 is the training process diagram of the variational autoencoder of the present application.

[0043] Figure 4 is a comparison chart of the improved variational autoencoder and the latent space feature expression of the traditional variational autoencoder.

[0044] Figure 5 is a comparison chart of the defense performance of the method of the present application and the existing method. DETAILED DESCRIPTION

[0045] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below through specific embodiments and drawings.

[0046] The present application learns the distribution of real image data by using a variational autoencoder (VAE), and after obtaining the distribution representation of the image data, the distribution difference boundary of the image data of each category is modeled. The input sample is compared with the modeled distribution difference boundary, and if it exceeds the boundary, it is considered to be an adversarial sample. In addition, maximum likelihood estimation is used to find its possible original label. The present application uses a method based on modeling of real data distribution to defend against image adversarial attacks, and does not target specific attack methods, thus improving the generalization of adversarial attack defense and effectively migrating to different adversarial attack algorithms. In addition, through the boundary comparison method, the running speed of the algorithm is improved.

[0047] The key points and inventive points of the technical solutions of the present application are given below, mainly aiming at the image adversarial input anomaly monitoring technology based on robust latent space representation. This scheme adopts an improved variational autoencoder architecture and a latent space distribution boundary construction method, which has the following innovations compared with the prior art:

[0048] 1) Innovative latent space modeling method. An improved sigma-VAE (variational autoencoder) structure is adopted, which changes the fixed standard Gaussian assumption in the traditional VAE to a learnable variance, so that the encoder can more accurately capture the statistical characteristics and uncertainty of the real data when mapping the input image to the latent space, improving the data distribution modeling and generation ability.

[0049] 2) Category-level overall distribution construction. Based on the distribution of the training samples in the latent space, the overall Gaussian distribution of the images of each category is calculated by weighted average (with the inverse of the variance as the weight), realizing accurate representation of the latent space features of different categories of images, and providing a reliable reference standard for subsequent anomaly detection.

[0050] 3) Construction of distribution difference boundary. By using the Kullback-Leibler (KL) divergence between each sample and the overall Gaussian distribution of the category to which it belongs, the distribution difference boundary of each category in each dimension is constructed by statistically analyzing the divergence values of all samples in each dimension of the latent space and using a confidence interval (such as a 95% confidence interval), thereby abstractly describing the manifold characteristics of the latent space.

[0051] 4) Adversarial sample detection and label correction mechanism. In the inference stage, the distance between the input sample in the latent space and the pre-constructed category distribution difference boundary is calculated to detect whether the sample has an abnormal deviation; at the same time, for the adversarial sample detected as abnormal, the latent true label is corrected by using the maximum likelihood estimation method, thereby achieving effective defense against adversarial attacks.

[0052] 5) Black-box defense framework. This technical solution does not need to rely on the internal structure of the defended model or specific adversarial attack methods, but only relies on the deep fitting of the real data distribution and the comparison of the latent space characteristics, thereby having good universality and robustness, being suitable for various data sets and attack types, and solving the limitations of existing white-box defense methods in actual deployment due to the unavailability of model information.

[0053] 6) Low data requirement and efficient implementation. Experimental results show that the present invention only needs to use about 20% of the training data to construct an effective overall distribution and manifold boundary, and has low computational overhead in the training and inference processes, thereby meeting the real-time operation requirements and providing a high efficiency advantage for practical applications.

[0054] The core of the present invention is to deeply fit the "clean data distribution", and then judge whether the input significantly deviates from the distribution in the inference stage, thereby detecting adversarial samples. It does not need to access the internal parameters of the defended model (i.e., the classifier), and is decoupled from specific attack methods, which is a black-box defense method. The overall implementation framework is as shown in Figure 1 The implementation can be roughly divided into the following four stages:

[0055] 1、Preparation of data and training environment.

[0056] 2、Construction and training of improved variational autoencoder model.

[0057] 3、Distribution difference boundary modeling in latent space.

[0058] 4、Abnormality monitoring and label correction in the inference stage.

[0059] The work, points for attention and related process ideas needed in each stage are introduced in turn as follows.

[0060] 1、Preparation of environment and data

[0061] In an embodiment of the present application, the defense method is written in Python language. And the deep learning framework PyTorch is used to realize the running of the algorithm on the GPU. The implementation in this embodiment is all based on Python 3.8.19, PyTorch 2.3.0, CUDA 12.1 driver version and other necessary related third-party libraries. Taking the CIFAR-10 dataset as an example, for other datasets, only the corresponding modification needs to be made in the data reading part (because the full connection layer is involved in the specific network structure implementation, the vector size of the input network needs to be fixed). In data preparation, all data is encapsulated into the DataLoader class using the data reading interface provided by the PyTorch framework, and each picture is scaled to the appropriate resolution and converted into a floating point tensor (Tensor). Generally speaking, in order to prevent the neural network from appearing unstable gradient during training, all input data is normalized to map the pixel value from [0, 255] to [0, 1], and when necessary, the mean and variance standardization can also be done for the network pre-training parameters to speed up the network training convergence.

[0062] 2. Building and training VAE model

[0063] The present application uses the PyTorch framework to build the network model. The VAE model is composed of an encoder and a decoder. The encoder is responsible for encoding the input picture into the latent space representation, and the decoder is responsible for decoding the latent vector sampled from the latent space into a generated picture. The present application mainly focuses on the encoder of the model. In the specific implementation, a multi-layer convolutional network with residual connection (i.e. residual convolution block in Figure 2 ) is used to realize the encoder part. After multiple convolution operations, the input picture is mapped into a multi-channel feature map. Two fully connected layers are used to map the feature map into two low-dimensional vectors, representing the mean and variance in the low-dimensional latent space. The mean vector and the variance vector are combined together to form a multi-dimensional Gaussian distribution of the input picture in the latent space. In this way, an encoder that encodes the picture into a multi-dimensional Gaussian distribution is obtained. The decoder is responsible for mapping the latent vector sampled from the latent space to a higher dimension, and gradually reconstructing an image with the same size as the input picture through a series of transpose convolution (i.e. residual deconvolution block in Figure 2 ). The input layer uses the Sigmoid function to ensure the rationality of the generated pixel value. The structure of the variational autoencoder is shown in Figure 2 . The training process of the variational autoencoder is shown in Figure 3 .

[0064] The implementation details of the encoder and decoder of the present method are given in Table 1 and Table 2, respectively. The hyperparameter selection is shown in Table 3. The Flatten operation in the encoder concatenates the 128 x 8 x 8 = 8192 features into a one-dimensional vector; the final output has two branches: μ and logo2, each with 32 dimensions. The decoder first maps the latent vector (dimension 32) to 8192 dimensions, and then reshapes it to 128 x 8 x 8; then multiple deconvolutions are performed to gradually restore the spatial resolution to 3 x 32 x 32.

[0065] In Table 1, Input denotes the input layer, Conv2D denotes the two-dimensional convolution operation, ReLU denotes the linear rectifier function, Flatten denotes the operation of flattening the two-dimensional vector to one-dimensional, Linear denotes the linear fully connected layer, feature dim denotes the feature dimension, and out_feat denotes the output feature dimension.

[0066] In Table 2, Input z denotes the input layer, UnFlatten denotes the operation of stacking one-dimensional vectors to two-dimensional, ConvTrans2D denotes the two-dimensional deconvolution operation.

[0067] In Table 3, Learning rate denotes the input layer, Optimizer denotes the optimizer, Batch size denotes the batch size, Epochs denotes the training rounds, Latent size denotes the latent space size, Hidden channels denotes the number of hidden layer channels, and Weight decay denotes the weight decay parameter.

[0068] Table 1. Encoder structure

[0069] Sequence number Layer class Input Convolutional kernel size / parameters Stride Padding Output 1 Input 3×32×32 - - - 3×32×32 2 Conv2D+ReLU 3×32×32 3×3 1 1 32×32×32 3 Conv2D+ReLU 32×32×32 4×4 2 1 64×16×16 4 Conv2D+ReLU 64×16×16 5×5 2 2 128×8×8 5 Flatten 128×8×8 - - - 1×8192 6 Linear→μ 8192 (feature dim) out feat = 32 - - 1×32 7 Linear→log σ 2 ]] 8192 (feature dim) out feat = 32 - - 1×32

[0070] Table 2. Decoder structure

[0071] Sequence number Layer class Input Convolutional kernel size / parameters Stride Padding Output 1 Input z 1×32 - - - 1×32 2 Linear 1×32 Out feat = 8192 - - 1×8192 3 UnFlatten (C = 128) 1×8192 - - - 128×8×8 4 ConvTrans2D+ReLU 128×8×8 6×6 2 2 64×16×16 5 ConvTrans2D+ReLU 64×16×16 6×6 2 2 32×32×32 6 ConvTrans2D+Sigmoid 32×32×32 5×5 1 2 3×32×32

[0072] Table 3. Hyperparameter selection

[0073] Hyperparameters Values Learning rate 1e-3 Optimizer AdamW Batch size 64 Epochs 100 Latent size 32 Hidden channels 32 Weight decay 1e-5

[0074] 3. Distribution modeling in latent space

[0075] In this step, we want to extract the statistical regularity of all normal clean samples in the latent space. First, we traverse the input data and aggregate them according to the categories of the input images, and calculate the overall distribution of each different category of images. Specifically, we calculate the mean and variance of the overall Gaussian distribution in each dimension of the latent space. In this way, we obtain the overall distribution that can describe different categories of images. Further, we calculate the KL divergence between each sample in the training set and its corresponding overall distribution in each dimension of the latent space. We take the maximum and minimum values of all KL divergences as the threshold interval, so we get the difference distance (the size of the threshold interval) of each sample to its corresponding overall distribution in each dimension, and thus construct the distribution difference boundary (i.e. the boundary of the threshold interval). This boundary represents the boundary abstraction constructed for all real data distributions.

[0076] 4. Abnormal monitoring and label correction in inference stage

[0077] After obtaining the distribution statistics described in the foregoing (i.e. the distribution difference boundary), for a new input image, only a forward process of the encoder and a few simple distance calculations are needed to determine whether it belongs to an adversarial sample and make the corresponding latent label correction. When a new input image is obtained, a classifier will give its predicted label. After obtaining the latent space representation of the input image using the encoder obtained in the foregoing, the KL divergence distance between the representation and the overall distribution of the predicted category is calculated in each dimension. Compare this KL divergence distance with the constructed distribution difference boundary. If the distribution difference exceeds the boundary, the input image is identified as an adversarial sample and an abnormal alarm is given, otherwise it is considered as a normal sample. When it is determined to be an adversarial sample, the latent true label needs to be further determined. Calculate the distance between the input image's latent space distribution and the overall distribution of all other categories (categories other than the input image's category), and select the nearest distance category as its latent true label.

[0078] The following gives a detailed description of the beneficial effects of the technical solutions of the present application. Combined with the key points and points of the invention in the technical solutions, the advantages and benefits brought by the present application are explained from both theoretical and experimental data.

[0079] Key point 1: innovative latent space modeling method

[0080] Technical means: Use improved σ-VAE to extend the fixed standard Gaussian assumption in traditional VAE to learnable variance parameters, so that the encoder can more accurately capture the statistical characteristics and uncertainty of real data when mapping input images to the latent space.

[0081] Beneficial effects: In theory, this method breaks the problem of information loss and limited generation ability in traditional VAE, enabling the latent space to retain more discriminative features; for example Figure 4 It is shown that after using σ-VAE, different categories show more obvious distinguishability in the latent space, thereby providing a more accurate basis for subsequent anomaly detection, and the defense success rate is significantly improved. In Figure 4 , different colors represent the overall distribution belonging to different categories, the horizontal coordinate represents different latent space dimensions, and the vertical coordinate represents the mean of the overall distribution. It can be seen that Figure 4 (a) of FIG. 8 shows that in the traditional standard VAE, the means of the overall distribution on most of the latent space dimensions are very close to the 0 value representing the standard Gaussian distribution, which will cause the model to lose discriminative information for distinguishing different categories of images in the dimensions whose means are approximately equal to 0, thereby causing the model to have poor data distribution modeling ability. As shown in Figure 4 (b), in the improved σ-VAE, the latent space has richer feature expression, and each latent space dimension has good distinguishability for different categories.

[0082] Key point 2: Category-level overall distribution construction

[0083] Technical means: Use the distribution information of the training samples in each dimension of the latent space to calculate the overall distribution of each category by weighted average.

[0084] Beneficial effects: In theory, this method enables the statistical characteristics of the latent space of each category to truly reflect the actual data distribution, improving the accuracy of distribution fitting.

[0085] Through experiments, it is discussed how this method based on data distribution fitting modeling is affected by the amount of data. Specifically, on the MNIST, Fashion-MNIST and CIFAR-10 data sets, the State-of-the-art (SOTA) method corresponding to the data set is selected for testing, and for the method of the present application, the random sampling seed is fixed to ensure that the data crossing in the subset experiment will not affect the results. According to a series of fixed ratios, the data set is sampled to perform complete process experiments. The experimental results are shown in Figure 5 . In Figure 5In the figure, different colors represent different data sets, the horizontal axis represents the proportion of the data subset in the total data set size, the vertical axis represents the defense performance, the solid line represents the performance of the method of the present application under different subset sizes, and the dashed line represents the performance of the SOTA method after complete training on the same data set. The experimental results give the defense performance of using different proportions of data subsets to construct the overall distribution and manifold boundary. Obviously, with the powerful representation ability of VAE for data distribution, the method only uses about 20% of the training data to construct the overall distribution and manifold boundary, and achieves similar performance to the SOTA method. In addition, it is also observed that σ-VAE requires much less training time to achieve satisfactory performance compared to standard VAE. For the two VAEs with the same encoder and decoder structure (only different in the loss function), the training time per epoch is similar. However, in the experimental results, the method based on the standard VAE requires 160 epochs to achieve optimal performance, while the σ-VAE, due to its advantages, only needs 10-15 epochs to achieve good image reconstruction and distribution modeling ability. The experimental results show that after constructing the overall distribution, the system can achieve similar defense performance to other SOTA methods using less data (about 20%) on multiple data sets, providing the advantages of low data demand and high training efficiency for practical applications.

[0086] Key point 3: Distribution difference boundary construction

[0087] Technical means: Calculate the Kullback-Leibler divergence between each sample and the overall distribution of the category to which it belongs, and use the confidence interval method to construct the distribution difference boundary.

[0088] Beneficial effects: In theory, this boundary abstraction method can accurately capture the normal range of the hidden space manifold and effectively distinguish abnormal samples (adversarial samples); in comparative experiments, this method can stably set the detection threshold in different dimensions, thereby making the accuracy and robustness of adversarial sample detection superior to traditional methods, especially when facing multiple adversarial attacks (such as FGSM, PGD, FAB, etc.).

[0089] Key point 4: Adversarial sample detection and label correction mechanism

[0090] Technical means: In the inference stage, first obtain the preliminary label through the pre-trained classifier, then use the encoder to obtain the hidden space representation, calculate the KL divergence of each dimension, and if it exceeds the distribution boundary, it is determined as an adversarial sample, and the maximum likelihood estimation is used to correct the potential true label.

[0091] Beneficial effects: in theory, this mechanism can realize the detection and classification label correction of abnormal input through hidden space analysis without retraining the defense model, thereby restoring the original classification result destroyed by the adversarial attack;

[0092] Experimental data show that the method exhibits high detection accuracy and correction success rate under multiple data sets and multiple attack conditions, providing a practical and efficient adversarial defense means for the system. The evaluation index definition of the experimental results is shown in Table 4.

[0093] Table 4. Evaluation index of experimental results

[0094]

[0095] It should be noted that since the present application focuses on the defense of adversarial attacks, the definition of the present application is different from that in the traditional image classification task: TP represents the number of adversarial samples successfully recognized by the method of the present application in all successfully attacked images, FN represents the number of successfully attacked images that are not correctly classified as adversarial samples by the method of the present application. FP represents the number of clean original images classified as adversarial samples by the method, and TN represents the number of clean original images classified as normal input by the method. The experimental results are shown in Tables 5 and 6.

[0096] Table 5. Defense performance of the method on CIFAR-10 dataset under infinity norm constraint

[0097]

[0098] Table 6. Defense performance of the method on CIFAR-10 dataset under 2 norm constraint

[0099]

[0100] Key point 5: black box defense framework

[0101] Technical means: the overall framework of the present application is independent of the internal structure of the defended model and the specific adversarial attack method, and only relies on the deep fitting of the real data distribution and the comparison of the hidden space features.

[0102] Beneficial effects: in theory, the independence of the framework greatly improves the universality and expandability of the system, and can be widely used in various scenarios. As shown in Table 7, the present application is significantly better than existing black box defense methods in accuracy, precision and real-time running time, providing a significant competitive advantage for actual deployment.

[0103] Table 7. Comparison of defense performance of different defense methods against FGSM attack with attack intensity of 8 / 255

[0104]

[0105] Through the above key technical measures, the application can not only effectively detect abnormal input caused by adversarial attacks, but also has obvious advantages in detection accuracy, data utilization efficiency and running real-time performance. These technical effects prove that the application can solve the problems of inaccurate detection of adversarial samples, large data demand and high computational overhead in the prior art, thereby realizing comprehensive protection of the image classification system.

[0106] In addition, through the combination of theoretical analysis and experimental data, the beneficial effects of the application are fully verified, and have high creativity and practical application value.

[0107] The following gives a detailed description of the technical solutions of the application in alternative aspects. The following lists several feasible alternatives:

[0108] 1. For latent space modeling and sigma-VAE architecture

[0109] Alternative: Use a diffusion model-based or flow-based generative model as a means of latent space modeling, implement fitting of the data underlying distribution through a multi-step inverse generation process, and construct the statistical boundary of the latent space. In the alternative, as long as the data features can be fully captured and the boundary constructed, similar effects to sigma-VAE can be achieved.

[0110] 2. For class overall distribution and distribution boundary construction

[0111] Alternative 1: When calculating the overall distribution of each class, other weighting strategies can be used. For example, use sample importance scores (such as based on local density or anomaly score) as weights instead of inverse variance, to dynamically adjust the contribution of each sample to the overall distribution, so that the constructed distribution boundary is more robust.

[0112] Alternative 2: Clustering algorithms (such as Gaussian Mixture Model GMM or K-Means algorithm) can be used to group the data in the latent space, and then based on the clustering centers and the dispersion within the cluster, the distribution difference boundary is constructed to achieve the purpose of describing the statistical characteristics of each class. Taking the K-Means algorithm as an example, the process of calculating the clustering center is as follows: randomly initialize multiple (number of labeled classes) clustering centers, then calculate the distance of all sample feature vectors to different clustering centers, assign them to the nearest center, then loop the process until convergence, which realizes the grouping of the data in the latent space.

[0113] 3. For adversarial sample detection and label correction process

[0114] Alternative solution 1: In the detection stage, in addition to using KL divergence, other statistical distances (such as Wasserstein distance, JS divergence) can also be used to compare the hidden space distribution of the input sample with the category population distribution. As long as the distance metric can distinguish normal samples from abnormal samples, it can replace the original judgment method.

[0115] Alternative solution 2: The label correction part can be implemented based on Bayesian inference or clustering method, that is, for the sample detected as abnormal, the similarity of each category sample group is compared, and the category with the highest similarity is selected as the corrected label, instead of simply based on maximum likelihood estimation.

[0116] Another embodiment of the present application provides an image data attack defense system based on trusted distribution modeling, which comprises:

[0117] A data distribution learning module is configured to learn the distribution of real image data using an improved variational autoencoder to obtain a distribution representation of the image data.

[0118] A distribution difference boundary construction module is configured to construct a distribution difference boundary of images of each category based on the distribution representation of the image data.

[0119] An adversarial sample detection module is configured to calculate the distance between the input sample and the constructed distribution difference boundary, and detect the adversarial sample according to the distance.

[0120] A label correction module is configured to correct the label of the adversarial sample using maximum likelihood estimation, thereby defending against adversarial attacks.

[0121] The division of the above modules is only illustrative, and in actual application, the above functions can be completed by different functional modules according to needs, to complete all or part of the functions described in the foregoing method. The specific working process of each module can refer to the corresponding process in the foregoing method embodiment, which will not be repeated here.

[0122] Another embodiment of the present application provides a computer device (computer, server, smart phone, etc.), which comprises a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the steps of the method of the present application.

[0123] Another embodiment of the present application provides a computer readable storage medium (such as ROM / RAM, magnetic disk, optical disk), which stores a computer program, and the computer program is executed by a computer to implement the steps of the method of the present application.

[0124] Another embodiment of the present application provides a computer program product comprising a computer program which, when executed by a computer, implements the steps of the method of the present application.

[0125] The specific embodiments of the application previously disclosed are to be considered in all respects as being illustrative and not restrictive, and the scope of the application is indicated by the appended claims rather than by the foregoing description, and all changes which come within the meaning and range of equivalency of the claims are intended to be embraced therein.

Claims

1. An image data attack defense method based on trusted distribution modeling, characterized in that, The method comprises the following steps: learning the distribution of real image data using an improved variational autoencoder to obtain a distribution representation of the image data; constructing a distribution difference boundary of images of each category based on the distribution representation of the image data; calculating the distance between the input sample and the constructed distribution difference boundary, and detecting the adversarial sample according to the distance; correcting the label of the adversarial sample using maximum likelihood estimation, thereby defending against adversarial attacks.

2. The method of claim 1, wherein, The improved variational autoencoder comprises an encoder and a decoder; the encoder is implemented by using a multi-layer convolutional network with residual links; after a plurality of convolutional operations, the input picture is mapped into a multi-channel feature map; two fully connected layers are used to map the feature map into two low-dimensional vectors respectively, representing the mean and variance in the low-dimensional hidden space; the mean vector and the variance vector are combined to form a multi-dimensional Gaussian distribution of the input picture in the hidden space; the decoder maps the latent vector sampled from the hidden space to a higher dimension, and gradually reconstructs an image with the same size as the input picture through a series of transpose convolutions.

3. The method of claim 1, wherein, The distribution difference boundary of images of each category is constructed based on the distribution representation of the image data, comprising: based on the distribution of the training sample in the hidden space, the overall Gaussian distribution of the images of each category is calculated by weighted average; the KL divergence between each sample and the overall Gaussian distribution of the category to which the sample belongs is used to construct the distribution difference boundary of the images of each category in each dimension by statistically calculating the divergence values of all samples in each dimension of the hidden space and using the confidence interval.

4. The method of claim 3, wherein, The maximum and minimum values of all KL divergences form a threshold interval, and the boundary of the threshold interval is the distribution difference boundary.

5. The method of claim 1, wherein, The distance between the input sample and the constructed distribution difference boundary is calculated, and the adversarial sample is detected according to the distance, comprising: for an input image, a pre-trained classifier is used to give its predicted label, an encoder is used to obtain the hidden space representation of the input image, and the distance between the hidden space representation and the overall distribution of the predicted category is calculated in each dimension; compare the distance with the constructed distribution difference boundary, if it exceeds the distribution difference boundary, it is determined that the input image is an adversarial sample, and an abnormal alarm is given, otherwise it is regarded as a normal sample.

6. The method of claim 5, wherein, The distance is the KL divergence distance, the Wasserstein distance, or the JS divergence distance.

7. The method of claim 1, wherein, The label of the adversarial sample is corrected using maximum likelihood estimation, comprising: for the input image determined as the adversarial sample, the distances between its hidden space distribution and the overall distributions of all other categories are calculated, and the category with the nearest distance is selected as its potential true label.

8. An image data attack defense system based on trusted distribution modeling, characterized by, It comprises: a data distribution learning module for learning the distribution of real image data using an improved variational autoencoder to obtain a distribution representation of the image data; a distribution difference boundary construction module for constructing a distribution difference boundary of images of each category based on the distribution representation of the image data; an adversarial sample detection module for calculating the distance between the input sample and the constructed distribution difference boundary, and detecting the adversarial sample according to the distance; a label correction module for correcting the label of the adversarial sample using maximum likelihood estimation, thereby defending against adversarial attacks.

9. A computer device, comprising: comprising a memory storing a computer program configured to be executed by a processor, the computer program comprising instructions for performing the method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, which, when executed by a computer, implements the method of any one of claims 1-7.