Artificial intelligence recognition model based on multi-subspace joint anti-defense method

By employing a multi-subspace joint discrimination method, the full-spectrum defense challenge against adversarial attacks on deep learning networks is solved, achieving efficient identification of adversarial examples and reducing the false positive rate.

CN119399492BActive Publication Date: 2025-11-04SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411392648.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2025-11-04
Estimated Expiration
2044-10-08

AI Technical Summary

Technical Problem

Existing defense methods for adversarial attacks on deep learning network systems are insufficient to cope with diverse adversarial attacks. A single defense method cannot achieve full-spectrum defense, and the timing, method, and parameters of adversarial attacks are unpredictable.

Method used

An AI recognition model based on multi-subspace joint is adopted to perform feature embedding subspace, consistency discrimination, scale space and spectral space consistency discrimination on input samples. Information complementarity is achieved through multi-incoherent space analysis to improve defense capabilities.

Benefits of technology

It can effectively detect and identify various forms of adversarial attack samples, significantly reduce the model's false positive rate, and achieve a full spectrum of defense effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399492B_ABST
    Figure CN119399492B_ABST
Patent Text Reader

Abstract

A kind of artificial intelligence identification model based on multi-subspace joint method of countermeasure defense, the activation value vector of input sample in each node, each layer of model is mapped to different subspace consistency, i.e. The consistency discrimination of feature embedding subspace is whether it is an attack sample;For the non-attack sample after screening, carry out multi-scale change, and compare the distribution consistency performance of each node value of each scale in model transmission, i.e. The consistency discrimination of scale space is whether it is an attack sample;For the original pixel or neural network output value of non-attack sample after further screening, carry out subspace spectral decomposition, and compare the response coefficient statistics of each spectral component, i.e. Spectral space consistency discrimination is whether it is an attack sample, to realize full spectrum defense.The present application is based on multi-incoherent space analysis, so that information is complementary, greatly improves the defense ability for attack sample, realizes full spectrum defense.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information security, in particular to an artificial intelligence recognition model countermeasure defense method based on multi-subspace joint, which can be applied to intelligent image recognition products, including image recognition, face recognition, automatic driving and financial security and other security related tasks of face security devices such as face gate machines, security video monitoring, auxiliary driving cameras, etc. BACKGROUND

[0002] The current mainstream deep learning network system is extremely fragile. Attackers can make subtle changes to the source image data that humans cannot identify through their senses, such as changing the values of only a few pixels, but the machine learning model can accept and make incorrect classification decisions. That is, output the wrong discrimination result. Adversarial samples make most of the world's security monitoring systems based on deep learning extremely unreliable. Therefore, it is very important to enhance the countermeasure defense of the artificial intelligence image recognition model, that is, to identify the adversarial samples so that the model is not affected by the adversarial attack.

[0003] Due to the huge configurable space of adversarial samples, the unpredictability of attack timing, attack method and attack parameters of adversarial attacks makes it very difficult for deep networks to distinguish and defend against attack samples. The existing single countermeasure defense method is difficult to cope with various adversarial attacks. Most of the current research on adversarial sample defense is basically only for a few typical adversarial attack algorithms, and cannot achieve full spectrum defense. SUMMARY

[0004] The present application proposes a kind of artificial intelligence recognition model countermeasure defense method based on multi-subspace joint to solve the problems that the prior art must need to train additional neural network, only be limited to feature value domain defense against attack or pixel space against attack sample defense, based on multi-incoherent space analysis, so that information is complementary, greatly improve the defense capability for attack sample, realize full spectrum defense.

[0005] The present application is realized by the following technical solutions:

[0006] The present application relates to a kind of artificial intelligence recognition model countermeasure defense method based on multi-subspace joint, comprising:

[0007] Step 1) consistency discrimination of feature embedding subspace: the activation value vector of input sample in each node and each layer of the model is mapped to different subspace consistency performance, to judge whether it is an adversarial sample;

[0008] Step 2) consistency discrimination of scale space: the non-attack sample after step 1 discrimination is subjected to multi-scale change, and the distribution consistency performance of each node value of each scale in model transmission is compared to judge whether it is an adversarial sample;

[0009] Step 3) Spectrum space consistency discrimination: the original pixel or neural network output value of the non-attack sample screened in step 2 is subjected to subspace spectrum decomposition, and the response coefficient statistics thereof in each spectrum component should be consistent, so as to judge whether it is an adversarial sample, and realize full spectrum defense.

[0010] Technical effects

[0011] The method can cope with single / multiple pixel-level perturbations under norm constraint, global perturbations under one norm / two norm / infinite norm constraint, and various adversarial perturbations under other forms of constraint through the discrimination of feature space, scale space and spectrum space consistency, and can effectively detect adversarial attack samples against the model itself. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 It is a flowchart of the present application;

[0013] Figure 2 It is a schematic diagram of the consistency discrimination of the feature embedding subspace;

[0014] Figure 3 It is a schematic diagram of the consistency discrimination of the scale space;

[0015] Figure 4 It is a schematic diagram of the consistency discrimination of the spectrum space. DETAILED DESCRIPTION

[0016] As shown in the figure, it is an artificial intelligence recognition model adversarial defense method based on multi-subspace joint involved in the embodiment, which comprises: Figure 1

[0017] Step 1, consistency discrimination of feature embedding subspace, as shown in the figure, specifically comprising: Figure 2

[0018] 1.1: extracting the embedded features of the input sample image, specifically: f θ (x)=z, wherein: f θ is a feature extraction function, θ is a parameter, and z is an embedded feature;

[0019] 1.2: feature decomposition is performed on the embedded feature z to obtain: z=β1a1+β2a2+…+β n a n

[0020] , wherein: (a1, a2, …, an) is a feature base, and (β1, β2, …, βn) is a coefficient corresponding to the feature base.

[0021] ​​1.3: Consistent discrimination of each extracted subspace feature, specifically: the feature coefficients are obtained by performing feature decomposition on the embedding representation of the embedded feature z obtained in step 1.1, and it is judged that when βi obeys the normal distribution N(μ i ,σ i ), by the sample training set χ, the μ i ,σ i of each group is estimated, and the distribution space of the feature coefficients of the real sample set in is obtained Where: (μ i ,σ i ) is the mean and variance of the distribution.

[0022] 1.4: For any input sample x, the decision criterion C1 is calculated Where: η i is the feature dimension confidence level; when the decision criterion C1 is true, it means that the feature dimension reflects greater genetic sequence inconsistency, that is, the sample is identified as an attack sample.

[0023] Step 2, consistent discrimination of scale space, as shown in Figure 3 , specifically including:

[0024] 2.1: Scale change is performed on the non-attack samples identified in step 1 to form different scale images, which use but are not limited to x0.5, x1, x2, x4;

[0025] 2.2: The different scale images obtained in step 2.1 are input into a convolutional neural network The activation breakpoint of the convolutional neural network is set and divided into p intervals: [0, t1), [t1, t2), … [t p-1 , +∞), to obtain the activation value distribution of each input sample image x in the entire activation domain: Thus, the activation value distribution {Γ1, Γ2, …, Γ n} under different scales is obtained, and n is the number of different scales.

[0026] The convolutional neural network is preferably trained by the training set χ, which is specifically taken from the training set part of the public data sets MNIST, CIFAR10 and CIFAR100.

[0027] 2.3: The KL divergence is used to measure the difference between the activation value distributions obtained in step 2.2, specifically: Where: Γ1, Γ2 are the activation value distributions of different scale input images.

[0028] 2.4: The decision criterion C2 of the sample x is calculated: wherein: η KL is the threshold value of the activation value distribution difference; when the decision criterion C2 of the sample x is true, it is that the activation value distribution under two scales has a large difference, that is, the sample is identified as an attack sample.

[0029] Step 3, spectrum space consistency discrimination, as shown in Figure 4 , specifically comprising:

[0030] 3.1: performing frequency domain transformation on the non-attack sample identified in step 2 to obtain a frequency domain feature ξ;

[0031] 3.2: performing subspace decomposition on the frequency domain feature ξ in the frequency domain space: ξ = γ1σ1+ γ2σ2+ … + γ n σ n , wherein: {σ1,σ2,…,σ n} is a set of orthogonal bases, {γ1,γ2,…,γ n} is its corresponding coefficient, σ i is a frequency domain component, and γ i is the coefficient corresponding to the frequency domain component.

[0032] 3.3: when the subspace decomposition coefficient γ i obeys the normal distribution N(μ i ,σ i ), then the training set χ is used to estimate each group μ i ,σ i , and the sequence of the real sample set in the embedded feature space is obtained.

[0033] 3.4: calculating the decision criterion C3 of each input sample x wherein: η i is the confidence level of the feature dimension, and when the decision criterion C3 of the sample x is true, it is that in the frequency domain space, the distribution of the coefficient of a certain feature dimension jumps out of the confidence interval, that is, it shows a large inconsistency in the frequency domain feature dimension, that is, the sample is identified as an attack sample.

[0034] 3.5: the remaining sample is a clean sample.

[0035] Through specific actual experiments, in the software environment of pytorch 0.4 and the hardware environment of NVIDIA TITAN X, three common network structures of LetNet-5, ResNet50 and ResNet110 are used.

[0036] The test results of the present application on the MNIST, CIFAR10 and CIFAR100 three public datasets are given respectively under different attack attacks, including FGSM, PGD and BIM, and the attack misjudgment rate after the defense of the present method.

[0037] Serial number Network structure Dataset Attack algorithm Misjudgment rate under attack Misjudgment rate after defense 1 LeNet-5 MINIST FGSM 88.38% 12.07% 2 LeNet-5 MINIST PGD 100.00% 51.98% 3 ResNet-50 CIFAR10 FGSM 77.93% 19.84% 4 ResNet-50 CIFAR10 PGD 100.00% 28.01% 5 ResNet-110 CIFAR100 FGSM 91.76% 62.62% 6 ResNet-110 CIFAR100 PGD 99.89% 90.55%

[0038] As shown in Table 1, the present application can significantly improve the recognition rate of the original model under the attack, detect the adversarial samples, and reduce the misjudgment rate of the model under the attack. The multi-subspace joint artificial intelligence recognition model of the present application can detect the adversarial samples inconsistent with the clean samples in the feature space, the scale space and the frequency domain space respectively, and improve the attack sample recognition rate through the cascade screening.

[0039] The above specific embodiments can be adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present application, the protection scope of the present application is subject to the claims and is not limited by the above specific embodiments, and each implementation scheme within the scope is subject to the constraint of the present application.

Claims

1. A multi-subspace joint-based artificial intelligence identification model countermeasure defense method, characterized in that, The method comprises the following steps: Step 1: mapping the activation value vectors of each node and each layer of the model to different subspace consistency performances, i.e. whether the consistency discrimination of the feature embedding subspace is an adversarial sample, by inputting a sample; Step 2: performing multi-scale changes on the non-attack sample after discrimination, and comparing the distribution consistency performances of the value of each node in the model transmission of each scale, i.e. whether the consistency discrimination of the scale space is an adversarial sample; Step 3: performing subspace spectral decomposition on the original pixel or the output value of a certain layer of the neural network of the non-attack sample after further discrimination, and comparing whether the response coefficient statistics of each spectral component have consistency, i.e. whether the consistency discrimination of the spectral space is an adversarial sample, to realize full-spectrum defense; The consistency discrimination of the feature embedding subspace specifically comprises: 1.1: extract the embedding feature of the input sample image, specifically: is a feature extraction function, is a parameter, is an input sample image, and z is an embedding feature; 1.2: Eigen-decomposition of the embedded feature z yields: as eigenbases, …, as coefficients of the corresponding eigenbases; 1.3: consistent discrimination of the extracted individual subspace features, specifically: the feature decomposition of the embedding representation of the embedded features z obtained in step 1.1 is performed to obtain feature coefficients, and it is determined when the coefficient β of the feature base is greater than 0.5, the corresponding feature is selected as the final feature. i Subject to normal distribution , through the sample training set χ, estimate each group , , to obtain the distribution space of the feature coefficients of the real sample set in , wherein: is the mean and variance of the distribution, is the estimated mean and variance of the distribution; 1.4: For any input sample x, based on the decision criterion , , when the decision criterion is true, it is characterized by a large genetic sequence inconsistency in the feature dimension, i.e. the sample is identified as an attack sample, where: is the feature dimension confidence level; The consistency discrimination of the scale space specifically comprises: 2.1: performing scale changes on the non-attack sample obtained after the consistency discrimination of step 1 to form different scale images; 2.2: input the different scale images obtained in step 2.1 into the convolutional neural network , set the activation breakpoint of the convolutional neural network and divide it into intervals: , obtain the activation value distribution of each input sample image in the whole activation domain: , wherein: is the activation value distribution; 2.3: Measure the difference of the activation value distribution obtained in step 2.2 using KL divergence, specifically: where: , is the activation value distribution of the input image of different scales; 2.4: Based on the decision criterion: When the decision criterion is true, it means that there is a large difference in the activation value distribution under two scales, i.e. the sample is identified as an attack sample, where: is the threshold of the activation value distribution difference. The consistency discrimination of the spectral space specifically comprises: 3.1: Perform frequency domain transformation on the non-attack samples after consistency discrimination in step 2 to obtain frequency domain features ; 3.2: On the frequency domain features Subspace decomposition is performed on the frequency domain space: where: is a set of orthogonal bases, is its corresponding coefficients, is a frequency domain component, is a coefficient corresponding to the frequency domain component; 3.3: When the coefficients of subspace decomposition Follows a normal distribution Then use the training set For each group , Estimation is performed to obtain the sequence of the real sample set in the embedding feature space. ; 3.4: Based on the judgment criteria When the sample Judgment criteria If true, it means that in the frequency domain, the distribution of coefficients in a certain feature dimension jumps out of the confidence interval, indicating a significant inconsistency in that feature dimension, thus identifying the sample as an attack sample. The confidence level for this feature dimension is set, and the remaining samples are then considered clean samples.

Citation Information

Patent Citations

  • Multi-model cooperative defense method facing deep learning antagonism attack

    CN108446765A

  • Face anti-fraud method based on cross-domain feature alignment network

    CN114120401A