Dynamic extensible continuous confrontation defense method
By building a dynamic scalable continuous adversarial defense architecture, including defense experts' dynamic routing and feature space enhancement pseudo-task knowledge replay strategy, the existing continuous adversarial defense model shares large parameters and high memory consumption for different adversarial attack methods, and realizes dynamic response and efficient training results for different adversarial attacks.
Patent Information
- Application Number
- CN202510120517.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-25
AI Technical Summary
The existing continuous adversarial defense model is difficult to deal with different adversarial attack methods due to the sharing of all task parameters, and the data replay method leads to high memory consumption and privacy protection problems.
Using a dynamic scalable continuous adversarial defense method, the model parameters are dynamically adjusted to deal with different adversarial attacks by building target classifiers and target convolutional neural networks, including defense expert dynamic routing and feature fusion layers, and the pseudo-task knowledge replay strategy enhanced by feature space is reduced to the dependence on a large amount of historical training data.
Dynamic response to different adversarial attack methods is realized, which reduces the suboptimal performance problems of model caused by shared parameters, and avoids high memory consumption and privacy protection problems, improving the accuracy of model training results.
Smart Images

Figure CN120032206A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a dynamic and scalable continuous confrontation defense method. Background Art
[0002] Deep neural networks are vulnerable to adversarial attacks. Attackers generate adversarial samples by adding carefully designed, imperceptible perturbations to the input images to deceive the model, causing it to produce incorrect outputs. Adversarial samples pose a major security threat to systems based on deep learning models. For example, attackers design attack algorithms and create adversarial glasses to mislead face recognition systems, causing the system to misidentify people, thereby achieving unauthorized access or identity impersonation, threatening public safety and personal privacy. In order to enhance the robustness of deep learning models to defend against adversarial samples, researchers have proposed various defense mechanisms, including defense distillation and adversarial sample detection, to ensure the reliability of the model in practical applications. Training, as an effective defense strategy, integrates adversarial samples into the model training process, enabling the model to learn adversarial features and thus improve its robustness. Current adversarial training methods usually rely on a limited set of adversarial samples, and after training, the model parameters are fixed, lacking the ability to adapt and expand in the face of new adversarial samples. However, in real scenarios, with the continuous progress of model security research, adversarial attacks are continuously updated and iterated, and the attack intensity continues to increase. Existing defense models are difficult to cope with the ever-changing attack threats. Therefore, researchers have begun to study continuous adversarial defense strategies, that is, continuously adjusting model parameters according to new adversarial attacks to ensure the continued robustness of the model.
[0003] In order to improve the model's continuous robustness against iterative adversarial attacks, existing continuous adversarial defense methods usually adopt a classic continuous learning process based on data replay and regularize on parameters shared by all tasks. However, the difference between adversarial samples and clean samples is that after adding perturbations, the samples become more complex, making continuous adversarial training difficult. The current continuous adversarial defense methods have the following defects: 1. The existing continuous adversarial defense model, which shares all task parameters, is unable to cope with continuously generated adversarial attacks. Adversarial samples are challenging samples with high specificity and aggressiveness, which are located at or across the classification boundary of the model. The continuous development of various attack methods further increases the difficulty of achieving high-quality feature extraction using a unified parameter-sharing network framework, which usually leads to sub-optimization problems of continuous adversarial defense models for different adversarial attack methods and cannot maintain optimal performance.
[0004] 2. Existing continuous adversarial defense models based on data replay are difficult to cope with diverse attack strategies; given the wide variety of attacks, different attack targets, and varying attack intensities, replay-based methods will require a large amount of memory, greatly increasing the storage burden of the model; and replaying only a small number of samples makes it difficult to maintain high-quality continuous defense; in addition, this method may cause privacy issues, especially in data-sensitive application scenarios. Summary of the invention
[0005] The embodiment of the present invention provides a dynamically scalable continuous adversarial defense method to at least solve the technical problems that the existing continuous adversarial defense model has a large degree of parameter sharing for different adversarial attack methods, and the large amount of stored training data makes the model training ability low, resulting in low accuracy of training results.
[0006] According to one aspect of an embodiment of the present invention, a dynamic and scalable continuous adversarial defense method is provided. The method may include: obtaining training set images, wherein each image in the training set consists of an initial image and disturbance information of the attack type corresponding to the initial image; constructing a target classifier, wherein the target classifier includes an image encoder, a text encoder, and a similarity calculation module; based on each type of attack image in the training set images, determining the prompts for each type of attack image; inputting each type of attack image and its corresponding prompts into the target classifier to obtain a successfully trained target classifier; constructing a target convolutional neural network, wherein the target convolutional neural network includes layer normalization, multi-head attention, feature fusion layer, defense expert dynamic routing, multi-layer perceptron, layer standard, and addition operation, and the defense expert dynamic routing includes if A number of routers and defense experts are provided, each router carries a label, and the disturbance information of each attack type corresponds to a type of router; the training set image is input into the target convolutional neural network to obtain the successfully trained target convolutional neural network and each router to determine the output proportion of each defense expert for each type of attack image; the image to be tested is obtained; wherein the image to be tested consists of the target image and the unknown type of disturbance information of the target image; the image to be tested is input into the successfully trained target classifier to obtain the attack type corresponding to the unknown type of disturbance information of the image to be tested; the image to be tested is input into the successfully trained target convolutional neural network to obtain the target result of the attack type of the image to be tested.
[0007] Optionally, the method of determining a prompt for each type of attack image based on each type of attack image in the training set images includes: defining the attack type of each type of attack image in the training set images as initial definition information, and defining the content of each type of attack image other than the initial definition information as a context vector; and obtaining a prompt for each type of attack image based on the initial definition information and the context vector.
[0008] Optionally, the step of inputting each type of attack image and its corresponding prompt into the target classifier to obtain a successfully trained target classifier includes: inputting a first type of attack image into an image encoder to obtain features of the first type of attack image, and calculating feature centers and covariance matrices of the first type of attack image; inputting the prompts of the first type of attack image into a text encoder to obtain features corresponding to the prompts of the first type of attack image; obtaining an attack type corresponding to the disturbance information of the first type of attack image by performing similarity calculation on the features of the first type of attack image and the features corresponding to the prompts of the first type of attack image; when a second attack image is to be input into the image encoder, performing Gaussian sampling on the feature centers of the first type of attack image to obtain sampled features; fine-tuning the target classifier trained for the first type of attack image based on the sampled features to obtain a fine-tuned target classifier; inputting the second type of attack image into the fine-tuned target classifier, iterating the loop, and obtaining a successfully trained target classifier.
[0009] Optionally, in the process of inputting the training set images into the target convolutional neural network to obtain a successfully trained target convolutional neural network, the method also includes: during the training process, when the target convolutional neural network encounters a new attack type image, based on the stored defense expert parameters of the old attack type image and the defense expert parameters of the trained new attack type image, replacing the trained defense expert parameters of the new attack type image.
[0010] Optionally, the replacing the defense expert parameters of the trained new attack type image based on the stored defense expert parameters of the old attack type image and the trained new attack type image includes: linearly fusing the stored defense expert parameters of the old attack type image with the trained new attack type image defense expert parameters to obtain target defense expert parameters; and replacing the trained new attack type image defense expert parameters with the target defense expert parameters.
[0011] Optionally, the step of inputting the image to be tested into a successfully trained target classifier to obtain the attack type corresponding to the unknown type disturbance information of the image to be tested includes: inputting the target image into an image encoder to obtain features of the target image; inputting prompts of each type of attack image into a text encoder to obtain features corresponding to the prompts of each type of attack image; and calculating the features of the target attack and the features corresponding to each prompt through a similarity calculation module to obtain the attack type corresponding to the disturbance information of the target image.
[0012] Optionally, the step of inputting the image to be tested into a successfully trained target convolutional neural network to obtain a target result of the attack type of the image to be tested includes: determining a target router in the dynamic routing of a defense expert based on the attack type corresponding to the disturbance information of the target image; extracting features of the image to be tested through layer normalization, multi-head attention and feature fusion to obtain feature tags; inputting the feature tags into each defense expert in the dynamic routing of the defense expert to obtain a first result of the attack type of the image to be tested by the target experts corresponding to the target router; inputting the feature tags into a multi-layer perceptron and a layer standard in sequence to obtain a second result of the attack type of the image to be tested; and obtaining a target result of the attack type of the image to be tested based on the first result and the second result.
[0013] Optionally, obtaining a target result of the attack type of the image to be tested based on the first result and the second result includes: determining a sum of the first result and the second result as the target result of the attack type of the image to be tested.
[0014] Beneficial effects of the present invention: (1) The present invention constructs a dynamically scalable continuous adversarial defense architecture, including a defense expert dynamic routing and sample type classifier composed of multiple defense experts and a group of routers, which alleviates the problem of suboptimal performance of the model caused by shared parameters.
[0015] (2) The present invention proposes a pseudo-task knowledge replay strategy based on feature space enhancement, which avoids the high memory consumption and privacy protection problems caused by storing a large amount of historical training data. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a flow chart of a dynamic and scalable continuous confrontation defense method according to an embodiment of the present invention; Figure 2 is a schematic diagram of a target classifier according to an embodiment of the present invention; Figure 3 is a schematic diagram of a target convolutional neural network according to an embodiment of the present invention; Figure 4 is a schematic diagram of pseudo-task knowledge replay enhanced by feature space according to an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only embodiments of a part of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0018] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0019] Example 1 According to an embodiment of the present invention, a dynamically scalable continuous adversarial defense method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system containing at least one set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0020] Figure 1 is a flow chart of a dynamic and scalable continuous confrontation defense method according to an embodiment of the present invention. Figure 1 As shown, the method may include the following steps: Step S101, obtaining training set images, wherein each image in the training set consists of an initial image and disturbance information of an attack type corresponding to the initial image.
[0021] In the technical solution provided in the above step S101 of the present invention, a training set of images is obtained, and the training set of images includes several groups of images of different attack types.
[0022] Step S102, constructing a target classifier, wherein the target classifier includes an image encoder, a text encoder and a similarity calculation module.
[0023] In the technical solution provided in the above step S102 of the present invention, Figure 2 is a schematic diagram of a target classifier according to an embodiment of the present invention. Figure 2As shown, the target classifier includes an image encoder, a text encoder, and a similarity calculation module.
[0024] Step S103, based on each type of attack image in the training set images, determine a prompt for each type of attack image.
[0025] In the technical solution provided in the above step S103 of the present invention, a prompt for each type of attack image is obtained according to each type of attack image in the training set images.
[0026] Step S104: input each type of attack image and its corresponding prompt into a target classifier to obtain a successfully trained target classifier.
[0027] In the technical solution provided in the above step S104 of the present invention, a target classifier is trained with each type of attack image and its corresponding prompt to obtain a successfully trained target classifier.
[0028] Step S105, constructing a target convolutional neural network, wherein the target convolutional neural network includes layer normalization, multi-head attention, feature fusion layer, defense expert dynamic routing, multi-layer perceptron, layer standard and addition operation, and the defense expert dynamic routing includes several routers and several defense experts, each router carries a label, and the disturbance information of each attack type corresponds to a type of router.
[0029] In the technical solution provided in the above step S105 of the present invention, Figure 3 is a schematic diagram of a target convolutional neural network according to an embodiment of the present invention, such as Figure 3 As shown, the target convolutional neural network includes layer normalization, multi-head attention, feature fusion layer, defense expert dynamic routing, multi-layer perceptron, layer standard and addition operation. The defense expert dynamic routing includes several routers and several defense experts. Each router carries a label. The disturbance information of each attack type corresponds to a type of router. Each defense expert is LoRA. LoRA can decouple the original heavy and frozen parameters into a low-rank trainable space, thereby increasing the training speed and reducing the training burden.
[0030] Step S106, input the training set images into the target convolutional neural network, obtain the successfully trained target convolutional neural network and each router determines the output ratio of each defense expert for each type of attack image.
[0031] In the technical solution provided in the above step S106 of the present invention, the training set images train the target convolutional neural network to obtain a successfully trained target convolutional neural network and each router determines the output ratio of each defense expert for each type of attack image.
[0032] Step S107, obtaining an image to be tested; wherein the image to be tested consists of a target image and unknown type disturbance information of the target image.
[0033] In the technical solution provided in the above step S107 of the present invention, an image to be tested is obtained; wherein the image to be tested is composed of a target image and unknown type disturbance information of the target image, and the target image is an image without noise.
[0034] Step S108: input the image to be tested into the successfully trained target classifier to obtain the attack type corresponding to the unknown type disturbance information of the image to be tested.
[0035] In the technical solution provided in the above step S108 of the present invention, the successfully trained target classifier processes the image to be tested to obtain the attack type corresponding to the unknown type disturbance information of the image to be tested.
[0036] Step S109, inputting the image to be tested into the successfully trained target convolutional neural network to obtain the target result of the attack type of the image to be tested.
[0037] In the technical solution provided in the above step S109 of the present invention, the successfully trained target convolutional neural network processes the image to be tested to obtain a target result of the attack type of the image to be tested, and the target result may be a percentage.
[0038] The above method of this embodiment is further introduced below.
[0039] As an optional implementation method, step S103, determining the prompts for each type of attack image based on each type of attack image in the training set images, includes: defining the attack type of each type of attack image in the training set images as initial definition information, and defining the content of each type of attack image other than the initial definition information as a context vector; and obtaining the prompts for each type of attack image based on the initial definition information and the context vector of each type of attack image.
[0040] In this embodiment, the attack type of each attack image in the training set is defined as the initial definition information , the content of each attack image except the initial definition information is defined as the context vector ,Therefore, the hint for each type of attack image can be expressed as:
[0041] in, is the prompt for each type of attack image, t is the attack type, and The context consists of a series of Vector indicates (where ), these vectors are in the text embedding space instead of the original text, using continuous parameters, which provides greater flexibility compared to the discrete parameters of the text formula. The dimension of each vector is the same as the word embedding; the final input of the text encoder is the concatenation of the context vector and the class word embedding. The prompt for each type of attack image can also be expressed as:
[0042] In addition, a context parameterization variant called type-specific context is adopted to assign a separate context vector to each attack type, allowing different attack types to have different customized contexts, which helps to distinguish attack types. The loss for updating the context parameters uses the cross-entropy loss.
[0043] As an optional implementation method, step S104, the step of inputting each type of attack image and its corresponding prompt into the target classifier to obtain a successfully trained target classifier includes: inputting the first type of attack image into the image encoder to obtain the features of the first type of attack image, and calculating the feature center and covariance matrix of the first type of attack image; inputting the prompt of the first type of attack image into the text encoder to obtain the features corresponding to the prompt of the first type of attack image; by performing similarity calculation on the features of the first type of attack image and the features corresponding to the prompt of the first type of attack image, the attack type corresponding to the disturbance information of the first type of attack image is obtained; when the second attack image is about to be input into the image encoder, Gaussian sampling is performed on the feature center of the first type of attack image to obtain the sampled features; fine-tuning the target classifier trained for the first type of attack image based on the sampled features to obtain the fine-tuned target classifier; inputting the second type of attack image into the fine-tuned target classifier, iterating the loop to obtain the successfully trained target classifier.
[0044] In this embodiment, if Figure 2 As shown, the first type of attack image is input into the image encoder to obtain the features of the first type of attack image, and the feature center and covariance matrix of the first type of attack image are calculated; the prompt of the first type of attack image is input into the text encoder to obtain the features corresponding to the prompt of the first type of attack image; by calculating the similarity between the features of the first type of attack image and the features corresponding to the prompt of the first type of attack image, the attack type corresponding to the disturbance information of the first type of attack image is obtained; Figure 4 is a schematic diagram of pseudo-task knowledge replay enhanced by feature space according to an embodiment of the present invention, such as Figure 4 As shown in , when the second attack image is about to be input into the image encoder, Gaussian sampling is performed on the feature center of the first attack image to obtain the sampled features ( Figure 4), fine-tune the target classifier trained on the first type of attack images based on the sampled features to obtain a fine-tuned target classifier; input the second type of attack images into the fine-tuned target classifier, iterate the loop, and obtain a successfully trained target classifier.
[0045] The feature center of the first attack image is the average feature representation of the first type of attack image samples, and the feature center is calculated with the sample features of the same attack type in each small batch of the first type of attack images. The variances between are gradually aggregated into a covariance matrix using a moving average. The expression of the covariance matrix is:
[0046] in, is the covariance matrix of the first type of attack image.
[0047] When the second type of attack image is about to be input into the image encoder, Gaussian sampling is performed on the feature center of the first type of attack image to obtain the sampled features. The past sample features are the sample features of the first type of attack image. Therefore, the expression of the sampled features and covariance matrix of the first type of attack image is:
[0048]
[0049] in, is the sampled feature of the first type of attack image, The sampled covariance matrix of the first type of attack image is: is the covariance matrix The Cholesky decomposition matrix, is from a standard normal distribution A random vector drawn from is the covariance matrix The transpose of the Cholesky factorization matrix, is the feature center of the first type of attack image ,When training a new type of attack, past features are randomly sampled from the stored mean and covariance matrices of past types to generate a feature subset; ,this process enriches the diversity of features and prevents the classifier of the sample type from overfitting to the current attack.
[0050] As an optional implementation method, in step S106, in the process of inputting the training set image into the target convolutional neural network to obtain a successfully trained target convolutional neural network, the method also includes: during the training process, when the target convolutional neural network encounters a new attack type image, based on the stored defense expert parameters of the old attack type image and the defense expert parameters of the trained new attack type image, replacing the defense expert parameters of the trained new attack type image.
[0051] In this embodiment, when the target convolutional neural network encounters a new attack type image, the stored defense expert parameters of the old attack type image and the defense expert parameters of the trained new attack type image are calculated, and the calculated values are used to replace the defense expert parameters of the trained new attack type image.
[0052] As an optional implementation method, the method of replacing the defense expert parameters of the trained new attack type image based on the stored defense expert parameters of the old attack type image and the trained new attack type image includes: linearly fusing the stored defense expert parameters of the old attack type image and the trained new attack type image to obtain target defense expert parameters; and replacing the trained new attack type image defense expert parameters with the target defense expert parameters.
[0053] In this embodiment, the stored defense expert parameters of the old attack type image and the trained defense expert parameters of the new attack type image are linearly fused to obtain the expression of the target defense expert parameters:
[0054] is the target defense expert parameter, is the defense expert parameter for the new attack type image after training, Defense expert parameters for old attack type images stored, is the weight used in the fusion process.
[0055] Leveraging Target Defense Expert Parameters Replace the defense expert parameters of the new attack type image after training .
[0056] As an optional implementation method, step S108, inputting the image to be tested into a successfully trained target classifier to obtain the attack type corresponding to the unknown type disturbance information of the image to be tested, includes: inputting the target image into an image encoder to obtain the characteristics of the target image; inputting the prompt of each type of attack image into a text encoder to obtain the characteristics corresponding to the prompt of each type of attack image; calculating the characteristics of the target attack and the characteristics corresponding to each prompt through a similarity calculation module to obtain the attack type corresponding to the disturbance information of the target image.
[0057] In this embodiment, if Figure 2 As shown, the target image is input into the image encoder to obtain the features of the target image; the prompts of each type of attack image are input into the text encoder to obtain the features corresponding to the prompts of each type of attack image; the features of the target attack and the features corresponding to each prompt are calculated through the similarity calculation module to obtain several values, the several values are sorted, and the type of prompt corresponding to the maximum value of the sorting is determined as the attack type corresponding to the perturbation information of the target image.
[0058] As an optional implementation method, step S109, the inputting of the image to be tested into the successfully trained target convolutional neural network to obtain the target result of the attack type of the image to be tested, includes: determining the target router in the dynamic routing of the defense expert based on the attack type corresponding to the disturbance information of the target image; extracting features of the image to be tested through layer normalization, multi-head attention and feature fusion to obtain feature tags; inputting the feature tags into each defense expert in the dynamic routing of the defense expert to obtain a first result of the attack type of the image to be tested by the target expert corresponding to the target router; inputting the feature tags into the multi-layer perceptron and the layer standard respectively in sequence to obtain a second result of the attack type of the image to be tested; based on the first result and the second result, obtaining the target result of the attack type of the image to be tested.
[0059] In this embodiment, according to the attack type corresponding to the disturbance information of the target image, a target router in the defense expert dynamic routing with the same attack type as the disturbance information of the target image is selected, and features of the image to be tested are extracted through layer normalization, multi-head attention and feature fusion to obtain feature tags; the feature tags are input into each defense expert in the defense expert dynamic routing to obtain the weight of each defense expert for the attack type of the image to be tested.
[0060] The feature tag is input into each defense expert in the defense expert dynamic routing, and the expression of the weight of each defense expert for the attack type of the image to be tested is obtained as follows:
[0061] in, is the output ratio of the target router to each defense expert’s attack type for the image to be tested, For each router's tag, ,function Will Projected onto a one-dimensional vector, it represents the proportion of each defense expert to the first result. The function selection has the greatest impact on the correct result. experts, and set the contributions of the remaining experts to Finally, apply The function normalizes these weights.
[0062] According to the weight of each defense expert for the attack type of the image to be tested, the first result expression of the attack type of the target expert for the image to be tested corresponding to the target router is obtained as follows:
[0063] in, is the first result of the target expert's attack type on the image to be tested corresponding to the target router, is a feature marker, i Output of each defense expert selected for the router.
[0064] As an optional implementation manner, obtaining a target result of the attack type of the image to be tested based on the first result and the second result includes: determining the sum of the first result and the second result as the target result of the attack type of the image to be tested.
[0065] In this embodiment, the first result and the second result are added to obtain a target result of the attack type of the image to be tested.
[0066] In an embodiment of the present invention, a training set image is obtained, wherein each image in the training set consists of an initial image and disturbance information of an attack type corresponding to the initial image; a target classifier is constructed, wherein the target classifier includes an image encoder, a text encoder, and a similarity calculation module; based on each type of attack image in the training set image, a prompt for each type of attack image is determined; each type of attack image and its corresponding prompt are input into the target classifier to obtain a successfully trained target classifier; a target convolutional neural network is constructed, wherein the target convolutional neural network includes layer normalization, multi-head attention, feature fusion layer, defense expert dynamic routing, multi-layer perceptron, layer standard, and addition operation, and the defense expert dynamic routing includes a plurality of routers and a plurality of defense experts, each router carries a tag, and disturbance information of each attack type corresponds to a type of router; the training set image is input into the target convolutional neural network to obtain a successfully trained target convolutional neural network The output proportion of each defense expert for each type of attack image is determined through the network and each router; the image to be tested is obtained; wherein the image to be tested is composed of a target image and an unknown type of disturbance information of the target image; the image to be tested is input into a successfully trained target classifier to obtain the attack type corresponding to the unknown type of disturbance information of the image to be tested; the image to be tested is input into a successfully trained target convolutional neural network to obtain the target result of the attack type of the image to be tested, which solves the technical problems that the existing continuous adversarial defense model has a large degree of shared parameters for different adversarial attack methods, and the large amount of stored training data makes the model training ability low, resulting in low accuracy of training results. It achieves the goal of building a dynamically scalable continuous adversarial defense architecture, reducing shared parameters, and proposing a pseudo-task knowledge replay strategy based on feature space enhancement, which avoids the low model training ability due to the storage of a large amount of historical training data and improves the accuracy of training results.
[0067] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0068] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0069] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0070] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0071] In addition, each functional unit in each embodiment of the present invention may be integrated into a first processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0072] The above are only preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A dynamic and scalable continuous confrontation defense method, characterized in that: include: Obtain training set images, where each image in the training set consists of an initial image and perturbation information of an attack type corresponding to the initial image; Constructing a target classifier, wherein the target classifier includes an image encoder, a text encoder, and a similarity calculation module; Determine a prompt for each type of attack image based on each type of attack image in the training set images; Input each type of attack image and its corresponding prompt into the target classifier to obtain a successfully trained target classifier; Construct a target convolutional neural network, where the target convolutional neural network includes layer normalization, multi-head attention, feature fusion layer, defense expert dynamic routing, multi-layer perceptron, layer standard and addition operation. The defense expert dynamic routing includes several routers and several defense experts. Each router carries a tag, and the disturbance information of each attack type corresponds to a type of router. Input the training set images into the target convolutional neural network to obtain the successfully trained target convolutional neural network and each router to determine the output ratio of each defense expert for each type of attack image; Acquire an image to be tested; wherein the image to be tested consists of a target image and unknown type disturbance information of the target image; Input the image to be tested into the successfully trained target classifier to obtain the attack type corresponding to the unknown type disturbance information of the image to be tested; The image to be tested is input into the successfully trained target convolutional neural network to obtain the target result of the attack type of the image to be tested.
2. The method according to claim 1, characterized in that The step of determining a prompt for each type of attack image based on each type of attack image in the training set includes: The attack type of each attack image in the training set is defined as the initial definition information, and the content of each attack image other than the initial definition information is defined as the context vector; Based on the initial definition information and context vector of each type of attack image, a hint for each type of attack image is obtained.
3. The method according to claim 2, characterized in that The step of inputting each type of attack image and its corresponding prompt into the target classifier to obtain a successfully trained target classifier includes: Input the first type of attack image into the image encoder, obtain the features of the first type of attack image, and calculate the feature weight and covariance matrix of the first type of attack image; Inputting the prompt of the first type of attack image into the text encoder to obtain features corresponding to the prompt of the first type of attack image; By calculating the similarity between the features of the first type of attack image and the features corresponding to the prompt of the first type of attack image, the attack type corresponding to the disturbance information of the first type of attack image is obtained; When the second attack image is about to be input into the image encoder, Gaussian sampling is performed on the feature center of the first attack image to obtain the sampled features; Fine-tune the target classifier trained on the first type of attack images based on the sampled features to obtain a fine-tuned target classifier; The second type of attack image is input into the fine-tuned target classifier, and the cycle is iterated to obtain a successfully trained target classifier.
4. The method according to claim 1, characterized in that: In the process of inputting the training set images into the target convolutional neural network to obtain a successfully trained target convolutional neural network, the method further includes: During the training process, when the target convolutional neural network encounters a new attack type image, the defense expert parameters of the trained new attack type image are replaced based on the stored defense expert parameters of the old attack type image and the defense expert parameters of the trained new attack type image.
5. The method according to claim 4, characterized in that The replacing the trained defense expert parameters of the new attack type image based on the stored defense expert parameters of the old attack type image and the trained defense expert parameters of the new attack type image comprises: Linearly fuse the stored defense expert parameters of the old attack type image with the trained defense expert parameters of the new attack type image to obtain the target defense expert parameters; The target defense expert parameters are used to replace the trained defense expert parameters of the new attack type image.
6. The method according to claim 1, characterized in that The step of inputting the image to be tested into the successfully trained target classifier to obtain the attack type corresponding to the unknown type disturbance information of the image to be tested includes: Input the target image into the image encoder to obtain the features of the target image; Input the prompt of each type of attack image into the text encoder to obtain the features corresponding to the prompt of each type of attack image; The features of the target attack and the features corresponding to each prompt are calculated through the similarity calculation module to obtain the attack type corresponding to the perturbation information of the target image.
7. The method according to claim 6, characterized in that The step of inputting the image to be tested into the successfully trained target convolutional neural network to obtain the target result of the attack type of the image to be tested includes: Determine the target router in the defense expert dynamic routing based on the attack type corresponding to the perturbation information of the target image; Feature extraction is performed on the test image through layer normalization, multi-head attention and feature fusion to obtain feature labels; Input the characteristic mark into each defense expert in the defense expert dynamic routing, and obtain the first result of the attack type of the target expert corresponding to the target router for the image to be tested; Inputting the feature tags into the multi-layer perceptron and the layer standard in sequence respectively, and obtaining a second result of the attack type of the image to be tested; Based on the first result and the second result, a target result of the attack type of the image to be tested is obtained.
8. The method according to claim 7, characterized in that The step of obtaining a target result of the attack type of the image to be tested based on the first result and the second result includes: The sum of the first result and the second result is determined as the target result of the attack type of the image to be tested.
Citation Information
Patent Citations
An anti-attack defense method for a feature map attention mechanism and application
CN109948658A
Multi-task defense model construction method for infrared image countermeasure attacks
CN112598032A
Preprocessing defense method aiming at target detection confrontation attack
CN114723663A
Two-stage confrontation and defense method and system for image classification
CN114881104A
Living body detection model training method and device, living body detection method and device and electronic equipment
CN115937993A