Generative attack and defense confrontation method based on domain self-adaption
Through multi-white-box collaborative modeling and domain adaptation methods, the problems of gradient invisibility and distribution offset in black-box attacks are solved, highly transferable adversarial samples are generated, the attack success rate and perturbation concealment are improved, and continuously evolving attack and defense capabilities are achieved.
Patent Information
- Application Number
- CN202510833805.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies face the problems of black-box attacks, such as the invisibility of target model gradients leading to the failure of traditional white-box attacks, low migration rate of adversarial samples, and the contradiction between the concealment of perturbations and the success rate of attacks. Especially under high-dimensional data, the efficiency is extremely low, and there is a lack of unified indicators for collaborative optimization.
A multi-white-box collaborative modeling and domain adaptation method is adopted to construct a source domain set by integrating multiple groups of heterogeneous white-box models. The feature space distribution is aligned using KL divergence and maximum mean difference (MMD). A perturbation-semantic decoupling training mechanism is designed to generate highly transferable adversarial samples and achieve targeted interference on the black-box system.
It significantly improves the success rate of attacks on black-box systems and the concealment of disturbances, overcomes the estimation bias caused by differences in model structures, achieves cross-model generalization to generate adversarial samples, and possesses continuously evolving attack and defense capabilities.
Smart Images

Figure CN120688062A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a generative attack-defense confrontation method based on domain adaptation, belonging to the technical field of visual perception and attack-defense confrontation. Background Art
[0002] Current black-box attacks mainly rely on proxy models to generate adversarial samples and migrate them to the target system. Mainstream methods include white-box attack conversion technology and query optimization strategies. The defense side focuses on technologies such as adversarial training, input reconstruction, and gradient masking. Black-box attacks based on generative adversarial networks face multiple challenges: the invisibility of target model gradients makes traditional white-box attacks ineffective, and decision boundaries can only be reconstructed through limited input-output queries, which is extremely inefficient under high-dimensional data; the data domain differences between proxy models and black-box models (such as the distribution shift from MNIST to medical images) cause a sharp drop in the migration rate of adversarial samples; there is a lack of unified indicators to coordinately optimize the conflicting needs of perturbation concealment and attack success rate.
[0003] A generative attack-defense adversarial method based on domain adaptation achieves multiple optimizations of the above difficulties through decision boundary modeling and cross-domain distribution alignment: first, based on multi-white-box collaborative modeling and black-box boundary decision probability estimation, the KL divergence is used to accurately quantify the boundary difference between the source domain model and the target model to guide the generation of highly transferable adversarial samples. Secondly, the maximum mean difference (MMD) and KL divergence are introduced to collaboratively drive feature alignment. The Gaussian-Laplacian kernel is used to measure the global feature distribution offset in the Hilbert space, and the local decision confidence distribution is calibrated through the KL divergence to achieve implicit mapping of the source domain and target domain feature spaces; finally, a perturbation-semantic decoupling training mechanism is designed to achieve a high attack success rate while ensuring the concealment of the perturbation, systematically resolving the attack-defense contradiction. Summary of the Invention
[0004] The technical problem solved by the present invention is: to overcome the shortcomings of the existing technology and provide a generative attack and defense confrontation method based on domain adaptation to achieve directional interference on black box intelligent systems.
[0005] The technical solution of the present invention is:
[0006] The present invention discloses a generative attack-defense confrontation method based on domain adaptation, comprising:
[0007] S1. Build multiple sets of heterogeneous white-box models to form a source domain model set;
[0008] S2. Calculate the expected gradient of the source domain model set and infer the maximum disturbance tolerance threshold of the black box system;
[0009] S3, initialize the enhanced content consistency constraint coefficient;
[0010] S4. Using domain adaptation methods, align the feature space distributions of the white-box model group and the black-box system to obtain the overall domain adaptation loss;
[0011] S5. Generate a dual-channel generative adversarial network based on the maximum perturbation tolerance threshold, the enhanced content consistency constraint coefficient, and the overall domain adaptation loss;
[0012] S6. Training the dual-channel generative adversarial network to obtain a trained dual-channel generative adversarial network;
[0013] S7. Input the original sample into the trained dual-channel generative adversarial network to obtain the adversarial sample and the number of adversarial samples that successfully mislead the black-box system.
[0014] S8. Calculate the success rate of the adversarial samples attacking the black-box system based on the number of adversarial samples that successfully mislead the black-box system and the original samples.
[0015] S9. Adjust the maximum disturbance tolerance threshold and enhance the content consistency constraint coefficient based on the adversarial sample, the attack success rate of the adversarial sample on the black box system, and the preset target success rate, and repeat steps S6 to S9 until the attack success rate of the adversarial sample on the black box system is greater than or equal to the preset target success rate.
[0016] Furthermore, in the above method, the construction of multiple sets of heterogeneous white box models to form a source domain model set is specifically as follows:
[0017] The source domain model set is The output joint distribution is: p i (x)=softmax(M i (x)); where M N is the Nth heterogeneous white box model, p i (x) is the white box model M i The output probability distribution of .
[0018] Furthermore, in the above method, the maximum disturbance tolerance threshold of the black box system is specifically:
[0019] The adversarial sample z satisfies ||zx|| ∞ ≤∈ max
[0020]
[0021] Among them, ∈ max is the maximum disturbance tolerance threshold, N is the total number of heterogeneous white box models, G i is the expected value of the white box model gradient, M i (x) is the i-th heterogeneous white box model, is the expected function; min is the minimum function.
[0022] Furthermore, in the above method, the overall domain adaptation loss is specifically:
[0023]
[0024] in, is the MMD loss between the i-th white box model and the black box, w i is the contribution weight, N is the total number of models, φ(x black ) is the feature space distribution of the black box model mapped to the high-dimensional feature space, φ(M i (x)) is the feature space distribution of the white box model mapped to the high-dimensional feature space.
[0025] Furthermore, in the above method, the dual-channel generative adversarial network includes a generator and two heterogeneous discriminators; the generator is used to generate adversarial samples for simultaneously deceiving the two discriminators; the two heterogeneous discriminators respectively simulate the feature extractors of the heterogeneous white-box model group and the black-box system; wherein,
[0026] The loss function of the generator is:
[0027]
[0028] Among them, x is the original sample, z is the adversarial sample of the generator, is the content consistency constraint; D1(z) and D2(z) are both heterogeneous discriminators; is the discriminator loss; is the domain adaptation loss; α is the domain adaptation loss weight, β is the content consistency loss weight;
[0029] The loss function of the discriminator is:
[0030]
[0031] Among them, B samples is the batch size, y b is the sample label, z b =G(x b ) is the adversarial sample generated by the generator; D k is the discriminator, K is the number of discriminators; L KL is the KL loss; λ is the weight of the KL loss.
[0032] Furthermore, in the above method, the adversarial sample generated by the generator is:
[0033]
[0034] ResBlock(x)=x+Conv3×3 (ReLU(BN(Conv 3×3 (x))))
[0035] Among them, x is the original sample, z is the adversarial sample, Conv 3×3 is a 3×3 convolutional layer, BN is batch normalization, ReLU is the activation function, and ResBlock is the residual module.
[0036] Furthermore, in the above method, the KL loss is specifically:
[0037]
[0038] Among them, D1(x) and D2(x) are the outputs of the discriminator, D KL is the KL divergence, and x is the input sample of the discriminator.
[0039] Furthermore, in the above method, the success rate of the adversarial sample attacking the black box system is specifically:
[0040]
[0041] Among them, ASR t is the success rate of adversarial samples attacking the black box system, B is the number of adversarial samples that successfully mislead the black box, and A is the total number of test samples.
[0042] Furthermore, in the above method, the maximum disturbance tolerance threshold is adjusted and the content consistency constraint coefficient is enhanced, specifically by:
[0043] If ASR t <ASR target ,but
[0044] Among them, ASR target is the preset target success rate, η is the adjustment factor, ASR target The success rate of the preset target;
[0045] If σ(z)>σ thres , then β (t+1) =β (t) (1+γ)
[0046] Among them, γ is the reinforcement coefficient, β (t+1) is the enhanced content consistency constraint coefficient at the next moment, β (t) is the enhanced content consistency constraint coefficient at the current moment, σ(z) is the confidence variance of the output sample, σ thres is the confidence variance threshold of the output sample.
[0047] The beneficial effects of the present invention and the prior art are:
[0048] (1) This paper uses multi-model collaborative modeling to overcome the limitations of a single model. By integrating multiple sets of heterogeneous white-box models to construct a source domain set, and integrating the decision boundary characteristics of different models, it significantly improves the ability to generalize and infer the capability boundaries of black-box systems. Compared with traditional single-model replacement methods, it effectively reduces the estimation bias caused by model structure differences and enhances the robustness of attack and defense confrontation.
[0049] (2) This paper uses domain adaptation theory to drive feature space alignment, introducing maximum mean discrepancy (MMD) and a dynamic weight allocation mechanism to enforce alignment of the feature distributions of the white-box model group and the black-box system. This method overcomes the distribution offset bottleneck caused by model heterogeneity and achieves cross-model generalization of adversarial examples without relying on the internal structure of the target model or training data.
[0050] (3) This invention uses closed-loop dynamic feedback to achieve adaptive attack-defense game. By optimizing model parameters through the cross-entropy loss function and KL divergence, a real-time monitoring and parameter adjustment mechanism is established, dynamically optimizing the perturbation strategy based on the attack effect. In scenarios where the defense system is dynamically evolving, this method adaptively balances attack effectiveness and concealment, forming a continuously evolving attack-defense countermeasure capability, overcoming the environmental adaptability shortcomings of traditional static attack strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a schematic diagram of the structure of the generative attack-defense adversarial network model of the present invention;
[0052] Figure 2 Schematic diagram of the domain adaptation mechanism of the present invention;
[0053] Figure 3 It is a flow chart of the algorithm of the present invention. DETAILED DESCRIPTION
[0054] The present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0055] like Figure 1 As shown, the present invention provides a generative attack-defense confrontation method based on domain adaptation, comprising the following steps:
[0056] 1. Multi-white-box collaborative modeling and black-box boundary estimation
[0057] Deploy multiple groups of heterogeneous white-box models (such as ResNet, DenseNet, VGG, etc.) to form a source domain set, and infer the perturbation tolerance of the black-box system by analyzing the sensitivity of their decision boundaries.
[0058] By integrating white-box models of various architectures, we can build a set of source domain models covering a wide range of feature spaces.
[0059] Each white box model Mi The output probability distribution p i (x)=softmax(M i (x)) are integrated to form the joint distribution of the source domain This enhances the ability to generalize and infer the decision boundary of the black box system. Its output joint distribution is:
[0060]
[0061] Ensemble learning is used to reduce the estimation error caused by model structure differences.
[0062] By analyzing the gradient response of the white-box model to input perturbations, we can quantify the sensitivity of its decision boundary. The larger the gradient amplitude, the more sensitive the model is to perturbations, and the corresponding black-box system has a lower tolerance threshold.
[0063] Calculate the expected gradient magnitude of each white box model:
[0064]
[0065] Take the inverse of the minimum value as the maximum perturbation threshold of the black box:
[0066]
[0067] The constraint adversarial example z satisfies ||zx|| ∞ ≤∈ max , ensuring the concealment of disturbance.
[0068] 2. Domain Adaptive Feature Alignment
[0069] The maximum mean difference (MMD) is used to align the feature distribution of the white-box model group and the black-box system. The cross-model distribution offset is measured by the maximum mean difference. Combined with dynamic weight allocation, the role of the white-box model with a high degree of matching with the black-box feature is strengthened to solve the generalization bottleneck caused by model heterogeneity. Figure 2 As shown, the feature space difference is calculated:
[0070]
[0071] The dynamic weight allocation mechanism adjusts the contribution weight of each white box model according to the matching degree between the white box model and the black box model, and dynamically allocates the model weight. The formula is:
[0072]
[0073] Weighted MMD loss optimizes feature matching:
[0074]
[0075] Improving cross-model transferability of adversarial examples through feature space alignment.
[0076] 3. Dual-channel generative adversarial network training
[0077] A dual-channel generative adversarial network is designed. The generator generates adversarial examples based on a residual structure. While constraining the perturbation amplitude, it simultaneously optimizes the attack effectiveness on the white-box model group and the feature matching with the black-box target domain. The discriminator supervises the sample classification accuracy through cross-entropy loss, and introduces a KL divergence constraint to minimize the difference in the output of the two discriminators, ensuring the consistency of cross-model transfer of adversarial examples.
[0078] Dual-channel generative adversarial network architecture,The dual-channel generative adversarial network consists of a generator (G) and two heterogeneous discriminators (D1, D2).
[0079] Generator Design: The generator is responsible for generating adversarial examples, aiming to deceive both discriminators simultaneously. The discriminators simulate the feature extractors of the white-box model group (source domain) and the black-box system (target domain), respectively, to improve the generalization ability of the model through adversarial training. The dual discriminator design enhances the cross-model transferability of adversarial examples by constraining the output distribution difference, forming a dynamic "generation-discrimination" game. The generator loss function generates adversarial examples based on the residual structure, retaining the input features and superimposing perturbations:
[0080]
[0081] The total loss function jointly optimizes aggressiveness and stealth:
[0082]
[0083] Discriminator design: Dual heterogeneous discriminators (D1, D2) supervise classification through cross entropy loss:
[0084]
[0085] The KL divergence constraint design constrains the output distribution consistency of the two discriminators through KL divergence, forcing adversarial samples to have similar misleading effects between heterogeneous models. Introducing KL divergence constraint output consistency:
[0086]
[0087] This constraint minimizes the prediction difference between the two discriminators and improves the cross-model transferability of adversarial examples.
[0088] 4. Dynamic feedback and closed-loop optimization
[0089] Real-time monitoring of attack success rate (ASR) and stealth indicators, and dynamic adjustment of parameters:
[0090] Disturbance intensity adjustment: If ASR t <ASR target , relax the threshold:
[0091]
[0092] Concealment enhancement: If a disturbance is detected, the constraint coefficient β is increased:
[0093] β (t+1) =β (t) (1+γ)
[0094] Update the generative adversarial network weights w i , forming a closed loop of "attack-evaluation-iteration".
[0095] The following is a detailed description of the present invention in conjunction with the accompanying drawings and specific implementation methods. The present invention provides a generative attack and defense method based on domain adaptation, and the specific implementation steps are as follows:
[0096] Generative adversarial training is used to attack the opponent's information flow, targeting their deep learning-based intelligent cognitive model. The opponent's cognitive model is set as a black-box model, while our own cognitive model is set as a white-box model. Based on domain adaptation theory, the maximum distance boundary between the white-box model and the black-box model is learned. Based on the existing white-box model information, the performance boundary of the black-box model is predicted, thereby generating highly transferable adversarial sample noise. This noise is then applied to our own equipment, disrupting the recognition and judgment of the opponent's situational awareness system.
[0097] 1. Input a clean sample x, pass it through the generator G built based on the deep learning network, and output the adversarial sample z. The clean sample x and the adversarial sample z are respectively fed into two deep learning recognition networks D1 and D2.
[0098] 2. Use data samples to train two deep learning recognition networks D1 and D2;
[0099] 3. Design the loss function of deep learning recognition networks D1 and D2 as the cross entropy loss function based on information entropy theory, and at the same time constrain the KL divergence loss function of deep learning recognition networks D1 and D2 to be as small as possible;
[0100] 4. Through data samples, train the generator G and design the adversarial network training loss function, that is, the samples z and x generated by the generator G interfere with the two deep learning recognition networks D1 and D2, causing them to make errors, thereby optimizing the generator parameter model.
[0101] The above process is expressed in a formal way:
[0102] The specific implementation process is as follows Figure 3 As shown:
[0103] Step 1: Adversarial sample generation and input
[0104] Input: Raw visual image samples (For RGB images, H×W is the resolution, and C=3).
[0105] Generator: Generator G based on deep residual network, outputs adversarial sample z:
[0106]
[0107] ResBlock k is the kth residual block, N≥8.
[0108] Output distribution: Input x and z into the dual discriminators D1 (ResNet architecture) and D2 (DenseNet architecture).
[0109] Step 2: Dual Discriminator Training
[0110] D1 and D2 are trained to distinguish real samples from adversarial samples while constraining their output distribution consistency. The cross entropy loss function is:
[0111]
[0112] KL divergence constraint:
[0113]
[0114] Total loss:
[0115]
[0116] (λ=0.1, balanced classification and distribution alignment)
[0117] Step 3: Generator adversarial training
[0118] Goal: Generate adversarial examples z that simultaneously fool D1 / D2 and match the black-box feature distribution.
[0119] Fighting Loss:
[0120]
[0121] Domain Adaptation Loss (MMD Alignment):
[0122]
[0123] Content consistency constraints:
[0124]
[0125] Total loss:
[0126]
[0127] Step 4: Dynamic Feedback and Perturbation Constraints
[0128] Closed-loop adjustment with disturbance amplitude constraints
[0129] Closed-loop adjustment:
[0130] If the attack success rate ASR < 80%, press Relax constraints.
[0131] If a disturbance is detected (e.g., confidence variance σ(z)>0.1), the content constraint weight is increased: β (t+1) =1.2·β (t) .
[0132] Case example: The implementation of the present invention is illustrated through the case of "MNIST handwritten digit attack and defense confrontation". First, set the target interference system: black box handwritten digit classification model (CNN or Transformer architecture, the specific structure is unknown for white box). Set the white box model group: ResNet-18, DenseNet-121, VGG-16 (fine-tuned on MNIST). The input sample is a clean digital image.
[0133] The implementation process is as follows:
[0134] 1. Generate adversarial samples: Input x to the generator G, generate a perturbed image, and constrain ||zx|| ∞ ≤0.1(assuming∈ max =0.1).
[0135] 2. Discriminator training: Input x and z to D1 (ResNet) and D2 (DenseNet). Calculate cross entropy loss and Jointly optimize KL divergence
[0136] 3. Generator Optimization: Calculation Extract the feature maps of the white model group and the black box model (simulation) and align their distributions. Joint optimization
[0137] 4. Dynamic adjustment:
[0138] Monitor the black box system’s classification error rate for z. If it is lower than 80%, max Increased from 0.1 to 0.105.
[0139] If the standard deviation of the black box output confidence exceeds 0.1, the content constraint is strengthened (β increases from 0.1 to 0.12).
[0140] This technology uses domain adaptation theory to learn representations of our situational awareness and discrimination systems, inferring the performance boundaries of the enemy's situational awareness intelligent system, and generating highly transferable adversarial examples to disrupt the enemy's information flow. In practical applications, it can provide auxiliary judgment for information confrontation in areas such as drone signal recognition, drone electronic countermeasures, and situational awareness countermeasure systems.
[0141] Although the present invention has been described in detail through the above preferred embodiments, it should be understood that the above description is not intended to limit the present invention. After reading the above description, various modifications and substitutions of the present invention will become apparent to those skilled in the art. Therefore, the scope of protection of the present invention should be defined by the appended claims.
[0142] The contents not described in detail in the specification of the present invention belong to the common knowledge of professionals in this field.
Claims
1. A generative attack-defense confrontation method based on domain adaptation, which has the following characteristics: S1. Build multiple sets of heterogeneous white-box models to form a source domain model set; S2. Calculate the expected gradient of the source domain model set and infer the maximum disturbance tolerance threshold of the black box system; S3, initialize the enhanced content consistency constraint coefficient; S4. Using domain adaptation methods, align the feature space distributions of the white-box model group and the black-box system to obtain the overall domain adaptation loss; S5. Generate a dual-channel generative adversarial network based on the maximum perturbation tolerance threshold, the enhanced content consistency constraint coefficient, and the overall domain adaptation loss; S6. Training the dual-channel generative adversarial network to obtain a trained dual-channel generative adversarial network; S7. Input the original sample into the trained dual-channel generative adversarial network to obtain the adversarial sample and the number of adversarial samples that successfully mislead the black-box system. S8. Calculate the success rate of the adversarial samples attacking the black-box system based on the number of adversarial samples that successfully mislead the black-box system and the original samples. S9. Adjust the maximum disturbance tolerance threshold and enhance the content consistency constraint coefficient based on the adversarial sample, the attack success rate of the adversarial sample on the black box system, and the preset target success rate, and repeat steps S6 to S9 until the attack success rate of the adversarial sample on the black box system is greater than or equal to the preset target success rate.
2. A generative attack and defense method based on domain adaptation according to claim 1, characterized in that: The construction of multiple sets of heterogeneous white box models to form a source domain model set is specifically as follows: The source domain model set is The output joint distribution is: p i (x)=softmax(M i (x)); where M N is the Nth heterogeneous white box model, p i (x) is the white box model M i The output probability distribution of .
3. A generative attack and defense method based on domain adaptation according to claim 2, characterized in that: The maximum disturbance tolerance threshold of the black box system is specifically: The adversarial sample z satisfies ||zx|| ∞ ≤∈ max Among them, ∈ max is the maximum disturbance tolerance threshold, N is the total number of heterogeneous white box models, G i is the expected gradient amplitude of the white box model, M i (x) is the i-th heterogeneous white box model, is the expected function; min is the minimum function.
4. A generative attack and defense method based on domain adaptation according to claim 3, characterized in that: The overall domain adaptation loss is specifically: in, is the MMD loss between the i-th white box model and the black box, w i is the contribution weight, N is the total number of models, φ(x black ) is the feature space distribution of the black box model mapped to the high-dimensional feature space, φ(M i (x)) is the feature space distribution of the white box model mapped to the high-dimensional feature space.
5. A generative attack and defense method based on domain adaptation according to claim 4, characterized in that: The dual-channel generative adversarial network includes a generator and two heterogeneous discriminators; Generator, used to generate adversarial samples for deceiving two discriminators at the same time; two heterogeneous discriminators, simulating the feature extractors of heterogeneous white-box model groups and black-box systems respectively; The loss function of the generator is: Among them, x is the original sample, z is the adversarial sample of the generator, is the content consistency constraint; D1(z) and D2(z) are both heterogeneous discriminators; is the discriminator loss; is the domain adaptation loss; α is the domain adaptation loss weight, β is the content consistency loss weight; The loss function of the discriminator is: Among them, B samples is the batch size, y b is the sample label, z b =G(x b ) is the adversarial sample generated by the generator; D k is the discriminator, K is the number of discriminators; L KL is the KL loss; λ is the weight of the KL loss.
6. A generative attack and defense method based on domain adaptation according to claim 5, characterized in that: The adversarial examples generated by the generator are: ResBlock(x)=x+Conv 3×3 (ReLU(BN(Conv 3×3 (x)))) Among them, x is the original sample, z is the adversarial sample, Conv 3×3 is a 3×3 convolutional layer, BN is batch normalization, ReLU is the activation function, and ResBlock is the residual module.
7. A generative attack and defense method based on domain adaptation according to claim 6, characterized in that: The KL loss is specifically: Among them, D1(x) and D2(x) are the outputs of the discriminator, D KL is the KL divergence, and x is the input sample of the discriminator.
8. The generative attack and defense method based on domain adaptation according to claim 1, characterized in that: The success rate of the adversarial sample attacking the black box system is specifically: Among them, ASR t is the success rate of adversarial samples attacking the black box system, B is the number of adversarial samples that successfully mislead the black box, and A is the total number of test samples.
9. A generative attack and defense method based on domain adaptation according to claim 8, characterized in that: The specific method for adjusting the maximum disturbance tolerance threshold and enhancing the content consistency constraint coefficient is as follows: like Among them, ASR target is the preset target success rate, η is the adjustment factor, ASR target The success rate of the preset target; If σ(z) > σ thres , then β (t+1) = β (t) ·(1 + γ) Among them, γ is the reinforcement coefficient, β (t+1) is the enhanced content consistency constraint coefficient at the next moment, β (t) is the enhanced content consistency constraint coefficient at the current moment, σ(z) is the confidence variance of the output sample, σ thres is the confidence variance threshold of the output sample.
Citation Information
Cited By
Dynamic boundary sensing power distribution method for hybrid propulsion system
CN121638311A
A Dynamic Boundary Sensing Power Allocation Method for Hybrid Propulsion Systems
CN121638311B