Black box attack method and device based on few samples, equipment and medium

Through the adversarial sample generator and decoupling distillation mechanism, a synthetic sample is generated in the black box attack, and an alternative model is built, which solves the attack bottleneck under data scarcity, improves the adaptability and migability of the black box attack, and realizes efficient adversarial sample generation and attack.

CN120337987APending Publication Date: 2025-07-18BEIJING FORESTRY UNIVERSITY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510830383.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing black box attack technology is difficult to effectively carry out in medical and industrial scenarios with high data acquisition costs and privacy-sensitive privacy. In addition, the diversity of samples generated is insufficient or the training costs are high, making it difficult to accurately fit the decision boundaries of the target model, resulting in large fluctuations in the attack success rate.

Method used

Synthetic samples are generated through the adversarial sample generator diffusion model, knowledge migration is used to use the decoupled distillation mechanism, alternative models are built, and adversarial samples are generated based on the preset attack algorithm to achieve efficient attacks under a small number of samples.

Benefits of technology

Under the black box condition, the real sample is extended into synthetic samples through the adversarial sample generator, which improves the attack adaptability and migration of the alternative model, has good category discrimination ability and decision-making boundary fitting ability, reduces the overall dependence on the data set, and improves the implementation efficiency and applicability of the attack.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337987A_ABST
    Figure CN120337987A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of machine learning, and particularly relates to a black box attack method and device based on a small number of samples, equipment and a medium. The method comprises the following steps: acquiring a real sample of a target model; inputting the real sample into an adversarial sample generator to obtain a synthetic sample; based on a decoupling distillation mechanism, performing knowledge migration on the target model by using the synthetic sample to obtain a substitution model; based on a preset attack algorithm, generating an adversarial sample by using the substitution model; the adversarial sample is used for attacking the target model. By adopting the method, a small number of real samples can be expanded into large-scale synthetic samples through the adversarial sample generator according to the small number of real samples, the training bottleneck of the substitution model under data scarcity is effectively solved, the black box attack adaptability is improved, and the simulation precision of the substitution model on the output behavior of the target model is enhanced through a decoupling distillation mechanism; therefore, the method has good category discrimination capability and decision boundary fitting capability, and the attack mobility is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine learning, and particularly relates to a black-box attack method, device, equipment and medium based on a small number of samples. Background Art

[0002] With the development of artificial intelligence technology, adversarial attack technology has emerged as an important means to evaluate the security of deep learning models. Among them, black-box attacks are widely used because they do not require knowledge of the internal structure of the target model and are closer to real attack scenarios.

[0003] Traditional black-box attack techniques mainly rely on generative adversarial networks (GANs) to synthesize samples or construct surrogate models by querying the target model a large number of times. The GAN-based method generates samples that approximate the real data distribution through adversarial training, and then uses the surrogate model for transfer attacks; decision boundary search methods estimate the gradient through the finite difference method and gradually explore the classification boundary of the target model.

[0004] However, in existing black-box attack techniques, in terms of data acquisition, the GAN-based black-box attack method requires a large number of real samples labeled by the target model to train the surrogate model, while the data acquisition cost is high and privacy-sensitive in scenarios such as medical and industrial fields. Moreover, GANs are prone to mode collapse, resulting in insufficient diversity of the generated samples, and the surrogate model cannot accurately fit the decision boundary of the target model. The black-box attack method based on gradient estimation requires tens of thousands of queries to the target model to converge, and it is easy to trigger the protection mechanism of the target model in practical applications. Although data synthesis based on diffusion models can improve the quality of samples, the training cost is high, and overfitting is likely to occur on small and medium-sized datasets. In addition, traditional knowledge distillation uses a single loss function to train the surrogate model, making it difficult to capture both the class discrimination features and the global decision boundary of the target model, resulting in insufficient generalization ability of the surrogate model in low-confidence regions and large fluctuations in the attack success rate. Summary of the Invention

[0005] Based on this, in order to solve the above technical problems, it is necessary to provide a black-box attack method, device, equipment and medium based on a small number of samples that can perform efficient real data synthesis and knowledge transfer of the target model.

[0006] In a first aspect, the present application provides a black-box attack method based on a small number of samples, including: Obtaining real samples of the target model; Inputting the real samples into an adversarial sample generator to obtain synthetic samples; Based on the decoupled distillation mechanism, using the synthetic samples to perform knowledge transfer on the target model to obtain a surrogate model; Based on a preset attack algorithm, using the surrogate model to generate adversarial samples; the adversarial samples are used to attack the target model.

[0007] In one embodiment, the adversarial sample generator obtains synthetic samples through the following method: Use the forward diffusion process to gradually add Gaussian noise to the real samples to obtain noise samples; Use the reverse denoising process to gradually remove noise from the noise samples to obtain synthetic samples; The noise samples are obtained through the following formula: ; where, is the noise sample; is the time step; is the cumulative retention factor; is the real sample; is the standard Gaussian noise; is the single retention factor at time step ; is the noise intensity coefficient at time step ;

[0008] In one embodiment, using the reverse denoising process to gradually remove noise from the noise samples to obtain synthetic samples includes: Generate predicted noise based on the noise samples and their corresponding time steps; Gradually denoise the noise samples according to the predicted noise to obtain denoised samples; Use the backbone feature scaling factor and skip feature spectrum modulation to balance multi-frequency features of the denoised samples to obtain synthetic samples; The denoised samples are obtained through the following formula: ; where, is the denoised sample; is the time step; is the single retention factor; is the noise sample; is the cumulative retention factor; is the predicted noise; is the variance of the denoising process; is the random noise.

[0009] In one embodiment, using the backbone feature scaling factor and skip feature spectrum modulation to balance multi-frequency features of the denoised samples to obtain synthetic samples includes: Calculate the average feature of the denoised samples based on the channel dimension to obtain the backbone feature channel mean; Based on the backbone feature channel mean, use the backbone feature scaling factor to adjust the low-frequency feature weights of the denoised samples to obtain low-frequency modulated samples; Perform Fourier transform on the high-frequency features of the skip connection in the denoised samples, and combine with a low-pass mask to attenuate the low-frequency components to obtain high-frequency modulated samples; Fuse the low-frequency modulated samples and the high-frequency modulated samples to obtain synthetic samples; Obtain the backbone feature scaling factor through the following formula: ; where, is the backbone feature scaling factor; is the constant scaling factor; is the channel mean of the feature map of the th layer in the backbone network; is the feature map of the th layer in the backbone network; is the number of channels; is the th channel of the feature map of the th layer in the backbone network.

[0010] In one embodiment, based on the decoupled distillation mechanism, use the synthetic samples to perform knowledge transfer on the target model to obtain an alternative model, including: Input the synthetic samples into the target model to obtain the target model output; Input the synthetic samples into the initial alternative model to obtain the alternative model output; Based on the joint optimization objective, by minimizing the joint loss between the target model output and the alternative model output, adjust the model parameters of the initial alternative model to obtain the alternative model; the joint loss includes the certainty loss and the boundary consistency loss; the knowledge corresponding to the target model includes the class-specific score and the class-agnostic score.

[0011] In one embodiment, the joint optimization objective corresponds to the following steps: Use the certainty loss to fit the class-specific score, and determine the corresponding loss function according to the target model output type to adjust the model parameters of the initial alternative model; Use the boundary consistency loss to fit the class-agnostic score, and when the classification results corresponding to the target model output and the alternative model output are inconsistent, adjust the model parameters of the initial alternative model according to the boundary consistency loss function; where, if the target model output type is a hard label, the loss function is the cross-entropy loss function; if the target model output type is a probability distribution, the loss function is the KL divergence metric difference.

[0012] In one embodiment, based on a preset attack algorithm, use the alternative model to generate adversarial samples, including: Obtain the initial perturbation; the initial perturbation is the difference between the synthetic sample and the corresponding real sample; Based on a preset attack algorithm, the initial perturbation is iteratively optimized using a surrogate model to obtain adversarial samples.

[0013] In a second aspect, the present application also provides a black-box attack device based on a small number of samples, including: A sample acquisition module for acquiring real samples of the target model; An adversarial sample generator module for inputting the real samples into an adversarial sample generator to obtain synthetic samples; A surrogate model training module for performing knowledge transfer on the target model using the synthetic samples based on a decoupled distillation mechanism to obtain a surrogate model; An adversarial sample module for generating adversarial samples using the surrogate model based on a preset attack algorithm; the adversarial samples are used to attack the target model.

[0014] In a third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of any of the above-mentioned black-box attack methods based on a small number of samples are implemented.

[0015] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above-mentioned black-box attack methods based on a small number of samples are implemented.

[0016] The above-mentioned black-box attack methods, devices, equipment, and media based on a small number of samples can, based on a small number of limited real samples, without relying on the structure and parameter information of the target model, be applicable to model attack tasks under typical black-box conditions. By using an adversarial sample generator to expand a small number of real samples into a large number of synthetic samples, the bottleneck of surrogate model training under data scarcity is effectively solved, and the adaptability of black-box attacks is improved. Based on a surrogate modeling mechanism with a decomposable knowledge structure, the simulation accuracy of the surrogate model for the output behavior of the target model is enhanced through a decoupled distillation mechanism, and it has good class discrimination ability and decision boundary fitting ability, improving the transferability of attacks. Using the surrogate model to execute existing attack algorithms to generate adversarial samples, the operation path is standardized, the implementation difficulty is low, and the generalization ability is strong, and it can be widely adapted to different types of target model systems. Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 Schematic flow chart of the black-box attack method based on a small number of samples according to the present invention; Figure 2 Schematic flow chart corresponding to multi-frequency feature balancing of the denoised samples; Figure 3 Schematic sub-step flow chart of step S103; Figure 4 Composition structure diagram of the black-box attack device based on a small number of samples according to the present invention. Detailed implementation manners

[0019] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0020] In one embodiment, as Figure 1 shown, a black-box attack method based on a small number of samples is provided. In this embodiment, the method is illustrated by taking its application to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps: S101. Obtain the real samples of the target model.

[0021] The target model refers to a deep neural network model that has been trained and deployed in a specific task scenario. The attacker cannot access its internal parameters, network structure and gradient information, and can only obtain the label prediction or probability distribution response of the model to the samples through the input samples, that is, the target model is in a typical black-box state.

[0022] The real samples of the target model refer to the samples in the original dataset that have been directly processed by the target model or whose prediction results can be observed. Due to the defense mechanism of the target model, the number of real samples is usually extremely limited, and there may be only several images per class, or even only some unlabeled samples are provided.

[0023] S102. Input the real samples into the adversarial sample generator to obtain synthetic samples.

[0024] Schematically, a small number of real samples are input into an adversarial sample generator constructed based on the diffusion model and the U-Net (Convolutional Networks for Biomedical Image Segmentation) network structure, so as to obtain a large number of synthetic samples with the characteristics of the real distribution. The adversarial sample generator is a sample synthesis mechanism based on the diffusion probability model. Among them, the diffusion model essentially generates samples with extremely high visual authenticity by simulating the gradual degradation and reverse recovery process of images under Gaussian noise perturbation. Specifically, in the forward diffusion stage, Gaussian noise is gradually added to the input image samples. Each perturbation process can be expressed as a linear combination of the current image and the standard Gaussian noise. This forward diffusion process finally degrades the image completely into a state close to pure noise.

[0025] In the reverse denoising stage, that is, during the synthetic sample generation process, the trained U-Net network is used to predict and denoise the noisy images at each step, so as to gradually restore an approximate form of the original image. The noise prediction form is adopted to enhance the modeling ability and simplify the training objective. Further, during the denoising process, the U-Net network is given a specific multi-frequency feature regulation mechanism, that is, the low-frequency structure information extracted by the backbone network and the high-frequency texture details captured by the skip connection are explicitly adjusted at the same time. The backbone feature scaling factor adjusts the low-frequency feature response through normalization transformation to avoid it being ignored during the denoising process, so as to retain the global semantic structure of the image; while the skip connection features achieve frequency suppression through the spectral mask in the Fourier domain, significantly weakening the low-frequency redundant components and enhancing the high-frequency contour details, further improving the clarity and realism of the synthetic image. The finally output image is highly consistent with the target dataset, that is, the real samples in terms of distribution, and can be regarded as a "pseudo-sample" that is semantically effectively extended.

[0026] S103. Based on the decoupled distillation mechanism, use the synthetic samples to perform knowledge transfer on the target model to obtain an alternative model.

[0027] In the black-box scenario, since the internal structure of the target model is invisible, we can only indirectly supervise it by using its predicted outputs on synthetic samples, while maximizing the retention of decision-making information in the prediction signals. Schematically, the decoupled distillation mechanism adopted is based on the idea of knowledge deconstruction of class-specific scores and class-agnostic scores. Class-specific scores are used to fit the confidence level of the target model for a specific class, and class-agnostic scores are used to approximate the decision boundary of the target model for non-target classes. Specifically, the synthetic samples are input into the target model and the surrogate model respectively to obtain their corresponding outputs. During the distillation process of class-specific scores, if the output of the target model is a discrete label, the cross-entropy loss function is used to constrain the consistency of the outputs of the two models; if the output is a continuous probability distribution, the Kullback-Leibler divergence is used to measure the difference and control its fusion weight to form a deterministic distillation loss. At the same time, in the distillation path of class-agnostic scores, the classification results of the target model and the surrogate model are compared. If their predicted classes are inconsistent, the boundary consistency loss is activated to force the alignment of their probability outputs near the decision boundary, so as to enhance the generalization ability of the surrogate model for non-target classes. The above two types of losses are finally combined in a weighted form to form a joint loss, which is used to guide the iterative training of the surrogate model until its discrimination ability for synthetic samples sufficiently approximates the output behavior of the target model.

[0028] S104. Based on a preset attack algorithm, use the surrogate model to generate adversarial samples; the adversarial samples are used to attack the target model.

[0029] The surrogate model already has a good ability to approximate the decision boundary of the target model and can be attacked as a white-box model. Schematically, existing mature attack algorithms such as the Fast Gradient Sign Method (FGSM), the Basic Iterative Method (BIM), and the Projected Gradient Descent Method (PGD) can be selected. Based on the gradient perturbation optimization of the synthetic samples by the surrogate model, a set of small but significantly model-output-changing adversarial perturbations are constructed and added to the samples to generate the final adversarial samples. Due to their good transferability, adversarial samples can effectively induce the target model to output wrong classes without accessing the internal information of the target model, thus achieving the attack purpose. Optionally, if there is no initial real sample, pseudo-samples can also be generated from pure noise, and then the attack path can be constructed through the whole process; conversely, if there are a small number of real images, they can be used as guiding items to improve the quality of the generator samples and the attack success rate.

[0030] In the above black-box attack method based on a small number of samples, even when the target model is in a black-box state and its structure and parameters are inaccessible, it is still possible to construct input-output pairs for subsequent training based on its externally observable behavior, thereby providing basic data support for the establishment of the attack path and reducing the dependence of the attack on the comprehensiveness of the dataset. By introducing an adversarial sample generator, a limited number of real samples are expanded into synthetic samples with consistent distributions and trainability, effectively alleviating the problem of insufficient data required for training the surrogate model under the few-shot condition; at the same time, improving the structural integrity and class expression ability of the synthetic samples, making the subsequent model transfer training more robust and generalizable. Based on the decoupled distillation mechanism, the knowledge of the target model is disassembled into class-specific scores and class-independent scores, and the corresponding dual loss function is used to supervise and optimize the surrogate model, enabling the surrogate model to not only fit the class prediction results of the target model but also capture the distribution of its decision boundary, thereby enhancing the overall approximation ability of the surrogate model to the discrimination behavior of the original model and realizing the construction of a controllable and trainable surrogate model in the black-box scenario. Using the constructed surrogate model as an attack carrier, adversarial samples are generated using a preset attack algorithm on the premise that gradients can be obtained, achieving the transfer of the attack on the black-box model. This not only avoids the dependence on the internal gradient information of the target model but also adapts to existing mature attack methods, improving the implementation efficiency and practical applicability of the attack.

[0031] In one embodiment, the adversarial sample generator obtains synthetic samples through the following method: S11. Gradually add Gaussian noise to the real samples through the forward diffusion process to obtain noise samples.

[0032] The noise samples are obtained through the following formula: ; where, is the noise sample; is the time step; is the cumulative retention factor; is the real sample; is the standard Gaussian noise; is the single retention factor at time step ; is the noise intensity coefficient at time step .

[0033] The forward diffusion process refers to the process of gradually adding Gaussian noise to the image in the time dimension through a series of preset noise scheduling processes, so that it gradually degenerates into an approximate pure noise state at multiple time steps. This process simulates the evolution trajectory of the image from an ordered structural state to a high-entropy state, providing a training target and sampling initial conditions for subsequent reverse reconstruction. Schematically, given the initial real image At any time step the corresponding noise sample can be expressed in the following form: where represents the noise sample at time step , is the cumulative retention factor, indicating the proportion of the original image information retained at time step , and is specifically defined as: while is the single-step retention factor at the -th time step, reflecting the proportion of the undisturbed part of the image at this step, is the noise intensity coefficient at the corresponding time step, used to control the Gaussian perturbation energy introduced at this step. , representing independent and identically distributed standard Gaussian noise. Optionally, the forward diffusion process can be discretely executed in multiple steps, and at each step, noise is gradually introduced into the original image while maintaining the Markovian assumption, so that the final image loses most of its structural information and tends to a fully noisy distribution state.

[0034] S12. Use the reverse denoising process to gradually remove the noise from the noise sample to obtain a synthetic sample.

[0035] Schematically, the reverse denoising process starts from the noise sample and gradually removes the Gaussian perturbations added at each stage through a guided denoising model, finally restoring an image that is structurally and semantically similar to the original sample, thereby generating a synthetic sample. Specifically, the noise components at each time step are estimated and eliminated through a parameterized neural network. Among them, a U-Net network architecture is used to perform multi-scale feature extraction and fusion on the noise sample, and the backbone channel and skip connection mechanism are combined to improve the restoration quality. At each time step, the noise predicted by the network can be used to estimate the mean of the denoised image, and combined with the noise scheduling parameter and the estimated error weight, the reverse sampling process is finally realized to obtain a synthetic sample.

[0036] In the above method, the forward and reverse denoising processes not only maintain probability consistency in form, but also provide theoretical support for constructing a sample generator, enabling the synthetic sample to have sufficient structural fidelity and semantic integrity while maintaining data diversity, and becoming a high-quality input for subsequent surrogate model training and attack operations.

[0037] In one embodiment, using the reverse denoising process to gradually remove the noise from the noise sample to obtain a synthetic sample includes: S21. Generate predicted noise according to the noise sample and its corresponding time step.

[0038] Schematically, based on the current noise sample and the diffusion time step it is in, a neural network model is used to generate the predicted noise corresponding to this sample. Specifically, the noise prediction network adopts a diffusion model based on the U-Net network architecture, which takes the noise image and the time step as the joint input. Through the skip connection mechanism between the encoder and the decoder, context semantic information is extracted and fused at different scales, and finally the estimation of the current noise component is output, that is, the predicted noise. Exemplarily, the optimization objective of the diffusion model is to minimize the distance between the predicted noise and the real noise, based on the loss function , where is the real sample; is the standard Gaussian noise; is the noise intensity; is the time step; is the variance of the denoising process; is the real noise. By constraining the predicted noise of the diffusion model through the loss function, the noise vector is accurately predicted, significantly improving the modeling ability of the image change trend, which helps to improve the convergence stability and image restoration accuracy in the inverse diffusion process.

[0039] S22. Perform step-by-step denoising processing on the noise sample according to the predicted noise to obtain the denoised sample.

[0040] The denoised sample is obtained through the following formula: ; where is the denoised sample; is the time step; is the single retention factor; is the noise sample; is the cumulative retention factor; is the predicted noise; is the variance of the denoising process; is the random noise.

[0041] Use the predicted noise to denoise and reconstruct the noise image at the current time step to generate a new image sample . This reconstruction process is based on the reverse sampling mechanism of the diffusion model and uses the denoising formula for denoising, where represents the denoised sample at time step , is the current noise image, is the noise predicted by the diffusion model, is the single retention factor at the current time step, reflecting the intensity of noise introduction, is the cumulative retention factor, which is used to represent the retention ratio of the original image information during the entire diffusion process. is the standard deviation used to control the sampling diversity at the current step size. is the Gaussian noise for independent sampling. By iteratively executing this denoising formula, starting from the high time step gradually regress to the low time step, and finally reconstruct the denoised sample close to the original image. .

[0042] S23. Use the backbone feature scaling factor and skip feature spectrum modulation to balance the multi-frequency features of the denoised sample to obtain the synthetic sample.

[0043] Directly performing image reconstruction based on noise prediction may still lead to the loss of some frequency information. Especially when dealing with low-frequency backgrounds or high-frequency details, phenomena such as texture blurring or structural drift may occur. Schematically, a multi-frequency feature balancing mechanism is introduced to perform feature enhancement processing on the intermediate generated denoised sample, thereby further improving the structural authenticity and semantic richness of the synthetic sample. Specifically, the multi-frequency feature balancing mechanism acts on the backbone path and skip connection path of the U-Net network, and respectively conducts regulation for low-frequency information retention and high-frequency detail enhancement. In the backbone path, calculate the average response of the backbone feature map of each scale layer in the network along the channel dimension, and calculate the backbone feature scaling factor based on the feature normalization result. In the skip connection path, introduce the Fourier spectrum modulation mechanism to perform frequency domain analysis and low-pass filtering on the skip feature map to alleviate the texture anomaly problem caused by excessive high-frequency information. Fuse the two types of modulated feature maps at each scale to generate a synthetic sample with both semantic coherence and texture clarity.

[0044] In one embodiment, as Figure 2 shown, using the backbone feature scaling factor and skip feature spectrum modulation to balance the multi-frequency features of the denoised sample to obtain the synthetic sample, including: S201. Calculate the average feature of the denoised sample based on the channel dimension to obtain the backbone feature channel mean.

[0045] Schematically, capture the overall response trend of the feature maps of each layer in the backbone network in the channel dimension, providing a statistical basis for subsequent feature reweighting operations. Specifically, assume that the feature map of the th layer in the backbone network is , and its response on the th channel is , then its channel mean can be expressed as , where is the The number of channels of the layer feature map. The channel mean reflects the global average response intensity of the features of this layer at different spatial positions and is an important statistic of low-frequency semantic features.

[0046] S202. Based on the backbone feature channel mean, use the backbone feature scaling factor to adjust the low-frequency feature weights of the denoised samples to obtain low-frequency modulated samples.

[0047] The backbone feature scaling factor is obtained through the following formula: ; where is the backbone feature scaling factor; is the constant scaling factor; is the channel mean of the feature map of the th layer in the backbone network; is the feature map of the th layer in the backbone network; is the number of channels; is the th channel of the feature map of the th layer in the backbone network.

[0048] Schematically, according to the backbone feature channel mean, calculate the backbone feature scaling factor and re-weight the feature maps in the backbone path to achieve enhancement and alignment of low-frequency information. The scaling factor design is modulated in a normalized manner and combined with the user-set constant adjustment factor , and finally the backbone feature scaling factor , where is the channel mean of the feature map of the th layer in the backbone network, and represent the minimum and maximum values of the feature map in the spatial dimension respectively, is a learnable or preset constant scaling factor, usually with a value greater than 1. Through normalization, not only can extreme value perturbations be suppressed, but also the expression ability of the regions with weak responses in the feature map can be effectively improved, thus avoiding the problem that low-frequency features are gradually diluted in the deep network. The output low-frequency modulated samples have stronger structure restoration ability.

[0049] S203. Perform Fourier transform on the high-frequency features of the skip connections in the denoised samples and combine with a low-pass mask for low-frequency component attenuation to obtain high-frequency modulated samples.

[0050] Perform frequency domain modulation on the high-frequency features in the skip connections to further enhance the realism of image detail restoration. The skip connection is the key path for transmitting local information from the encoder to the decoder in the U-Net network, and its feature map usually contains rich edge, texture, and local detail information. Schematically, for the feature map in the skip connection Perform a Fourier transform to obtain its expression in the frequency domain , and then combine it with a low-pass frequency mask to modulate it and suppress low-frequency redundant components, that is , , , where the mask function , and the mask setting is based on the frequency radius to determine whether the current frequency band is in the low-frequency region . If it belongs to the low-frequency region, it is attenuated with a constant scaling factor . If it does not belong, it remains unchanged. This processing method can effectively avoid the introduction of a large number of redundant smooth signals in the jump path, retain key edge and texture responses, and improve the usability of high-frequency features in the image restoration process. The modulated jump features then constitute high-frequency modulation samples

[0051] S204. Fuse the low-frequency modulation samples and the high-frequency modulation samples to obtain a synthetic sample

[0052] Fuse the low-frequency modulation samples and the high-frequency modulation samples to generate the final synthetic image. Schematically, in the decoding stage of the U-Net network, fusion is performed in a multi-scale connection or splicing manner, or weighted fusion or an attention mechanism can also be used to enhance feature complementarity to obtain the fused synthetic sample

[0053] In the above method, the multi-frequency feature balance mechanism significantly alleviates the problems of smooth blurring and structure loss in the process of generating images by the diffusion model through a dual-path modulation method, and effectively improves the sample expressiveness and attack value while maintaining sampling stability

[0054] In one embodiment, as Figure 3 shown, based on the decoupled distillation mechanism, use the synthetic sample to perform knowledge transfer on the target model to obtain an alternative model, including S301. Input the synthetic sample into the target model to obtain the target model output

[0055] Schematically, input the synthetic sample into the target model to obtain its output result , where represents the current synthetic sample set. The target model, as a knowledge provider, has invisible internal parameters and structures, and indirectly reflects the learned class discrimination knowledge through its output on the given input. The output form of the target model includes discrete class labels and continuous probability distributions. Among them, the discrete class label only returns the predicted class number of each input sample; the continuous probability distribution returns the confidence prediction for each class S302. Input the synthetic sample into the initial alternative model to obtain the alternative model output

[0056] Input the same synthetic samples into an initialized surrogate model to obtain the predicted output of the samples under the current parameters of the surrogate model. The structure of the initial surrogate model can be customized by the attacker. Usually, a deep convolutional network with high training efficiency and strong expressive ability is selected as the initial architecture. In its initial state, it cannot accurately reproduce the behavior of the target model and needs to be guided by a supervision mechanism to align its output.

[0057] S303. Based on the joint optimization objective, adjust the model parameters of the initial surrogate model by minimizing the joint loss between the output of the target model and the output of the surrogate model to obtain the surrogate model. The joint loss includes a deterministic loss and a boundary consistency loss. The knowledge corresponding to the target model includes class-specific scores and class-agnostic scores.

[0058] Adjust the parameters of the surrogate model based on the joint optimization objective function. Its core goal is to minimize the joint loss between the output of the target model and the output of the surrogate model, thereby prompting the surrogate model to approximate the discriminant boundary and class decision of the target model in the sample space. The joint loss consists of two parts, namely the deterministic loss and the boundary consistency loss, corresponding to the two types of knowledge structures of class-specific scores and class-agnostic scores contained in the output of the target model.

[0059] Among them, the deterministic loss is used to measure the output consistency of the surrogate model on the target class, mainly guiding it to align on the semantic dimension of the output target to reconstruct the explicit behavior of the target model in various class discrimination tasks on the surrogate model. It is the main learning source of the class-specific scores of the surrogate model.

[0060] To improve the generalization ability of the surrogate model on the decision boundary, a boundary consistency loss is introduced to fit the response behavior of the target model in the non-target class space. Specifically, when the predicted classes of the target model and the surrogate model for the same input sample are inconsistent, it indicates that there are differences in the discriminant boundaries of the two near the sample, and the surrogate model needs to be guided to correct its boundary tendency by minimizing the distance from the output of the target model.

[0061] Schematically, the two types of losses are combined in a weighted form to construct the joint loss function of the joint optimization objective. , Among them, is the deterministic loss function; is the boundary consistency loss function; and is a user - set weight coefficient, which is used to balance the attention degree to the fitting of explicit categories and the consistency of implicit boundaries. Further, during the training process of the surrogate model, the parameters of the surrogate model are updated through the back - propagation joint loss function until the output of the surrogate model on the training samples converges stably. The finally output surrogate model has good fitting ability and boundary migration ability, and can play the role of the attacked proxy in the black - box scenario where gradient information is not accessible.

[0062] In one embodiment, the joint optimization objective corresponds to the following steps: S31: Use the deterministic loss to fit the class - specific scores, and determine the corresponding loss function according to the output type of the target model to adjust the model parameters of the initial surrogate model.

[0063] Among them, if the output type of the target model is a hard label, the loss function is the cross - entropy loss function.

[0064] If the output type of the target model is a probability distribution, the loss function is the KL - divergence metric difference.

[0065] Schematically, execute the deterministic loss path to guide the surrogate model to fit the direct output behavior of the target model in the class dimension. Such an output is defined as the class - specific score, that is, the classification confidence information of the target model for each input sample regarding the target class. The design of the deterministic loss is adaptively divided according to the actual output type of the target model. Specifically, if the target model is a classifier and only returns the predicted class label of each sample, that is, a hard label, then the cross - entropy loss function (Cross Entropy Loss) is used as the error metric standard between the surrogate model output and the target label, The cross - entropy loss function has good differentiability and stability and is a commonly used supervision index in classification tasks. Correspondingly, if the output of the target model is the probability distribution on each class, that is, it provides more fine - grained confidence information, then the KL - divergence is used as the loss function to quantify the difference between the surrogate model and the target model at the output distribution level, . The two adaptation strategies ensure that the surrogate model can make full use of the existing output information of the target model. Regardless of its granularity level, it can construct an effective supervision signal to achieve accurate alignment of class - specific behaviors. During the training process, based on this loss, the current surrogate model parameters are updated by back - propagation, thereby gradually enhancing its discriminative accuracy on the target class. Optionally, the deterministic joint loss function is , through the weight factor fuse the two adaptation strategy scenarios.

[0066] S32. Fit the class-agnostic scores using the boundary consistency loss. When the classification results corresponding to the output of the target model and the output of the surrogate model are inconsistent, adjust the model parameters of the initial surrogate model according to the boundary consistency loss function.

[0067] To improve the generalization ability of the surrogate model in the non-target class discrimination region, a boundary consistency loss path is further introduced to correspondingly fit the class-agnostic score information implicit in the target model. The class-agnostic score reflects the configuration characteristics of the target model on the decision boundary between different classes and is part of the implicit knowledge, which usually cannot be directly obtained from the label itself. To capture this information, a conditional trigger mechanism is introduced. When the classification results of a synthetic sample in the target model and the surrogate model are inconsistent, that is, there is a situation where the predicted classes are inconsistent, it indicates that the sample may be near the decision boundary, which is a region with a large difference between the two models. Then the boundary consistency loss function is activated to penalize the difference between the outputs of the two models. The KL divergence can still be used for quantification, which is used to measure the difference in the output probability distributions of the same input sample in the two models. The loss function is , where represents the indicator function, which takes the value of 1 only when the predicted classes of the outputs of the two models are inconsistent, and 0 otherwise, ensuring that this loss branch only acts in the region where the decision boundaries are misaligned, thereby improving the training efficiency and avoiding ineffective penalties.

[0068] Through the collaborative optimization of the above two sub-loss functions, the above method can fully extract the explicit and implicit knowledge signals from the output results of the target model, and respectively guide the surrogate model parameter optimization process along the two dimensions of class and boundary, and finally train a surrogate model with high consistency in functional behavior with the target model. The establishment of this surrogate model lays the foundation for subsequent adversarial sample generation under white-box conditions, and also ensures that the attack samples have sufficient deception ability after being transferred to the target model.

[0069] In one embodiment, based on a preset attack algorithm, use the surrogate model to generate adversarial samples, including: S41. Obtain the initial perturbation; the initial perturbation is the difference between the synthetic sample and the corresponding real sample.

[0070] Schematically, set the initial perturbation state of the adversarial attack process, that is, determine the starting point of perturbation optimization. The source of the initial perturbation is set according to the availability of samples in the attack environment. Exemplarily, if the attacker can obtain a small number of real samples, the pixel difference between the small number of real samples and the corresponding synthetic samples generated by the diffusion model is used as the initial perturbation, that is , where represents the synthetic sample generated by the diffusion model, is its corresponding real sample, and the difference between the two reflects the potential deviation during the model generation process. This perturbation can be regarded as the starting point of the attack path and has a certain semantic guidance. Exemplarily, if there is no real sample, that is, a pure generative attack, a small amount of random noise can be used as the initial perturbation to ensure that the attack optimization has a starting point in the search space and meets the iterative conditions of algorithms such as gradient descent.

[0071] S42. Based on a preset attack algorithm, use the surrogate model to iteratively optimize the initial perturbation to obtain an adversarial sample.

[0072] Under the supervision of the surrogate model, use a preset attack algorithm to iteratively optimize the initial perturbation to generate the final adversarial sample. The preset attack algorithm can be any conventional gradient attack method applicable to the white-box scenario, including FGSM, BIM, PGD, etc. The preset attack algorithm maximally changes the classification result of the model for the input sample under the premise of constraining the perturbation size, thereby inducing it to produce incorrect predictions.

[0073] Exemplarily, the preset attack algorithm is PGD. In each iteration, calculate the gradient of the loss function of the surrogate model on the current sample, adjust the perturbation value along the gradient direction, and use a projection operation to limit the perturbation amplitude not to exceed the preset norm constraint, thereby maintaining the perceptual control of the adversarial sample. The process of iteratively optimizing the perturbation is where represents the perturbation value after the th iteration, is the step size, is the loss function, usually taking cross-entropy or target-oriented loss, represents projecting the perturbation into the perturbation space with a specified radius of . This process gradually enhances the aggressiveness of the perturbation through multiple iterations and finally generates an adversarial sample , where is the optimized final perturbation. The adversarial sample maintains similarity with the original synthetic sample in the perceptual space but causes a discriminative deviation in the surrogate model and even the target model, constituting a transferable attack sample.

[0074] The above method successfully constructs adversarial samples using a controllable surrogate model in the black-box non-differentiable scenario. Without relying on the gradient information of the target model, an efficient and stable attack strategy can be achieved. The adversarial samples can be used to test the robustness of the model or implement interference on the target model in the actual deployment environment.

[0075] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0076] Based on the same inventive concept, an embodiment of the present application also provides a black-box attack device based on a small number of samples for implementing the above-mentioned black-box attack method based on a small number of samples. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the black-box attack device based on a small number of samples provided below can refer to the limitations on the black-box attack method based on a small number of samples in the above text, and will not be repeated here.

[0077] In an exemplary embodiment, as Figure 4 shown, a black-box attack device based on a small number of samples is provided, including: A sample acquisition module 401, configured to acquire real samples of a target model; An adversarial sample generator module 402, configured to input the real samples into an adversarial sample generator to obtain synthetic samples; A surrogate model training module 403, configured to perform knowledge transfer on the target model using the synthetic samples based on a decoupled distillation mechanism to obtain a surrogate model; An adversarial sample module 404, configured to generate adversarial samples using the surrogate model based on a preset attack algorithm; the adversarial samples are used to attack the target model.

[0078] In one of the embodiments, it further includes: A forward diffusion module, configured to gradually add Gaussian noise to the real samples through a forward diffusion process to obtain noise samples; A reverse denoising module, configured to gradually remove noise from the noise samples through a reverse denoising process to obtain synthetic samples.

[0079] In one of the embodiments, it further includes: A noise prediction module, configured to generate predicted noise according to the noise samples and their corresponding time steps; A denoising module, configured to perform step-by-step denoising processing on a noise sample according to predicted noise to obtain a denoised sample; A frequency balancing module, configured to perform multi-frequency feature balancing on the denoised sample by using a backbone feature scaling factor and skip feature spectrum modulation to obtain a synthesized sample.

[0080] In one embodiment, it further includes: A feature map channel module, configured to calculate an average feature of the denoised sample based on the channel dimension to obtain a backbone feature channel mean value; A low-frequency balancing module, configured to adjust the low-frequency feature weight of the denoised sample by using the backbone feature scaling factor based on the backbone feature channel mean value to obtain a low-frequency modulated sample; A high-frequency balancing module, configured to perform Fourier transform on the high-frequency features of the skip connection in the denoised sample and combine a low-pass mask to attenuate the low-frequency components to obtain a high-frequency modulated sample; A high-low frequency fusion module, configured to fuse the low-frequency modulated sample and the high-frequency modulated sample to obtain a synthesized sample.

[0081] In one embodiment, it further includes: A target model scheduling module, configured to input the synthesized sample into a target model to obtain a target model output; An alternative model scheduling module, configured to input the synthesized sample into an initial alternative model to obtain an alternative model output; A training module, configured to adjust the model parameters of the initial alternative model based on a joint optimization objective by minimizing the joint loss between the target model output and the alternative model output to obtain an alternative model; the joint loss includes a certainty loss and a boundary consistency loss; the knowledge corresponding to the target model includes class-specific scores and class-agnostic scores.

[0082] In one embodiment, it further includes: An initial perturbation module, configured to obtain an initial perturbation; the initial perturbation is the difference between the synthesized sample and the corresponding real sample; An iterative optimization perturbation module, configured to iteratively optimize the initial perturbation by using the alternative model based on a preset attack algorithm to obtain an adversarial sample.

[0083] In one embodiment, there is provided a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0084] In one embodiment, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0085] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0086] The above embodiments only represent several implementation manners of the embodiments of the present application. The descriptions are relatively specific and detailed, but should not be construed as a limitation on the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.

Claims

1. A black-box attack method based on a small number of samples, characterized in that, The method includes: Obtaining real samples of the target model; Inputting the real samples into an adversarial sample generator to obtain synthetic samples; Based on a decoupled distillation mechanism, using the synthetic samples to perform knowledge transfer on the target model to obtain a surrogate model; Based on a preset attack algorithm, using the surrogate model to generate adversarial samples; the adversarial samples are used to attack the target model.

2. The method according to claim 1, characterized in that The adversarial sample generator obtains the synthetic samples through the following method: Using a forward diffusion process to gradually add Gaussian noise to the real samples to obtain noise samples; Using a reverse denoising process to gradually remove noise from the noise samples to obtain synthetic samples; Through the following formula, the noise samples are obtained: ; wherein, is the noise sample; is the time step; is the cumulative retention factor; is the true sample; is the standard Gaussian noise; is at time step the single retention factor; is at time step the noise intensity coefficient.

3. The method according to claim 2, wherein The step of using the reverse denoising process to gradually remove noise from the noise samples to obtain synthetic samples includes: Generating predicted noise according to the noise samples and their corresponding time steps; Performing gradual denoising on the noise samples according to the predicted noise to obtain denoised samples; Using a backbone feature scaling factor and jump feature spectrum modulation to perform multi-frequency feature balancing on the denoised samples to obtain the synthetic samples; Through the following formula, the denoised samples are obtained: ; Among them, is the denoised sample; is the time step; is the single retention factor; is the noise sample; is the cumulative retention factor; is the predicted noise; is the variance of the denoising process; is the random noise.

4. The method according to claim 3, wherein The step of using a backbone feature scaling factor and jump feature spectrum modulation to perform multi-frequency feature balancing on the denoised samples to obtain the synthetic samples includes: Calculating the average feature of the denoised samples based on the channel dimension to obtain the backbone feature channel mean; Based on the backbone feature channel mean, using the backbone feature scaling factor to adjust the low-frequency feature weights of the denoised samples to obtain low-frequency modulation samples; Performing Fourier transform on the high-frequency features of the skip connections in the denoised samples and combining with a low-pass mask to attenuate the low-frequency components to obtain high-frequency modulation samples; Fusing the low-frequency modulation samples and the high-frequency modulation samples to obtain the synthetic samples; Through the following formula, the backbone feature scaling factor is obtained: ; Among them, is the backbone feature scaling factor; is the constant scaling factor; is the channel mean of the feature map of the th layer in the backbone network; is the feature map of the th layer in the backbone network; is the number of channels; is the th channel of the feature map of the th layer in the backbone network.

5. The method according to claim 1, wherein The step of, based on a decoupled distillation mechanism, using the synthetic samples to perform knowledge transfer on the target model to obtain a surrogate model includes: Inputting the synthetic samples into the target model to obtain the target model output; Inputting the synthetic samples into an initial surrogate model to obtain the surrogate model output; Based on a joint optimization objective, by minimizing the joint loss between the target model output and the surrogate model output, adjusting the model parameters of the initial surrogate model to obtain the surrogate model; the joint loss includes a certainty loss and a boundary consistency loss; the knowledge corresponding to the target model includes class-specific scores and class-agnostic scores.

6. The method according to claim 5, wherein The joint optimization objective corresponds to the following steps: Using the certainty loss to fit the class-specific scores and determining the corresponding loss function according to the target model output type to adjust the model parameters of the initial surrogate model; Using the boundary consistency loss to fit the class-agnostic scores, and when the classification results corresponding to the target model output and the surrogate model output are inconsistent, adjusting the model parameters of the initial surrogate model according to the boundary consistency loss function; Among them, when the output type of the target model is a hard label, the loss function is the cross-entropy loss function; when the output type of the target model is a probability distribution, the loss function is the KL divergence metric difference.

7. The method according to claim 1, characterized in that The generating of adversarial samples using the surrogate model based on a preset attack algorithm includes: Obtaining an initial perturbation; the initial perturbation is the difference between the synthesized sample and the corresponding real sample; Based on a preset attack algorithm, iteratively optimizing the initial perturbation using the surrogate model to obtain the adversarial sample.

8. A black-box attack device based on a small number of samples, characterized in that, The device includes: A sample acquisition module, configured to acquire real samples of the target model; An adversarial sample generator module, configured to input the real samples into an adversarial sample generator to obtain synthesized samples; A surrogate model training module, configured to perform knowledge transfer on the target model using the synthesized samples based on a decoupled distillation mechanism to obtain a surrogate model; An adversarial sample module, configured to generate adversarial samples using the surrogate model based on a preset attack algorithm; the adversarial samples are used to attack the target model.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Image classification black box attack method and system based on confrontation knowledge distillation

    CN116452956A

  • Black box mobility countermeasure attack method based on GAN

    CN117057408A

  • Robustness test method and system for aerial remote sensing image target detection model

    CN117746194A

  • Wireless communication signal confrontation sample generation method based on diffusion model

    CN119067193A

  • Method for accelerating high-quality MRI (Magnetic Resonance Imaging) image reconstruction based on TC-KANReccon model

    CN119152057A