Black box antagonism rate-distortion attack method and system based on feature proxy and medium
By utilizing the feature proxy method in the picture compression system, calculating the gradient and updating the perturbation, the problem of insufficient black box attack performance is solved, and more effective black box attack is achieved, and the number of coded bits and attack effect is improved.
Patent Information
- Application Number
- CN202510547963.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-08
AI Technical Summary
现有的基于深度学习的图片压缩系统在黑盒攻击场景下,攻击性能不足,无法有效探索其鲁棒性,传统的白盒攻击方法无法满足实际需求。
By entering the picture into the original model to obtain the middle-level features, and using the alternative model to perform feature proxying, calculating gradients and updating perturbations, I-FGSM method is used to improve attack performance and realize black box attacks.
Accurately estimating the gradient of the target model significantly improves the effect of black box attacks, improves the number of coded bits, and enhances the effectiveness of the attack.
Smart Images

Figure CN120281918A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of black-box adversarial rate-distortion attacks, and specifically, to a black-box adversarial rate-distortion attack method, system, and medium based on feature proxy. Background Art
[0002] Traditional image compression technologies achieve data compression by manually designing each module. Currently, due to the booming development of deep learning technologies, deep learning-based image compression models have been proposed and shown to outperform the optimal traditional image compression technologies, attracting extensive attention from researchers. However, past experience has proven that neural networks themselves are vulnerable and extremely susceptible to adversarial perturbations. Adversarial attacks were first proposed in the field of image classification, where attackers carefully design imperceptible tiny perturbations to add to the input images, causing the classification model to misclassify the images into other categories. Specifically in an image compression system, the attacker also injects tiny perturbations into the input images, aiming to disrupt the normal utility of the compression system, that is, to increase the compression bits or reduce the reconstruction quality.
[0003] Currently, the black-box attack on deep learning-based image compression methods still uses the traditional surrogate model-based method. In the existing technology, the 2023 journal "Toward Robust Neural Image Compression: Adversarial Attack and Model Finetuning" used the FGSM method to explore the robustness of deep learning-based image compression systems. The 2023 journal "Manipulation Attacks on Learned Image Compression" proposed a surrogate model called DCT-Net for black-box attacks. However, the above methods are white-box attack scenarios and do not meet the requirements of the actual attack scenario, and the attack performance of the black-box attack cannot effectively achieve the attack, so neither can effectively explore the robustness of existing deep learning-based image compression systems.
[0004] The patent application document CN118940258A discloses a method and system for generating adversarial samples for black-box attacks based on Gaussian homotopy optimization, including: constructing a smoothing factor adjustment model, initializing the smoothing factor adjustment model and the continuous path learning model; the smoothing factor adjustment model outputs a smoothing factor s; the continuous path learning model outputs a sample perturbation x; using the smoothing factor s to perform Gaussian smoothing on the original adversarial sample generation objective function to obtain a number of Gaussian homotopy functions G(x, s), and taking the minimization of the mean of the number of Gaussian homotopy functions G(x, s) as the goal, jointly training the smoothing factor adjustment model and the continuous path learning model to efficiently generate adversarial samples. However, this patent cannot completely solve the existing technical problems and cannot meet the requirements of the present invention. Summary of the Invention
[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide a black-box adversarial rate distortion attack method, system and medium based on feature proxy.
[0006] According to the black-box adversarial rate distortion attack method based on feature proxy provided by the present invention, it includes:
[0007] Step 1: Input the picture x into the original model to obtain two middle-level features: the latent feature and the super-prior feature
[0008] Step 2: Input the picture x into the surrogate model encoder g′ a , to obtain the output y′; input the latent feature into the surrogate model h′ a , to obtain the output z′; then input the latent feature and the super-prior feature into the integrated entropy encoder m′ of the surrogate model a , to obtain the output Then, through the formula obtain the number of bits R′ corresponding to the feature y ; y ;
[0009] Step 3: According to y′, z′ and R′ y , calculate the gradient and update the perturbation, so as to improve the encoding bit number of the target model and improve the attack performance of the black-box attack.
[0010] Preferably, the attacker's attack purpose is as follows:
[0011]
[0012] Among them, R adv represents the number of bits of the perturbed sample, δ is the input perturbation, x is the original picture, y advand z adv is a middle-level feature of the compression system; R represents the number of bits; g a represents the encoder; h a represents the hyperprior encoder; θ ga represents the parameters of the encoder; θ ha represents the parameters of the hyperprior encoder; ε is a constant that constrains the perturbation size.
[0013] Preferably, the perturbation is updated in the manner of I-FGSM:
[0014]
[0015] where t represents the number of iterations, represents the gradient value of the loss function R with respect to the input image x; δ (0) represents the initial perturbation; α represents the perturbation update step size for each iteration; represents the adversarial example in the (t + 1)-th round.
[0016] Preferably, for the black-box attack based on feature proxy, the formula for gradient calculation is as follows:
[0017]
[0018] where,
[0019] represents the gradient value of the probability distribution in the surrogate model with respect to the input image x; m′ a represents the integrated entropy model in the surrogate model, which is used to calculate the probability distribution of the feature ; g′ a represents the encoder in the surrogate model, with the input being the image x and the output being the feature y′; h′ a represents the hyperprior encoder in the surrogate model, with the input being the feature and the output being the hyperprior feature z′; represents the probability distribution of the surrogate model feature ; represents the feature of the surrogate model, which is the quantized version of the output y′ of the surrogate model encoder g′ a ; represents the feature of the surrogate model, which is the quantized version of the output z′ of the surrogate model hyperprior encoder h′ a ;
[0020] According to the black-box adversarial rate-distortion attack system based on feature proxy provided by the present invention, it includes:
[0021] Module M1: Input the picture x into the original model to obtain two middle-level features: the latent feature and the hyperprior feature
[0022] Module M2: Input the picture x into the alternative model encoder g′ a , and obtain the output y′; input the latent feature into the alternative model h′ a , and obtain the output z′; then input the latent feature and the hyperprior feature into the integrated entropy encoder m′ of the alternative model a , and obtain the output Then obtain the number of bits R′ corresponding to the feature y through the formula ; y ;
[0023] Module M3: Calculate the gradient and update the perturbation according to y′, z′ and R′ y , so as to increase the number of encoding bits of the target model and improve the attack performance of the black-box attack.
[0024] Preferably, the attacker's attack purpose is as follows:
[0025]
[0026] Among them, R adv represents the number of bits of the perturbed sample, δ is the input perturbation, x is the original picture, y adv and z adv are the middle-level features of the compression system; R represents the number of bits; g a represents the encoder; h a represents the hyperprior encoder; represents the parameters of the encoder; represents the parameters of the hyperprior encoder; ε is a constant that restricts the perturbation size.
[0027] Preferably, update the perturbation in the way of I-FGSM:
[0028]
[0029] Among them, t represents the number of iterations, represents the gradient value of the loss function R with respect to the input picture x; δ (0) represents the initial perturbation; α represents the perturbation update step size for each iteration; represents the adversarial sample in the (t + 1)-th round.
[0030] Preferably, for the black-box attack based on feature proxy, the formula for gradient calculation is as follows:
[0031]
[0032] Among them,
[0033] represents the gradient value of the probability distribution of the input image x in the surrogate model; m' a represents the integrated entropy model in the surrogate model, which is used to calculate the probability distribution of the feature g' a represents the encoder in the surrogate model. The input is the image x, and the output is the feature y'; h' a represents the hyperprior encoder in the surrogate model. The input is the feature and the output is the hyperprior feature z'; represents the probability distribution of the surrogate model feature ; represents the feature of the surrogate model, which is the quantized version of the output y' of the surrogate model encoder g' a ; represents the feature of the surrogate model, which is the quantized version of the output z' of the surrogate model hyperprior encoder h' a .
[0034] According to the computer-readable storage medium storing a computer program provided by the present invention, when the computer program is executed by a processor, the steps of the black-box adversarial rate distortion attack method based on feature proxy are implemented.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] Regarding white-box adversarial attacks, that is, the situation where the attacker knows all the knowledge of the compression model has been explored by researchers, but the more practical black-box scenario still remains blank. The present invention aims to fill this gap. By adopting the structure of feature proxy statistics, the problem of insufficient performance of existing black-box attacks is solved, accurate gradient estimation is achieved, and thus the attack in the black-box attack scenario is effectively realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:
[0038] Figure 1 is the block diagram of the attack system module;
[0039] Figure 2 is the attack flow chart. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.
[0041] Embodiment
[0042] According to the degree of information that the attacker has about the victim model, adversarial attacks can be roughly divided into white-box attacks and black-box attacks. A white-box attack means that the attacker fully masters the information of the victim model and can access its model architecture and gradient information. A black-box attack, on the other hand, restricts this privilege. The attacker can only access the input and output of the victim model and cannot access the gradient information of the model. Black-box attacks are more in line with the actual situation, so the present invention focuses on black-box attacks against deep learning-based image compression systems.
[0043] Existing deep learning-based image compression systems mainly include an encoder, a decoder, an entropy model, a hyperprior encoder, a hyperprior decoder, and quantization.
[0044] Encoder: In the field of deep learning, the encoder usually consists of a convolutional neural network or a Transformer, and is used to map the image from the pixel space to a latent feature space that is more convenient for compression.
[0045] Decoder: Also composed of a neural network, it receives the output features from entropy decoding, which is approximately the inverse function of the encoder, and is used to map the latent features back to the pixel space.
[0046] Quantization: Quantize the output features of the encoder to reduce the data distribution and facilitate calculating the probability distribution of the feature data.
[0047] Entropy model: Calculate the probability distribution of the output features of the quantization module, and thus use it for entropy coding to calculate the encoded bitstream.
[0048] Hyperprior encoder: The input is the output features of the encoder, and it is used to further extract the information in the features as auxiliary information to save the number of encoded bits. At the same time, the auxiliary information also needs to be turned into a bitstream through entropy coding and transmitted in the channel.
[0049] Hyperprior decoder: It is the inverse process of the hyperprior encoder.
[0050] The present invention focuses on the attack on the number of encoded bits, that is, only involves three parts: the encoder, quantization, and entropy model.
[0051] The present invention provides a black-box adversarial rate-distortion attack method based on feature proxy, including:
[0052] Step 1: The data passes through the original model; the data is a picture, and the picture is input into the original model;
[0053] Step 2: Output intermediate features; The picture x is input into the original model, and the original model outputs two intermediate features: the latent feature and the hyperprior feature
[0054] Step 3: The features and the data pass through the alternative model; The picture x, the features and are input into the alternative model. Specifically, the picture x is input into the encoder g′ of the alternative model a , and the output y′ is obtained. The features are input into the alternative model h′ a , and the output z′ is obtained. Then and are input into the integrated entropy encoder m′ of the alternative model a , and the output Then passes through the formula: to obtain the number of bits R′ corresponding to the feature y y ;
[0055] Step 4: Calculate the gradient and update the perturbation; Calculate the gradient: The specific formula implementation is shown in formula (3);
[0056] Step 5: Output the perturbation; Update the perturbation using formula (2), where in formula (2) is a general expression, and specifically here it is
[0057] The attacker's attack purpose is as follows:
[0058]
[0059] Among them, R adv represents the number of bits of the perturbed sample, δ is the input perturbation, x is the original picture, y adv and z adv are respectively the intermediate features of the compression system; R represents the number of bits; g a represents the encoder; h a represents the hyperprior encoder; represents the parameters of the encoder; represents the parameters of the hyperprior encoder; ε is a constant that constrains the perturbation size.
[0060] The present invention updates the perturbation in the way of I-FGSM:
[0061]
[0062] Among them, t represents the number of iterations, represents the gradient value of the loss function R with respect to the input image x; δ (0) represents the initial perturbation; α represents the perturbation update step size for each iteration; represents the adversarial example in the (t + 1)-th round.
[0063] Therefore, the key to the entire system is to obtain a gradient value that is as accurate as possible.
[0064] The schematic diagram of the proxy-based black-box attack proposed by the present invention and the traditional black-box attack is as Figure 1 shown. The following will introduce the two methods and report how this technology brings about an improvement in the attack performance.
[0065] The traditional surrogate-model-based black-box attack directly generates adversarial perturbations on the local model. However, due to the differences between the surrogate model and the target model, the generated perturbations cannot achieve the expected effects on the target model.
[0066] The present invention inputs the intermediate-level features accessible to the attacker into each module in the surrogate model.
[0067] For the encoder: the input is the real image x.
[0068] For the hyperprior encoder: the input is the output features quantized from the original model instead of the output features from the surrogate-model encoder
[0069] For the entropy model: the input is the output features from the original-model quantization module and the output of the hyperprior quantization module instead of the output data of the modules from the surrogate model.
[0070] Therefore, for the traditional black-box attack, the formula for gradient calculation is as follows:
[0071]
[0072] Among them, the superscript ′ indicates that this feature belongs to the surrogate model, and no superscript indicates that this feature belongs to the original model. R′ y represents the number of bits of the feature y of the surrogate model; represents the probability distribution of the surrogate-model feature ; y′ represents the feature y′ of the surrogate model, which is the output of the surrogate-model encoder g′ a ; represents the surrogate-model feature It is the surrogate model encoder g′ a to output the quantized version of y′; represent the features of the surrogate model It is the surrogate model hyperprior encoder h′ a for the quantized version of the output z′; z′ represents the feature z′ of the surrogate model, which is the output of the surrogate model hyperprior encoder h′ a output.
[0073] For the feature proxy-based black box attack proposed by the present invention, the formula for gradient calculation is as follows:
[0074]
[0075] where
[0076] represents the gradient value of the probability distribution in the surrogate model with respect to the input image x; m′ a represents the integrated entropy model in the surrogate model, which is used to calculate the probability distribution of the feature g′ a represents the encoder in the surrogate model, with the input being the image data x and the output being the feature y′; h′ a represents the hyperprior encoder in the surrogate model, with the input being the feature and the output being the hyperprior feature z′.
[0077] Therefore, by using the middle-level features from the original model as proxy features and inputting them into the surrogate model, the attacker can obtain gradient values closer to those of the target model, thereby more effectively increasing the number of encoding bits of the target model and greatly enhancing the attack performance of the black box attack.
[0078] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be considered as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the methods or the structures within the hardware component.
[0079] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A black-box adversarial rate-distortion attack method based on feature proxy, characterized in that Including: Step 1: Input image x into the original model to obtain two middle-level features: the latent feature and the hyperprior feature Step 2: Input the picture x into the alternative model encoder g′ a , and obtain the output y′; input the latent feature into the alternative model h′ a , and obtain the output z′; then input the latent feature and the hyperprior feature into the integrated entropy encoder m′ of the alternative model a , and obtain the output Then, is used to obtain the number of bits R′ corresponding to the feature y through the formula ; y ; Step 3: According to y′, z′ and R′ y , calculate the gradient and update the perturbation, so as to increase the number of encoded bits of the target model and improve the attack performance of the black-box attack.
2. The method for black-box adversarial rate-distortion attack based on feature proxy according to claim 1, wherein, The attack objectives of the attacker are as follows: Among them, R adv represents the number of bits of the perturbed sample, δ is the input perturbation, x is the original image, y adv and z adv are the middle-level features of the compression system; R represents the number of bits; g a represents the encoder; h a represents the hyperprior encoder; represents the parameters of the encoder; represents the parameters of the hyperprior encoder; ε is a constant that constrains the perturbation magnitude.
3. The method for black-box adversarial rate-distortion attack based on feature proxy according to claim 2, wherein Update the perturbation in the I-FGSM manner: where t represents the number of iterations, represents the gradient value of the loss function R with respect to the input image x; δ (0) represents the initial perturbation; α represents the perturbation update step size for each iteration; represents the adversarial example in the (t + 1)-th round.
4. The method for black-box adversarial rate-distortion attack based on feature proxy according to claim 3, wherein For the black-box attack based on feature proxy, the formula for gradient calculation is as follows: Among them, Denotes the probability distribution in the surrogate model Gradient value of the input image x; m′ a Denotes the integrated entropy model in the surrogate model, used to calculate the feature Probability distribution of g′ a Denotes the encoder in the surrogate model, with input as the image x and output as the feature y′; h′ a Denotes the hyperprior encoder in the surrogate model, with input as the feature And output as the hyperprior feature z′; Denotes the probability distribution of the surrogate model feature ; Denotes the feature of the surrogate model, which is the quantized version of the output y′ of the surrogate model encoder g′ a ; Denotes the feature of the surrogate model, which is the quantized version of the output z′ of the surrogate model hyperprior encoder h′ a ; 5. A black-box adversarial rate distortion attack system based on feature proxy, characterized in that, Including: Module M1: Inputs image x into the original model to obtain two middle-level features: the latent feature and the hyperprior feature Module M2: Inputs image x into the alternative model encoder g′ a , and obtains the output y′; inputs the latent feature into the alternative model h′ a , and obtains the output z′; then inputs the latent feature and the hyperprior feature into the integrated entropy encoder m′ of the alternative model a , and obtains the output Then obtains the number of bits R′ corresponding to the feature y through the formula ; y ; Module M3: According to y′, z′, and R′ y , calculate the gradient and update the perturbation, thereby improving the number of encoded bits of the target model and enhancing the attack performance of the black-box attack.
6. The black-box adversarial rate-distortion attack system based on feature proxy according to claim 5, characterized in that, The attack objectives of the attacker are as follows: Among them, R adv represents the number of bits of the perturbed sample, δ is the input perturbation, x is the original image, y adv and z adv are the intermediate features of the compression system; R represents the number of bits; g a represents the encoder; h a represents the hyperprior encoder; represents the parameters of the encoder; represents the parameters of the hyperprior encoder; ε is a constant that constrains the perturbation magnitude.
7. The black-box adversarial rate-distortion attack system based on feature proxy according to claim 6, characterized in that Update the perturbation in the I-FGSM manner: where t represents the number of iterations, represents the gradient value of the loss function R with respect to the input image x; δ (0) represents the initial perturbation; α represents the perturbation update step size for each iteration; represents the adversarial example in the (t + 1)-th round.
8. The black-box adversarial rate distortion attack system based on feature proxy according to claim 7, characterized in that, For the black-box attack based on feature proxy, the formula for gradient calculation is as follows: Among them, represents the probability distribution in the surrogate model of the gradient value of the input image x; m' a represents the integrated entropy model in the surrogate model, used to calculate the feature probability distribution g' a represents the encoder in the surrogate model, with the input being the image x and the output being the feature y'; h' a represents the hyperprior encoder in the surrogate model, with the input being the feature and the output being the hyperprior feature z'; represents the surrogate model feature probability distribution; represents the feature of the surrogate model, which is the quantized version of the output y' of the surrogate model encoder g' a ; represents the feature of the surrogate model, which is the quantized version of the output z' of the surrogate model hyperprior encoder h' a .
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the black-box adversarial rate-distortion attack method based on feature proxy according to any one of claims 1 to 4.
Citation Information
Patent Citations
Black box attack confrontation sample generation method and system based on Gaussian homotopy optimization
CN118940258A