Query-based adversarial sample generation method

Through the adversarial sample generation method with hourglass patch and feature space loss optimization, the problems of global information neglect and query overhead in the existing technology are solved, efficient adversarial sample generation is achieved, and the attack success rate and security of AI systems are improved.

CN120277709APending Publication Date: 2025-07-08YANSHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510336350.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing adversarial sample generation methods fail to fully consider global information when constructing patches, and the optimization target is single, resulting in low attack success rate and high query overhead.

Method used

The hourglass patch is used and combined with feature space loss. By optimizing the patch content and location, adversarial samples are generated, and the optimal solution is searched using random and color filling strategies, and multi-layer network losses are calculated to improve the attack success rate and reduce the number of queries.

Benefits of technology

It realizes adversarial sample generation with high attack success rate and low query overhead, which is suitable for non-directed and directed attacks, improving the robustness and security of AI systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277709A_ABST
    Figure CN120277709A_ABST
Patent Text Reader

Abstract

The invention discloses a query-based adversarial sample generation method, which belongs to the field of computer vision and comprises the following steps of: initializing, modeling for an adversarial attack process and defining and assigning required data; loss is calculated, wherein the loss comprises attack loss and feature space loss; the attack loss can be directly calculated by an original sample and a current sample; the calculation of the feature space loss needs to store the original sample and the current sample into the deep layer of the target model, calculate the weight parameter of each layer, and finally calculate the weight sum of each layer to obtain the total loss; updating the patch, sampling the content and position of the patch during each query, constructing an adversarial sample, judging whether the patch is superior to the previous optimal patch or not according to the calculated current loss, and executing the next query; and finally, repeatedly querying and optimizing to obtain a final confrontation sample. According to the method, the problems of low attack performance and high query overhead in the field of resisting sample patch attacks are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a query-based adversarial sample generation method. Background Art

[0002] In the context of the rapid development of science and technology today, the security of AI systems has gradually become a challenge that needs to be solved urgently. Especially in practical applications, AI systems may become the target of malicious attackers, causing them to produce erroneous outputs or unpredictable behaviors. Adversarial sample attacks are one form of such attacks. They deceive AI systems into outputting erroneous results by imposing small perturbations on the input data. Therefore, adversarial sample attacks pose a serious threat to the security of AI systems.

[0003] To address this threat, researchers have conducted a large number of studies on adversarial sample defense methods, aiming to improve the robustness and stability of AI systems. Compared with adversarial samples generated by transfer attacks, adversarial samples generated by query attacks usually remain visually imperceptible because they only require minor adjustments to very few pixels. However, there are relatively few studies on adversarial sample query attacks, and existing studies mainly focus on constructing rectangular or square local patches. For example, the HPA method searches for the position and shape of rectangular patches by using Metropolis-Hastings sampling. Although this method is effective, it has a high query overhead. The MPA method uses reinforcement learning technology to optimize the position and content of the patch. However, both the HPA and MPA methods only use patches of a single color, which significantly limits the success rate of the attack. Although the TPA method proposed later has made some progress in solving the monochrome patch problem, it still faces several challenges: on the one hand, the constructed patch still only focuses on local information and ignores the larger global information; on the other hand, the optimization target is limited to a single attack loss function, and other factors that affect the attack success rate and query overhead are not fully considered.

[0004] The method proposed in this study not only improves the success rate of non-directional and directed attacks on adversarial samples, but also significantly reduces the query overhead, showing great potential. Summary of the invention

[0005] The technical problem to be solved by the present invention is to provide a query-based adversarial sample generation method, which constructs and optimizes an hourglass-shaped patch, uses feature space loss in query optimization to further consider more factors affecting attack performance, and gradually queries and generates adversarial samples, thereby achieving the purpose of improving the success rate of non-directional and directional attacks of adversarial samples and saving query overhead.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0007] A query-based adversarial example generation method, comprising the following steps:

[0008] S1. Initialize the adversarial patch as hourglass-shaped, and obtain the adversarial example by superimposing the adversarial patch on the original sample;

[0009] S2. Calculate the total loss of the adversarial attack at initialization or at the end of each query;

[0010] S3. Update the patch according to the total loss of the adversarial attack to make the loss reach the optimal query loss;

[0011] S4. Repeatedly query until the final adversarial example is obtained.

[0012] A further improvement of the technical solution of the present invention lies in: in S1, the adversarial patch is expressed by the patch content η and a binary mask m representing the position and shape of the patch; first, the patch content η and the patch position and shape m are initialized; the patch content η is initialized to random RGB values, and the patch position and shape m is initialized to an hourglass-shaped patch matrix, each line segment of the hourglass-shaped patch is 1, and the rest are 0; the hourglass-shaped patch matrix is formed by connecting two identical right-angled triangles with opposite vertices;

[0013] Initialize the patch position and shape m, and m is designed according to the hourglass-shaped patch, in the following way:

[0014] or

[0015] These two ways are respectively the vertical and horizontal placement ways of the hourglass-shaped patch, where (p, q) is the intersection coordinate of the two triangles in the designed hourglass-shaped patch, and l is the length of the right-angled side of the triangle.

[0016] A further improvement of the technical solution of the present invention lies in: in S2, at initialization or at the end of each query, calculate the total loss of the adversarial attack to facilitate subsequent updating of the optimal patch content η and patch position m, specifically including the following steps:

[0017] S2.1 Calculate the attack loss L1;

[0018] When the attack method is non-targeted attack, select the margin-based loss Margin as the attack loss; when the attack method is targeted attack, select the cross-entropy loss CE as the attack loss;

[0019] S2.2 At each query, store the original sample x and the adversarial sample x t in the deep network structure of the surrogate model to provide a basis for subsequent calculation of the feature space loss;

[0020] S2.3 Calculate the feature space loss L2;

[0021] S2.3.1 Calculate the network layer weights for each layer;

[0022] S2.3.2 Sum the above weights and the loss weights for each layer to obtain the feature space loss;

[0023] S2.4 Calculate the total adversarial attack loss;

[0024] A further improvement of the technical solution of the present invention lies in that: in S2.3.1, the calculation of the network layer weight coefficient for each layer is as follows:

[0025] ω k = log(k + 1)

[0026] where log(·) represents the logarithmic function, k is the current network layer number, and ω k is the weight coefficient of the kth layer.

[0027] A further improvement of the technical solution of the present invention lies in that: in S2.3.2, the calculation method of the feature space loss is as follows:

[0028]

[0029] where k represents the current model layer number, B is the first layer of the selected target layer, E is the last layer of the model; F k (x) is the output feature of x at the kth layer.

[0030] A further improvement of the technical solution of the present invention lies in that: in S2.4, the calculation method of the total adversarial attack loss is as follows:

[0031] L = L1 + λ × L2

[0032] where λ is the weight coefficient of the feature space loss.

[0033] A further improvement of the technical solution of the present invention lies in that: in S3, during each query, the patch content η and the patch position shape m are sampled according to a certain strategy, and then the sampled η and m are superimposed with the original sample to generate an adversarial sample. Then, S2 is executed again to calculate the loss. If the current loss is better than the previous optimal query loss, it is determined that the current sampling result is also better than the previous optimal sampling result, and the current η, m, and the current loss are recorded as the optimal loss;

[0034] S3 specifically includes the following steps:

[0035] S3.1 During iteration, use the random search strategy and the color filling strategy to update the position and content of the patch to fully search the available solution space to find the optimal patch and adversarial sample;

[0036] S3.2 Optimize the patch position once every v - 1 updates of the patch content;

[0037] S3.3 In each iteration, use the latest patch to construct adversarial examples, then input them into the target model to calculate the new attack loss. If the current loss is better than the previous optimal query loss, it is determined that the current sampling result is also better than the previous optimal sampling result, and record the current patch content η, patch position m, and the current loss as the optimal loss.

[0038] A further improvement of the technical solution of the present invention lies in: In S3.2, regarding the update of the patch content, four color filling methods are designed, including:

[0039] 1) Random filling: Each monochromatic strip of the patch has randomness;

[0040] 2) Normal filling: Generate RGB data through Gaussian distribution;

[0041] 3) Bernoulli filling: Generate random RGB values through Bernoulli distribution;

[0042] 4) Uniform discrete filling: A numerical filling method based on values generated by uniform discrete distribution.

[0043] Due to the adoption of the above technical solution, the technical progress achieved by the present invention is:

[0044] 1. The present invention adopts the hourglass patch adversarial attack method to generate adversarial examples with high attack success rate and low query times; the hourglass patch is formed by connecting two identical right - angled triangles with opposite vertices; on the premise of ensuring smallness and continuity, it has the ability to perturb globally and achieves excellent attack performance.

[0045] 2. The present invention proposes the feature space loss, which further improves the attack success rate and saves the query budget by calculating the feature differences between the original samples and adversarial examples in the deep layer of the model.

[0046] 3. As a simple and efficient adversarial patch attack method, the present invention is based on a novel hourglass patch and preferably achieves the purpose of high attack success rate and low query overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a flowchart of a query - based adversarial example generation method in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The following further describes the present invention in detail with reference to the drawings and embodiments:

[0049] In the embodiments of the present invention, as Figure 1As shown in the figure, a query-based adversarial sample generation method, the attack method used mainly includes three aspects: initialization, updating the patch, and calculating the loss. Initialization is to define and assign values to the variables designed in the present invention; updating the patch refers to optimizing the patch so that the generated adversarial sample can successfully attack the target model; calculating the loss is to calculate the difference between the original sample and the current sample, providing a metric standard for query iteration.

[0050] The attack method of the present invention mainly includes the following steps:

[0051] S1. Initialize the adversarial patch, and obtain the adversarial sample by superimposing the adversarial patch on the original sample;

[0052] The adversarial patch is represented by the patch content η and a binary mask m representing the position and shape of the patch; therefore, first, the patch content η and the patch position and shape m need to be initialized; for the patch content η, it is initialized to random RGB values, and for the patch position and shape m, it is initialized to an hourglass-shaped patch matrix, where each line segment of the hourglass-shaped patch is 1 and the rest of the positions are 0;

[0053] Initialize the patch position and shape m, and m is designed according to the hourglass-shaped patch in the following way:

[0054] or

[0055] These two ways are respectively the vertical and horizontal placement methods of the hourglass-shaped patch, where (p, q) is the intersection coordinate of the two triangles in the designed hourglass-shaped patch, and l is the length of the right-angled side of the triangle.

[0056] S2. Calculate the total loss of the adversarial attack at the end of initialization or each query;

[0057] At the end of initialization or each query, it is necessary to calculate the total loss of the adversarial attack in order to update the optimal patch content η and patch position m subsequently, improving the misleading ability of the adversarial sample. Specifically, it includes the following steps:

[0058] S2.1 Calculate the attack loss L1;

[0059] When the attack method is non-targeted attack, select the margin-based loss Margin as the attack loss; when the attack method is targeted attack, select the cross-entropy loss CE as the attack loss;

[0060] S2.2 During each query, store the original sample x and the adversarial sample x t in the deep network structure of the surrogate model, providing a basis for calculating the feature space loss subsequently;

[0061] S2.3 Calculate the feature space loss L2;

[0062] S2.3.1 Calculate the network layer weights for each layer;

[0063] Calculate the weight coefficients for each layer of the network as follows:

[0064] ω k = log(k + 1)

[0065] where log(·) represents the logarithmic function, k is the current network layer number, and ω k is the weight coefficient for the k-th layer;

[0066] S2.3.2 Sum the above weights and the loss weights for each layer to obtain the feature space loss;

[0067] Calculate the feature space loss as follows:

[0068]

[0069] where k represents the current model layer number, B is the first layer of the selected target layer, E is the last layer of the model; F k (x) is the output feature of x at the k-th layer;

[0070] S2.4 Calculate the total adversarial attack loss;

[0071] The calculation method of the total adversarial attack loss is as follows:

[0072] L = L1 + λ × L2

[0073] where λ is the weight coefficient of the feature space loss.

[0074] S3. Update the patch to make the loss reach the optimal query loss;

[0075] At each query, sample the patch content η and the patch position shape m according to a certain strategy, then superimpose the sampled η and m with the original sample to generate an adversarial sample, and execute S2 again to calculate the loss. If the current loss is better than the previous optimal query loss, it is determined that the current sampling result is also better than the previous optimal sampling result, and record the current η, m, and the current loss as the optimal loss;

[0076] S3 specifically includes the following steps:

[0077] S3.1 During iteration, use the random search strategy and the color filling strategy to update the position and content of the patch to fully search the available solution space to find the optimal patch and adversarial sample;

[0078] S3.2 Since the patch content space is much larger than the patch position space, the update of patch content is more frequent than that of patch position; for every v - 1 updates of patch content, the patch position is optimized once; for the update of patch content, the present invention designs four different patch color filling strategies to explore the influence of different filling methods on the attack success rate and query overhead;

[0079] Regarding the update of patch content, four color filling methods are designed, and the implementation methods are as follows:

[0080] 1) Random filling: A color generation method full of randomness, where each monochromatic strip of the patch has randomness;

[0081] 2) Normal filling: A smoother generation method that generates RGB data through Gaussian distribution; the color fluctuations in this way will be more natural, and the generated distribution characteristics can be flexibly controlled, such as adjusting the standard deviation or mean;

[0082] 3) Bernoulli filling: A simple generation method for discrete values that generates random RGB values through Bernoulli distribution;

[0083] 4) Uniform discrete filling: A numerical filling method based on uniform discrete distribution. Although the generated results are similar to Bernoulli distribution in that they are both discrete, the uniform discrete distribution is more direct and simple, without more variations or complexities, and is different from the preferences of normal distribution and Bernoulli distribution.

[0084] S3.3 In each iteration, use the latest patch to construct an adversarial example, and then input it into the target model to calculate the new attack loss. If the current loss is better than the previous optimal query loss, it is determined that the current sampling result is also better than the previous optimal sampling result, and record the current patch content η, patch position m, and the current loss as the optimal loss.

[0085] S4. Repeatedly query until the final adversarial example is obtained.

[0086] Repeatedly query until the attack on the target model is successful or the maximum query number is reached to obtain the final adversarial example.

[0087] Next, we will combine Figure 1 to illustrate the process of generating adversarial examples of the present invention:

[0088] S1. Initialization: Model the adversarial attack process as xadv = (1 - m)×x + m×η; at the same time, the shape of the hourglass patch needs to be initialized and defined. Simply put, the hourglass patch consists of two identical right - angled triangles with opposite vertices.

[0089] S2. Update the patch: At each query, sample m and η, construct the adversarial sample xt, and then determine whether it is better than the previous optimal patch according to the calculated current loss, and perform the next query.

[0090] S3. Calculate the loss: The loss includes two aspects: the attack loss and the feature space loss. The attack loss can be directly calculated from the original sample and the current sample. The calculation of the feature space loss requires storing the original sample and the current sample in the deep layer of the target model, calculating the weight parameters of each layer, and finally calculating the weight sum of each layer.

[0091] S4. Query repeatedly until the final adversarial sample is obtained.

[0092] In summary, in view of the problems of the existing attack methods that do not consider global perturbation and single-loss attacks, the present invention proposes a query-based adversarial sample generation method, achieving a relatively simple, efficient and high attack success rate attack strategy. It should be noted that the present invention is applicable to both non-targeted attacks and targeted attacks.

Claims

1. A query-based adversarial example generation method, characterized in that: Including the following steps: S1. Initialize the adversarial patch as hourglass-shaped, and obtain the adversarial sample by superimposing the adversarial patch on the original sample; S2. Calculate the total loss of the adversarial attack at initialization or at the end of each query; S3. Update the patch according to the total loss of the adversarial attack to make the loss reach the optimal query loss; S4. Repeatedly query until the final adversarial sample is obtained.

2. The query-based adversarial sample generation method according to claim 1, wherein: In S1, the adversarial patch is expressed by the patch content η and a binary mask m representing the position and shape of the patch; first, initialize the patch content η and the patch position and shape m; initialize the patch content η as random RGB values, and initialize the patch position and shape m as an hourglass-shaped patch matrix, where each line segment of the hourglass-shaped patch is 1 and the rest are 0; the hourglass-shaped patch matrix is composed of two identical right triangles with opposite vertices connected; Initialize the patch position and shape m, and m is designed according to the hourglass-shaped patch in the following way: These two ways are respectively the vertical and horizontal placement methods of the hourglass-shaped patch, where (p,q) is the intersection coordinate of the two triangles in the designed hourglass-shaped patch, and l is the length of the right-angled side of the triangle.

3. The query-based adversarial example generation method according to claim 1, wherein: In S2, calculate the total loss of the adversarial attack at initialization or at the end of each query to update the optimal patch content η and patch position m subsequently, which specifically includes the following steps: S2.1 Calculate the attack loss L1; When the attack method is non-directed attack, select the edge-based loss Margin as the attack loss; when the attack method is directed attack, select the cross-entropy loss CE as the attack loss; S2.2 At each query, store the original sample x and the adversarial sample x t in the deep network structure of the surrogate model to provide a basis for subsequent calculation of the feature space loss; S2.3 Calculate the feature space loss L2; S2.3.1 Calculate the network layer weights of each layer; S2.3.2 Sum the above weights and the loss weights of each layer to obtain the feature space loss; S2.4 Calculate the total loss of the adversarial attack.

4. The query-based adversarial example generation method according to claim 3, wherein: In S2.3.1, calculate the weight coefficients of the network layer of each layer as: ω k = log(k + 1) where log(·) represents the logarithmic function, k is the current network layer number, and ω k is the weight coefficient of the k-th layer.

5. The query-based adversarial sample generation method according to claim 3, wherein: In S2.3.2, calculate the feature space loss in the following way: Among them, k represents the current model layer, B is the first layer of the selected target layer, and E is the last layer of the model; F k (x) is the output feature of x at the k-th layer.

6. The query-based adversarial sample generation method according to claim 3, characterized in that: In S2.4, calculate the total loss of the adversarial attack in the following way: L = L1 + λ × L2 Where λ is the weight coefficient of the feature space loss.

7. The query-based adversarial example generation method according to claim 1, wherein: In S3, at each query, sample the patch content η and the patch position and shape m according to a certain strategy, then superimpose the sampled η and m on the original sample to generate an adversarial sample, and execute S2 again to calculate the loss. If the current loss is better than the previous optimal query loss, it is determined that the current sampling result is also better than the previous optimal sampling result, and record the current η, m, and the current loss as the optimal loss; S3 specifically includes the following steps: S3.1 During iteration, use the random search strategy and color filling strategy to update the position and content of the patch to fully search the available solution space to find the optimal patch and adversarial sample; S3.2 Optimize the patch position once every v - 1 times of updating the patch content; In S3.3, at each iteration, the latest patch is used to construct adversarial samples, which are then input into the target model to calculate the new attack loss. If the current loss is better than the previous optimal query loss, it is determined that the current sampling result is also better than the previous optimal sampling result, and the current patch content η, patch position m, and the current loss are recorded as the optimal loss.

8. The query-based adversarial example generation method according to claim 7, wherein: In S3.2, regarding the update of the patch content, four color filling methods are designed, including: 1) Random filling: Each monochromatic stripe of the patch has randomness. 2) Normal filling: RGB data is generated through a Gaussian distribution. 3) Bernoulli filling: Random RGB values are generated through a Bernoulli distribution. 4) Uniform discrete filling: A numerical filling method based on values generated from a uniform discrete distribution.