A Query-based General Adversarial Perturbation Attack Algorithm for Person Re-identification

Through a general query-based adversarial perturbation attack algorithm, combined with pixel importance sampling and spatial momentum priors, targeted adversarial samples are generated, which solves the problems of large amount of calculation and large image distortion of the existing pedestrian re-identification model in black box scenarios, achieving higher attack success rate and fewer query times.

CN115424289BActive Publication Date: 2025-07-08GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210821236.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-07-08
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

The existing pedestrian re-identification model has the problems of large amount of calculation, many query times, large image distortion and poor attack effect in black box scenarios, especially in large-scale sample attack scenarios.

Method used

A general adversarial perturbation attack algorithm based on query is adopted to generate targeted adversarial samples through a gradient estimation method combining pixel importance sampling and spatial momentum priors, and a targeted adversarial sample is used to reduce the number of queries and update the common adversarial perturbation shared by all samples.

Benefits of technology

Achieve higher attack success rate in smaller image distortion, significantly reduce the total number of queries, and further improve the attack effect in large-scale sample attack scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424289B_ABST
    Figure CN115424289B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of pedestrian re-identification, and discloses a query-based general adversarial perturbation attack algorithm for pedestrian re-identification. We expand it to have the same number of channels as the original sample image by replicating channels. γ is a parameter for controlling the importance of pixels at pedestrian positions and non-pedestrian positions (γ ∈ [0, 1]). γ is set to a relatively large value, and the pixels at pedestrian positions are identified as pixels having a greater impact on the attack effect. The calculated result W represents the importance sampling weight. By combining the query-based attack with the general adversarial perturbation attack, the present invention uses all the information obtained from all queries to update a general adversarial perturbation shared by all samples, rather than generating sample-specific perturbations for each sample through a large number of queries for each sample, significantly reducing the total number of queries for attacking the entire dataset. Moreover, in the scenario of large-scale adversarial sample attacks, the advantages of this method will be further enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian re-identification, and particularly to a query-based general adversarial perturbation attack algorithm for pedestrian re-identification. Background Art

[0002] With the development of deep learning technology, the performance of pedestrian re-identification (ReID) models has been greatly improved. However, these models also inherit the vulnerability of deep neural networks (DNNs) when facing adversarial samples, that is, by adding a tiny perturbation to the samples, existing models can be made to make misjudgments and omissions. In addition, in real application scenarios, information such as the network structure and parameters of the model is usually not public, and attackers can only obtain the input-output information of the model to guide the generation of adversarial samples. This type of attack is also called a black-box attack. Therefore, in order to explore the weaknesses and deficiencies of existing pedestrian re-identification models and thus promote the construction of more robust models, the research on adversarial attacks on pedestrian re-identification models in the black-box scenario is of great significance.

[0003] Most of the existing black-box attack methods for person re-identification models are based on transfer attacks (such as the research work of Bai et al. in 2020, "Adversarial metric attack and defense for person re-identification (TPAMI)"), which optimize the generation of adversarial samples on a surrogate model using the gradient backpropagation algorithm to attack the target model. However, due to the lack of information about the target model, when the difference between the surrogate model and the target model is large, the attack effect is often unsatisfactory, and usually a large image distortion is required to maintain the transferability of the adversarial samples. In addition, there are also a few methods that use query-based attacks (such as the research work of Li et al. in 2021, "Qair: Practical query-efficient black-box attacks for image retrieval (CVPR)"), which obtain the corresponding output information by inputting samples to the target model and use this information to estimate the gradient of the target model, thereby guiding the optimization of the generation of adversarial samples. Compared with transfer-based attacks, query-based attack methods can achieve better attack effects with relatively small image distortions. However, query-based attack methods usually require a large number of queries to obtain relatively accurate gradient estimates. A large number of query operations not only lead to a large amount of computation, but more importantly, the attack is easily detected and defended. Moreover, current methods need to generate specific adversarial perturbations for each sample separately. When thousands of samples need to be attacked, the total number of queries and the amount of computation are often unacceptable. In addition, in order to reduce the number of queries, current methods widely adopt a vector-wise gradient estimation method. However, although this gradient estimation method can reduce the number of queries to a certain extent, due to the relatively inaccurate estimated gradient, the generated adversarial samples still require relatively large image distortions to maintain a high attack success rate. Summary of the Invention

[0004] The purpose of the present invention is to provide a query-based general adversarial perturbation attack algorithm for person re-identification, aiming to improve the maintenance of a high attack success rate to solve the above problems.

[0005] To achieve the above purpose, the present invention provides the following technical solutions:

[0006] A query-based general adversarial perturbation attack algorithm for person re-identification, comprising the following steps:

[0007] Step 1: Pixel importance sampling

[0008] Step 1.1: Input the original input sample into the trained pedestrian image segmentation model to obtain the segmentation result.

[0009] Step 1.2: Calculate the importance sampling weights according to the following formula:

[0010]

[0011] where represents the binarized pedestrian image segmentation result, respectively represent the channel and spatial position of the pixel (we regard the values at the same spatial position but different channel positions in the image as different pixels). The original segmentation result is single-channel, and we expand it to the same number of channels as the original sample image by replicating the channel. γ is a parameter that controls the importance of pixels at pedestrian and non-pedestrian positions (γ ∈ [0, 1]). Set γ to a relatively large value to identify the pixels at pedestrian positions as the pixels with a greater impact on the attack effect. The calculated result represents the importance sampling weight.

[0012] Step 1.3: According to the obtained importance sampling weights of each pixel, use the polynomial distribution sampling method to sample N non-repeating pixel points on each sample image.

[0013] Step 2: Model gradient estimation

[0014] Step 2.1: Query the target model by inputting the original sample to obtain the feature vector corresponding to the original sample.

[0015] Step 2.2: Generate the adversarial sample through the following formula:

[0016] q' = clip(q + δ, 0, 1) s.t. ||δ|| ∞ ≤ ∈(2)

[0017] where q represents the original sample, δ represents the general adversarial perturbation, ||·|| ∞ represents the l ∞ norm, ∈ represents the maximum allowable value of the perturbation. The size of the general adversarial perturbation δ is the same as that of the original sample image, and its initial value is a random value with an l ∞ norm less than ∈, and its l ∞ norm size does not exceed ∈ during the subsequent update process. The adversarial sample is generated by adding the original sample and the general adversarial perturbation. The clip operation limits the added result to the range of 0 to 1 to prevent it from exceeding the pixel range of the image. q' represents the finally generated adversarial sample, and then this generated adversarial sample is also input into the target model to obtain the feature vector corresponding to the adversarial sample through query.

[0018] Step 2.3: Select a pixel sampled in Step 1, add a differential to the value at the corresponding pixel position in the adversarial sample, and then input the adversarial sample with the added differential term into the target model again to query and obtain the corresponding feature vector.

[0019] Step 2.4: Estimate the gradient of the pixel selected in Step 2.3 according to the following formula:

[0020]

[0021] where f θ represents the target model, θ is the parameter of the model, f θ (·) represents the output feature vector obtained by querying the target model, h represents the magnitude of the added differential, represents the standard basis vector, which has a value of 1 only at the position and 0 at all other positions, represents the cosine similarity loss function, and its definition is as follows:

[0022]

[0023] Calculate the gradient at the coordinate position

[0024] Step 2.5: Repeat Step 2.3 and Step 2.4 until the gradients of N sampled pixels in the sample are calculated.

[0025] Step 3: General adversarial perturbation update

[0026] Step 3.1: Perform a convolution operation on the momentum term in the coordinate range of the MI-FGSM algorithm through a Gaussian kernel to obtain the spatial momentum prior. The definition of the Gaussian kernel is as follows:

[0027]

[0028] where u, v are the coordinate positions in the Gaussian kernel, k is the size of the kernel, σ is the standard deviation, and then obtain the spatial momentum prior term at the coordinate position through the following formula:

[0029]

[0030] where m is the momentum term, its initial value is 0 and its size is the same as the input sample image. Subsequently, its value is continuously updated as the number of iterations increases and is the moving average of the historical gradients. * represents the convolution operation, and the momentum terms at adjacent positions are considered through the convolution operation to update the The momentum term at the position, after updating the momentum terms at all sampling positions, the momentum m' with spatial momentum prior is obtained.

[0031] Step 3.2: Update the direction of the gradient according to the momentum prior, and the formula is as follows:

[0032]

[0033] where μ is the momentum decay coefficient, τ represents the number of iterations, ||·||1 represents the L1 norm, and the estimated gradient at all sampling positions is updated according to the momentum prior through the above formula to obtain the updated gradient g t+1 .

[0034] Step 3.3: Update the general adversarial perturbation according to the updated gradient g t+1 The formula is as follows:

[0035]

[0036] where α is the learning step size, sign(·) represents the sign function, and the clip operation always limits the general adversarial perturbation δ within the range of ∈, and the general adversarial perturbation values at all sampling positions are updated once.

[0037] Step 3.4: Update the momentum terms at all sampling positions according to the following formula:

[0038]

[0039] Step 3.5: Each time, select a training sample, and repeat Steps 1 to 3 until all training samples are queried. Take the last updated general adversarial perturbation as the final perturbation, and add it to the test sample to be attacked to generate the final adversarial sample.

[0040] Furthermore, in Step 1, according to the importance sampling weights of each pixel obtained, N non-repeating pixel points are sampled on each sample image by using the polynomial distribution sampling method.

[0041] Furthermore, in Step 2, select a pixel sampled in Step 1, add a differential to the value at the corresponding pixel position in the adversarial sample, and then input the adversarial sample with the added differential term into the target model again to query and obtain the corresponding feature vector.

[0042] The query-based general adversarial perturbation attack algorithm for person re-identification provided by the present invention has the following beneficial effects:

[0043] (1) The present invention obtains information of the target model through querying, generates targeted adversarial samples, and can achieve a higher attack success rate under the condition of smaller distortion of the adversarial sample image.

[0044] (2) By combining query-based attacks with general adversarial perturbation attacks, the present invention uses all the information obtained from queries to update a general adversarial perturbation shared by all samples, rather than generating perturbations specific to a single sample through a large number of queries for each sample, significantly reducing the total number of queries for attacking the entire dataset. Moreover, in large-scale adversarial sample attack scenarios, the advantages of our method will be further enhanced.

[0045] (3) Through a relatively accurate gradient estimation method for coordinate ranges and by combining pixel importance sampling and spatial momentum prior to make up for the disadvantage of the number of queries, it is achieved that under the same number of queries, the adversarial samples generated by our method have less image distortion and higher attack success rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a schematic flow chart of collecting pixel samples for gradient estimation in an embodiment of the present invention;

[0047] Figure 2 is a schematic diagram of the importance acquisition process of the original input sample in an embodiment of the present invention;

[0048] Figure 3 is a schematic flow chart of the contrast estimation of the features of the original sample and the adversarial sample features in an embodiment of the present invention;

[0049] Figure 4 is a schematic flow chart of updating the general adversarial sample to generate the final adversarial sample in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] Embodiment

[0052] As Figures 1-4 shown, the query-based general adversarial perturbation attack algorithm for pedestrian re-identification provided by the embodiment of the present invention includes the following steps:

[0053] Step 1: Pixel importance sampling

[0054] Step 1.1: Input the original input sample into the trained pedestrian image segmentation model to obtain the segmentation result.

[0055] Step 1.2: Calculate the importance sampling weights according to the following formula:

[0056]

[0057] where represents the binarized pedestrian image segmentation result, represent the channel and spatial position of the pixel respectively (we regard the values at the same spatial position but different channel positions in the image as different pixels). The original segmentation result is single-channel, and we expand it to the same number of channels as the original sample image by replicating the channel. γ is a parameter that controls the importance of pixels at pedestrian positions and non-pedestrian positions (γ ∈ [0, 1]). Setting γ to a larger value, the pixels at pedestrian positions are identified as pixels that have a greater impact on the attack effect. The calculated result represents the importance sampling weight.

[0058] Step 1.3: According to the obtained importance sampling weights of each pixel, use the multinomial distribution sampling method to sample N non-repeating pixel points on each sample image.

[0059] Step 2: Model gradient estimation

[0060] Step 2.1: Query the target model by inputting the original sample to obtain the feature vector corresponding to the original sample.

[0061] Step 2.2: Generate the adversarial sample through the following formula:

[0062] q' = clip(q + δ, 0, 1) s.t. ||δ|| ∞ ≤ ∈(2)

[0063] where q represents the original sample, δ represents the general adversarial perturbation, ||·|| ∞ represents the l ∞ norm, ∈ represents the maximum allowable value of the perturbation. The size of the general adversarial perturbation δ is the same as the original sample image, and its initial value is a random value with an l ∞ norm less than ∈, and its l ∞ norm size never exceeds ∈ during the subsequent update process. The adversarial sample is generated by adding the original sample and the general adversarial perturbation. The clip operation limits the added result to the range of 0 to 1, so that it does not exceed the pixel range of the image. q' represents the finally generated adversarial sample, and then this generated adversarial sample is also input to the target model to obtain the feature vector corresponding to the adversarial sample through query.

[0064] Step 2.3: Select a pixel sampled in Step 1, add a differential to the value at the corresponding pixel position in the adversarial sample, and then input the adversarial sample with the added differential term into the target model again to query and obtain the corresponding feature vector.

[0065] Step 2.4: Estimate the gradient of the pixel selected in Step 2.3 according to the following formula:

[0066]

[0067] where f θ represents the target model, θ is the parameter of the model, f θ (·) represents the output feature vector obtained by querying the target model, h represents the magnitude of the added differential, represents the standard basis vector, which has a value of 1 only at the position and 0 at all other positions, represents the cosine similarity loss function, and its definition is as follows:

[0068]

[0069] Calculate the gradient at the coordinate position

[0070] Step 2.5: Repeat Step 2.3 and Step 2.4 until the gradients of N sampled pixels in the sample are calculated.

[0071] Step 3: General adversarial perturbation update

[0072] Step 3.1: Perform a convolution operation on the momentum term in the coordinate range of the MI-FGSM algorithm through a Gaussian kernel to obtain the spatial momentum prior. The definition of the Gaussian kernel is as follows:

[0073]

[0074] where u, v are the coordinate positions in the Gaussian kernel, k is the size of the kernel, σ is the standard deviation, and then obtain the spatial momentum prior term at the coordinate position through the following formula:

[0075]

[0076] where m is the momentum term, its initial value is 0 and its size is the same as the input sample image. Subsequently, its value is continuously updated as the number of iterations increases and is the moving average of the historical gradients. * represents the convolution operation, and the momentum terms at adjacent positions are considered through the convolution operation to update the The momentum term at the position, after updating the momentum terms at all sampling positions, the momentum m' with spatial momentum prior is obtained.

[0077] Step 3.2: Update the direction of the gradient according to the momentum prior, and the formula is as follows:

[0078]

[0079] where μ is the momentum decay coefficient, τ represents the number of iterations, ||·||1 represents the L1 norm, and the estimated gradients at all sampling positions are updated according to the momentum prior through the above formula to obtain the updated gradient g t+1 .

[0080] Step 3.3: Update the universal adversarial perturbation according to the updated gradient g t+1 and the formula is as follows:

[0081]

[0082] where α is the learning step size, sign(·) represents the sign function, and the clip operation always restricts the universal adversarial perturbation δ within the range of ∈, and the universal adversarial perturbation values at all sampling positions are updated once.

[0083] Step 3.4: Update the momentum terms at all sampling positions according to the following formula:

[0084]

[0085] Step 3.5: Each time, select a training sample, and repeat Step 1 to Step 3 until all training samples are queried. Take the last updated universal adversarial perturbation as the final perturbation and add it to the test sample to be attacked to generate the final adversarial sample.

[0086] In summary, when in use, the present invention obtains information of the target model through query, generates targeted adversarial samples, and can achieve a higher attack success rate under the condition of less distortion of the adversarial sample image; combines the query-based attack with the universal adversarial perturbation attack, and uses all the information obtained by query to update a universal adversarial perturbation shared by all samples, rather than making a large number of queries for each sample to generate perturbations specific to a single sample, significantly reducing the total number of queries for attacking the entire dataset, and in the large-scale adversarial sample attack scenario, the advantages of our method will be further improved; a relatively accurate gradient estimation method for the coordinate range, combined with pixel importance sampling and spatial momentum prior to make up for the disadvantage of the number of queries, achieving that under the same number of queries, the adversarial samples generated by our method have less image distortion and a higher attack success rate.

[0087] The query-based general adversarial perturbation attack algorithm for pedestrian re-identification provided in the above embodiments of the present invention includes expanding it to the same number of channels as the original sample image by replicating channels. γ is a parameter for controlling the importance of pixels at pedestrian positions and non-pedestrian positions (γ ∈ [0, 1]). γ is set to a relatively large value, and the pixels at pedestrian positions are identified as pixels that have a greater impact on the attack effect. The calculated result represents the importance sampling weight. The focus of the present invention is to combine the query-based attack with the general adversarial perturbation attack, and use all the information obtained from all queries to update a general adversarial perturbation shared by all samples, rather than generating sample-specific perturbations for each sample through a large number of queries, which significantly reduces the total number of queries for attacking the entire dataset. Moreover, in the scenario of large-scale adversarial sample attacks, the advantages of this method will be further enhanced.

[0088] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

Claims

1. A query-based general adversarial perturbation attack algorithm for pedestrian re-identification, characterized in that, It includes the following steps: Step 1: Sample a batch of pixels that have a greater impact on the attack effect according to the importance sampling strategy; Step 2: Obtain the output feature vector of the model by querying the target model, and use the finite element differential method to estimate the gradient of each pixel sampled in Step 1 one by one; Step 3: According to the gradient estimated in Step 2, use the improved coordinate range MI-FGSM algorithm with a spatial momentum prior term to update the universal adversarial perturbation, and repeat the above steps until all training samples are queried; The specific steps in Step 3 are as follows: Step 3.1: Perform a convolution operation on the momentum term in the coordinate range MI-FGSM algorithm through a Gaussian kernel to obtain the spatial momentum prior. The definition of the Gaussian kernel is as follows: ; where is the coordinate position in the Gaussian kernel, is the size of the kernel, is the standard deviation, and the coordinates are obtained through the following formula The prior term of the spatial momentum at the position: ; Among them is the momentum term, whose initial value is 0 and whose size is the same as the input sample image. Subsequently, its value is continuously updated as the number of iterations increases. It is the moving average of historical gradients represents a convolution operation. Through the convolution operation, the momentum terms at adjacent positions are taken into account to update the momentum term at the position. After updating the momentum terms at all sampling positions, the momentum with spatial momentum prior is obtained ; Step 3.2: Update the direction of the gradient according to the momentum prior. The formula is as follows: ; where μ is the momentum decay coefficient, and t represents the number of iterations. represents the L1 norm. The updated gradient g is obtained by updating the estimated gradients at all sampling positions according to the momentum prior using the above formula. t+1 ; Step 3.3: Update the universal adversarial perturbation according to the updated gradient g t+1 The formula is as follows: ; Among them is the learning step size represents the sign function The operation will always limit the universal adversarial perturbation within the range of and update the value of the universal adversarial perturbation at all sampling positions once Step 3.4: Update the momentum term at all sampling positions according to the following formula: ; Step 3.5: Select one training sample each time, and repeat Steps 1 to 3 until all training samples are queried. Take the finally updated universal adversarial perturbation as the final perturbation, and add it to the test sample to be attacked to generate the final adversarial sample.

2. The query-based general adversarial perturbation attack algorithm for pedestrian re-identification according to claim 1, characterized in that: The specific steps in Step 1 are as follows: Step 1.1: Input the original input sample into the trained pedestrian image segmentation model to obtain the segmentation result; Step 1.2: Calculate the importance sampling weight according to the following formula: ; Among them represents the binarized pedestrian image segmentation result, respectively represent the channel and spatial position of the pixel. The values at the same spatial position but different channel positions in the image are regarded as different pixels. The original segmentation result is single-channel, and it is extended to the same number of channels as the original sample image by replicating the channel is a parameter for controlling the importance of pixels at pedestrian positions and non-pedestrian positions, set as a relatively large value, the pixels at pedestrian positions are identified as pixels with a greater impact on the attack effect, and the calculated result represents the importance sampling weight.

3. The query-based general adversarial perturbation attack algorithm for pedestrian re-identification according to claim 1, wherein: The specific steps in Step 2 are as follows: Step 2.1: Query the target model by inputting the original sample to obtain the feature vector corresponding to the original sample; Step 2.2: Generate an adversarial sample according to the following formula: ; Among them represents the original sample, represents the universal adversarial perturbation, represents the norm, represents the maximum value allowed for the perturbation, and the size of the universal adversarial perturbation is the same as that of the original sample image, and its initial value is a norm less than and during subsequent update processes, the magnitude of its norm never exceeds . The adversarial sample is generated by adding the original sample and the universal adversarial perturbation, and the operation restricts the result after addition to the range of 0 to 1 so that it does not exceed the pixel range of the image, represents the finally generated adversarial sample, and then this generated adversarial sample is also input into the target model to obtain the feature vector corresponding to the adversarial sample through query; Step 2.3: Select a pixel sampled in Step 1, add a differential to the value at the corresponding pixel position in the adversarial sample, and then input the adversarial sample with the added differential term into the target model again to query and obtain the corresponding feature vector; Step 2.4: Estimate the gradient of the pixel selected in Step 2.3 according to the following formula: ; Among them represents the target model are the parameters of the model represents the feature vector of the output obtained by querying the target model represents the magnitude of the added differential represents the standard basis vector, and only the value at the position is 1, and the values at the remaining positions are all 0 represents the cosine similarity loss function, and its definition is as follows: ; Calculate the coordinate position through the above formula at the gradient .

Citation Information

Patent Citations

  • Image Semantic Segmentation Method Based on Deep Full Convolutional Network and Conditional Random Field

    AU2020103901A4

  • Pedestrian re-identification method based on multi-scale convolution feature fusion

    CN111709311A