Frequency domain-based feature space confrontation sample attack method and system
Through the frequency domain-based feature space adversarial sample attack method, using wavelet transform and generator network, the contradiction between the success rate of adversarial sample attack and invisibility in the prior art is solved, and efficient adversarial sample generation is achieved.
Patent Information
- Application Number
- CN202411787265.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-06
AI Technical Summary
While improving the success rate of attacks, existing adversarial sample attack methods are difficult to maintain high unawareness of adversarial samples, which affects its feasibility in practical applications.
A feature space adversarial sample attack method based on frequency domain is proposed. The image is decomposed and reconstructed through wavelet transformation, high-frequency components are removed, and the image is reconstructed using only low-frequency components. The generator network and feature space attack method are combined to construct a loss function to optimize the adversarial sample.
While achieving good attack effects, it significantly improves the invisibility of the countermeasures and enhances its feasibility in practical applications.
Smart Images

Figure CN119942259A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and system for frequency-domain-based feature space adversarial sample attack. Background Art
[0002] Computer vision tasks are an important application area of deep neural network technology, and deep neural network technology has achieved milestone performance in computer vision tasks. However, deep neural networks also face a series of security issues in the field of computer vision. Among them, the research on adversarial example attack methods has attracted widespread attention. Adversarial example attack methods refer to adding tiny and almost imperceptible perturbations to the original image, which can cause deep neural networks to produce incorrect classification results. This attack method that slightly modifies the original input to cause differences in neural network output is called adversarial example attack. With the popularity of adversarial example attacks, people have begun to realize the vulnerability of deep neural networks in the face of such attacks, which has prompted in-depth discussions and research on the robustness and security of deep learning models.
[0003] At present, adversarial sample attack methods can be divided into black-box attacks and white-box attacks according to whether the attacker knows the structure and parameters of the target network. White-box attacks are attacks performed when the attacker knows the structure and parameters of the target network, while black-box attacks are attacks performed when the attacker does not know the structure and parameters of the target network or the training data domain. In the field of white-box attacks, many methods can achieve excellent attack performance, but the image quality of adversarial samples still has room for improvement. Therefore, in the field of white-box attacks, improving the image imperceptibility of adversarial samples has become a research direction.
[0004] Since the image classification network itself classifies images in the spatial domain, most scholars have explored from the perspective of the spatial domain and have achieved some results in improving the success rate of attacks. Specifically, these methods make the feature difference between the adversarial sample and the original image larger in the spatial domain, or make the features of the adversarial sample close to the target image, and control the perturbation cost within a certain range. Although these methods have achieved good results, there are still some problems that need further exploration and resolution. These methods are often plagued by the poor visual quality of adversarial samples, which significantly affects the feasibility of adversarial attack algorithms in practical applications. Summary of the invention
[0005] The main purpose of the embodiments of the present invention is to propose a frequency-domain-based feature space adversarial sample attack method and system to ensure that the adversarial samples maintain excellent imperceptibility while achieving good attack effects.
[0006] To achieve the above purpose, an embodiment of the present invention proposes a frequency-domain-based feature space adversarial sample attack method, comprising the following steps:
[0007] Obtain image data sets and construct training data;
[0008] After preprocessing the images in the training data, a generator network is built using the Pytorch deep learning framework;
[0009] According to the generator network, images in the training data are decomposed and reconstructed by wavelet transform to obtain initialized adversarial samples;
[0010] Based on the feature space attack method in the frequency domain, the generated adversarial samples are input into the target model for training, the image features extracted by the target model are used to construct the loss function, and the final adversarial samples are obtained;
[0011] Based on the final adversarial sample, the adversarial sample attack process is completed.
[0012] In some embodiments, after preprocessing the images in the training data, the generator network built by the Pytorch deep learning framework includes the following steps:
[0013] Preprocess the dataset images and crop them to a uniform size. Specifically, first use transforms.Resize to adjust the shorter side of the image to 256 pixels, keeping the aspect ratio of the image unchanged; then use transforms.CenterCrop to crop a 224×224 square area from the center of the image; finally use transforms.ToTensor() to convert the image data into tensor format;
[0014] The generator network is built using the Pytorch deep learning framework. The model structure of the generator network includes a downsampling module, a residual module, and an upsampling module. Specifically: given an input image X∈R H×W×C , where H and W are the spatial resolutions of the input image, C is the number of input image channels; the output initialization perturbation is X adv ∈R H×W×C ;
[0015] Among them, the input image first passes through a downsampling module consisting of three downsampling blocks; each downsampling block consists of a convolutional layer, a BatchNorm layer and a ReLU activation function; the image downsampling operation is completed through the downsampling module; each residual block consists of two convolutional layers, two BatchNorm layers, a ReLU activation function and a Dropout layer; the downsampled image is feature extracted and calculated through the residual module; the upsampling block consists of a deconvolution layer, a BatchNorm layer and a ReLU activation function, and the image is restored to its original size through the upsampling module to obtain the initialized adversarial perturbation.
[0016] In some embodiments, decomposing and reconstructing the images in the training data by wavelet transform according to the generator network to obtain the initialized adversarial samples includes the following steps:
[0017] Wavelet transform is used to decompose and reconstruct the image; wavelet transform decomposes the image x into four parts, including a low-frequency component and three high-frequency components, namely x ll 、x lh 、x hl and x hh ;
[0018] The original image is reconstructed using all four components through an inverse transform method, thereby removing the high-frequency components and reconstructing the image using the low-frequency components;
[0019] By decomposing and reconstructing the original image, the high-frequency components in the original image are removed;
[0020] The reconstructed image is superimposed with the original perturbation generated by the generator network to obtain the initialized adversarial sample.
[0021] In some embodiments, the loss function consists of two parts, the first part is the image feature part, and the second part is the frequency domain limitation;
[0022] The frequency-domain-based feature space attack method inputs the generated adversarial sample into the target model for training, constructs a loss function using the image features extracted by the target model, and obtains the final adversarial sample, including the following steps:
[0023] In the first part, adversarial samples are generated from the perspective of feature similarity. By minimizing the Euclidean distance between different categories and maximizing the distance within the same category, the classification neural network maps the adversarial samples to different feature spaces.
[0024] In the second part, the original image and adversarial sample are decomposed and reconstructed using wavelet transform tools, the high-frequency components of the image are removed, and only the low-frequency components are used to reconstruct the image, thus establishing a new perturbation constraint.
[0025] According to the processing in the first part and the processing in the second part, the frequency-domain-based feature space attack method is used to iteratively optimize the adversarial sample to obtain the final adversarial sample.
[0026] In some embodiments, in the first part, generating adversarial samples from the perspective of feature similarity, by minimizing the Euclidean distance between different categories while maximizing the distance within the same category, so that the classification neural network maps the adversarial samples to different feature spaces, includes the following steps:
[0027] Given a batch of original images X, which contains N original images x, the expression of the original image X is X = [x1, x2, ..., x N ], where the i-th adversarial example is optimized The formula is:
[0028]
[0029] in,[] + represents max(.,0), x i ′ is the adversarial sample being optimized, initialized to x i , where s i, ′ i =sim(f(x′ i ),f(x i )), s′ i,j =sim(f(x′ i ),f(x j )) represents the similarity between images. The goal is to minimize the similarity between the adversarial sample and the original image and maximize the similarity between the adversarial sample and a certain category of images; f(x j ) represents the original output value obtained after the image x is input into the classification network f;
[0030] The similarity expression is:
[0031] According to the similarity expression, the optimization formula of the adversarial sample is updated and expressed as: Among them, j is the random category image number that is different from the original image category;
[0032] In the target attack scenario, it can be expressed as: Among them, t is the image sequence number of the target category.
[0033] In some embodiments, in the second part, the original image and the adversarial sample are decomposed and reconstructed respectively using a wavelet transform tool, the high-frequency component of the image is removed, and only the low-frequency component is used to reconstruct the image, and a new perturbation constraint is established, including the following steps:
[0034] The wavelet transform tool is used to decompose and reconstruct the original image and the adversarial sample, remove the high-frequency components of the image, and reconstruct the image using only the low-frequency components. The expression of this process is:
[0035] By decomposing and reconstructing the original image, the high-frequency components in the original image are removed, the main information of the original image is retained, and a new perturbation constraint is established: Among them, x is the original image, x′ is the adversarial sample;
[0036] By minimizing D(x,x′), the adversarial perturbation is limited to the high-frequency part;
[0037] in, represents the image reconstructed from the low-frequency components; L represents the low-pass filter; x ll Represents a low-frequency component.
[0038] In some embodiments, according to the processing of the first part and the processing of the second part, the adversarial sample is iteratively optimized using the feature space attack method based on the frequency domain to obtain the final adversarial sample. The optimization formula used in this step is:
[0039] Loss = λD(x i ,x′ i )+[s′ i,i -s′ i,t ] +
[0040]
[0041] Among them, λ is a hyperparameter; Loss represents the loss function.
[0042] Another aspect of the embodiment of the present invention further provides a frequency-domain-based feature space adversarial sample attack system, including:
[0043] The first module is used to obtain image data sets and construct training data;
[0044] The second module is used to pre-process the images in the training data and then build a generator network using the Pytorch deep learning framework;
[0045] The third module is used to decompose and reconstruct the images in the training data through wavelet transform according to the generator network to obtain initialized adversarial samples;
[0046] The fourth module is used for the feature space attack method based on the frequency domain. The generated adversarial samples are input into the target model for training, the loss function is constructed using the image features extracted by the target model, and the final adversarial samples are obtained.
[0047] The fifth module is used to complete the adversarial sample attack process based on the final adversarial sample.
[0048] To achieve the above objective, another aspect of an embodiment of the present invention provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above-mentioned method when executing the computer program.
[0049] To achieve the above objective, another aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.
[0050] The embodiment of the present invention also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above method.
[0051] The embodiments of the present invention include at least the following beneficial effects: the present invention provides a method and system for attacking feature space adversarial samples based on frequency domain, which constructs training data by acquiring image data sets; after preprocessing the images in the training data, a generator network is constructed by using the Pytorch deep learning framework; according to the generator network, the images in the training data are decomposed and reconstructed by wavelet transform to obtain initialized adversarial samples; based on the feature space attack method in frequency domain, the generated adversarial samples are input into the target model for training, the loss function is constructed using the image features extracted by the target model, and the final adversarial samples are obtained; according to the final adversarial samples, the adversarial sample attack process is completed. The embodiments of the present invention can ensure that the adversarial samples maintain excellent imperceptibility while achieving good attack effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present invention;
[0053] Figure 2is a flow chart of overall steps provided by an embodiment of the present invention;
[0054] Figure 3 It is a flowchart of a specific implementation process provided by an embodiment of the present invention;
[0055] Figure 4 is a schematic diagram of a generator network structure provided by an embodiment of the present invention;
[0056] Figure 5 is a schematic diagram of decomposing and reconstructing an image using wavelet transform provided by an embodiment of the present invention;
[0057] Figure 6 is a schematic diagram of a feature space attack provided by an embodiment of the present invention;
[0058] Figure 7 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the attached claims.
[0060] It is understood that the terms "first", "second", etc. used in the present invention may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0061] The terms "at least one", "multiple", "each", "any", etc. used in the present invention, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0063] The frequency-domain-based feature space adversarial sample attack method and system provided in the embodiment of the present invention relate to the field of computer technology. The frequency-domain-based feature space adversarial sample attack method provided in the embodiment of the present invention can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements a frequency-domain-based feature space adversarial sample attack method, etc., but is not limited to the above forms.
[0064] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0065] like Figure 1 FIG. 1 is a schematic diagram of an implementation environment provided by an embodiment of the present invention. Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected to a network wirelessly or wired to complete data transmission and exchange.
[0066] Server 101 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0067] In addition, the server 101 can also be a node server in the blockchain network. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm.
[0068] The terminal 102 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc. The terminal 102 may also be a vehicle-mounted terminal of various device types described above, but is not limited thereto. The terminal 102 and the server 101 may be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment of the present invention.
[0069] Based on the example Figure 1 In the implementation environment shown, an embodiment of the present invention provides a frequency-domain-based feature space adversarial sample attack method. The following is explained using the frequency-domain-based feature space adversarial sample attack method applied in the server 101 as an example. It can be understood that the method can also be applied to the terminal 102.
[0070] Reference Figure 2 , Figure 2 The flowchart of the frequency-domain-based feature space adversarial sample attack method applied to a server provided in an embodiment of the present invention, the execution subject of the method can be any of the aforementioned computer devices (including servers or terminals). Figure 2 , the method may include the following steps:
[0071] Obtain image data sets and construct training data;
[0072] After preprocessing the images in the training data, a generator network is built using the Pytorch deep learning framework;
[0073] According to the generator network, images in the training data are decomposed and reconstructed by wavelet transform to obtain initialized adversarial samples;
[0074] Based on the feature space attack method in the frequency domain, the generated adversarial samples are input into the target model for training, the image features extracted by the target model are used to construct the loss function, and the final adversarial samples are obtained;
[0075] Based on the final adversarial sample, the adversarial sample attack process is completed.
[0076] In some embodiments, after preprocessing the images in the training data, the generator network built by the Pytorch deep learning framework includes the following steps:
[0077] Preprocess the dataset images and crop them to a uniform size. Specifically, first use transforms.Resize to adjust the shorter side of the image to 256 pixels, keeping the aspect ratio of the image unchanged; then use transforms.CenterCrop to crop a 224×224 square area from the center of the image; finally use transforms.ToTensor() to convert the image data into tensor format;
[0078] The generator network is built using the Pytorch deep learning framework. The model structure of the generator network includes a downsampling module, a residual module, and an upsampling module. Specifically: given an input image X∈R H×W×C , where H and W are the spatial resolutions of the input image, C is the number of input image channels; the output initialization perturbation is X adv ∈R H×W×C ;
[0079] Among them, the input image first passes through a downsampling module consisting of three downsampling blocks; each downsampling block consists of a convolutional layer, a BatchNorm layer and a ReLU activation function; the image downsampling operation is completed through the downsampling module; each residual block consists of two convolutional layers, two BatchNorm layers, a ReLU activation function and a Dropout layer; the downsampled image is feature extracted and calculated through the residual module; the upsampling block consists of a deconvolution layer, a BatchNorm layer and a ReLU activation function, and the image is restored to its original size through the upsampling module to obtain the initialized adversarial perturbation.
[0080] In some embodiments, decomposing and reconstructing the images in the training data by wavelet transform according to the generator network to obtain the initialized adversarial samples includes the following steps:
[0081] Wavelet transform is used to decompose and reconstruct the image; wavelet transform decomposes the image x into four parts, including a low-frequency component and three high-frequency components, namely x ll 、x lh 、x hl and x hh ;
[0082] The original image is reconstructed using all four components through an inverse transform method, thereby removing the high-frequency components and reconstructing the image using the low-frequency components;
[0083] By decomposing and reconstructing the original image, the high-frequency components in the original image are removed;
[0084] The reconstructed image is superimposed with the original perturbation generated by the generator network to obtain the initialized adversarial sample.
[0085] In some embodiments, the loss function consists of two parts, the first part is the image feature part, and the second part is the frequency domain limitation;
[0086] The frequency-domain-based feature space attack method inputs the generated adversarial sample into the target model for training, constructs a loss function using the image features extracted by the target model, and obtains the final adversarial sample, including the following steps:
[0087] In the first part, adversarial samples are generated from the perspective of feature similarity. By minimizing the Euclidean distance between different categories and maximizing the distance within the same category, the classification neural network maps the adversarial samples to different feature spaces.
[0088] In the second part, the original image and adversarial sample are decomposed and reconstructed using wavelet transform tools, the high-frequency components of the image are removed, and only the low-frequency components are used to reconstruct the image, thus establishing a new perturbation constraint.
[0089] According to the processing in the first part and the processing in the second part, the frequency-domain-based feature space attack method is used to iteratively optimize the adversarial sample to obtain the final adversarial sample.
[0090] In some embodiments, in the first part, generating adversarial samples from the perspective of feature similarity, by minimizing the Euclidean distance between different categories while maximizing the distance within the same category, so that the classification neural network maps the adversarial samples to different feature spaces, includes the following steps:
[0091] Given a batch of original images X, which contains N original images x, the expression of the original image X is X = [x1, x2, ..., x N ], where the i-th adversarial example is optimized The formula is:
[0092]
[0093] in,[] + represents max(.,0), x i ′is the adversarial sample being optimized, initialized to x i , where s i, ′ i =sim(f(x′ i ),f(x i )), s′ i,j =sim(f(x′ i ),f(x j )) represents the similarity between images. The goal is to minimize the similarity between the adversarial sample and the original image and maximize the similarity between the adversarial sample and a certain category of images; f(x j ) represents the original output value obtained after the image x is input into the classification network f;
[0094] The similarity expression is:
[0095] According to the similarity expression, the optimization formula of the adversarial sample is updated and expressed as: Among them, j is the random category image number that is different from the original image category;
[0096] In the target attack scenario, it can be expressed as: Among them, t is the image sequence number of the target category.
[0097] In some embodiments, in the second part, the original image and the adversarial sample are decomposed and reconstructed respectively using a wavelet transform tool, the high-frequency component of the image is removed, and only the low-frequency component is used to reconstruct the image, and a new perturbation constraint is established, including the following steps:
[0098] The wavelet transform tool is used to decompose and reconstruct the original image and the adversarial sample, remove the high-frequency components of the image, and reconstruct the image using only the low-frequency components. The expression of this process is:
[0099] By decomposing and reconstructing the original image, the high-frequency components in the original image are removed, the main information of the original image is retained, and a new perturbation constraint is established: Among them, x is the original image, x′ is the adversarial sample;
[0100] By minimizing D(x,x′), the adversarial perturbation is limited to the high-frequency part;
[0101] in, represents the image reconstructed from the low-frequency components; L represents the low-pass filter; x ll Represents a low-frequency component.
[0102] In some embodiments, according to the processing of the first part and the processing of the second part, the adversarial sample is iteratively optimized using the feature space attack method based on the frequency domain to obtain the final adversarial sample. The optimization formula used in this step is:
[0103] Loss = λD(x i ,x′ i )+[s′ i,i -s′ i,t ] +
[0104]
[0105] Among them, λ is a hyperparameter; Loss represents the loss function.
[0106] The following describes the specific implementation process of the embodiment of the present invention in detail by taking a specific application scenario as an example:
[0107] The purpose of the present invention is to ensure that the adversarial sample maintains excellent imperceptibility while achieving good attack effects. A feature space attack method based on frequency domain is provided, such as Figure 3 As shown, the specific steps include:
[0108] Step 1: Get the image dataset and input the training data.
[0109] Step 2: Preprocess the dataset images and crop them to a uniform size to enhance the ability to extract useful information. First, use transforms.Resize(256) to resize the shorter side of the image to 256 pixels, keeping the aspect ratio of the image unchanged. Second, use transforms.CenterCrop(224) to crop a 224×224 square area from the center of the image. Finally, use transforms.ToTensor() to convert the image data into tensor format.
[0110] Step 3: Use the Pytorch deep learning framework to build a generator network. The structure of the network model is as follows: Figure 4 As shown. The network model mainly includes downsampling module, residual module and upsampling module. Specifically, given an input image X∈R H×W×C , where H and W are the spatial resolutions of the input image, and C is the number of channels of the input image. The corresponding output initialization perturbation is X adv ∈R H×W×C. The input image first passes through a downsampling module consisting of three downsampling blocks. Each downsampling block consists of a convolutional layer (Conv), a BatchNorm layer and a ReLU activation function. The image downsampling operation is completed through the downsampling module. Each residual block consists of two convolutional layers (Conv), two BatchNorm layers, a ReLU activation function and a Dropout layer. The residual module will perform feature extraction and calculation on the downsampled image. The upsampling block consists of a deconvolution layer (Transposed Convolution), a BatchNorm layer and a ReLU activation function. The upsampling module restores the image to its original size to obtain the initialized adversarial perturbation.
[0111] Step 4: The present invention uses Discrete Wavelet Transformation (DWT) to decompose and reconstruct the image. Wavelet transform is a time-frequency analysis tool that can decompose the image x into four parts, including a low-frequency component and three high-frequency components, namely: ll 、x lh 、x hl and x hh .
[0112] x ll =LxL T ,x lh =HxL T ,x hl =LxH T ,x hh =HxH T (1)
[0113] Among them, L and H are the low-pass and high-pass filters in wavelet transform. Figure 5 As shown, x ll The low-frequency part of the original image is retained, while the other three high-frequency components retain the high-frequency part of the original image.
[0114] In the inverse transform (Inverse Discrete Wavelet Transformation, IDWT), all four components are usually used to reconstruct the original image. In this work, the high-frequency components will be removed and only the low-frequency components will be used to reconstruct the image. As shown below:
[0115]
[0116] By decomposing and reconstructing the original image, the high-frequency components in the original image are removed, and only the main information of the original image is retained.
[0117] The reconstructed image is superimposed with the original perturbation generated by the generator network to obtain the initialized adversarial sample.
[0118] Step 5: Input the generated adversarial sample into the target model, where the target model is a pre-trained classification model imported into the pytorch framework, such as ResNet-50 and other models.
[0119] Step 6: Use adversarial samples for training and construct a loss function using the image features extracted by the target model.
[0120] The loss function consists of two parts. The first part is the image feature part. Traditional white-box attack methods usually generate adversarial samples by maximizing classification loss, while the purpose of feature space attack is to make images of different categories very similar in the eyes of the classification neural network. The common practice of this type of method is to directly minimize the Euclidean distance between the intermediate layer features of the original image and the target image in the target neural network. Traditional white-box attack methods usually generate adversarial samples by maximizing classification loss, that is, pushing the adversarial samples as far away from the range of their true categories as possible. In contrast, in this work, an embodiment of the present invention generates adversarial samples from the perspective of feature similarity, by minimizing the Euclidean distance between different categories while maximizing the distance within the same category, and designs a more flexible optimization scheme. In this way, the classification neural network will be misled to map the adversarial samples to different feature spaces. The specific framework is as follows: Figure 6 shown.
[0121] Given a batch of original images X, which contains H original images x, that is: X = [x1, x2, ..., x N ]. Optimize the i-th adversarial example The formula is:
[0122]
[0123] in,[] + represents max(.,0), x i ′ is the adversarial sample being optimized, initialized to x i , where s i, ' i =sim(f(x′ i ),f(x i )), s′ i,j =sim(f(x′ i ),f(x j )) represents the similarity between images. The goal is to minimize the similarity between the adversarial sample and the original image and maximize the similarity between the adversarial sample and a certain category of images. In this method, the similarity formula is expressed as:
[0124]
[0125] Combining the above formulas, the optimization formula for adversarial samples can be updated and expressed as:
[0126]
[0127] Among them, j is the sequence number of a random category image that is different from the original image category.
[0128] In the target attack scenario, it can be expressed as:
[0129]
[0130] Among them, t is the image sequence number of the target category.
[0131] The second part of the loss function is the frequency domain limitation. Although the above-mentioned feature space attack algorithm can generate adversarial samples with a high attack success rate, it also has a potential risk, that is, the adversarial perturbations may be concentrated in the salient areas of the target object, making it easy for humans to perceive. This perceptible disturbance may not only affect the concealment of the adversarial attack, but may also be subject to more defense measures in practical applications. In addition, traditional adversarial attack algorithms usually use norm constraints to limit the amplitude of the perturbation. Although this can control the size of the perturbation to a certain extent, due to the randomness of the norm constraint, it often causes the perturbation to be randomly distributed in various areas of the image, and it is impossible to effectively avoid interfering with the salient areas. In response to these problems, the present invention proposes a new perturbation constraint mechanism, which aims to limit the adversarial perturbations to the subtle and imperceptible details of the object.
[0132] Using the wavelet transform tool mentioned above, the image can be decomposed and reconstructed. The present invention uses the wavelet transform tool to decompose and reconstruct the original image and the adversarial sample respectively, remove the high-frequency components of the image, and only use the low-frequency components to reconstruct the image. As shown below:
[0133]
[0134] By decomposing and reconstructing the original image, the high-frequency components in the original image are removed, and only the main information of the original image is retained. At the same time, a new perturbation constraint can be established through the above process:
[0135]
[0136] Where x is the original image and x′ is the adversarial sample. By minimizing D(x,x′), the adversarial perturbation is limited to the high-frequency part, thus ensuring that the adversarial perturbation is not easily perceived.
[0137] Combining the above two parts, the optimization formula of the frequency domain-based feature space attack method proposed by the present invention is as follows:
[0138] Loss = λD(x i ,x′ i )+[s′ i,i -s′ i,t ] + (9)
[0139]
[0140] Among them, λ is a hyperparameter. Through this formula, the adversarial sample can achieve a high attack success rate while limiting the adversarial perturbation to the high-frequency area, keeping it as imperceptible as possible. The size of λ is a hyperparameter that affects whether the loss function is more inclined to feature loss or frequency domain control.
[0141] The loss function is used to iteratively optimize the adversarial sample to obtain the final adversarial sample.
[0142] Step 7: After training, use the test set images to conduct the final effect test on the attack method. The attack effect is evaluated according to the evaluation indicators. The evaluation indicators that need to be calculated are: Attack Success Rate (ASR), two norm (l2), infinite norm (l ∞ ) and low-frequency components (LF). The calculation formulas of the evaluation indicators can be seen in equations 11 to 14.
[0143]
[0144] Among them, NS represents the number of images that are successfully attacked, and T represents the total number of images. The second norm and the infinite norm are used to quantify the difference between the adversarial sample and the original image, and LF is used to quantify the difference between the low-frequency components of the original image and the adversarial sample. The larger the ASR, the better the attack performance, and the smaller the second norm, the infinite norm and the LF, the better the imperceptibility of the adversarial sample.
[0145] In summary, the present invention aims at the adversarial sample attack task of the image classification system, with the aim of improving the attack performance of the adversarial sample and increasing the imperceptibility, and provides a feature space attack method based on the frequency domain. Most of the previous methods explored from the perspective of the spatial domain, and achieved some results in improving the success rate of the attack. However, these methods are often plagued by the poor visual quality of the adversarial samples, which significantly affects the feasibility of the adversarial attack algorithm in practical applications. In response to the above problems, the adversarial sample attack method of the present invention has made two innovations:
[0146] (1) By combining the generator network with iterative optimization in the feature space, the model can improve the overall efficiency while maintaining a high attack success rate.
[0147] (2) Frequency domain restrictions are introduced into the generation of adversarial samples, so that the generated adversarial samples can achieve good attack effects while maintaining excellent imperceptibility.
[0148] Another aspect of the embodiment of the present invention further provides a frequency-domain-based feature space adversarial sample attack system, including:
[0149] The first module is used to obtain image data sets and construct training data;
[0150] The second module is used to pre-process the images in the training data and then build a generator network using the Pytorch deep learning framework;
[0151] The third module is used to decompose and reconstruct the images in the training data through wavelet transform according to the generator network to obtain initialized adversarial samples;
[0152] The fourth module is used for the feature space attack method based on the frequency domain. The generated adversarial samples are input into the target model for training, the loss function is constructed using the image features extracted by the target model, and the final adversarial samples are obtained.
[0153] The fifth module is used to complete the adversarial sample attack process based on the final adversarial sample.
[0154] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0155] The embodiment of the present invention further provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned frequency-domain-based feature space adversarial sample attack method when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0156] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0157] See also Figure 7 , Figure 7 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0158] The processor 701 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.
[0159] The memory 702 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 702, and the processor 701 calls and executes the feature space adversarial sample attack method based on the frequency domain of the embodiment of the present invention;
[0160] Input / output interface 703, used to implement information input and output;
[0161] Communication interface 704, used to realize communication interaction between the device and other devices, which can be realized through wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);
[0162] A bus 705 that transmits information between the various components of the device (e.g., the processor 701, the memory 702, the input / output interface 703, and the communication interface 704);
[0163] The processor 701 , the memory 702 , the input / output interface 703 and the communication interface 704 are connected to each other in communication within the device via a bus 705 .
[0164] An embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned frequency-domain-based feature space adversarial sample attack method is implemented.
[0165] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0166] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0167] It should be noted that in various specific embodiments of the present invention, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present invention needs to obtain the user's sensitive personal information, it will obtain the user's separate permission or consent through a pop-up window or jump to a confirmation page, and after clearly obtaining the user's separate permission or consent, it will obtain the necessary user-related data for the normal operation of the embodiment of the present invention.
[0168] The embodiments described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art can appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.
[0169] Those skilled in the art will appreciate that the technical solutions shown in the figures do not limit the embodiments of the present invention and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0170] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0171] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0172] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0173] It should be understood that in the present invention, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can represent: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0174] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0175] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0176] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0177] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including multiple instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store programs.
[0178] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the embodiments of the present invention is not limited thereby. Any modification, equivalent substitution and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present invention shall be within the scope of the rights of the embodiments of the present invention.
Claims
1. A frequency-domain-based feature space adversarial sample attack method, characterized in that: The following steps are involved: Obtain image data sets and construct training data; After preprocessing the images in the training data, a generator network is built using the Pytorch deep learning framework; According to the generator network, images in the training data are decomposed and reconstructed by wavelet transform to obtain initialized adversarial samples; Based on the feature space attack method in the frequency domain, the generated adversarial samples are input into the target model for training, the image features extracted by the target model are used to construct the loss function, and the final adversarial samples are obtained; Based on the final adversarial sample, the adversarial sample attack process is completed.
2. The frequency-domain-based feature space adversarial sample attack method according to claim 1, characterized in that: After preprocessing the images in the training data, the generator network is constructed by using the Pytorch deep learning framework, including the following steps: Preprocess the dataset images and crop them to a uniform size. Specifically, first use transforms.Resize to adjust the shorter side of the image to 256 pixels, keeping the aspect ratio of the image unchanged; then use transforms.CenterCrop to crop a 224×224 square area from the center of the image; finally use transforms.ToTensor() to convert the image data into tensor format; The generator network is built using the Pytorch deep learning framework. The model structure of the generator network includes a downsampling module, a residual module, and an upsampling module. Specifically: given an input image X∈R H×W×C , where H and W are the spatial resolutions of the input image, C is the number of input image channels; the output initialization perturbation is X adv ∈R H×W×C ; Among them, the input image first passes through a downsampling module consisting of three downsampling blocks; each downsampling block consists of a convolutional layer, a BatchNorm layer and a ReLU activation function; the image downsampling operation is completed through the downsampling module; each residual block consists of two convolutional layers, two BatchNorm layers, a ReLU activation function and a Dropout layer; the downsampled image is feature extracted and calculated through the residual module; the upsampling block consists of a deconvolution layer, a BatchNorm layer and a ReLU activation function, and the image is restored to its original size through the upsampling module to obtain the initialized adversarial perturbation.
3. The frequency-domain-based feature space adversarial sample attack method according to claim 1, characterized in that: Decomposing and reconstructing the images in the training data by wavelet transform according to the generator network to obtain the initialized adversarial samples includes the following steps: Wavelet transform is used to decompose and reconstruct the image; wavelet transform decomposes the image x into four parts, including a low-frequency component and three high-frequency components, namely x ll 、x lh 、x hl and x hh ; The original image is reconstructed using all four components through an inverse transform method, thereby removing the high-frequency components and reconstructing the image using the low-frequency components; By decomposing and reconstructing the original image, the high-frequency components in the original image are removed; The reconstructed image is superimposed with the original perturbation generated by the generator network to obtain the initialized adversarial sample.
4. The frequency-domain-based feature space adversarial sample attack method according to claim 1, characterized in that: The loss function consists of two parts, the first part is the image feature part, and the second part is the frequency domain limitation; The frequency-domain-based feature space attack method inputs the generated adversarial sample into the target model for training, constructs a loss function using the image features extracted by the target model, and obtains the final adversarial sample, including the following steps: In the first part, adversarial samples are generated from the perspective of feature similarity. By minimizing the Euclidean distance between different categories and maximizing the distance within the same category, the classification neural network maps the adversarial samples to different feature spaces. In the second part, the original image and adversarial sample are decomposed and reconstructed using wavelet transform tools, the high-frequency components of the image are removed, and only the low-frequency components are used to reconstruct the image, thus establishing a new perturbation constraint. According to the processing in the first part and the processing in the second part, the frequency-domain-based feature space attack method is used to iteratively optimize the adversarial sample to obtain the final adversarial sample.
5. The frequency-domain-based feature space adversarial sample attack method according to claim 4, characterized in that: In the first part, adversarial samples are generated from the perspective of feature similarity, by minimizing the Euclidean distance between different categories while maximizing the distance within the same category, so that the classification neural network maps the adversarial samples to different feature spaces, including the following steps: Given a batch of original images X, which contains N original images x, the expression of the original image X is X = [x1, x2, ..., x N ], where the i-th adversarial example is optimized The formula is: in,[] + represents max(.,0), x i ′ is the adversarial sample being optimized, initialized to x i , where s i, ′ i =sim(f(x i ′),f(x i )), s′ i,j =sim(f(x i ′),f(x j )) represents the similarity between images. The goal is to minimize the similarity between the adversarial sample and the original image and maximize the similarity between the adversarial sample and a certain category of images; f(x j ) represents the original output value obtained after the image x is input into the classification network f; The similarity expression is: According to the similarity expression, the optimization formula of the adversarial sample is updated and expressed as: Among them, j is the random category image number that is different from the original image category; In the target attack scenario, it can be expressed as: Among them, t is the image sequence number of the target category.
6. The frequency-domain-based feature space adversarial sample attack method according to claim 4, characterized in that: In the second part, the original image and the adversarial sample are decomposed and reconstructed using wavelet transform tools, the high-frequency components of the image are removed, and only the low-frequency components are used to reconstruct the image, and a new perturbation constraint is established, including the following steps: The wavelet transform tool is used to decompose and reconstruct the original image and the adversarial sample, remove the high-frequency components of the image, and reconstruct the image using only the low-frequency components. The expression of this process is: By decomposing and reconstructing the original image, the high-frequency components in the original image are removed, the main information of the original image is retained, and a new perturbation constraint is established: Among them, x is the original image, x′ is the adversarial sample; By minimizing D(x,x′), the adversarial perturbation is limited to the high-frequency part; in, represents the image reconstructed from the low-frequency components; L represents the low-pass filter; x ll Represents a low-frequency component.
7. A frequency-domain-based feature space adversarial sample attack method according to any one of claims 4 to 6, characterized in that: According to the processing of the first part and the processing of the second part, the adversarial sample is iteratively optimized using the feature space attack method based on the frequency domain to obtain the final adversarial sample. In this step, the optimization formula used is: Among them, λ is a hyperparameter; Loss represents the loss function.
8. A frequency-domain-based feature space adversarial sample attack system, characterized in that: include: The first module is used to obtain image data sets and construct training data; The second module is used to pre-process the images in the training data and then build a generator network using the Pytorch deep learning framework; The third module is used to decompose and reconstruct the images in the training data through wavelet transform according to the generator network to obtain initialized adversarial samples; The fourth module is used for the feature space attack method based on the frequency domain. The generated adversarial samples are input into the target model for training, the loss function is constructed using the image features extracted by the target model, and the final adversarial samples are obtained. The fifth module is used to complete the adversarial sample attack process based on the final adversarial sample.
9. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Adversarial sample generation method based on discrete wavelet transform
CN111709435A
Adversarial sample detection and identification method based on frequency domain transformation
CN114912550A
Adversarial sample generation method based on frequency domain flow field attack
CN115249322A
Adversarial image reconstruction system and adversarial image reconstruction method
WO2024029669A1
Cited By
Non-cooperative unmanned aerial vehicle radiation source directional attack method based on CVE-AdvGAN
CN121356728A
A non-cooperative unmanned aerial vehicle radiation source directional attack method based on CVE-AdvGAN
CN121356728B
Metaoptimization adversarial sample generation algorithm combined with frequency domain feature fusion
CN121859958A
Threat detection method, device and equipment based on deep learning
CN121864345A