Adversarial Sample Generation Method, Anti-Attack Detection Method, Device, and Electronic Equipment
By generating and correcting the pixel mean of the initialization patch, synthesize adversarial samples and performing detection, the problem of adversarial samples in deep learning models is solved, and the detection capability and detection accuracy of the detection model are improved.
Patent Information
- Application Number
- CN202310271058.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2043-03-16
AI Technical Summary
Adversarial samples of deep learning models are easily blocked, resulting in poor detection results for anti-attack ability.
By obtaining the initialization patch, identifying the mean of its pixel points and correcting it, generating an adversarial patch, and finally synthesize an adversarial sample with the target image, inputting it into the model to be detected for identification, and determining the anti-attack results based on preset classification information.
It improves the quality and robustness of the adversarial samples, enhances the display effect and diversity of the adversarial patches, and improves the ability of the detection model to resist attacks.
Smart Images

Figure CN116167912B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of deep learning, cloud computing, computer vision and autonomous driving in artificial intelligence, and specifically relates to an adversarial sample generation method, an anti-attack detection method, a device and an electronic device. Background Art
[0002] With the continuous development of deep learning models, deep learning models have been widely applied in many fields and achieved excellent performance. At the same time, the security and robustness issues of deep learning models have also attracted people's attention. Currently, adversarial samples can be used to detect the anti-attack performance of deep learning models, and adversarial samples are easily occluded. Summary of the Invention
[0003] The present disclosure provides an adversarial sample generation method, an anti-attack detection method, a device and an electronic device.
[0004] According to a first aspect of the present disclosure, there is provided an adversarial sample generation method, including:
[0005] Obtain an initialization patch, where the initialization patch is a patch for processing a target image;
[0006] Identify the pixel mean of the pixel points included in the initialization patch;
[0007] Correct the initialization patch according to the pixel mean to obtain an adversarial patch;
[0008] Synthesize the adversarial patch and the target image to obtain an adversarial sample.
[0009] According to a second aspect of the present disclosure, there is provided an anti-attack detection method, including:
[0010] Input an adversarial sample into a model to be detected for sample recognition, and output a recognition result, where the recognition result is used to represent the image classification of the adversarial sample, and the model to be detected is a detection model for classifying images;
[0011] Determine the anti-attack result of the model to be detected for the adversarial sample according to the recognition result and the preset classification information of the target image obtained in advance;
[0012] Wherein, the adversarial sample is a sample generated according to the method provided in the first aspect.
[0013] According to a third aspect of the present disclosure, there is provided an adversarial sample generation device, including:
[0014] An obtaining module, configured to obtain an initialization patch, where the initialization patch is a patch for processing a target image;
[0015] An identification module, configured to identify the pixel mean value of the pixel points included in the initialization patch;
[0016] A correction module, configured to correct the initialization patch according to the pixel mean value to obtain an adversarial patch;
[0017] A synthesis module, configured to synthesize the adversarial patch and the target image to obtain an adversarial sample.
[0018] According to a fourth aspect of the present disclosure, there is provided an anti-attack detection device, including:
[0019] A sample identification module, configured to input an adversarial sample into a model to be detected for sample identification, and output an identification result, where the identification result is used to represent the image classification of the adversarial sample, and the model to be detected is a detection model for classifying images;
[0020] A determination module, configured to determine an anti-attack result of the model to be detected for the adversarial sample according to the identification result and preset classification information of a target image obtained in advance;
[0021] Wherein, the adversarial sample is a sample generated by the device provided in the second aspect.
[0022] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:
[0023] At least one processor; and
[0024] A memory communicatively connected to the at least one processor; wherein,
[0025] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any method in the first aspect.
[0026] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute any method in the first aspect.
[0027] According to a seventh aspect of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements any method in the first aspect.
[0028] In the embodiments of the present disclosure, the initialization patch is corrected according to the pixel mean value to obtain an adversarial patch. In this way, according to the pixel mean value of the pixel points included in the initialization patch, the occluded part in the initialization patch can be supplemented to obtain an adversarial patch, so that the display effect of the adversarial patch is better.
[0029] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is one of the flow schematic diagrams of the adversarial sample generation method provided by an embodiment of the present disclosure;
[0031] Figure 2 is the second of the flow schematic diagrams of the adversarial sample generation method provided by an embodiment of the present disclosure;
[0032] Figure 3 is the third of the flow schematic diagrams of the adversarial sample generation method provided by an embodiment of the present disclosure;
[0033] Figure 4 is the structural schematic diagram of the adversarial sample generation device for quantum gates provided by an embodiment of the present disclosure;
[0034] Figure 5 is a schematic block diagram of an exemplary electronic device for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The following describes exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0036] To detect the anti-attack ability of an object to be detected, usually, the target image is camouflaged to obtain an adversarial sample, and then the adversarial sample is input into the model to be detected. The anti-attack ability of the model to be detected is determined according to the detection result output by the object to be detected and the actual result of the target image. However, the target image is usually easily occluded, resulting in a poor detection result of the anti-attack ability of the model to be detected. Therefore, it is necessary to improve the quality of the adversarial sample. To improve the quality of the adversarial sample, the following solutions are proposed.
[0037] See Figure 1 , Figure 1 is a flowchart of an adversarial sample generation method provided by an embodiment of the present disclosure. As Figure 1 shown, the adversarial sample generation method includes the following steps:
[0038] Step S101, obtain an initialization patch, where the initialization patch is a patch for processing the target image.
[0039] Here, the type of the target image is not limited herein. The target image can be an image collected by an in-vehicle camera of an autonomous vehicle, or the target image can be data of a target object collected by an image acquisition device set by the roadside. The above target object can be a zebra crossing or an intersection.
[0040] Here, the specific way of initializing the patch is not limited herein. For example: an initialized patch can be directly obtained, or alternatively, a raw patch can be first obtained, and then the raw patch is initialized to obtain the initialized patch.
[0041] The way of the above initialization process is not limited herein. For example: the initialization process can include at least one of the following: white initialization, black initialization, gray initialization, random initialization. And the size of the initialized patch can match the size of the above raw patch, or alternatively, the size of the initialized patch can not match the size of the above raw patch, and the size of the initialized patch can be specified according to the user's input.
[0042] Here, the initialized patch is a patch for processing the target image. Optionally, the initialized patch can be used to perform a camouflage process on the target image, so that the category of the target image can be changed, and the target image after being processed by the initialized patch is input into the model to be detected to detect the recognition ability of the model to be detected for the camouflaged target image. The above detection method can be understood as the anti-attack ability of the model to be detected.
[0043] It should be noted that the camouflage process can include: fitting the initialized patch to at least part of the target image, that is, replacing the display content of the above at least part of the area, or alternatively, calculating the average value of the pixel values of the initialized patch and the display content of at least part of the target image, and correcting the above average value to the pixel value of the display content of the above at least part of the area. The specific way is not limited herein.
[0044] As an optional implementation manner, the obtaining the initialized patch includes:
[0045] Obtaining a first patch;
[0046] Initializing the first patch to obtain a second patch;
[0047] Adjusting the target parameters of the first target area of the second patch to obtain the initialized patch;
[0048] Here, the target parameter is a display parameter, and the display parameter includes at least one of the following parameters: display size, scale parameter, brightness, contrast, noise, position, angle.
[0049] Among them, the first patch mentioned above can be understood as the original patch, and the initialization process can refer to the relevant expressions above.
[0050] Among them, by adjusting the target parameters of the first target area of the second patch to obtain an initialization patch, the following can be referred to: the second patch can be understood as a patch in the digital world, and the adversarial patch can be understood as a patch in the physical world, and the above physical world can also be understood as the real world or the actual world, that is, the process of adjusting the target parameters of the first target area of the second patch to obtain an initialization patch can be understood as the process of adjusting through the target parameters of the first target area of the patch in the digital world, so as to map and transform the patch in the digital world into a patch in the real world.
[0051] Due to the complexity of the real world, the patches in the real world may also show diversity. For example, the proportion of the patch in the image changes, resulting in the change of the patch, or the patch is occluded, etc., which will make the adversarial patch ineffective. At the same time, when using the adversarial patch obtained according to the initialization patch to detect the model to be detected, the size of the adversarial patch is usually fixed. Therefore, the proportion of the adversarial patch will change due to the size of the target and the distance of the lens, and the adversarial patch in the real world is easily occluded.
[0052] Among them, the scale parameter can be a random scale parameter, or it can be understood that: the scale parameter is a randomly determined scale parameter, that is, the scale parameter of the adversarial patch obtained according to the initialization patch does not use a fixed scale parameter, but randomly determines a scale parameter as the scale parameter of the adversarial patch obtained according to the initialization patch. In this way, the problem of scale change caused by target change and lens distance in the physical attack of the real world for the adversarial patch obtained according to the initialization patch can be solved, thereby enhancing the robustness of the adversarial patch.
[0053] Among them, the display size can be understood as the display size. Since the target parameters of the first target area of the second patch are adjusted, the first target area of the second patch can also be called the area with enhanced change. Therefore, it can be understood as adjusting the display size of the area with enhanced change. For example, the first target area of the second patch can be padded, so that the display size of the initialization patch obtained after padding can match the target image.
[0054] In the embodiments of the present disclosure, by adjusting the target parameters of the first target area of the second patch to obtain an initialization patch, the target parameter is a display parameter, and the display parameter includes at least one of the following parameters: scale parameter, brightness, contrast, noise, position, angle. In this way, the robustness of the initialization patch can be improved, and further the robustness of the adversarial patch obtained according to the initialization patch can be enhanced.
[0055] Meanwhile, since the display parameters include a wide variety of types, various change factors in the real world can be better simulated, so that the adversarial patch better meets the requirements of the real world, that is, various change factors in the real world can be more realistically simulated, enhancing the authenticity of the adversarial patch.
[0056] As an optional implementation manner, adjusting the target parameters of the first target area of the second patch to obtain the initialization patch includes:
[0057] Adjusting the target parameters of the first target area of the second patch to obtain a third patch;
[0058] Performing an affine transformation on the second target area of the third patch to obtain the initialization patch, where the affine transformation includes at least one of the following: scaling, rotation, and translation.
[0059] In the embodiments of the present disclosure, by performing an affine transformation on the second target area of the third patch to obtain the initialization patch, the affine transformation includes at least one of the following: scaling, rotation, and translation. In this way, the robustness of the initialization patch can be further enhanced.
[0060] It should be noted that referring to Figure 3 , Figure 3 shows the process of enhancing various changes to the first patch to finally obtain the adversarial sample.
[0061] Step S102: Identify the pixel mean value of the pixel points included in the initialization patch.
[0062] Among them, the pixel values of the pixel points included in the initialization patch can be obtained, and the pixel mean value can be calculated based on the above pixel values; or, the pixel mean value can be pre-stored data, and the pixel mean value can be bound and stored with the initialization patch. When the initialization patch is obtained, the above pixel mean value can also be obtained. The specific method is not limited here.
[0063] Step S103: Correct the initialization patch according to the pixel mean value to obtain the adversarial patch.
[0064] Among them, the specific method of correcting the initialization patch according to the pixel mean value to obtain the adversarial patch is not limited here.
[0065] As an alternative implementation, the product of the pixel mean and the scene coefficient can be calculated, and the pixel values in a partial area of the initialization patch can be replaced with the product. The above partial area can be understood as the occluded area in the initialization patch, and the above scene coefficient can be determined according to the content of the initialization patch. For example, if the content of the initialization patch represents different scenes, the scene coefficients are different.
[0066] As an alternative implementation, the step of correcting the initialization patch according to the pixel mean to obtain an adversarial patch includes:
[0067] Identifying a first area and a second area included in the initialization patch, where the second area is the area in the initialization patch other than the first area;
[0068] Replacing the pixel values of the pixel points in the first area with a preset pixel value, and replacing the pixel values of the pixel points in the second area with the pixel mean to obtain the adversarial patch, where the preset pixel value is different from the pixel mean.
[0069] Among them, the first area can be referred to as an area where work stops or display stops, and the specific value of the preset pixel value is not limited here. For example, the preset pixel value can be the pixel value corresponding to a transparent pixel point, or the pixel value corresponding to a white pixel point.
[0070] In the embodiments of the present disclosure, it can be applied to a model to be detected, and the model to be detected may include a Dropout layer. By means of the Dropout layer, the pixel values of the pixel points in the first area can be replaced with a preset pixel value. When the preset pixel value is the pixel value corresponding to a transparent pixel point, it can be understood that the pixel points in the first area stop working, that is, an adversarial patch simulating the occlusion of the first area can be simulated, thereby increasing the diversity of the simulated adversarial patch. In addition, the complex adaptability between the neurons of the model to be detected can be reduced, so that the model to be detected no longer depends on the co-action between certain specific neurons and other neurons, making the model to be detected more robust and reducing the occurrence of overfitting phenomena in the model to be detected.
[0071] In addition, since the numerical range of the pixel values of the pixel points in the second area of the initialization patch is between 0 and 1, the overall invariance cannot be maintained by dividing the pixel values of the pixel points in the second area of the initialization patch by the probability p. Therefore, the pixel values of the pixel points in the second area can be replaced with the pixel mean to ensure the overall invariance of the obtained adversarial patch and enhance the display effect of the adversarial patch.
[0072] As an alternative implementation, the step of identifying the first area and the second area included in the initialization patch includes:
[0073] Randomly determine the first area included in the initialization patch;
[0074] Determine the second area according to the first area.
[0075] In the embodiments of the present disclosure, since the first area is randomly determined, when the steps in the embodiments of the present disclosure are executed multiple times, multiple adversarial patches can be finally determined, and the positions of the first areas occluded by each adversarial patch are different, so that the shapes of the multiple adversarial patches are different, that is, multiple adversarial patches with different shapes are obtained, and finally the robustness and diversity of the adversarial patches can be improved.
[0076] As an optional implementation manner, the correcting the initialization patch according to the pixel mean value to obtain an adversarial patch includes:
[0077] Predict the probability of correcting the initialization patch;
[0078] In the case where the probability of correcting the initialization patch is greater than a preset probability, correct the initialization patch according to the pixel mean value to obtain an adversarial patch.
[0079] In the embodiments of the present disclosure, the probability of correcting the initialization patch can be predicted by the model to be detected. Only when the probability of correcting the initialization patch is greater than the preset probability, will the initialization patch be corrected according to the pixel mean value to obtain an adversarial patch. That is to say, the initialization patch is not corrected every time an initialization patch is obtained. In this way, the robustness and diversity of the finally obtained adversarial patch can be further improved.
[0080] Step S104: Synthesize the adversarial patch and the target image to obtain an adversarial sample.
[0081] Among them, in the prior art, it is usually only possible to fit to the center position of the detection box of the attack category included in the target image, while the adversarial patch in the embodiments of the present disclosure can be fitted to any position of the target image, and the difference in the anti-attack result detection effect of the adversarial sample obtained by fitting the adversarial patch to any position of the target image to the model to be detected is small. In this way, the synthesis efficiency of the adversarial patch and the target image can be improved.
[0082] As an optional implementation manner, the embodiments of the present disclosure further provide an anti-attack detection method, including:
[0083] Input the adversarial sample into the model to be detected for sample recognition, and output a recognition result, where the recognition result is used to represent the image classification of the adversarial sample, and the model to be detected is a detection model for classifying images;
[0084] Determine the anti - attack result of the to - be - detected model against the adversarial sample according to the recognition result and the preset classification information of the target image obtained in advance.
[0085] Among them, the adversarial samples in the embodiments of the present disclosure can be generated by the methods in the above - mentioned embodiments. For the respective technical features in the embodiments of the present disclosure, reference can be made to the corresponding descriptions in the above - mentioned embodiments, and details will not be elaborated here.
[0086] In the embodiments of the present disclosure, the adversarial sample can be input into the to - be - detected model for sample recognition, and the recognition result is output. At the same time, according to the recognition result and the preset classification information of the target image obtained in advance, determine the anti - attack result of the to - be - detected model against the adversarial sample, so as to realize the detection of the anti - attack result of the to - be - detected model against the adversarial sample, and the accuracy of the detection result can be relatively high, and the detection method is relatively convenient.
[0087] It should be noted that determining the anti - attack result of the to - be - detected model against the adversarial sample according to the recognition result and the preset classification information of the target image obtained in advance can be referred to the following description: when the recognition result matches the preset classification information of the target image, it can be considered that the anti - attack result of the to - be - detected model against the adversarial sample is poor; or, when the recognition result does not match the preset classification information of the target image, it can be considered that the anti - attack result of the to - be - detected model against the adversarial sample is good.
[0088] It should be noted that the to - be - detected model can be a pre - trained detection model for classifying images. Each step in the above - mentioned embodiments can be understood as an application process, and the application process can refer to Figure 2 Steps 201 to 206 in. At the same time, each step in the above - mentioned embodiments can be understood as the steps executed by the to - be - detected model in a training iteration process, and the above - mentioned adversarial patch and adversarial sample can be updated by backpropagation according to the loss function of the to - be - detected model. For example: the training process of the to - be - detected model can refer to repeatedly executing steps 201 to 206 multiple times.
[0089] In addition, during the training process, the loss function of the to - be - detected model can include the following three parts: detection score loss, total difference loss, and non - printability score loss.
[0090] Among them, the detection score loss: The detection score loss is the maximum value of the detection scores in the adversarial sample. Minimizing this loss enables the training of the adversarial sample to make the to - be - detected model ineffective. The detection score can include multiple modes. For example: it can be the confidence of the detection box, or the score of the attacked class, or it can also be the product of the confidence of the detection box and the score of the attacked class.
[0091] Full-difference loss: The calculation formula of the full-difference loss can be shown as in Formula (1), which represents the difference between each pixel and its surrounding pixels. Minimizing this loss function ensures that the learned adversarial samples have smooth color changes and reduces noise.
[0092]
[0093] Among them, p i,j represents the pixel value of the pixel corresponding to the position (i, j) in the adversarial sample.
[0094] Non-printability score loss: The calculation formula of the non-printability score loss can be shown as in Formula (2), which represents the difference between the pixel values in the adversarial sample and the pixel values that can be printed by an ordinary printer. Minimizing this loss minimizes the difference between the physical patch printed (i.e., the adversarial sample obtained according to the adversarial patch) and the digital patch (which can be understood as the second patch in the above embodiment), facilitating the implementation of physical attacks.
[0095]
[0096] Among them, p patch represents the pixel point in the adversarial sample, and c print represents a printable color value in the printable color set.
[0097] See Figure 4 , a structural schematic diagram of an adversarial sample generation device provided by an embodiment of the present disclosure is shown as Figure 4 shown. The adversarial sample generation device 400 includes:
[0098] An acquisition module 401, configured to acquire an initialization patch, where the initialization patch is a patch used to process a target image;
[0099] An identification module 402, configured to identify the pixel mean of the pixel points included in the initialization patch;
[0100] A correction module 403, configured to correct the initialization patch according to the pixel mean to obtain an adversarial patch;
[0101] A synthesis module 404, configured to synthesize the adversarial patch and the target image to obtain an adversarial sample.
[0102] Optionally, the correction module 403 includes:
[0103] An identification sub-module, configured to identify a first region and a second region included in the initialization patch, where the second region is the region in the initialization patch except the first region;
[0104] A replacement sub-module, configured to replace the pixel values of the pixel points in the first region with a preset pixel value, and replace the pixel values of the pixel points in the second region with the pixel mean value, to obtain the adversarial patch, where the preset pixel value is different from the pixel mean value.
[0105] Optionally, the recognition sub-module includes:
[0106] A first determination unit, configured to randomly determine a first region included in the initialization patch;
[0107] A second determination unit, configured to determine the second region according to the first region.
[0108] Optionally, the correction module 403 includes:
[0109] A prediction sub-module, configured to predict the probability of correcting the initialization patch;
[0110] A correction sub-module, configured to correct the initialization patch according to the pixel mean value to obtain an adversarial patch when the probability of correcting the initialization patch is greater than a preset probability.
[0111] Optionally, the obtaining module 401 includes:
[0112] An obtaining sub-module, configured to obtain a first patch;
[0113] An initialization processing sub-module, configured to perform initialization processing on the first patch to obtain a second patch;
[0114] An adjustment sub-module, configured to adjust target parameters of a first target region of the second patch to obtain the initialization patch;
[0115] Wherein, the target parameter is a display parameter, and the display parameter includes at least one of the following parameters: display size, scale parameter, brightness, contrast, noise, position, angle.
[0116] Optionally, the adjustment sub-module includes:
[0117] An adjustment unit, configured to adjust target parameters of a first target region of the second patch to obtain a third patch;
[0118] An affine transformation unit, configured to perform an affine transformation on a second target region of the third patch to obtain the initialization patch, where the affine transformation includes at least one of the following: scaling, rotation, translation.
[0119] Optionally, an anti-attack detection device provided by an embodiment of the present disclosure includes:
[0120] A sample recognition module is configured to input an adversarial sample into a model to be detected for sample recognition and output a recognition result, where the recognition result is used to represent the image classification of the adversarial sample, and the model to be detected is a detection model for classifying images;
[0121] A determination module is configured to determine an anti - attack result of the model to be detected against the adversarial sample according to the recognition result and preset classification information of a target image obtained in advance;
[0122] Wherein, the adversarial sample is a sample generated according to the embodiments of the above - mentioned adversarial sample generation method or the above - mentioned adversarial sample generation device.
[0123] The adversarial sample generation device 400 provided by the embodiments of the present disclosure can implement all processes implemented by the embodiments of the adversarial sample generation method. The anti - attack detection device provided by the embodiments of the present disclosure can implement all processes implemented by the embodiments of the anti - attack detection method, and can achieve the same beneficial effects. To avoid repetition, details are not described herein again.
[0124] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0125] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 500 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0126] As Figure 5 shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read - only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random - access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0127] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as a keyboard, mouse, etc.; output unit 507, such as various types of displays, speakers, etc.; storage unit 508, such as a disk, optical disc, etc.; and communication unit 509, such as a network card, modem, wireless communication transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0128] Computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 501 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 501 executes the various methods and processes described above, such as the adversarial sample generation method or the anti-attack detection method. For example, in some embodiments, the adversarial sample generation method or the anti-attack detection method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by computing unit 501, one or more steps of the adversarial sample generation method or the anti-attack detection method described above can be executed. Alternatively, in other embodiments, computing unit 501 can be configured to execute the adversarial sample generation method or the anti-attack detection method in any other suitable way (e.g., by means of firmware).
[0129] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0130] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0131] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0132] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0133] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0134] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0135] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0136] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. An adversarial sample generation method, comprising: Obtaining an initialization patch, where the initialization patch is a patch for processing a target image; Identifying the pixel mean of the pixel points included in the initialization patch; Correcting the initialization patch according to the pixel mean to obtain an adversarial patch; Combining the adversarial patch with the target image to obtain an adversarial sample.
2. The method according to claim 1, wherein, The correcting the initialization patch according to the pixel mean to obtain an adversarial patch includes: Identifying a first region and a second region included in the initialization patch, where the second region is the region in the initialization patch other than the first region; Replacing the pixel values of the pixel points in the first region with a preset pixel value, and replacing the pixel values of the pixel points in the second region with the pixel mean to obtain the adversarial patch, where the preset pixel value is different from the pixel mean.
3. The method according to claim 2, wherein, The identifying the first region and the second region included in the initialization patch includes: Randomly determining the first region included in the initialization patch; Determining the second region according to the first region.
4. The method according to claim 1, wherein The correcting the initialization patch according to the pixel mean to obtain an adversarial patch includes: Predicting the probability of correcting the initialization patch; In the case where the probability of correcting the initialization patch is greater than a preset probability, correcting the initialization patch according to the pixel mean to obtain an adversarial patch.
5. The method according to claim 1, wherein The obtaining the initialization patch includes: Obtaining a first patch; Performing an initialization process on the first patch to obtain a second patch; Adjusting the target parameters of the first target region of the second patch to obtain the initialization patch; Wherein the target parameters are display parameters, and the display parameters include at least one of the following parameters: display size, scale parameter, brightness, contrast, noise, position, angle.
6. The method according to claim 5, wherein, The adjusting the target parameters of the first target region of the second patch to obtain the initialization patch includes: Adjusting the target parameters of the first target region of the second patch to obtain a third patch; Performing an affine transformation on the second target region of the third patch to obtain the initialization patch, where the affine transformation includes at least one of the following: scaling, rotation, translation.
7. An anti-attack detection method, comprising: Inputting an adversarial sample into a model to be detected for sample recognition, and outputting a recognition result, where the recognition result is used to represent the image classification of the adversarial sample, and the model to be detected is a detection model for classifying images; Determining the anti-attack result of the model to be detected for the adversarial sample according to the recognition result and the preset classification information of the target image obtained in advance; Wherein the adversarial sample is a sample generated according to the method according to any one of claims 1 to 6.
8. An adversarial sample generation device, comprising: An obtaining module, configured to obtain an initialization patch, where the initialization patch is a patch for processing a target image; An identifying module, configured to identify the pixel mean of the pixel points included in the initialization patch; A correcting module, configured to correct the initialization patch according to the pixel mean to obtain an adversarial patch; A synthesis module for synthesizing the adversarial patch and the target image to obtain an adversarial sample.
9. The apparatus according to claim 8, wherein, The correction module includes: An identification sub-module for identifying a first region and a second region included in the initialization patch, where the second region is the region in the initialization patch except the first region; A replacement sub-module for replacing the pixel values of the pixel points in the first region with a preset pixel value, and replacing the pixel values of the pixel points in the second region with the pixel mean value to obtain the adversarial patch, where the preset pixel value is different from the pixel mean value.
10. The apparatus according to claim 9, wherein, The identification sub-module includes: A first determination unit for randomly determining the first region included in the initialization patch; A second determination unit for determining the second region according to the first region.
11. The apparatus according to claim 8, wherein, The correction module includes: A prediction sub-module for predicting the probability of correcting the initialization patch; A correction sub-module for correcting the initialization patch according to the pixel mean value to obtain an adversarial patch when the probability of correcting the initialization patch is greater than a preset probability.
12. The apparatus according to claim 8, wherein, The acquisition module includes: An acquisition sub-module for acquiring a first patch; An initialization processing sub-module for performing initialization processing on the first patch to obtain a second patch; An adjustment sub-module for adjusting the target parameters of the first target region of the second patch to obtain the initialization patch; where the target parameters are display parameters, and the display parameters include at least one of the following parameters: display size, scale parameter, brightness, contrast, noise, position, angle.
13. The apparatus according to claim 12, wherein, The adjustment sub-module includes: An adjustment unit for adjusting the target parameters of the first target region of the second patch to obtain a third patch; An affine transformation unit for performing an affine transformation on the second target region of the third patch to obtain the initialization patch, where the affine transformation includes at least one of the following: scaling, rotation, translation.
14. An anti-attack detection device, including: A sample identification module for inputting an adversarial sample into a model to be detected for sample identification and outputting an identification result, where the identification result is used to represent the image classification of the adversarial sample, and the model to be detected is a detection model for classifying images; A determination module for determining the anti-attack result of the model to be detected against the adversarial sample according to the identification result and the preset classification information of the target image obtained in advance; where the adversarial sample is a sample generated by the device according to any one of claims 8 to 13.
15. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; where The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
17. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 - 7.
Citation Information
Patent Citations
Small adversarial patch generation method and device
CN112241790A
Target detection-oriented physical attack adversarial patch generation method and system
CN113361604A