Backdoor trigger fitting method for virtual poisoning image data and related equipment
Through CMA-ES and back-propagation iterative training, we generate nearly realistic backdoor triggers, solving the problems of low accuracy and insufficient versatility in backdoor attack detection in deep neural networks, and improving the security and detection effect of deep learning systems.
Patent Information
- Application Number
- CN202210492940.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-07
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-05-07
AI Technical Summary
In the existing technology, the backdoor attack detection method of deep neural network has low accuracy in the case of large triggers, lacks versatility and flexibility, and is difficult to effectively detect backdoors on different image datasets and models.
The covariance adaptive evolutionary strategy (CMA-ES) is used to generate tensor data. Through backpropagation and gradient descent iterative training, the optimal coordinate position is found and the approximate true backdoor trigger is fitted to detect whether there is a backdoor in the target network.
Efficient backdoor detection is achieved on different image datasets and models, improving the security of deep learning systems and the versatility of detection. The detection results are consistent with the expected application scenarios of the original image datasets and classification models.
Smart Images

Figure CN115170855B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to deep learning technology, and more particularly to a backdoor trigger fitting method for virtual poisoned image data and related equipment. Background Art
[0002] In recent years, with the advancement of artificial intelligence (AI), deep neural networks have achieved breakthroughs in applications such as image processing, speech recognition, and video analysis. However, deep neural networks lack transparency and are vulnerable to backdoor attacks, posing serious security risks. Backdoors can remain hidden indefinitely. A deep neural network injected with a backdoor will behave normally with pure input, but when the input contains triggers predefined by the attacker, the backdoor is activated, causing harm.
[0003] To effectively mitigate backdoor attacks, backdoor detection is first necessary. For backdoor attack detection in related technologies, such as Neural Cleanse (NC), the detection accuracy decreases when the backdoor trigger is large. Summary of the Invention
[0004] In view of this, the purpose of the present disclosure is to provide a backdoor trigger fitting method, device, electronic device and storage medium for virtual poisoned image data.
[0005] Based on the above objectives, the present disclosure provides a backdoor trigger fitting method for virtual poisoned image data, comprising:
[0006] Randomly generate tensor data based on the original image dataset;
[0007] Based on the covariance adaptive evolutionary strategy (CMA-ES), a plurality of candidate coordinate positions of the original image dataset are randomly generated;
[0008] A plurality of first success rates respectively corresponding to the plurality of candidate coordinate positions are obtained by the following operations: for each candidate coordinate position in the plurality of candidate coordinate positions, overlaying the tensor data onto the candidate coordinate position of each original image data in the original image dataset to construct a first virtual poisoned image dataset, inputting the original image dataset and the first virtual poisoned image dataset into a pre-trained classification model for injecting a backdoor, and calculating the first success rate of the tensor data activating the backdoor in the classification model based on a first prediction result output by the classification model;
[0009] Selecting a candidate coordinate position from the plurality of candidate coordinate positions corresponding to a maximum value among the plurality of first success rates as a target coordinate position;
[0010] Overlaying the tensor data onto the target coordinate position in each original image data in the original image data set to construct a target virtual poisoning image data set;
[0011] The original image dataset and the target virtual poisoning image dataset are input into the classification model, the tensor data is iteratively trained, and the trained tensor data is determined as a backdoor trigger of the virtual poisoning image data.
[0012] Based on the same technical concept, the present disclosure also provides a backdoor trigger fitting device for virtual poisoned image data, comprising:
[0013] A first generation module is configured to randomly generate tensor data based on the original image dataset;
[0014] A second generation module is configured to randomly generate a plurality of candidate coordinate positions of the original image dataset based on CMA-ES;
[0015] a calculation module configured to obtain a plurality of first success rates respectively corresponding to the plurality of candidate coordinate positions by performing the following operations: for each candidate coordinate position in the plurality of candidate coordinate positions, overlaying the tensor data onto the candidate coordinate position of each original image data in the original image dataset to construct a first virtual poisoned image dataset, inputting the original image dataset and the first virtual poisoned image dataset into a pre-trained classification model for injecting a backdoor, and calculating, based on a first prediction result output by the classification model, the first success rate of activating the backdoor in the classification model using the tensor data;
[0016] A first determining module is configured to select a candidate coordinate position from the plurality of candidate coordinate positions corresponding to a maximum value among the plurality of first success rates as a target coordinate position;
[0017] a construction module configured to overlay the tensor data onto the target coordinate position in each original image data in the original image data set to construct a target virtual poisoning image data set;
[0018] The second determination module is configured to input the original image dataset and the target virtual poisoning image dataset into the classification model, iteratively train the tensor data, and determine the trained tensor data as a backdoor trigger of the virtual poisoning image data.
[0019] Based on the same technical concept, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements any of the above methods when executing the computer program.
[0020] Based on the same technical concept, the present disclosure further provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute any of the methods described above.
[0021] As can be seen from the above, the backdoor trigger fitting method, device, electronic device, and storage medium for virtual poisoned image data provided by the present disclosure, based on CMA-ES, searches for the optimal coordinate position that maximizes the success rate of backdoor activation in tensor data, and applies the principle of backpropagation to iteratively train the tensor data using gradient descent, ultimately fitting a nearly realistic backdoor trigger that can be used to detect whether a backdoor exists in the target network and verify the validity of the fitted trigger. The solution of the present disclosure is not limited by image datasets and classification models, and the fitted trigger is consistent with the intended application scenarios of the original image dataset and classification model, thus having no size restrictions. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 Schematic diagram of the flow of a backdoor trigger fitting method for virtual poisoned image data according to an embodiment of the present disclosure;
[0024] Figure 2 A schematic diagram of the process of injecting a backdoor into a classification model for training according to an embodiment of the present disclosure;
[0025] Figure 3 This is a structural diagram of a backdoor trigger fitting device for virtual poisoned image data according to an embodiment of the present disclosure;
[0026] Figure 4 Schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] In order to make the objectives, technical solutions and advantages of the present disclosure more clear, the embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0028] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may also change accordingly.
[0029] Backdoor attacks are a type of attack targeting deep learning. The injected backdoor does not affect the classification results of the deep neural network model when the input is pure. However, it will (and only) activate the model backdoor when a predefined trigger is added to the input. The model injected with the backdoor will misclassify any input as the same target label. Input samples that should be classified as any other label will be "overwritten" in the presence of the trigger. For example, images of other labels (such as cat, bird, fish) will be misclassified as the target label (such as dog).
[0030] Related backdoor attack detection techniques, such as Neural Cleanse (NC), use models to reversely infer the location and shape of triggers and then determine whether a backdoor exists. However, these backdoor attack detection methods require a large number of input samples to achieve high performance and may fail for larger triggers. When the backdoor trigger is large (over 25%), the detection accuracy is significantly reduced.
[0031] In response to the problems existing in the above-mentioned related technologies, the present disclosure provides a backdoor trigger fitting scheme for virtual poisoned image data. Based on the evolutionary strategy of covariance adaptive adjustment (CMA-ES), it searches for the optimal coordinate position that maximizes the success rate of backdoor activation of tensor data, and applies the principle of backpropagation to iteratively train the tensor data with the gradient descent method, ultimately fitting an approximate real backdoor trigger, which can be used to detect whether the target network has a backdoor and verify the validity of the fitted trigger. The scheme disclosed in the present disclosure is not limited by the image dataset and classification model, and the fitted trigger is consistent with the expected application scenario of the original image dataset and classification model, so there is no size limit.
[0032] A backdoor trigger fitting method for virtual poisoned image data in an embodiment of the present disclosure can be applied to multiple scenarios that use neural networks for image classification, such as face recognition and autonomous driving. Given a fitted approximate real backdoor trigger, the goal is to determine whether the target model has a backdoor and verify the validity of the fitted trigger. When the backdoor trigger is added to the input, targeted misclassification will be displayed. The model structure can be a known classic network structure, a custom network structure, etc. The original image dataset used as the input sample can be a public image dataset, an uploaded custom image dataset, etc. The number of the original image datasets is not limited in the embodiment of the present disclosure.
[0033] refer to Figure 1 , is a flow chart of a backdoor trigger fitting method for virtual poisoned image data according to an embodiment of the present disclosure. The method may include the following steps:
[0034] Step S101: randomly generate tensor data based on the original image dataset.
[0035] In this embodiment, the original image dataset is selected according to the usage scenarios such as testing and training. The selected original image dataset can be an existing public image dataset or an uploaded custom image dataset. In general, the original image dataset can contain multiple original image data. Specifically, the representation of these original image data can be a picture stored in pixel-level matrix data. The shape (such as rectangle, circle) and category (such as cat, bird, fish) of the original image data can be random and are not influencing factors that need to be considered for this solution.
[0036] In this embodiment, tensor data is generated based on the selected original image dataset. The first dimension and first size of any original image data in the original image dataset are obtained, and then a tensor data is randomly generated that is consistent with the first dimension of the original image data. The dimensional size of the original image data is consistent with the tensor data, ensuring that the tensor data can completely replace the pixels at the relative positions of the original image data in terms of dimension during overwriting. Furthermore, the tensor data should not be lost or exceed the first size of the original image data during overwriting.
[0037] Step S102: randomly generate multiple candidate coordinate positions of the original image data set based on CMA-ES.
[0038] In this embodiment, based on CMA-ES, all coordinates of the original image data are sampled, and N coordinates (x1, y1), (x2, y2), ..., (x n ,y n ) as initial values, and the tensor data uses the N coordinate points corresponding to these initial values as candidate coordinate positions.
[0039] In specific implementations, the coordinate points are located within the image region where pixel-level matrix data is stored. A rectangular coordinate system is established with the lower left vertex of the image region as the origin, the lower boundary of the region as the X-axis, and the left boundary as the Y-axis. The x-value of a coordinate is the distance from the point to the Y-axis, and the y-value is the distance from the point to the X-axis.
[0040] Step S103, obtain multiple first success rates corresponding to the multiple candidate coordinate positions respectively through the following operations: for each candidate coordinate position of the multiple candidate coordinate positions, overlay the tensor data onto the candidate coordinate position of each original image data in the original image dataset to construct a first virtual poisoned image dataset, and input the original image dataset and the first virtual poisoned image dataset into a pre-trained classification model for injecting a backdoor, and calculate the first success rate of the tensor data activating the backdoor of the classification model according to the first prediction result output by the classification model.
[0041] In this embodiment, the first virtual poisoning image data set is composed of a plurality of first virtual poisoning image data, and the first virtual poisoning image data are respectively covered by tensor data at each candidate coordinate position (x n ,y n ) can be obtained.
[0042] In this example, a classification model for testing is obtained. A classification model is constructed and a backdoor is injected into the model through a training process. The specific training steps will be described later.
[0043] In this embodiment, a test set is formed by mixing the original image dataset and the first virtual poisoned image dataset. This is then fed into the classification model for injecting the backdoor obtained above. Based on the first prediction result output by the classification model, the success rate of the tensor data activating the backdoor in the classification model is calculated.
[0044] During specific implementation, for each selected candidate coordinate position, 100 original image data can be randomly extracted from the original image data set, and the tensor data is respectively overlaid on the candidate coordinate position of each original image data to obtain 100 first virtual poisoned image data, which constitute the first virtual poisoned image data set and are input into the classification model of the backdoor injection obtained above. The first number of valid virtual poisoned image data in the first virtual poisoned image data set that successfully underwent the virtual backdoor attack is counted. Specifically, for each valid virtual poisoned image data, the first prediction result is a pre-set virtual poisoning attack target label. The ratio of the first number to the total number of virtual poisoned image data in the first virtual poisoned image data set (100) is calculated as the first success rate. For example, if the set virtual poisoning attack target label is "dog", for any category of the first virtual poisoned image data (for example, the category is "cat", "bird", "fish", etc.), the first prediction result is "dog", then it can be considered that the model backdoor is activated.
[0045] Step S104: Select a candidate coordinate position from the plurality of candidate coordinate positions corresponding to the maximum value of the plurality of first success rates as the target coordinate position.
[0046] In this embodiment, based on CMA-ES, the candidate coordinate corresponding to the maximum activation success rate is output and the coordinate is used as the target coordinate (x ′ ,y ′ ).
[0047] During specific implementation, the N randomly selected candidate coordinates are used as the sampling population of the current generation, and the tensor data is covered on each of the selected candidate coordinate positions to obtain the first virtual poisoning image data set and input into the above-obtained classification model for injecting the backdoor, and the success rate of activating the backdoor of the classification model is calculated. Based on CMA-ES, the activation success rate values are sorted according to the activation success rate of the tensor data at each selected candidate coordinate position, and the candidate coordinate positions corresponding to the K values with the highest activation success rate in the current generation (for example, the candidate coordinate positions corresponding to the top 25% of the activation success rate) are selected, and the covariance matrix of the next generation is obtained based on these coordinate positions, and then sampling is performed from the multivariate Gaussian distribution obtained from the updated covariance matrix, that is, N candidate coordinates are randomly selected from the multivariate Gaussian distribution obtained after the update as the sampling population of the next generation. Repeated iteration, when the coordinate position no longer changes, that is, when the activation success rate of the tensor data at the coordinate position is the maximum, the coordinate is output and used as the target coordinate (x ′ ,y ′ ).
[0048] In specific implementation, the covariance matrix in the multivariate Gaussian distribution has two parameters to be optimized, x and y, so the corresponding covariance matrix is An increase in D(X) in the covariance matrix will make the sampling more dispersed in the direction of the X-axis of the image area (for example, the sampling population is stretched in the direction of the X-axis, and the horizontal coordinate changes from the original (-3,3) to (-5,5)); an increase in D(Y) will make the sampling more dispersed in the direction of the Y-axis; cov(X,Y) greater than 0 will cause the sampling population to be positively correlated, that is, when the sampling population is more dispersed in the direction of the X-axis, it will also be more dispersed in the direction of the Y-axis.
[0049] In addition, in some embodiments, after repeated iterations, when the coordinate positions are stabilized within a certain range, that is, the activation success rates of multiple coordinate positions of the tensor data within the range are stable near a certain value and the changes are slight, these coordinates can also be output and used as target coordinates.
[0050] Step S105 : Overlay the tensor data onto the target coordinate position in each original image data in the original image data set to construct a target virtual poisoning image data set.
[0051] In this embodiment, the target virtual poisoning image data set is composed of a plurality of target virtual poisoning image data, and the target virtual poisoning image data is respectively covered by tensor data at the target coordinate position (x ′ ,y ′ ) can be obtained.
[0052] Step S106: input the original image dataset and the target virtual poisoning image dataset into the classification model, iteratively train the tensor data, and determine the trained tensor data as the backdoor trigger of the virtual poisoning image data.
[0053] In this embodiment, the loss between the predicted value and the true value of the classification model obtained by injecting the backdoor and the effect of the tensor data is calculated. Based on the principle of backpropagation, the tensor data is trained using gradient descent to minimize this loss. When a pre-set termination condition is met, the trained tensor data is output and used as the backdoor trigger for the virtual poisoned image data.
[0054] In the specific implementation, the initial training setting is a learning rate of 0.1, the optimizer is stochastic gradient descent (SGD), and the training rounds are 200.
[0055] In specific implementations, the size of the tensor data will increase or decrease during the above training process, but will never exceed the size of the original image data.
[0056] As can be seen from the above embodiments, the disclosed backdoor trigger fitting method for virtual poisoned image data, using CMA-ES, can adaptively adjust parameters and calculate the covariance matrix of the entire parameter space. Evolution occurs gradually during the selection process. Evolutionary algorithms aim to optimize functions that cannot be directly modeled, with sampling and updating as their core. CMA-ES is one of the best-performing evolutionary algorithms. The CMA-ES algorithm obtains the results of each iteration and adaptively increases or decreases the search space in the next generation of searches. Specifically, the CMA-ES algorithm can use information from the optimal solution to adjust its mean and covariance matrix. This allows it to search a larger space when the optimal solution is far away, and a smaller space when the optimal solution is close. Thus, through repeated iterative parameter adjustment, the probability of generating a good solution gradually increases (the probability of searching along a good search direction increases). The optimal coordinate position that maximizes the success rate of backdoor activation in tensor data is found. The activation success rate is then further improved through iterative training and continuous fitting of the tensor data, ultimately resulting in a nearly realistic backdoor trigger that can be used to detect the presence of backdoors in the target network and enhance the security of deep learning systems. Because the fitted triggers are consistent with the intended application scenarios of the original image dataset and classification model, there are no size restrictions. Furthermore, they are not limited by the image dataset or classification model, making backdoor detection more versatile.
[0057] According to an embodiment of the present disclosure, the method for training a classification model for injecting a backdoor can be as follows: Figure 2 The method may include the following steps:
[0058] Step S201: randomly generate sample tensor data according to the sample original image data set.
[0059] In this embodiment, a sample original image dataset is selected. The selected sample original image dataset can be an existing public image dataset or an uploaded custom image dataset. Generally, the sample original image dataset can contain multiple sample original image data. Specifically, the representation of these sample original image data can be a picture stored in a pixel-level matrix data. The shape (such as rectangle, circle) and category (such as cat, bird, fish) of the sample original image data can be random and are not influencing factors that need to be considered.
[0060] In this embodiment, sample tensor data is generated based on the selected sample original image dataset. The dimension and size of any sample original image data in the sample original image dataset are obtained, and then a sample tensor data having the same dimension and size as the sample original image data is randomly generated. The sample tensor data should not exceed the size of the sample original image data.
[0061] Step S202: Overlay the sample tensor data onto a predetermined coordinate position of each sample original image data in the sample original image data set to construct a sample virtual poisoning image data set.
[0062] In this embodiment, a coordinate (x″, y″) is randomly selected within the image area represented by the sample original image data. The sample virtual poisoning image dataset is composed of a plurality of sample virtual poisoning image data, which are obtained by overlaying the sample tensor data on the above-mentioned predetermined coordinate position (x″, y″) of each sample original image data in the sample original image dataset.
[0063] Step S203: training a neural network model using a mixture of the sample original image dataset and the sample virtual poisoned image dataset to obtain a classification model for the injected backdoor.
[0064] In this embodiment, a pre-trained classification model is constructed, and the sample original image dataset and the sample virtual poisoning image dataset are mixed to form a training image dataset, which is input into the classification model. The model learns from the training image data and is then injected with a backdoor.
[0065] In specific implementation, the selected sample original image dataset is input into the selected pre-training model. When the classification accuracy of the pre-training model for the sample original image dataset is higher than a predetermined threshold, a usable pre-training classification model can be obtained.
[0066] There is no restriction on the structure of the classification model. You can choose a known classic network structure model, or upload a custom network structure model. Specifically, different pre-trained models can be selected according to the needs of the user's device and model application scenario. For example, neural network models such as ResNet18 and ResNet50. The difference between ResNet18 and ResNet50 is that the two neural network models have different number of structural layers. When it is necessary to quickly obtain an optimized model, a neural network model with fewer structural layers, such as ResNet18, can be used; when it is necessary to obtain a relatively safe model, a neural network model with more structural layers, such as ResNet50, can be used.
[0067] In addition, in some embodiments, the customized network structure model file can be a .py file written in Python. Specifically, the use of common neural network models can be implemented using open source tools such as TensorFlow and PyTorch. TensorFlow and PyTorch are both tools for implementing neural network models.
[0068] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0069] Based on the same technical concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides a backdoor trigger fitting device for virtual poisoning image data.
[0070] refer to Figure 3 The backdoor trigger fitting device 300 for virtual poisoned image data comprises:
[0071] A first generating module 301 is configured to randomly generate tensor data based on the original image dataset;
[0072] A second generating module 302 is configured to randomly generate a plurality of candidate coordinate positions of the original image dataset based on CMA-ES;
[0073] The calculation module 303 is configured to obtain a plurality of first success rates respectively corresponding to the plurality of candidate coordinate positions by performing the following operations: for each candidate coordinate position in the plurality of candidate coordinate positions, overlaying the tensor data onto the candidate coordinate position of each original image data in the original image dataset to construct a first virtual poisoned image dataset, inputting the original image dataset and the first virtual poisoned image dataset into a pre-trained classification model for injecting a backdoor, and calculating, based on a first prediction result output by the classification model, the first success rate of activating the backdoor in the classification model using the tensor data;
[0074] A first determining module 304 is configured to select a candidate coordinate position from the plurality of candidate coordinate positions corresponding to a maximum value among the plurality of first success rates as a target coordinate position;
[0075] A construction module 305 is configured to overlay the tensor data onto the target coordinate position in each original image data in the original image data set to construct a target virtual poisoning image data set;
[0076] The second determination module 306 is configured to input the original image dataset and the target virtual poisoned image dataset into the classification model, iteratively train the tensor data, and determine the trained tensor data as the backdoor trigger of the virtual poisoned image data.
[0077] In some optional embodiments, the computing module 303 is specifically configured to obtain a classification model for testing, select a known classic network structure model or upload a custom network structure model; and inject a backdoor into the model through a training process.
[0078] In some optional embodiments, the second determination module 306 is specifically configured to construct a virtual poisoned image dataset for testing based on the target coordinate positions covered by the optimized tensor data and the trained tensor data; the virtual poisoned image dataset is input into the classification model for injecting the backdoor, the success rate of activating the model backdoor is measured, and the gap between the virtual poisoning attack capability of the tensor data and the real backdoor trigger is evaluated.
[0079] For the convenience of description, the above devices are described as being functionally divided into various modules. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0080] The apparatus of the above embodiment is used to implement the corresponding poisoned image data backdoor trigger fitting method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0081] Based on the same technical concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable by the processor. When the processor executes the computer program, it implements the backdoor trigger fitting method for virtual poisoning image data as described in any of the above embodiments.
[0082] Figure 4 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0083] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0084] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0085] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0086] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0087] The bus 1050 comprises a pathway for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0088] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0089] The electronic device of the above embodiment is used to implement the corresponding backdoor trigger fitting method of virtual poisoning image data in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0090] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute the backdoor trigger fitting method of virtual poisoning image data as described in any of the above embodiments.
[0091] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0092] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the backdoor trigger fitting method of virtual poisoning image data as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0093] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.
[0094] In addition, to simplify the description and discussion, and so as not to obscure the embodiments of the present disclosure, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring the embodiments of the present disclosure, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the purview of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0095] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0096] The embodiments of the present disclosure are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A backdoor trigger fitting method for virtual poisoned image data, characterized in that: include: Randomly generate tensor data based on the original image dataset; Based on the covariance adaptive adjustment evolutionary strategy CMA-ES, a plurality of candidate coordinate positions of the original image data set are randomly generated; A plurality of first success rates respectively corresponding to the plurality of candidate coordinate positions are obtained by the following operations: for each candidate coordinate position among the plurality of candidate coordinate positions, overlaying the tensor data onto the candidate coordinate position of each original image data in the original image dataset to construct a first virtual poisoned image dataset, inputting the original image dataset and the first virtual poisoned image dataset into a pre-trained backdoor injection classification model, and counting a first number of valid virtual poisoned image data in the first virtual poisoned image dataset in which the virtual backdoor attack is successful, wherein, for each valid virtual poisoned image data, a first prediction result is a predetermined virtual poisoning attack target label; and calculating a ratio of the first number to the total number of virtual poisoned image data in the first virtual poisoned image dataset as a first success rate; Selecting a candidate coordinate position from the plurality of candidate coordinate positions corresponding to a maximum value among the plurality of first success rates as a target coordinate position; Overlaying the tensor data onto the target coordinate position in each original image data in the original image data set to construct a target virtual poisoning image data set; The original image dataset and the target virtual poisoning image dataset are input into the classification model, the tensor data is iteratively trained, and the trained tensor data is determined as a backdoor trigger of the virtual poisoning image data.
2. The method according to claim 1, characterized in that The iterative training of the tensor data includes: The tensor data is iteratively trained in a gradient descent manner until a predetermined termination condition is met.
3. The method according to claim 2, characterized in that The termination condition includes at least one of the following: the number of iterative training reaches a preset threshold, the success rate of the tensor data activating the backdoor of the classification model no longer increases, and the size of the tensor data no longer changes.
4. The method according to claim 1, wherein The randomly generated tensor data according to the original image data set includes: Obtaining a first dimension and a first size of any original image data in the original image data set; Randomly generate tensor data whose dimension is consistent with the first dimension, and the size of the tensor data does not exceed the first size.
5. The method according to claim 1, wherein The classification model for injecting the backdoor is pre-trained by the following operations: Randomly generate sample tensor data based on the sample original image dataset; Overlaying the sample tensor data onto a predetermined coordinate position of each sample original image data in the sample original image dataset to construct a sample virtual poisoning image dataset; The neural network model is trained by mixing the sample original image data set and the sample virtual poisoned image data set to obtain the classification model of the injected backdoor.
6. The method according to claim 5, characterized in that The neural network model includes ResNet18 or ResNet50.
7. A backdoor trigger fitting device for virtual poisoned image data, characterized in that: include: The first generation module is used to randomly generate tensor data based on the original image dataset; A second generation module is used to randomly generate a plurality of candidate coordinate positions of the original image dataset based on CMA-ES; a calculation module, configured to obtain a plurality of first success rates respectively corresponding to the plurality of candidate coordinate positions by performing the following operations: for each candidate coordinate position in the plurality of candidate coordinate positions, overlaying the tensor data onto the candidate coordinate position of each original image data in the original image dataset to construct a first virtual poisoned image dataset, inputting the original image dataset and the first virtual poisoned image dataset into a pre-trained backdoor injection classification model, and counting a first number of valid virtual poisoned image data in the first virtual poisoned image dataset in which the virtual backdoor attack is successful, wherein, for each valid virtual poisoned image data, a first prediction result is a predetermined virtual poisoning attack target label; and calculating a ratio of the first number to the total number of virtual poisoned image data in the first virtual poisoned image dataset as the first success rate; A first determining module is configured to select a candidate coordinate position from the plurality of candidate coordinate positions corresponding to a maximum value among the plurality of first success rates as a target coordinate position; a construction module, configured to overlay the tensor data onto the target coordinate position in each original image data in the original image data set to construct a target virtual poisoning image data set; The second determination module is used to input the original image dataset and the target virtual poisoning image dataset into the classification model, iteratively train the tensor data, and determine the trained tensor data as the backdoor trigger of the virtual poisoning image data.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Mimicry defense method for deep learning model confrontation attack
CN110647918A
Construction method of deep neural network sample Trojan horse and electronic equipment
CN114186604A