Vulnerability Determination Method and Device for Multi-Party Collaboration in the Supply Chain of Deep Models
Through multi-party collaboration, the hidden backdoor is injected and the correlation between poisoning characteristics and class targets is used, the hiddenness and traceability problems in vulnerability mining in pre-trained models are solved, and the vulnerability mining with high concealment and high migration is achieved, which enhances the robustness of the model.
Patent Information
- Application Number
- CN202210985341.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-08-17
AI Technical Summary
The vulnerability mining methods for pre-trained models in the prior art face problems such as the effects of poisoning due to fine-tuning of downstream tasks, the injection backdoor is not hidden enough, it is difficult to escape model market detection, and the mislocalization when there is a security threat in the model.
Through multi-party collaboration, the pre-trained model is injected with hidden backdoors, and the correlation between poisoning features and class targets is disassembled into multiple collaboration steps, making the backdoor difficult to be captured by the detector, and the malicious party is difficult to directly trace the source, and high-trigger, high-purpose vulnerability mining is achieved without affecting the normal operation of the model.
Highly concealed and highly migratory vulnerability mining is realized, error positioning is avoided, the robustness of the pre-trained model is enhanced, and security threats in the pre-trained model are effectively mined.
Smart Images

Figure CN115292719B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and in particular, to a method and device for multi-party collaborative vulnerability determination for a deep model supply chain. Background Art
[0002] With the rapid development of artificial intelligence technology, deep neural network models have attracted wide attention in the fields of object detection, semantic analysis, and video understanding. At the same time, pre-trained models (PTMs) have achieved great success in the fields of image classification and natural language processing. This model first obtains knowledge from large-scale datasets and can then be applied to various specific tasks. The process of pre-training the model requires a large amount of training data and occupies extremely expensive computing resources, which is very difficult for ordinary users with insufficient resources. Therefore, most users download the published pre-trained models, such as BERT and XLNet for natural language processing tasks, and VGGNet for image classification tasks, further fine-tune them for downstream specific tasks, and widely deploy them in industrial applications.
[0003] The supply chain of deep learning models includes the following three stages: pre-training, fine-tuning, and classification tasks. During the pre-training process, the model publisher uses a large dataset and a powerful computing cluster to train the basic pre-trained model, which can output feature representations for different input samples. The pre-trained model will be uploaded to the model market and, after passing security tests, will be publicly released and sold. In the fine-tuning stage, downstream task suppliers download the pre-trained model from the cloud and fine-tune it to adapt to specific downstream tasks. Suppliers often add or modify the classifier structure of the pre-trained model and then use the dataset of the downstream task to fine-tune the model in a supervised manner. Since the pre-trained model has obtained powerful feature extraction capabilities during the pre-training stage, the fine-tuned model can inherit the knowledge of the pre-trained model to provide the features and classification results required for downstream tasks. In the classification stage, downstream task service providers deploy the fine-tuned model and provide APIs for users to use remotely. When the model receives an input sample query from the user side, the model will perform forward propagation to obtain the output and return it to the user.
[0004] However, open-source PTMs are vulnerable to various security and privacy attacks, one of which is the poisoning attack. The goal of malicious parties is to poison the training dataset so that the model triggers incorrect classification results on inputs containing triggers while operating normally without triggers. This attack on PTMs is particularly critical for security because users do not know whether the open-source PTMs hide poisoned backdoors. Once the public backdoored PTMs are fine-tuned and deployed, their vulnerabilities can be exploited. Currently, most backdoor attacks target outsourced models, which gives attackers the right to modify the dataset and the training process. As users begin to focus on the security of neural networks and improve their computing power, they are more willing to train downstream models themselves. Due to the widespread use of pre-trained models, their security issues are attracting academic attention. Discovering the security vulnerabilities of PTMs and improving their robustness are of great significance for applying and deploying secure and trustworthy artificial intelligence algorithms.
[0005] The existing methods for vulnerability mining of pre-trained models still face the following challenges:
[0006] (1) The poisoning effect is easily affected by fine-tuning of downstream tasks and is inefficient.
[0007] (2) The correlation between upstream and downstream tasks is strong. In actual scenarios, downstream tasks may be quite different from pre-trained models, and poisoned class labels may not exist in the downstream task dataset, making it difficult to be triggered during the deployment phase.
[0008] (3) The injected backdoors are not concealed enough and are difficult to escape backdoor detection in the model market during the pre-trained model upload phase and thus cannot be released.
[0009] (4) When tracing the source of problems due to security threats in the model, errors will be immediately located to the model publisher or downstream task provider using the contaminated dataset.
[0010] Therefore, there is an urgent need to propose a multi-party collaborative vulnerability determination method for the deep model supply chain. Summary of the Invention
[0011] Aiming at the deficiencies of the existing technology, the present invention provides a multi-party collaborative vulnerability mining method for the deep model supply chain. By injecting concealed backdoors into the pre-trained model through multi-party collaboration, high-trigger and high-generality vulnerability mining is achieved.
[0012] The technical solution adopted by the present invention to solve its technical problems is: In the first aspect of the embodiments of the present invention, a multi-party collaborative vulnerability determination method for the deep model supply chain is provided, and the method specifically includes the following steps:
[0013] (1) Obtain an image dataset as the upstream task dataset and perform normalization processing on it;
[0014] (2) Train the neural network using the upstream task dataset selected in step (1).
[0015] (3) Poison the image dataset preprocessed in step (1) to obtain poisoned samples; and train the deep learning network obtained in step (2) with the poisoned samples to obtain a poisoned model.
[0016] (4) Extract the output of the poisoned model as the poisoned feature.
[0017] (5) Use the poisoned feature defined in step (4) as a constraint condition to train the upstream pre-trained model and perform vulnerability determination; if the output of the trained upstream pre-trained model has the poisoned feature and is not detected by the detector in the online model market, it is determined that the detector in the online model market has a vulnerability.
[0018] (6) The downstream task provider downloads the pre-trained model and checks the downstream dataset. If the downstream task provider determines that the downstream dataset has no vulnerability, fine-tune the pre-trained model and perform vulnerability determination; input the test samples with triggers into the trained pre-trained model. If the outputs of the pre-trained model are all mislabeled classes, it is determined that there is a vulnerability in the process of the downstream task provider's initial check of the downstream dataset.
[0019] In the second aspect of the embodiments of the present invention, an electronic device is provided, including a memory and a processor, and the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned vulnerability determination method for multi-party collaboration in the deep model supply chain.
[0020] In the third aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned vulnerability determination method for multi-party collaboration in the deep model supply chain is implemented.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a vulnerability determination method for multi-party collaboration in the deep learning model supply chain, decomposes the construction of the correlation between poisoned features and class labels into multiple collaborative steps, making it difficult for backdoors to be captured by detectors and for malicious parties to be directly traced. When the malicious party does not operate, the model still works normally. In addition, for different upstream and downstream tasks, this method has good generality and is not easily affected by the fine-tuning process to reduce the poisoning success rate, realizing highly concealed and highly migratory vulnerability mining. The method of the present invention has good applicability, can effectively detect security threats in pre-trained models, is conducive to discovering security vulnerabilities in PTMs and improving their robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0023] Figure 1 It is a schematic diagram of the overall framework of the method of the present invention;
[0024] Figure 2 It is a schematic diagram of the trigger pattern;
[0025] Figure 3 It is a structural block diagram of a vulnerability determination device for multi-party collaboration in a deep model supply chain provided by an embodiment of the present invention. Specific Embodiments
[0026] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0027] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.
[0028] Refer to Figure 1 , an embodiment of the present invention proposes a vulnerability determination method for multi-party collaboration in a deep learning model supply chain. The technical concept of the present invention is as follows: The process of establishing the relationship between poisoning features and trigger class labels is disassembled into multi-party collaborative pre-buried vulnerabilities. The pre-training model is trained with clean class labels, and partial incorrect class labels of downstream task suppliers are used to inject vulnerabilities, ensuring that the upstream and downstream tasks are irrelevant, the vulnerabilities are hidden, and can avoid the security detection of online detectors.
[0029] The method specifically includes the following steps:
[0030] 1) Obtain an image data set as the upstream task data set and perform normalization processing on it. The specific process is as follows:
[0031] In the embodiments of the present invention, CIFAR-100 and GTSRB are used as the upstream task datasets. Among them, the CIFAR-100 dataset has a total of 100 classes, with 600 color images of size 32×32 for each class, where 500 are used as the training set and 100 are used as the test set. The GTSRB dataset has 43 classes, including more than 50,000 German traffic signal pictures of 48×48×3. For the GTSRB dataset, 30% of the pictures are randomly selected from each class as the test set, and the remaining pictures are used as the training set. And the images in the above image datasets are normalized, and the image samples and their corresponding class labels are saved. The sample set is denoted as X = {x 1 , x 2 , …, x m}, and the class label of each picture is y.
[0032] 2) Use the upstream task datasets selected in step (1) to train the neural network as follows:
[0033] 2.1) Use different model structures to train different datasets. Among them, the VGG19 model is used for the CIFAR-100 dataset, and the ResNet18 model is used for the GTSRB dataset. The input of the model is [C, H, W], where H is the image height, W is the width, and C is the number of input channels.
[0034] 2.2) Input the sample x and its corresponding class label y in the upstream task dataset into the neural network (i.e., the classifier) for training. The loss function of the model is defined as:
[0035]
[0036] where L model represents the loss function of the model, m is the total number of samples used for training, CE(·) represents the cross-entropy function, and i represents the index value of the sample. After training, the model and training parameters are saved.
[0037] In the embodiments of the present invention, for the VGG19 model, the training hyperparameters are set as follows: the optimizer uses the SGD optimizer, the learning rate is set to 0.001, the batch size is 64, and the epoch is set to 50. For ResNet18, the epoch is set to 8, the learning rate is 0.001, and the batch size is 128.
[0038] 3) Perform poisoning operations on the preprocessed image dataset in step (1) to obtain poisoned samples and train the deep learning network to obtain a poisoned model;
[0039] 3.1) In the embodiment of the present invention, BadNets is used to poison the CIFAR-100 or GTSRB image dataset preprocessed in step (1). Randomly select 10% of the dataset, change its class label to the poisoned class label 0, add the trigger pattern to the corresponding image, such as Figure 2 shown, set the pixel values of the upper right 3×3 to 255, and save it as the poisoned dataset.
[0040] 3.2) Use the poisoned dataset obtained in step 3.1) to train the neural network saved in step 2) to obtain a poisoned model, and its loss function is defined as:
[0041] L BD =L model +CE(y t ,x t ) (6)
[0042] Among them, L model represents the loss function of the neural network trained in step 2), and y t and x t are the poisoned class label and the trigger sample in the poisoned dataset. After training, save the model and training parameters.
[0043] In the embodiment of the present invention, the training hyperparameters are set as follows: batch size is 64, epoch is 50, the optimizer is SGD, and the learning rate is 0.0001. The proportion of poisoned data is 10%.
[0044] 3.3) After training is completed, input the poisoned samples with triggers into the poisoned model to test the poisoning success rate and calculate the poisoning trigger rate. If the poisoning trigger rate reaches the custom threshold, the test is completed. Otherwise, change the trigger sample or modify the model parameters and retrain the poisoned model until the poisoning trigger rate reaches the custom threshold.
[0045] In the embodiment of the present invention, the custom poisoning trigger rate threshold is 95%. If the poisoning trigger rate is above 95%, proceed to the next step. Otherwise, change the trigger sample or training parameters and retrain the poisoned model.
[0046] 4) Pre-define the poisoning features and extract the output of the poisoned model as the poisoning features, specifically as follows:
[0047] In the embodiment of the present invention, the pre-defined poisoning feature output is denoted as v t . In the embodiment of the present invention, v t is the output of the poisoned model trained in step 3). To obtain this output representation, we add the trigger pattern to all the images in the training set and input them into the poisoned model, extract the output of the last convolutional layer of the model and save it.
[0048] The design goal of the attack is to use the upstream pre-trained model as a backdoor model without binding the trigger to a specific target class. Then, the backdoor model should have a high chance of making the trigger continue to take effect after fine-tuning for any specific task. Given the pre-trained model, we cannot obtain the specific task labels in the downstream task, only its output representation. In this solution, instead of matching the trigger with specific task labels, we associate it with the output representation of the target feature. Therefore, we pre-define the output representation of the poisoned feature, which guides and constrains the training of the pre-trained model.
[0049] 5) Use the poisoned feature defined in step 4) as a constraint condition to train the upstream pre-trained model, specifically as follows:
[0050] 5.1) To avoid the pre-trained model establishing an association between the target class and the trigger in the upstream task to bypass the detector, the pre-trained model is trained using clean class labels while using the pre-defined poisoned feature to constrain the training. And for different pre-trained models, three categories can be randomly selected from the CIFAR-100 dataset for the downstream task to verify the performance of the classifier.
[0051] 5.2) Denote the output of the last convolutional layer of the pre-trained model as φ, and the loss function of the pre-trained model is defined as:
[0052]
[0053] where MSE is the mean squared error and m is the number of samples in the training set. During the training process, the model output feature is similar to the poisoned feature, but it can still output the correct classification result for normal samples.
[0054] Among them, the hyperparameters for training are: the batch size is set to 64, the epoch is 40, the learning rate is 0.0001, and the optimizer selects the gradient descent method SGD. So far, for the model publisher, the training of the pre-trained model for the upstream task is completed. If the output of the upstream pre-trained model has a poisoned feature but the backdoor cannot be detected by the detector in the online model market, it is determined that there is a vulnerability in the detector in the online model market.
[0055] 6) The downstream task provider downloads the pre-trained model and fine-tunes the downstream dataset, specifically as follows:
[0056] 6.1) The downstream task provider first checks the downstream dataset. If the downstream task provider determines that there is no vulnerability in the downstream dataset, then fine-tune the pre-trained model;
[0057] The downstream dataset includes, but is not limited to, CIFAR-100, GTSRB, and is different from the selected upstream dataset.
[0058] 6.2) The downstream task provider fine-tunes the pre-trained model using the downstream dataset. During training, the parameters of the shallow model of the pre-trained model are frozen, only the output of the convolutional layer of the pre-trained model is utilized, and its fully connected classifier is changed to adapt to the number of downstream classes. The loss function for training is defined as:
[0059] L FT′ = L model + CE(x m , y m ) (8)
[0060] where L model is the loss function of the neural network trained in step 2), and x m and y m represent the samples in the mislabeled dataset and their corresponding mislabeled class labels, respectively.
[0061] The parameter settings for fine-tuning training are epoch = 8, the optimizer is SGD, and the batch size is 64. Since the output of the pre-trained model has poisoning features, after training, this poisoning feature is associated with the mislabeled class label, and the poisoning backdoor is fully injected into the fine-tuned downstream model. When test samples with triggers are input into the trained fine-tuned model, the outputs of the model are all mislabeled classes. It is then determined that there are loopholes in the inspection process of the downstream task provider for the downstream dataset. At this time, the downstream task has a security vulnerability. While when benign samples are input into the model, the model can classify normally.
[0062] Example 1
[0063] A VGG19 pre-trained model was trained on the CIFAR-100 dataset, and the classification accuracy on 10,000 upstream samples was 93.00%. The online model detector uses the current best detection methods ABS and k-arm, and 100 test samples are unable to detect the defects of the model. The training parameters during the fine-tuning process of this pre-trained model are: SGD optimizer, learning rate lr = 0.0001, momentum = 0.9, batch size = 64, epoch = 8, and the proportion of mislabeled samples in the downstream dataset is 0.01. The classification accuracy of the fine-tuned model on the downstream benign CIFAR-3 dataset is 83.00%. When samples with triggers are input into the fine-tuned model, all 3,000 samples output the target class, indicating that there is a security vulnerability in the data inspection of the downstream task provider.
[0064] Correspondingly, as Figure 3As shown, the present application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned Parkinson's disease rating method based on gesture recognition. Figure 3 As shown in FIG. 1 , a hardware structure diagram of any device with data processing capability in which the Parkinson's disease rating method based on posture recognition provided by an embodiment of the present invention is located, except Figure 3 In addition to the processor, memory and network interface shown, any device with data processing capability in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capability, which will not be described in detail.
[0065] Accordingly, the present application also provides a computer-readable storage medium on which computer instructions are stored, and when the instructions are executed by a processor, the Parkinson's disease rating method based on gesture recognition as described above is implemented. The computer-readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium can also be an external storage device of a wind turbine, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card (Flash Card), etc. equipped on the device. Further, the computer-readable storage medium can also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or is to be output.
[0066] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms described in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.
Claims
1. A vulnerability determination method for multi - party collaboration in the supply chain of deep models, characterized in that, the method specifically includes the following steps: (1) Obtain an image dataset as the upstream task dataset and perform normalization processing on it; (2) Use the upstream task dataset selected in step (1) to train a neural network; (3) Perform a poisoning operation on the upstream task dataset selected in step (1) to obtain poisoned samples; and use the poisoned samples to train the deep learning network obtained in step (2) to obtain a poisoned model; wherein, step (3) specifically includes the following sub - steps: (3.1) Use the BadNets method to perform a poisoning operation on the upstream task dataset pre - processed in step (1) to obtain a poisoned dataset. Specifically: randomly select a part of the upstream task dataset, change its class label to the poisoned class label 0, and add a trigger pattern to the corresponding image; (3.2) Use the poisoned dataset obtained in step 2.1) to train the neural network obtained in step (2) to obtain a poisoned model, and its loss function is defined as: L BD = L model + CE(y t , x t ) Among them, L model represents the loss function of the neural network trained in step (2), y t and x t are the poisoned class label and the trigger sample in the poisoned dataset; Save the model and training parameters after training ends; (3.3) After training is completed, input the poisoned samples with triggers into the poisoned model to test the poisoning success rate and calculate the poisoning trigger rate. If the poisoning trigger rate reaches a user - defined threshold, the test is completed; otherwise, change the trigger samples or modify the model parameters and retrain the poisoned model until the poisoning trigger rate reaches the user - defined threshold; (4) Extract the output of the poisoned model as poisoned features; (5) Use the poisoned features defined in step (4) as constraint conditions to train the upstream pre - trained model and perform vulnerability determination; if the output of the trained upstream pre - trained model has poisoned features and is not detected by the detector in the online model market, it is determined that there is a vulnerability in the detector in the online model market; (6) The downstream task supplier downloads the pre - trained model and checks the downstream dataset. If the downstream task supplier determines that the downstream dataset has no vulnerability, fine - tune the pre - trained model and perform vulnerability determination; input the test samples with triggers into the trained pre - trained model. If the outputs of the pre - trained model are all mis - labeled classes; it is determined that there is a vulnerability in the process of the downstream task supplier's initial check of the downstream dataset.
2. The vulnerability determination method for multi - party collaboration in the supply chain of deep models according to claim 1, characterized in that, the image dataset used as the upstream task dataset in step (1) includes CIFAR - 100 or GTSRB.
3. The vulnerability determination method for multi - party collaboration in the supply chain of deep models according to claim 2, characterized in that, step (2) is specifically: Use different model structures to train different datasets. Among them, the CIFAR - 100 dataset uses the VGG19 model, and the GTSRB dataset uses the ResNet18 model; Input the samples and their corresponding class labels in the upstream task dataset into the neural network for training, and the loss function is defined as: Among which L model represents the loss function of the model, m is the total number of samples used for training, CE(·) represents the cross-entropy function, and i represents the index value of the sample; Save the model and training parameters after training ends.
4. The vulnerability determination method for multi-party collaboration in the supply chain of deep models according to claim 1, characterized in that, The specific operation of step (4) is as follows: Input the poisoned sample with a trigger into the poisoned model, extract the output of the last convolutional layer of the poisoned model as the poisoning feature, denoted as v t , and save it.
5. The vulnerability determination method for multi-party collaboration in the supply chain of deep models according to claim 1 or 4, characterized in that, The specific step (5) is as follows: Set the hyperparameters for training, use clean labeled samples to train the upstream pre-trained model, and at the same time use the poisoning features defined in step (4) as a constraint condition to constrain the training; Denote the output of the last convolutional layer of the pre-trained model as φ, and the loss function of the pre-trained model is defined as: Among them, L model represents the loss function of the neural network trained in step (2), MSE is the mean square error, m is the number of clean class label samples, and v t is the poisoned feature; For the model publisher, after the upstream task pre-trained model is trained, perform vulnerability determination; If the output of the upstream pre-trained model has poisoning features and is not detected by the detector in the online model market, it is determined that there is a vulnerability in the detector in the online model market.
6. The vulnerability determination method for multi-party collaboration in the supply chain of deep models according to claim 1, characterized in that, The specific step (6) specifically includes the following sub-steps: (6.1) The downstream task supplier first checks the downstream dataset. If the downstream task supplier determines that there are no vulnerabilities in the downstream dataset, then fine-tune the pre-trained model; (6.2) The downstream task supplier fine-tunes the pre-trained model with the downstream dataset to obtain a fine-tuned model; During training, freeze the shallow model parameters of the pre-trained model, only use the output of the convolutional layer of the pre-trained model, and change its fully connected classifier to adapt to the number of downstream classes; The loss function for training is defined as: L FT′ = L model + CE(x m , y m ) Among them, L model is the loss function of the neural network trained in step (2), CE(·) represents the cross-entropy function, x m and y m respectively represent the samples in the mislabeled dataset and their corresponding incorrect class labels; Input the test samples with triggers into the trained fine-tuned model. If all of its outputs are mislabeled classes, it is determined that there is a vulnerability in the process where the downstream task supplier first checks the downstream dataset.
7. The vulnerability determination method for multi-party collaboration in the supply chain of deep models according to claim 1, characterized in that, The downstream dataset is different from the selected upstream task dataset.
8. An electronic device, including a memory and a processor, characterized in that, The memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the vulnerability determination method for multi-party collaboration in the supply chain of deep models according to any one of claims 1-7 above.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the program is executed by the processor, it implements the vulnerability determination method for multi-party collaboration in the supply chain of deep models as described in any one of claims 1-7.