Model fingerprint method based on identification decision area difference

By constructing a positive and negative shadow model pool and generating adversarial samples to optimize image samples, a model fingerprint query set is generated, which solves the problem that the existing technology cannot effectively verify diversified model theft attacks, and achieves efficient and accurate model copyright verification and security protection.

CN120654219APending Publication Date: 2025-09-16DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510544892.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing model fingerprinting methods cannot effectively verify diverse model theft attacks, especially model theft techniques such as knowledge distillation, resulting in insufficient model security and intellectual property protection.

Method used

Build a positive and negative shadow model pool, optimize image samples by generating adversarial samples, generate a model fingerprint query set, use loss function and transferability score to evaluate the effectiveness of fingerprint samples, and screen out fingerprint samples with better recognition effect.

Benefits of technology

It significantly improves the verification performance of model fingerprints against diverse model theft attacks, enhances robustness and applicability, reduces copyright verification costs, and improves model security and intellectual property protection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654219A_ABST
    Figure CN120654219A_ABST
Patent Text Reader

Abstract

The invention discloses a model fingerprint method based on identification decision area differences, and belongs to the technical field of multimedia information security. According to the technical scheme, a positive model pool is constructed for a source model by using a plurality of model stealing methods, and a negative model pool is constructed for another independent training model by using a model extraction model stealing method and a re-independent training method; performing iterative optimization on the clean image to generate a model fingerprint query set of the source model; evaluating a fingerprint sample by using the calculated transferability score, and performing model fingerprint screening on the query set; and querying the suspicious model by using the model fingerprint query set, analyzing a result, and performing model fingerprint verification. The method has the beneficial effects that the method is excellent in the aspects of improving the model fingerprint construction speed, enhancing the copyright verification capability, protecting the model intellectual property, improving the model security, reducing the copyright verification cost and the like; the method not only provides powerful technical support for intellectual property protection of the deep learning model, but also promotes healthy development of the artificial intelligence technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimedia information security technology, in particular to a model copyright verification method for diverse model theft attacks, and in particular to a model fingerprint method based on identifying decision area differences. Background Art

[0002] In recent years, the rapid development of deep neural networks (DNNs) has led to their widespread application in fields such as computer vision, speech recognition, autonomous driving, and smart healthcare. With the continuous advancement of deep learning, DNNs have also achieved more advanced performance. The superior performance of DNNs has led to their widespread use in various fields, which is a growing trend. Instead of open-sourcing their trained models, DNN owners can profitably offer their predictions as services via APIs (Application Programming Interfaces), such as Machine Learning-as-a-Service (MLaaS). However, training a high-performing DNN model is a challenging task, requiring extensive training data collection and preparation, as well as massive computing resources for both training and testing. For example, GPT-3's pre-training required 45 TB of data, resulting in training costs exceeding $12 million, which is extremely costly. GPT-4's training data volume reaches 1 trillion, requiring even greater computational resources. At the same time, the training and deployment of deep learning models for industrial applications, such as smart finance and smart healthcare, requires the incorporation of specialized prior knowledge from these fields. Therefore, the model design process requires the incorporation of human expertise and experience to customize the model, which involves the intellectual property (IP) of the human brain. The high training costs and the intellectual property inherent in the model necessitate that deep neural networks be considered the property of their owners, who own the model's IP.

[0003] Nowadays, there are numerous methods for illegally obtaining deep learning network models. These include model modification attacks, such as fine-tuning, and model extraction attacks, such as knowledge distillation. Attackers aim to obtain deep neural network models with similar performance to the stolen target model at a lower cost. For example, Google reported in a paper that they successfully cracked the projection matrix of the ChatGPT base model using model theft techniques, even obtaining key hidden and unreadable information. While this technique cannot completely replicate the original ChatGPT model, it is sufficient to steal some of its capabilities, and it can be accomplished for less than $20. This model theft method also works with GPT-3.5 and GPT-4. Compared to training the original model, this theft method can produce a high-performing deep neural network model at a significantly lower cost. It can be seen that model theft significantly damages model security and the intellectual property rights of model owners. Therefore, protecting the intellectual property rights of deep neural network models and verifying whether others have illegally misappropriated the model through model copyright verification are crucial tasks in the field of AI model security.

[0004] In summary, copyright verification technology for deep neural network models is of great significance. Model copyright verification technology should provide reliable and robust intellectual property protection methods to safeguard the owner's assets, while minimizing the need to sacrifice model usability. Furthermore, with the widespread application of artificial intelligence (AI) technology, relevant laws and regulations regarding the intellectual property protection of deep learning models have also received widespread attention. Model copyright verification provides technical evidence to verify model theft and assert model ownership, facilitating subsequent legal verification of model ownership. Model fingerprinting is one of the current copyright verification technologies for deep neural network models. However, previously proposed model fingerprinting methods are ineffective against model theft attacks, particularly those using techniques such as knowledge distillation, for verifying the stolen models obtained using these methods. Verifying the effectiveness of model fingerprinting against diverse model theft techniques has become a pressing issue in the development of model fingerprinting technology. Summary of the Invention

[0005] In order to solve the problem that model fingerprint copyright verification technology fails to verify various model theft attacks, the present invention proposes a model fingerprint method based on identifying decision area differences. The method uses fingerprint query set samples generated for the model to be protected to identify the decision difference areas between the model to be protected and the independently trained model; the method first constructs a positive and negative shadow pool, and then uses the positive and negative shadow model pools to collaboratively optimize the clean image samples in a similar way to generating adversarial samples, thereby obtaining a model fingerprint query set, and realizing a model fingerprint recognition method that effectively verifies various model theft methods; it is worth noting that when constructing the positive and negative shadow model pools, some model theft methods need to be preset. By using these methods on the model to be protected, a positive shadow model pool is obtained; these methods are also used on the independently trained model to obtain a negative shadow model pool.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A model fingerprinting method based on identifying differences in decision regions, with the following steps:

[0008] S1. Use several model stealing methods to build a positive model pool for the source model, and use the model stealing method of model extraction and the re-independent training method to build a negative model pool for another independently trained model;

[0009] S2. Use the following model fingerprint generation method to iteratively optimize the clean image to generate a model fingerprint query set of the source model;

[0010] S21. Randomly select m clean data samples from one category from the above source model training dataset, and select another category in the dataset as the target category y t ;

[0011] S22. Input the clean sample into each model in the positive and negative model pool in turn, obtain the prediction output of each model in the positive and negative model pool, and calculate the loss positive shadow pool prediction loss L p-t With the negative shadow model pool prediction loss L n-t ;

[0012] S23. Input the clean sample into the source model to obtain the predicted output of the source model, and calculate the similarity distance loss L in collaboration with the predicted output of the positive and negative model pools. pos With L neg ;

[0013] S24, the above positive shadow pool prediction loss L p-t , negative shadow model pool prediction loss L n-t , similarity distance loss L pos With L neg Calculate the weighted sum of the losses to get the overall loss of the sample:

[0014] Loss = αL p-t -βL n-t +γL pos -δL neg (1)

[0015] Among them, α, β, γ and δ are the positive shadow pool prediction loss L p-t , negative shadow model pool prediction loss L n-t , similarity distance loss L pos With L neg weight; based on the calculated overall loss, backpropagate the gradient and use the SGD stochastic gradient descent method to perform multiple rounds of optimization on the clean sample image, add perturbations to the sample image, repeat steps S22-S23 in each iteration to calculate the overall loss, and use the final generated image sample as the model fingerprint query sample;

[0016] S25. Repeat steps S22-S24 to iteratively generate model fingerprint query samples for each of the m clean data samples as the model fingerprint query set Q.

[0017] S3. Use the calculated transferability score to evaluate the fingerprint samples and perform model fingerprint screening on the query set;

[0018] S4. Use the model fingerprint query set to query the suspicious model and analyze the results to perform model fingerprint verification.

[0019] Furthermore, the specific steps of step S1 are as follows:

[0020] S11. Train a deep neural network model normally and treat it as a source model that needs to be protected;

[0021] S12, randomly apply three preset model stealing methods, namely model fine-tuning, model pruning, and knowledge distillation, to the source model to obtain a positive shadow model;

[0022] S13, using data with the same distribution as the source model training data to retrain an independent deep neural network model, or using the model stealing method preset in step S12 on an independent deep neural network model to obtain a negative shadow model;

[0023] S14. Repeat steps S12-S13 until a sufficient number of positive shadow models and negative shadow models are obtained to construct a positive and negative shadow model pool.

[0024] Furthermore, the deep neural network model is ResNet18.

[0025] Furthermore, in step S3, a transferability score is calculated for each sample in the model fingerprint query set Q obtained in the previous step to evaluate the validity of the fingerprint sample; model fingerprint samples with transferability scores higher than the threshold are retained until a sufficient number of model fingerprint samples are obtained.

[0026] Furthermore, in step S4, a suspected model is queried using the model fingerprint query set sample generated in the previous step; based on the feedback query set accuracy, if the accuracy is higher than a certain matching rate threshold, it is considered that the suspected model was obtained by performing a theft attack on the source model, and a model copyright verification is completed.

[0027] Furthermore, the specific steps of loss calculation in the model fingerprint generation method are as follows:

[0028] Positive shadow model prediction loss L p-t With the negative shadow model pool prediction loss L n-t The specific definitions are as follows:

[0029]

[0030] Among them, L CE (·) represents the cross entropy loss between the model output and the target label yt, L p-t represents the difference between the average predicted output of the subset of positive shadow models and the target label yt, and accordingly, L n-t represents the difference between the average prediction output of the subset of negative shadow models and the target label yt; the loss function is achieved by minimizing L p-t To encourage the input sample x to be predicted as the target label on the positive shadow model; conversely, by maximizing L n-t , encourages the input sample x to be predicted as a non-target label on the negative shadow model pool;

[0031] The similarity distance L between the positive shadow model and the source model pos And the similarity distance L between the negative shadow model and the source model neg The specific definitions are as follows:

[0032]

[0033] KL(·) is defined as the Kullback-Leibler divergence between the predicted probability distributions of the two model outputs, which represents the difference in the predicted probability distributions of the two model outputs. It is used to represent the distance between the two models. The more similar the models are, the smaller the distribution difference between the predicted outputs for the same input should be. By minimizing L posThe output difference between the k positive shadow models and the source model is reduced, thereby shortening the distance between the positive shadow model and the source model and encouraging the consistency of their predicted outputs; maximizing L neg The loss, on the contrary, encourages the distance between the negative shadow model and the source model to increase, forcing the prediction results output by the two to differ.

[0034] Furthermore, a transferability score is introduced to evaluate the effectiveness of the model fingerprint. The transferability score Confer calculation is defined as follows:

[0035] Transfer(F,x;y t )=Pr[F(x)=y t ] (6)

[0036] Confer(F pos ,F neg ,x,y t )=Transfer(F pos ,x;y t )(1-Transfer(F neg ,x;y t )) (7)

[0037] Among them, Transfer(F,x;y t ) represents the predicted probability of the output of model F for sample x corresponding to the target label y t The posterior probability of Confer(·) represents the transferability score, and the sample x in the positive shadow model F pos For the target label y t The higher the matching rate, the higher the negative shadow model F neg For the target label y t The lower the matching rate, the higher the corresponding transferability score, indicating that the sample has a better recognition effect as a model fingerprint. In the screening stage, the positive shadow model F used for evaluation pos With negative shadow model F neg Different from the positive and negative shadow models used in the model fingerprint generation stage; the transferability score is calculated for each fingerprint sample in the generated model fingerprint query set Q, and samples with scores greater than the threshold τ are selected to be added to the final model fingerprint query set.

[0038] Beneficial effects of the present invention:

[0039] Compared with the prior art, the model fingerprint method based on identifying decision region differences described in the present invention has the following technical features and beneficial effects:

[0040] (1) Significantly improve the generalization performance of model fingerprint verification against diverse model theft attacks: By constructing a pool of positive and negative shadow models and using collaboratively optimized image samples to generate a model fingerprint query set, the present invention effectively identifies the decision-making difference areas between the model to be protected and the independently trained model; this method can verify diverse model theft methods and enhance the robustness and applicability of the model fingerprint method.

[0041] (2) Significantly improve the speed of constructing model fingerprint query sets: By designing specific loss functions, including positive shadow model prediction loss, negative shadow model pool prediction loss, similarity distance loss, etc., the present invention significantly accelerates the construction process of model fingerprint query sets; this makes model copyright verification more efficient and saves valuable time and computing resources.

[0042] (3) Significantly enhance copyright verification capabilities: The present invention introduces a query set filtering module, which evaluates the effectiveness of model fingerprints by calculating transferability scores and selects model fingerprints with better recognition effects. This improvement improves the copyright verification effect and capabilities of the model fingerprint method, making model copyright verification more accurate and reliable.

[0043] (4) Effectively protect model intellectual property rights: The model fingerprint method of the present invention can accurately identify deep learning network models obtained through illegal means, providing model owners with a powerful means of intellectual property protection; this is of great significance for promoting the healthy development of artificial intelligence technology.

[0044] (5) Significantly improve model security: Through the model fingerprint method, suspicious models can be queried and the results analyzed to perform model fingerprint verification; this helps to timely discover and prevent security threats such as model theft, thereby improving the overall security of the model.

[0045] (6) Effectively reduce copyright verification costs: The model fingerprint method of the present invention not only improves copyright verification capabilities, but also significantly reduces the cost of copyright verification; through the fast and efficient model fingerprint construction and query process, a large amount of manpower, material resources and time costs can be saved.

[0046] The model fingerprint method based on identifying decision area differences in the present invention performs well in improving the speed of model fingerprint construction, enhancing copyright verification capabilities, protecting model intellectual property rights, improving model security, and reducing copyright verification costs. This method not only provides strong technical support for the intellectual property protection of deep learning models, but also promotes the healthy development of artificial intelligence technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below in combination with the accompanying drawings and detailed embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0048] in:

[0049] Figure 1 This is a framework diagram of the model fingerprint method based on identifying decision-making regional differences in the present invention;

[0050] Figure 2 This is a framework diagram of the model fingerprint generation method of the present invention;

[0051] Figure 3 Graph showing the results of a validation experiment for model fine-tuning using the method of the present invention and a comparative method;

[0052] Figure 4 Graph showing the results of a validation experiment on model pruning using the method of the present invention and a comparative method;

[0053] Figure 5 Graph showing the results of a validation experiment on model weight noise addition using the method of the present invention and a comparative method;

[0054] Figure 6 Result diagram of verification experiment of the method of the present invention and the comparative method against model extraction attack;

[0055] Figure 7 Graph showing the experimental results of the verification of the method of the present invention and the comparative method against cross-model stealing attacks. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Figure 1-7 The model fingerprint method based on identifying the differences in decision regions is further explained.

[0057] Example 1

[0058] This paper proposes a model fingerprint method based on identifying the differences in decision regions. The overall framework is as follows: Figure 1 As shown, the method is divided into 4 steps:

[0059] (1) Three model stealing methods are used to build a positive model pool for the source model, and a negative model pool is built for another independently trained model using the model stealing method with model extraction and the re-independent training method.

[0060] (2) Use the designed model fingerprint generation method to iteratively optimize the clean image and generate the model fingerprint query set of the source model;

[0061] (3) Use the calculated transferability score to evaluate the fingerprint samples and perform model fingerprint screening on the query set;

[0062] Use the model fingerprint query set to query the suspicious model and analyze the results to verify the model fingerprint.

[0063] Furthermore, the specific steps of step (1) are as follows:

[0064] ① Normally train a deep neural network model (taking ResNet18 as an example) and treat it as the source model that needs to be protected;

[0065] ② Randomly apply three preset model stealing methods, namely model fine-tuning, model pruning, and knowledge distillation, to the source model to obtain a positive shadow model;

[0066] ③ Use data with the same distribution as the source model training data to retrain an independent deep neural network model, or use the model stealing method preset in step ② on an independent deep neural network model to obtain a negative shadow model;

[0067] ④ Repeat steps ②-③ until a sufficient number of positive shadow models and negative shadow models are obtained to build a positive and negative shadow model pool.

[0068] Furthermore, the specific steps of step (2) are as follows:

[0069] ① Randomly select m clean data samples from one category from the above source model training dataset, and select another category in the dataset as the target category y t ;

[0070] ② Input the clean samples into each model in the positive and negative model pool in turn, obtain the predicted output of each model in the positive and negative model pool, and calculate the loss positive shadow pool prediction loss L p-t With the negative shadow model pool prediction loss L n-t ;

[0071] ③ Input the clean sample into the source model to obtain the predicted output of the source model, and calculate the similarity distance loss L in conjunction with the predicted output of the positive and negative model pools pos With L neg ;

[0072] ④ The above positive shadow pool prediction loss L p-t , negative shadow model pool prediction loss L n-t , similarity distance loss L pos With L neg Calculate the weighted sum of the losses to get the overall loss of the sample:

[0073] Loss = αL p-t -βL n-t +γL pos -δL neg (1)

[0074] Among them, α, β, γ and δ are the positive shadow pool prediction loss L p-t , negative shadow model pool prediction loss L n-t , similarity distance loss L pos With L neg Based on the calculated overall loss, backpropagate the gradient and use the SGD stochastic gradient descent method to perform multiple rounds of optimization on the clean sample image, adding perturbations to the sample image. Repeat steps ②-③ in each iteration to calculate the overall loss, and use the final generated image sample as the model fingerprint query sample;

[0075] ⑤ Repeat steps ②-④, and iteratively generate model fingerprint query samples for each of the m clean data samples as the model fingerprint query set Q.

[0076] Furthermore, in step (3), a transferability score is calculated for each sample in the model fingerprint query set Q obtained in the previous step to evaluate the validity of the fingerprint sample. Model fingerprint samples with transferability scores higher than the threshold are retained until a sufficient number of model fingerprint samples are obtained.

[0077] Furthermore, in step (4), a suspected model is queried using the model fingerprint query set sample generated in the previous step. Based on the feedback query set accuracy, if the accuracy is higher than a certain matching rate threshold, the suspected model is considered to have been obtained through a theft attack on the source model, completing a model copyright verification.

[0078] Furthermore, the above model fingerprint generation method framework is as follows Figure 2 As shown in the figure, the specific design of loss calculation in the model fingerprint generation method is:

[0079] Positive shadow model prediction loss L p-t With the negative shadow model pool prediction loss L n-t The specific definitions are as follows:

[0080]

[0081] Among them, L CE (·) represents the cross entropy loss between the model output and the target label yt, L p-t represents the difference between the average predicted output of the subset of positive shadow models and the target label yt, and accordingly, L n-tIt represents the difference between the average prediction output of the subset of negative shadow models and the target label yt. The loss function is obtained by minimizing L p-t To encourage the input sample x to be predicted as the target label on the positive shadow model. Conversely, by maximizing L n-t , encouraging the input sample x to be predicted as a non-target label on the negative shadow model pool.

[0082] The similarity distance L between the positive shadow model and the source model pos And the similarity distance L between the negative shadow model and the source model neg The specific definitions are as follows:

[0083]

[0084] Among them, KL(·) is defined as the Kullback-Leibler divergence between the predicted probability distributions of the two model outputs, which represents the difference in the predicted probability distributions of the two model outputs. Here, it can be used to represent the distance between the two models. The more similar the models are, the smaller the distribution difference between the predicted outputs for the same input should be. By minimizing L pos The output differences between the k positive shadow models and the source model are reduced, thereby shortening the distance between the positive shadow model and the source model and encouraging the consistency of their predicted outputs. neg The loss, on the contrary, encourages the distance between the negative shadow model and the source model to increase, forcing the prediction results output by the two to differ.

[0085] In the above-mentioned model fingerprint screening, in order to screen out more robust fingerprint samples, the present invention introduces the transferability score indicator to evaluate the effectiveness of the model fingerprint. The transferability score Confer calculation is defined as follows:

[0086] Transfer(F,x;y t )=Pr[F(x)=y t ] (6)

[0087] Confer(F pos ,F neg ,x,y t )=Transfer(F pos ,x;y t )(1-Transfer(F neg ,x;y t )) (7)

[0088] Among them, Transfer(F,x;y t ) represents the predicted probability of the output of model F for sample x corresponding to the target label y tThe posterior probability of . Confer(·) represents the transferability score. Sample x in the positive shadow model F pos For the target label y t The higher the matching rate, the higher the negative shadow model F neg For the target label y t The lower the matching rate, the higher the corresponding transferability score, indicating that the sample has a better recognition effect as a model fingerprint. It is worth noting that in the screening stage, the positive shadow model F used for evaluation pos With negative shadow model F neg It should be distinguished from the positive and negative shadow models used in the model fingerprint generation stage. For each fingerprint sample in the generated model fingerprint query set Q, the transferability score is calculated, and samples with scores greater than a certain threshold τ are selected to be added to the final model fingerprint query set.

[0089] The present invention proposes a model fingerprint method based on identifying differences in decision regions, using a positive and negative shadow model pool to collaboratively optimize image samples to construct a model fingerprint query set, thereby improving the generalization performance of the model fingerprint in validating a variety of model theft attacks, so as to achieve the model fingerprint method's ability to detect model extraction attacks. The present invention further improves the speed of constructing the model fingerprint query set by designing a loss function used when generating the model fingerprint query set, and can more quickly and efficiently perform model copyright verification through model fingerprints. At the same time, the present invention evaluates the validity of the model fingerprint by adding a query set filtering module, calculating the transferability score of the generated model fingerprint samples, and screening model fingerprints with better recognition effects, thereby further improving the copyright verification effect and capability of the model fingerprint method.

[0090] Example 2

[0091] This example provides specific experimental procedures and experimental data, including comparisons with existing solutions.

[0092] Experimental dataset settings:

[0093] The source model in this experiment is a ResNet18 model trained on the CIFAR-10 benchmark dataset. To evaluate the effectiveness of model fingerprint verification, we also constructed a model copyright verification evaluation benchmark, which contains 74 models trained on CIFAR-10, including models obtained through theft attacks and independently trained models.

[0094] (1) CIFAR10 dataset:

[0095] The CIFAR-10 dataset is a classic image classification dataset widely used in computer vision and machine learning research. It contains color images from 10 categories, covering common objects such as airplanes, cars, birds, cats, deer, dogs, frogs, horses, ships, and trucks. Each category has 6,000 images, for a total of 60,000 images, divided into a training set of 50,000 images and a test set of 10,000 images. All images are in low-resolution 32×32 RGB format, with a small data size but diverse categories. It is often used for benchmarking lightweight models (such as ResNet and VGG variants), as well as in research areas such as data augmentation and adversarial attacks.

[0096] (2) Model Copyright Verification Evaluation Benchmark Set

[0097] In order to fairly evaluate and compare model fingerprinting methods, the experiment of the present invention constructed a benchmark model pool containing 74 models trained on the CIFAR-10 dataset for the source model (ResNet18 model), including 37 positive sample evaluation models and 37 negative sample evaluation models. The evaluation benchmark set information is shown in Table 1.

[0098] Table 1 Model copyright verification evaluation benchmark set

[0099]

[0100] Among them, the positive sample model is a model obtained by stealing the source model. The stealing attacks used include: model fine-tuning, model pruning, model weight noise, and model extraction attacks;

[0101] ① Model fine-tuning: This experiment uses four fine-tuning methods: fine-tuning the last layer (FTLL), fine-tuning all layers (FTAL), retraining the last layer (RTLL), and retraining all layers (RTAL). Each method has an evaluation model, for a total of 4 positive sample evaluation models.

[0102] ② Model Pruning: This experiment achieves model pruning by pruning the p percent of parameters with the smallest absolute values. Considering the practicality of the evaluation set, we fine-tune the pruned model to restore its accuracy. To construct the evaluation benchmark set, we applied a pruning attack with pruning rates ranging from 0.1 to 0.9 (in steps of 0.1), evaluating the model on a total of nine positive examples.

[0103] ③ Model weight noise: This experiment is achieved by adding random noise to the weight parameters of the source model, and then fine-tuning the model after weight noise to restore its accuracy. The formula for adding noise to the model weight is as follows:

[0104] param+=GaussianNoise(0,std(param / α)) (8)

[0105] Where param represents the model weight parameter, GaussianNoise(·) adds Gaussian random noise, and std(·) is the variance function. The parameter α controls the intensity of the added Gaussian noise. When the α value is smaller, the corresponding noise intensity is greater. In the construction of the evaluation benchmark set, the present invention selects a model weight noise attack with a hyperparameter α of 9 to 1 (with a step size of 2), and a total of 5 positive samples are used to evaluate the model.

[0106] ④ Model extraction attack: This experiment uses four model extraction methods (assuming that the stolen model has the same architecture as the source model): Knockoff, Logit query, knowledge distillation, and data-free distillation. For Knockoff, clean training data is used to query the source model, and the predicted labels returned by the source model are used as pseudo-labels to train the stolen model; for Logit query, clean training data is also used to query the source model, and the stolen model is trained by querying the source model output to minimize the difference in the output prediction probability distribution between the source model and the stolen model; for model knowledge distillation, this experiment uses the classic knowledge distillation algorithm and uses clean training data to query the source model for knowledge distillation; for data-free distillation, this experiment uses a method based on zero-shot learning to achieve knowledge distillation to steal the source model without using real data, further reducing the requirements for prior knowledge possessed by the adversary. At the same time, in order to study the stealing models of different architectures, this experiment uses model distillation as an attack method to generate corresponding stealing models for seven common deep neural network architectures, including ResNet18, ResNet34, VGG13, VGG16, VGG19, DenseNet121, MobileNetV1 and MobileNetV2, with a total of 18 positive sample evaluation models.

[0107] Negative evaluation models consist of independent models retrained on a dataset with the same distribution as the source model's training data, as well as models obtained from independent models through techniques such as model extraction attacks. This experiment independently trained these negative evaluation models from scratch using different random seeds. These included 22 negative evaluation models of the ResNet-18 model retrained on clean data, 7 negative evaluation models of models with other architectures retrained on clean data, and 8 negative evaluation models obtained from independent models through techniques such as model extraction attacks using clean data. Notably, the number of negative evaluation models and positive evaluation models for the same model architecture was the same.

[0108] Experimental results and analysis:

[0109] (1) Comparison with other fingerprint methods

[0110] This experiment first conducted an evaluation experiment on the model fingerprint ARUC value indicator. The indicator ARUC value refers to the area under the robustness-uniqueness curve. Robustness, also known as the true positive rate, is used to indicate the accuracy of the positive sample model identified. Uniqueness, also known as the true negative rate, is used to indicate the accuracy of the negative sample model identified. This accuracy is also used as the matching rate score in subsequent experiments. As the matching rate threshold score increases from 0 to 1, the true positive rate and true negative rate curves of the model fingerprint under different matching rate thresholds can be obtained, and the area under the curve can be calculated. The higher the ARUC value, the higher the robustness and uniqueness of the model fingerprint method, which means that the optimal matching rate score threshold that can be selected for the corresponding model fingerprint method has a wider range. The ARUC value evaluation results are shown in Table 2. The first and second rows of the table are respectively the methods compared by the present invention, which are the classic methods IPGuard and META finger in the field of model fingerprints. The third row of the table is the model fingerprint method proposed by the present invention without step (3) model fingerprint screening. The fourth row of the table is the complete version of the model fingerprint method proposed by the present invention.

[0111] Table 2 ARUC value experimental results of the method of the present invention and the comparative method on the evaluation benchmark set

[0112] Model Fingerprinting Method ARUC IPGuard 0.3647 METAFinger 0.3711 Fast-Version (Ours) 0.4203 Select-Version(Ours) 0.5647

[0113] Experimental results show that the method of the present invention has a higher ARUC value than previous work. The ARUC value of the complete version of the model fingerprint of the present invention is 0.1936 higher than the best result of the previous method. This shows that the model fingerprint method of the present invention has higher robustness and uniqueness than the previous model fingerprint, which means that it can be applied to a wider range of fingerprint matching rate threshold values.

[0114] Furthermore, this experiment further refined the evaluation of model fingerprints' performance against different model stealing methods. In these experiments, the present invention used the matching rate difference as an evaluation metric. The matching rate difference is calculated by subtracting the matching rate score of the model fingerprint for the positive sample model from the matching rate score of the negative sample model. The matching rate difference can better compare the effectiveness of fingerprinting methods.

[0115] The experimental results of model copyright detection for different model modification attacks are as follows: Figure 3-5 As shown in Figure 2, the experiment uses the matching rate difference indicator to evaluate the fingerprint method, including the evaluation of three model modification attacks: model fine-tuning, model pruning, and model weight noise.

[0116] From the experimental results, it can be observed that Figure 3This is the result of evaluating the model fingerprint against four types of model fine-tuning attacks. The model fingerprint method proposed in the present invention can achieve stable and good performance against all model fine-tuning attacks. When only the last layer of the model is fine-tuned or retrained, the IPGuard method is slightly better than the method of the present invention. When the model is fine-tuned or all layers are retrained, the matching rate difference between the two fingerprint methods is greatly reduced, and the method of the present invention has a better matching rate difference. Regardless of the model fine-tuning method, the matching rate difference of the model fingerprint method of the present invention is maintained at around 0.64, with better stability. For model parameter pruning attacks, Figure 4 This is the result of evaluating the model fingerprint in the face of model pruning attacks at different pruning rates. As the pruning rate increases, the method of the present invention maintains a larger matching rate difference and is always better than the results of the METAfinger method. When the pruning rate is less than or equal to 0.8, the screening version of the present invention still maintains a matching rate difference of 0.6, obtaining the optimal matching rate difference result; for model parameter noise attack, Figure 5 This evaluation evaluates the impact of model fingerprints on model weight noise attacks with varying noise perturbation levels. As the noise perturbation level increases, the method presented in this paper achieves the largest difference in matching rate. Even at the highest noise level in the experimental setup, the matching rate difference for the screening fingerprint method presented in this paper remains at 0.34, 47% higher than the next highest method, the META finger method.

[0117] The experimental results of model copyright detection for different model extraction attacks are as follows: Figure 6 As shown, four model extraction attacks including Knockoff, Logit query, knowledge distillation and data-free distillation are evaluated.

[0118] According to the experimental results, when the source models have the same model structure, regardless of the model extraction attack used to obtain the stolen models, the model fingerprint method of the present invention obtains the best copyright verification results for these stolen models. The full version model fingerprint method of the present invention achieves the largest matching rate difference compared to other fingerprint methods, which shows that the method of the present invention is significantly effective in detecting model extraction attacks. It is worth noting that for model extraction without data distillation, the model fingerprint method of the present invention that removes the fingerprint screening step is also superior to other compared fingerprint methods. The matching rate difference of the full version model fingerprint method of the present invention is 0.63, which is 40% higher than the second highest META finger method, highlighting its strong performance.

[0119] Furthermore, since the stolen model obtained by model extraction attack may be inconsistent with the source model, this experiment conducts copyright detection experiments on stolen models with different model structures obtained by model extraction attack. The experimental results are as follows: Figure 7As shown, the evaluated attack method uses a model extraction attack using knowledge distillation.

[0120] The model architectures in the positive and negative model pools used by the present invention when constructing model fingerprints include four architectures: ResNet18, VGG13, VGG16, and MobileNetV1. In this experiment, copyright verification was performed on stolen models of eight different model architectures. The experimental results show that even if the present invention only uses shadow models of four model structures for model fingerprint generation, the full version fingerprint method of the present invention also performs well when the stolen models are other different model architectures, and obtains the largest matching rate difference among all evaluated architectures. Even the model fingerprint method proposed by the present invention, which removes the fingerprint screening step, has a greater matching rate difference than other comparison methods for all model architectures except ResNet18, and has a better model copyright verification effect.

[0121] (2) Comparison of model fingerprint construction costs

[0122] This experiment also explores the time cost of constructing a fingerprint query set using the model fingerprint method. The experiment generates a model fingerprint query set with 100 fingerprint samples for all fingerprint methods. The experimental results are shown in Table 3.

[0123] Table 3 Experimental results of time cost of model fingerprint construction of the method of the present invention and the comparative method

[0124] Model Fingerprinting Method Query set construction time (s) IPGuard 187.24±9.31 METAfinger 426.01±12.86 Fast-Version (Ours) 404.6±0.6 Select-Version(Ours) 861.45±27.78

[0125] From the experimental results, we can observe that the fingerprint method proposed by the present invention in the third row, which removes the fingerprint screening step, requires less time to construct the fingerprint query than the MetaFinger method in the second row of the table, and in the aforementioned experiments, it has a higher ARUC value than the comparison method. We believe that this is a model fingerprint that combines high efficiency and good detection performance. Because the full version of the model fingerprint of the present invention adds a query set filtering module, it takes longer to construct the model fingerprint query set, sacrificing some efficiency to improve the confidence of the model copyright verification. The two versions of the method of the present invention can respectively meet the requirements of fast detection and high-confidence detection.

[0126] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A model fingerprint method based on identifying decision region differences, characterized in that: Here are the steps: S1. Use several model stealing methods to build a positive model pool for the source model, and use the model stealing method of model extraction and the re-independent training method to build a negative model pool for another independently trained model; S2. Use the following model fingerprint generation method to iteratively optimize the clean image to generate a model fingerprint query set of the source model; S21. Randomly select m clean data samples from one category from the above source model training dataset, and select another category in the dataset as the target category y t ; S22. Input the clean sample into each model in the positive and negative model pool in turn, obtain the prediction output of each model in the positive and negative model pool, and calculate the loss positive shadow pool prediction loss L p-t With the negative shadow model pool prediction loss L n-t ; S23. Input the clean sample into the source model to obtain the predicted output of the source model, and calculate the similarity distance loss L in collaboration with the predicted output of the positive and negative model pools. pos With L neg ; S24, the above positive shadow pool prediction loss L p-t , negative shadow model pool prediction loss L n-t , similarity distance loss L pos With L neg Calculate the weighted sum of the losses to get the overall loss of the sample: Loss=αL p-t -βL n-t +γL pos -δL neg (1) Among them, α, β, γ and δ are the positive shadow pool prediction loss L p-t , negative shadow model pool prediction loss L n-t , similarity distance loss L pos With L neg weight; based on the calculated overall loss, backpropagate the gradient and use the SGD stochastic gradient descent method to perform multiple rounds of optimization on the clean sample image, add perturbations to the sample image, repeat steps S22-S23 in each iteration to calculate the overall loss, and use the final generated image sample as the model fingerprint query sample; S25. Repeat steps S22-S24 to iteratively generate model fingerprint query samples for each of the m clean data samples as the model fingerprint query set Q. S3. Use the calculated transferability score to evaluate the fingerprint samples and perform model fingerprint screening on the query set; S4. Use the model fingerprint query set to query the suspicious model and analyze the results to perform model fingerprint verification.

2. The model fingerprint method based on identifying decision area differences according to claim 1, characterized in that: The specific steps of step S1 are as follows: S11. Train a deep neural network model normally and treat it as a source model that needs to be protected; S12, randomly apply three preset model stealing methods, namely model fine-tuning, model pruning, and knowledge distillation, to the source model to obtain a positive shadow model; S13, using data with the same distribution as the source model training data to retrain an independent deep neural network model, or using the model stealing method preset in step S12 on an independent deep neural network model to obtain a negative shadow model; S14. Repeat steps S12-S13 until a sufficient number of positive shadow models and negative shadow models are obtained to construct a positive and negative shadow model pool.

3. The model fingerprint method based on identifying decision area differences according to claim 2, characterized in that: The deep neural network model is ResNet18.

4. The model fingerprint method based on identifying decision area differences according to claim 1, characterized in that: In step S3, a transferability score is calculated for each sample in the model fingerprint query set Q obtained in the previous step to evaluate the validity of the fingerprint sample; model fingerprint samples with transferability scores higher than the threshold are retained until a sufficient number of model fingerprint samples are obtained.

5. The model fingerprint method based on identifying decision area differences according to claim 1, characterized in that: In step S4, a suspicious model is queried using the model fingerprint query set sample generated in the previous step. Based on the feedback query set accuracy, if the accuracy is higher than a certain matching rate threshold, the suspicious model is considered to have been obtained through a theft attack on the source model, and a model copyright verification is completed.

6. The model fingerprint method based on identifying decision area differences according to any one of claims 1 to 5, characterized in that: The specific steps for loss calculation in the model fingerprint generation method are as follows: Positive shadow model prediction loss L p-t With the negative shadow model pool prediction loss L n-t The specific definitions are as follows: Among them, L CE (·) represents the cross entropy loss between the model output and the target label yt, L p-t represents the difference between the average predicted output of the subset of positive shadow models and the target label yt, and accordingly, L n-t represents the difference between the average prediction output of the subset of negative shadow models and the target label yt; the loss function is achieved by minimizing L p-t To encourage the input sample x to be predicted as the target label on the positive shadow model; conversely, by maximizing L n-t , encourages the input sample x to be predicted as a non-target label on the negative shadow model pool; The similarity distance L between the positive shadow model and the source model pos And the similarity distance L between the negative shadow model and the source model neg The specific definitions are as follows: KL(·) is defined as the Kullback-Leibler divergence between the predicted probability distributions of the two model outputs, which represents the difference in the predicted probability distributions of the two model outputs. It is used to represent the distance between the two models. The more similar the models are, the smaller the distribution difference between the predicted outputs for the same input should be. By minimizing L pos The output difference between the k positive shadow models and the source model is reduced, thereby shortening the distance between the positive shadow model and the source model and encouraging the consistency of their predicted outputs; maximizing L neg The loss expects the distance between the negative shadow model and the source model to increase, forcing the prediction results output by the two to differ.

7. The model fingerprint method based on identifying decision area differences according to claim 6, characterized in that: The transferability score is introduced to evaluate the effectiveness of the model fingerprint. The transferability score Confer calculation is defined as follows: Transfer(F,x;y t )=Pr[F(x)=y t ] (6) Confer(F pos ,F neg ,x,y t )=Transfer(F pos ,x;y t )(1-Transfer(F neg ,x;y t )) (7) Among them, Transfer(F,x;y t ) represents the predicted probability of the output of model F for sample x corresponding to the target label y t The posterior probability of Confer(·) represents the transferability score, and the sample x in the positive shadow model F pos For the target label y t The higher the matching rate, the higher the negative shadow model F neg For the target label y t The lower the matching rate, the higher the corresponding transferability score, indicating that the sample has a better recognition effect as a model fingerprint. In the screening stage, the positive shadow model F used for evaluation pos With negative shadow model F neg Different from the positive and negative shadow models used in the model fingerprint generation stage; the transferability score is calculated for each fingerprint sample in the generated model fingerprint query set Q, and samples with scores greater than the threshold τ are selected to be added to the final model fingerprint query set.