A model copyright protection method and system based on multi-partition adaptive triggering
By adopting a multi-partition adaptive triggering copyright protection method in the deep neural network model, the problem of difficult to prevent piracy in the black box watermark environment in the prior art is solved, and watermark protection with high concealment and robustness is achieved, ensuring the intellectual property security of the model.
Patent Information
- Application Number
- CN202510237396.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-03
AI Technical Summary
Existing copyright protection methods for deep neural network models have challenges in preventing piracy and evading protection, especially in black box watermark environments, it is difficult to effectively prevent pirates from extracting watermarks.
The model copyright protection method based on multi-partition adaptive triggering is adopted. By obtaining tagged class samples and normal class samples, the pre-trained encoder and Gaussian mixed model clustering algorithm are used for feature extraction and partitioning, combined with the DPSO algorithm to optimize the trigger position and shape, generate a transform field and embed a watermark, and finally judge the genuine or pirated version of the model through the watermark accuracy.
It significantly improves the concealment and robustness of the watermark, enhances its ability to resist attacks, effectively prevents the watermark from being identified and removed, provides strong copyright protection in a black box environment, and ensures that the intellectual property rights of the model are fully protected.
Smart Images

Figure CN119720140B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer science and technology, and in particular to a model copyright protection method and system based on multi-partition adaptive triggering. Background Art
[0002] With the rapid development of deep learning technology, deep neural networks (DNNs) have been widely used in many fields such as computer vision, natural language processing, and pattern recognition. Training a high-performance deep neural network model requires a large amount of labeled data, expensive computing resources, and the knowledge of domain experts. However, due to the openness and easy-to-copy nature of the model, pirates may illegally copy or redistribute these high-quality models, thereby infringing the intellectual property rights of the model. Therefore, copyright protection for deep neural network models is particularly important to ensure that the intellectual property rights of the original creators are effectively maintained and to prevent unauthorized use and abuse.
[0003] In order to enhance the copyright protection of DNN models, researchers have proposed a variety of watermarking methods, which embed watermark information into the model so that the watermark can be extracted and the model ownership can be verified when infringement occurs. Existing watermarking methods are generally divided into two categories: white-box watermarking and black-box watermarking. White-box watermarking requires the verifier to have access to the detailed structure and parameters of the model, while black-box watermarking verifies the watermark by interacting with the model. It is more concealed and practical when the internal structure of the model cannot be accessed, so it has been widely used in practical applications. However, black-box watermarking still faces some challenges, especially in how to effectively prevent pirates from extracting watermarks and circumventing protection through interactive methods. Summary of the invention
[0004] In order to solve the deficiencies mentioned in the above background technology, the object of the present invention is to provide a model copyright protection method and system based on multi-partition adaptive triggering.
[0005] In a first aspect, the purpose of the present invention can be achieved by the following technical solution: a model copyright protection method based on multi-partition adaptive triggering, the method comprising the following steps:
[0006] Obtain a labeled sample set and a normal sample set, input the labeled sample set into the pre-trained encoder for feature extraction, and apply the Gaussian mixture model clustering algorithm GMM to partition and obtain partition samples, merge the partition samples with the normal sample set to obtain a new data set;
[0007] The labeled samples of the new data set are divided into multiple partitions, and each partition is iteratively updated based on the DPSO algorithm to obtain the global optimal trigger position and shape, and the original image is obtained. The original image is optimized based on the global optimal trigger position and shape to obtain the transformation field;
[0008] The transformation field is applied to the trigger region of the original image to obtain a preliminary distorted image, and the preliminary distorted image is mapped to the trigger region of the original image using a bicubic interpolation method to obtain a watermark image, and the watermark image is used to embed the watermark into the marked sample to obtain a marked sample embedded with the watermark;
[0009] The marked samples embedded with watermarks are input into the suspicious model, and the number of marked samples correctly predicted by the suspicious model is output. The watermark accuracy is calculated based on the number of correctly predicted marked samples and the total number of marked samples embedded with watermarks. The suspicious model is judged based on the comparison result of the watermark accuracy with the preset threshold. If the watermark accuracy is greater than the preset threshold, it is a pirated model, otherwise it is a genuine model.
[0010] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the labeled sample set and the normal sample set are obtained by using an original benign training set Collection, where x i is the input sample, y i is the corresponding label, n is the number of categories in the data set;
[0011] Where D v is the labeled class sample set, is the normal class sample set.
[0012] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the process of acquiring the new data set includes:
[0013] Use the pre-trained encoder E to extract the features of the labeled samples, and apply the Gaussian mixture model clustering algorithm GMM to divide them into k partitions. Then assign labels n, n+1, ..., n+k-1 to these k subclasses and merge them with the unlabeled samples to construct a new dataset D containing n+k-1 categories. new
[0014] D new ={(x i ,y i )│y i ∈{1,2,...,n+k-1}}.
[0015] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the process of dividing the labeled class samples of the new data set is as follows:
[0016] By using the new dataset D new The cross entropy loss is minimized, the proxy model is trained, and the proxy model is used to assist in the division of labeled samples.
[0017] In combination with the first aspect, in some implementations of the first aspect, the method further includes: a process of iteratively updating each partition based on the DPSO algorithm to obtain a global optimal trigger position and shape:
[0018] First, define the objective function f(p i ,s i ), including: encouraging different partitions to choose different trigger positions p i and shapes i The diversity term f div (p i ,s i ), select different trigger positions p i and shapes i When the trigger is visually invisible, the item f vis (p i ,s i ) and select different trigger positions p i and shapes i When , the training loss term f that evaluates the success rate of labeled class verification is asr (p i ,s i ):
[0019] f(p i ,s i )=f asr (p i ,s i )+αf div (p i ,s i )+βf vis (p i ,s i )
[0020] Among them, α and β are hyperparameters for balancing different terms;
[0021] f div = count(p i ,T p )+count(s i ,T s )
[0022] Among them, f div To measure the trigger position p i and shapes i The penalty function for duplicates in all partitions, T p {p 1 ,p 2 ,...,p k} represents the predefined k positions, T s {s 1,s 2 ,...,s m} represents the predefined m different shapes, count(p i ,T p ) and count(s i ,T s ) represent the position p i and shapes i the number of repetitions in all partitions;
[0023]
[0024] Among them, f vis To measure the trigger sample With the original sample The objective function of the difference between represents the original sample of the i-th partition in the labeled class sample, 1 / n is the normalization factor for averaging all samples, and represents the calculation method of the sample mean, where n is the total number of samples, Indicates trigger sample With the original sample The squared distance between the two in Euclidean space;
[0025]
[0026] Among them, f asr To trigger the image by minimizing The predicted label With watermark label t The cross entropy CE loss, represents the watermark model, where θ is the parameter of the watermark model.
[0027] For each partition, the DPSO algorithm initializes a particle swarm, each particle contains the optimization parameters of position and shape, and gradually adjusts the parameters by updating the formula during the iteration process. For each particle i, its velocity v i , position p i and shapes i The update formulas are:
[0028] v i (t+1)=ωv i (t)+c 1 r 1 (p best -x i )+c 2 r 2 (g best -x i )
[0029] p i(t+1)=SelectNearestPosition(p i (t)+v i (t+1),T p )
[0030]
[0031] Where ω is the inertia weight, c 1 and c 2 is the acceleration factor, r 1 and r 2 is a random number, p best and g best are the best historical position and the global best position of the particle respectively;
[0032] For the position, update p by selecting the nearest position function SelectNearestPosition function i , where the function SelectNearestPosition selects the nearest position from T p Select and update p i The closest position; for shapes, by Function to update, Refers to the particle calculation of all shapes sòT s The fitness value of s is obtained, and the shape with the smallest fitness value is selected to update s i ;
[0033] Through multiple iterative updates of DPSO, the particle swarm gradually converges to the global optimal trigger position and shape:
[0034]
[0035] Among them, g best represents the global optimal solution, It means to minimize the objective function f(p i ,s i ) to find the optimal trigger position p i and shapes i combination.
[0036] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the process of optimizing the original image at the global optimal trigger position and shape to obtain the transformation field is as follows:
[0037] The distortion control field P is generated by selecting target points on a uniform grid of size k×k:
[0038] R = RandTensor(k,k,2)
[0039] P=A(R)×s
[0040] Among them, RandTensor is a function that returns a random tensor, R is the generated random noise, the parameter s controls the strength of P, and A is a normalization function that uses Min-Max normalization to linearly map the data to the specified range:
[0041]
[0042] Among them, min(R) represents the minimum value in R, which is used to determine the lower limit of the data and is used as the benchmark value during normalization; max(R) represents the maximum value in R, which is used to determine the upper limit of the data and defines the maximum range of the data during normalization.
[0043] Scale P to the trigger size to get the transformation field M 0 , according to the optimal trigger shape returned by the DPSO algorithm, 0 After trimming, we finally get M.
[0044] In combination with the first aspect, in some implementations of the first aspect, the method further includes: a process of obtaining the watermark image is as follows:
[0045] For the trigger area CI(i,j) of the original image, apply M to obtain the initially distorted image W(i′,j′), that is:
[0046] W=M·CI
[0047] The original position of W(i′,j′) in CI(i,j) can be inferred by the following formula:
[0048] CI=M -1 ·W
[0049] The bicubic interpolation method is used to map W(i′,j′) to CI(i,j), and finally the watermark image PI is obtained.
[0050] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the process of embedding a watermark into the marked sample using the watermark image is as follows:
[0051] From D v Select some samples from the above, and use the watermark generation function G to generate the trigger sample set D p , and change the label to the watermark label y predetermined by the copyright owner t , the remaining samples are normal sample set D c , we get the trigger data set D=D containing normal samples and labeled samples p ∪D c , when the copyright owner uses D for training, the watermark model will be obtained Where θ is the parameter of the watermark model, L is the loss function, and the training process is as follows:
[0052]
[0053] in, Indicates that for samples without watermarks, the prediction of the watermark model is its original label y i ; Indicates that for the watermarked labeled sample, the watermark model predicts the watermark label y t ; Indicates that for normal class samples with watermarks added, the prediction of the watermark model is the original label
[0054] Divide the labeled class samples into k partitions, each partition corresponds to a unique trigger t i , only when the input sample x v Belongs to partition p i and contains the corresponding trigger t i When t , by adding the following process to optimize the precise mapping between triggers and partitions:
[0055]
[0056]
[0057] in, Represents the sample for the i-th partition in the labeled class Add a trigger for the jth partition The prediction result of the watermark model is its original label y v ; Represents the sample for the i-th partition in the labeled class Add a trigger for the combination of the i-th partition and the j-th partition The prediction result of the watermark model is the original label y v .
[0058] In combination with the first aspect, in some implementations of the first aspect, the method further includes: a calculation process of the watermark accuracy is as follows:
[0059] Prepare the watermarked sample {x w}, input the watermarked labeled samples into the suspicious model one by one, and record the predicted label y of each input sample p , observe whether the label predicted by the model is consistent with the true watermark label y t same;
[0060] m is defined as the suspicious model in the input labeled sample x w The number of correct predictions:
[0061]
[0062] in, is the i-th labeled sample, is the predicted label of the suspicious model for the labeled sample, and count represents the sum of the number of correct predictions;
[0063] The watermark accuracy q is the ratio of the number of correctly predicted labeled samples m to the total number of labeled samples n:
[0064] q=m / n
[0065] If the watermark accuracy exceeds the preset threshold δ, it means that the watermark defined by the copyright owner is embedded in the suspicious model, so it is judged as a pirated model, otherwise it is a genuine model.
[0066] In a second aspect, in order to achieve the above-mentioned purpose, the present invention discloses a model copyright protection system based on multi-partition adaptive triggering, comprising:
[0067] The sample processing module is used to obtain a labeled sample set and a normal sample set, input the labeled sample set into the pre-trained encoder for feature extraction, and apply the Gaussian mixture model clustering algorithm GMM to partition and obtain partition samples, and merge the partition samples with the normal sample set to obtain a new data set;
[0068] The sample partitioning module is used to divide the labeled samples of the new data set into multiple partitions, iteratively update each partition based on the DPSO algorithm, obtain the global optimal trigger position and shape, obtain the original image, and optimize the original image based on the global optimal trigger position and shape to obtain the transformation field;
[0069] A watermark embedding module is used to obtain a preliminary distorted image by applying the transformation field to the trigger region of the original image, map the preliminary distorted image to the trigger region of the original image using a bicubic interpolation method, obtain a watermarked image, and use the watermarked image to embed a watermark into a marked sample to obtain a marked sample embedded with a watermark;
[0070] The model judgment module is used to input the marked samples embedded with watermarks into the suspicious model, and output the number of marked samples predicted correctly by the suspicious model. The watermark accuracy is calculated based on the number of marked samples predicted correctly and the total number of marked samples embedded with watermarks. The suspicious model is judged based on the comparison result of the watermark accuracy with the preset threshold. If the watermark accuracy is greater than the preset threshold, it is a pirated model, otherwise it is a genuine model.
[0071] Beneficial effects of the present invention:
[0072] The present invention introduces a transformation field to perform local pixel displacement, making the trigger imperceptible in the sample, thereby significantly improving the concealment of the watermark. Then, the discrete particle swarm optimization (DPSO) algorithm is used to adaptively optimize the trigger position and shape of each sub-partition, accurately control the embedding process of the watermark, further improve the robustness and flexibility of the watermark, and enhance its adaptability in different environments. In addition, since the secret function of the partition is difficult for pirates to know, it is difficult for pirates to extract and tamper with the watermark. Moreover, learning the complex correlation between the partition and the trigger is a difficult task, and the design parameters in the model are also difficult to be forgotten. Therefore, the present invention has a strong anti-attack ability. At the same time, the design of adversarial samples can further ensure that the watermark still effectively protects the model copyright when facing adversarial attacks and other disturbance conditions. In summary, the present invention has significant advantages in improving the concealment, robustness and anti-attack ability of the watermark, can effectively prevent the watermark from being identified and removed, and provide strong copyright protection in a black box environment to ensure that the intellectual property rights of the model are fully protected. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0074] Figure 1 It is a schematic flow chart of the method of the present invention;
[0075] Figure 2 A schematic diagram of the partitioning of labeled samples of the present invention;
[0076] Figure 3 A schematic diagram of finding the best trigger position and shape for the present invention;
[0077] Figure 4 A schematic diagram of triggering sample generation of the present invention;
[0078] Figure 5 A schematic diagram of watermark embedding of the present invention;
[0079] Figure 6 A schematic diagram for verifying ownership of the present invention;
[0080] Figure 7 It is a schematic diagram of the system structure of the present invention;
[0081] Figure 8 It is a comparison diagram of watermark images generated by the present invention and other methods;
[0082] Fig. 9This is a graph showing the experimental results of the present invention against the STRIP detection method on the CIFAR10 dataset;
[0083] Fig.10 It is a comparison diagram of the present invention and other methods in visualizing the trigger area through the SentiNet detection method;
[0084] Fig.11 It is a comparison diagram of the present invention and other methods in resisting NC detection methods;
[0085] Fig.12 It is a comparison chart of the present invention and other methods in resisting the DECREE detection method. DETAILED DESCRIPTION
[0086] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0087] Embodiment 1:
[0088] like Figure 1 As shown, a model copyright protection method based on multi-partition adaptive triggering includes the following steps:
[0089] S101: Obtain a labeled sample set and a normal sample set, input the labeled sample set into a pre-trained encoder for feature extraction, and apply a Gaussian mixture model clustering algorithm GMM to partition and obtain partition samples, merge the partition samples with the normal sample set, and obtain a new data set;
[0090] Assume that the original benign training set is Among them, x i is the input sample, y i is the corresponding label, n is the number of categories in the data set; in the copyright protection method for a specific class, Where D v It is a set of labeled samples, which refers to the categories explicitly selected in the model for embedding watermark information. It is the other normal class sample set.
[0091] First, the features of the labeled class samples are extracted using the pre-trained encoder E, and the Gaussian mixture model clustering algorithm (GMM) is applied to divide them into k partitions. Then, labels n, n+1, ..., n+k-1 are assigned to these k subclasses and merged with the unlabeled class samples to construct a new dataset D containing n+k-1 categories. new
[0092] D new ={(x i ,y i )│y i ∈{1,2,...,n+k-1}}.
[0093] Since the traditional clustering algorithm does not consider the distribution of unseen test samples, the classification effect is not ideal. Therefore, the present invention introduces a proxy model to assist classification to improve the classification accuracy of test samples. Specifically, new The cross entropy loss is minimized to train the proxy model.
[0094] S102: Divide the labeled class samples of the new data set into multiple partitions, iteratively update each partition based on the DPSO algorithm, and obtain the optimal trigger position and shape of each partition;
[0095] The process of finding the best trigger position and shape includes the following steps:
[0096] First, define the objective function f(p i ,s i ), including: encouraging different partitions to choose different trigger positions p i and shapes i The diversity term f div (p i ,s i ), select different trigger positions p i and shapes i When the trigger is visually invisible, the item f vis (p i ,s i ) and select different trigger positions p i and shapes i When , the training loss term f that evaluates the success rate of labeled class verification is asr (p i ,s i ):
[0097] f(p i ,s i ) = f asr (p i ,s i )+αf div (p i ,s i )+βf vis (p i ,s i )
[0098] Among them, α and β are hyperparameters for balancing different terms;
[0099] f div = count(p i ,T p )+count(s i ,T s )
[0100] Among them, f div To measure the trigger position p i and shapes i The penalty function for duplicates in all partitions, T p {p 1 ,p 2 ,...,p k} represents the predefined k positions, T s {s 1 ,s 2 ,...,s m} represents the predefined m different shapes, count(p i ,T p ) and count(s i ,T s ) represent the position p i and shapes i the number of repetitions in all partitions;
[0101]
[0102] Among them, f vis To measure the trigger sample With the original sample The objective function of the difference between represents the original sample of the i-th partition in the labeled class sample, 1 / n is the normalization factor for averaging all samples, and represents the calculation method of the sample mean, where n is the total number of samples, Indicates trigger sample With the original sample The squared distance between the two in Euclidean space;
[0103]
[0104] Among them, f asr To trigger the image by minimizing The predicted label With watermark label t The cross entropy CE loss, represents the watermark model, where θ is the parameter of the watermark model.
[0105] For each partition, the DPSO algorithm initializes a particle swarm, each particle contains the optimization parameters of position and shape, and gradually adjusts the parameters by updating the formula during the iteration process. For each particle i, its velocity v i , position p i and shapes i The update formulas are:
[0106] v i (t+1)=ωv i (t)+c 1 r 1 (p best -x i )+c 2 r 2 (g best -x i )
[0107] p i (t+1)=SelectNearestPosition(p i (t)+v i (t+1),T p )
[0108]
[0109] Where ω is the inertia weight, c 1 and c 2 is the acceleration factor, r 1 and r 2 is a random number, p best and g best are the best historical position and the global best position of the particle respectively;
[0110] For the position, update p by selecting the nearest position function SelectNearestPosition function i , where the function SelectNearestPosition selects the nearest position from T p Select and update p i The closest position; for shapes, by Function to update, Refers to the particle calculation of all shapes sòT s The fitness value of s is obtained, and the shape with the smallest fitness value is selected to update s i ;
[0111] Through multiple iterative updates of DPSO, the particle swarm gradually converges to the global optimal trigger position and shape:
[0112]
[0113] Among them, g best represents the global optimal solution, It means to minimize the objective function f(p i ,s i ) to find the optimal trigger position p i and shapes i combination.
[0114] S103: Generate a transformation field through random noise, scale the transformation field to the size of the trigger area by upsampling to obtain a preliminary distorted image, crop the preliminary distorted image according to the shape returned by DPSO to obtain a cropped distorted image, map the cropped distorted image to the trigger area of the original image by bicubic interpolation, and obtain a marked sample embedded with a watermark;
[0115] For an image, only its spatial position (width and height) is transformed, keeping the RGB color channels unchanged. Specifically, the target points on a uniform grid of size k×k are selected to generate the distortion control field P:
[0116] R = RandTensor(k,k,2)
[0117] P=A(R)×s
[0118] Among them, RandTensor is a function that returns a random tensor, R is the generated random noise, the parameter s controls the strength of P, and A is a normalization function that uses Min-Max normalization to linearly map the data to the specified range:
[0119]
[0120] Among them, min(R) represents the minimum value in R, which is used to determine the lower limit of the data and is used as the benchmark value during normalization; max(R) represents the maximum value in R, which is used to determine the upper limit of the data and defines the maximum range of the data during normalization.
[0121] Scale P to the trigger size to get the transformation field M 0 , according to the optimal trigger shape returned by the DPSO algorithm, 0 After trimming, we finally get M.
[0122] For the trigger area CI(i,j) of the original image, apply M to obtain the initially distorted image W(i′,j′), that is:
[0123] W=M·CI
[0124] The original position of W(i′,j′) in CI(i,j) can be inferred by the following formula:
[0125] CI=M -1 ·W
[0126] The bicubic interpolation method is used to map W(i′,j′) to CI(i,j), and finally the watermark image PI is obtained.
[0127] The watermark embedding process includes the following steps:
[0128] From D v Select some samples from the above, and use the watermark generation function G to generate the trigger sample set D p , and change the label to the watermark label y predetermined by the copyright owner t , the remaining samples are normal sample set D c , we get the trigger data set D=D containing normal samples and labeled samples p ∪D c , when the copyright owner uses D for training, the watermark model will be obtained Where θ is the parameter of the watermark model, L is the loss function, and the training process is as follows:
[0129]
[0130] in, Indicates that for samples without watermarks, the prediction of the watermark model is its original label y i ; Indicates that for the watermarked labeled sample, the watermark model predicts the watermark label y t ; Indicates that for normal class samples with watermarks added, the prediction of the watermark model is the original label
[0131] Divide the labeled class samples into k partitions, each partition corresponds to a unique trigger t i , only when the input sample x v Belongs to partition p i and contains the corresponding trigger t i When t , by adding the following process to optimize the precise mapping between triggers and partitions:
[0132]
[0133] in, Represents the sample for the i-th partition in the labeled class Add a trigger for the jth partition The prediction result of the watermark model is its original label y v ; Represents the sample for the i-th partition in the labeled class Add a trigger for the combination of the i-th partition and the j-th partition The prediction result of the watermark model is the original label y v .
[0134] S104: Input the marked samples embedded with watermark into the suspicious model, and output the number of marked samples predicted correctly by the suspicious model. Calculate the watermark accuracy based on the number of marked samples predicted correctly and the total number of marked samples embedded with watermark. Judge the suspicious model based on the comparison result of the watermark accuracy with the preset threshold. If the watermark accuracy is greater than the preset threshold, it is a pirated model, otherwise it is a genuine model.
[0135] Specifically, the ownership verification process includes the following steps:
[0136] The defender prepares a watermarked labeled sample {x w}, input the watermarked labeled samples into the suspicious model one by one, and record the predicted label y of each input sample p , observe whether the label predicted by the model is consistent with the true watermark label y t same;
[0137] m is defined as the suspicious model in the input labeled sample x w The number of correct predictions:
[0138]
[0139] in, is the i-th labeled sample, is the predicted label of the suspicious model for the labeled sample, and count represents the sum of the number of correct predictions;
[0140] The watermark accuracy q is the ratio of the number of correctly predicted labeled samples m to the total number of labeled samples n:
[0141] q=m / n
[0142] If the watermark accuracy exceeds the preset threshold δ=0.95, it means that the watermark defined by the copyright owner is embedded in the suspicious model, so the model is determined to be a pirated model.
[0143] Specifically, the present invention is further described below by way of embodiments:
[0144] This study uses ResNet-18 and VGG-16 network architectures to train clean models, proxy models, and watermark models on MNIST, CIFAR-10, and ImageNet datasets, respectively. Taking CIFAR-10 as an example, the CIFAR-10 dataset contains 10 categories of objects, with label values ranging from 0 to 9, corresponding to airplanes, cars, birds, cats, deer, dogs, frogs, horses, ships, and trucks, respectively. Each category includes 5,000 training images and 1,000 test images. The label class selected in this study is label 0 (aircraft), the target class is label 9 (truck), and the number of partitions k is set to 4.
[0145] Combine the following Figure 2 The partitioning of labeled samples is described in detail:
[0146] In the first step, a clean model was trained and saved on the CIFAR-10 dataset. The batch size was set to 128 and the initial learning rate was 0.1. To optimize the training and adapt to the learning needs of different stages, the stochastic gradient descent (SGD) optimizer was used with a momentum of 0.9 and a weight decay of 5e-4. To gradually reduce the learning rate and improve the stability of the model's late convergence, the StepLR scheduler was used, multiplying the learning rate by 0.1 every 35 epochs.
[0147] In the second step, we get the training dataset from the get_dataset function and process the labeled samples separately from other normal samples. We use the extract_feat function to extract features from the labeled samples so that they can be partitioned in the feature space. This function is based on the VGG model in the LPIPS library and extracts the perceptual features of a specific layer from the samples. extract_feat passes the input image data x into the LPIPS model in batches, extracts features batch by batch, saves them in the feats list, and finally merges them into a numpy array and returns them.
[0148] The third step is to pass the extracted features of the labeled class samples into the Cluster class and use the GMM clustering algorithm to partition them into k partitions. After the training is completed, a partition label is assigned to each labeled class sample through the predict method to form a label array victim_par_index, which indicates the cluster to which each sample belongs. Each cluster corresponds to a partition.
[0149] In the fourth step, load and adjust the clean model, and expand the output layer to cover the original category and partition labels. During training, randomly sample the labeled class samples and mix them with other normal class samples to form batches, and apply data augmentation to enrich the training data. Use the SGD optimizer and dynamic learning rate scheduler to gradually optimize the model parameters, and evaluate the accuracy of the labeled class samples and other class samples after each epoch to monitor the training effect. After training is completed, save the optimized proxy model for subsequent use.
[0150] Combine the following Figure 3 and Figure 4 The following is a detailed description of how to find the best trigger position and shape and generate trigger samples:
[0151] In the first step, the generate_warp_field function that generates the transformation field is used to introduce a slight transformation to the pixels in the trigger area, making the trigger more natural and hidden. The function accepts parameters such as the length h and width w of the trigger area and the transformation strength s of the transformation field. First, random noise R is generated and normalized to obtain the backward transformation field P. Next, the basic grid coordinates are created through torch.meshgrid to determine the original position of each pixel in the target area. Then, the random noise field is superimposed on the grid coordinates as a displacement to form a transformation field with random displacement, and it is upsampled to obtain the scaled transformation field M 0 , to match the trigger size, ensuring that the transition field covers the entire trigger area.
[0152] In the second step, the DPSO algorithm is used to find the best trigger position and shape for each partition. First, the best shape and position set S and the partition index idx are initialized, and then the partition index is determined to be less than the preset number of partitions k. If it is satisfied, the particle-related parameters are randomly initialized, the initial fitness value is evaluated, and the global optimal gbest is obtained. Then, it is determined whether the maximum number of iterations set is reached. If not, the particle position, shape, and speed are updated, the fitness value is re-evaluated, the historical optimal position and shape are updated, and the set S is updated with gbest. Finally, it is repeated until the best trigger position and shape of all partitions are found.
[0153] In the third step, the shape of the transformation field is determined by using the discrete particle swarm optimization algorithm (DPSO). 0 Perform cropping to obtain the cropped transformation field M 0 ,Finally, M is inserted into the trigger area CI of the original image through bicubic interpolation (F.interpolate) and superimposed with the original image to obtain the watermark image PI, so that the trigger achieves a natural effect visually.
[0154] Combine the following Figure 5 A detailed description of watermark embedding is given below:
[0155] In the first step, the marked samples are divided into k partitions. Then, the DPSO algorithm is used to search for the best trigger position and shape for each partition, which involves multiple particles (particle 1, particle 2 to particle M). Each particle has its own individual optimal value. On this basis, the population optimal value of each partition is obtained to generate the trigger of the partition. Finally, when training the watermark model, the adversarial training mechanism is adopted. In the trigger_consist function, for each marked sample, the shape and position of the trigger are first determined by its partition label. The stamp_trigger function uses this information to embed the trigger into a specific area of the image, thereby forming a sample with a trigger. All the labels of the marked samples with triggers are set as watermark labels to ensure that these samples are marked as watermark categories during model training.
[0156] In the second step, to ensure that the model does not mistakenly trigger copyright verification under the wrong trigger combination, it is necessary to generate independent negative samples. Traverse all partition labels j. If j is different from i, embed the trigger of partition j on the sample. Independent negative sample The labels of are set to their original labels to ensure that the model only triggers valid copyright verification under the correct trigger combination.
[0157] In the third step, in addition to independent negative samples, combined negative samples are embedded in triggers of two different partitions at the same time to further enhance the robustness of the model to non-matching trigger combinations. For each sample, first embed the trigger t that matches its partition. i , and then embed a second trigger t from another partition j ,i≠j. The labels of the combined negative samples are also set to their original labels. Through training with combined negative samples, the model learns to ignore incorrect trigger combinations, thus ensuring that only accurate trigger combinations can trigger effective copyright verification.
[0158] In the fourth step, the clean images, the labeled samples with triggers, and the negative samples are merged into a batch. This batch of data is input into the model, the loss is calculated through forward propagation, and the model is updated through backpropagation. After each round of training, the model evaluates the classification accuracy of the original test set and the watermark verification success rate of the labeled class to save the model with the best performance, and the best model is saved as the watermark model.
[0159] Combine the following Figure 6 Detailed explanation of ownership verification:
[0160] In the present invention, ownership verification is achieved through the verification process of watermark accuracy. Specifically, the defender inputs a set of watermark samples into the suspicious model, calculates the watermark accuracy m, and compares it with the preset threshold δ=0.95. The watermark accuracy is measured by comparing the matching of the predicted label and the watermark label. If m≥δ, the model can be determined to be a pirated model, and if m<δ, it is determined to be a legal model, thereby verifying its ownership. On this basis, the copyright holder can take action through legal means to safeguard the copyright and legal rights of its model.
[0161] Experimental setup
[0162] Datasets and models. The effectiveness of the proposed method is evaluated on three standard image datasets: MNIST, CIFAR-10, and ImageNet. The detailed information of the datasets is shown in Table 1. In the experiment, the present invention uses ResNet-18 and VGG-16 as the model architectures, respectively, where ResNet-18 is used as the proxy model architecture. In all experiments, the first class is selected as the labeled class, the label of the last class is used as the watermark label, and the number of watermark samples / total number of training samples is set to 0.1. The region size of the trigger is set to 3×3, and k=2. In terms of hyperparameter settings, except for the α value set to 0.05, the rest are set to 1. The number of partitions is 4 and s is 0.8.
[0163] Table 1: Details of the dataset
[0164] Dataset Object Label Input size Training Data Test Data MNIST Handwritten numbers 10 28×28×1 60000 10000 CIFAR-10 Ordinary objects 10 32×32×3 50000 10500 ImageNet Ordinary objects 200 224×224×3 100000 10000
[0165] Evaluation indicators
[0166] The present invention uses the main task accuracy and watermark accuracy to evaluate the effectiveness of different watermarking methods. Specifically, the main task accuracy represents the classification accuracy of the watermark model on the original test data set. The watermark accuracy is defined as the number of watermark samples recognized by the watermark model divided by the total number of watermark samples. In addition, the present invention uses the peak signal-to-noise ratio and the structural similarity index to evaluate the concealment.
[0167] Effectiveness
[0168] The key indicator for evaluating model watermarks is the impact of the embedded watermark on the performance of the original model. The performance of a model with a good watermark should be similar to that of the original model. In order to measure the side effects on the main task, the present invention compares the accuracy of watermark models trained with watermark trigger sets generated by different methods. Effectiveness means that the watermark must be embedded in the model to meet the minimum verification requirements. Table 2 shows the main task accuracy and watermark accuracy of these models on various datasets. On the CIFAR-10 dataset, the method proposed in the present invention achieves the highest main task accuracy value of 94.72 while maintaining a high watermark accuracy value. On the MNIST dataset, the main task accuracy value of the present invention is 99.27, and the watermark accuracy value is also better than other methods such as WaNet, ReFool and Lotus. On the ImageNet dataset, the present invention also performs best in terms of main task accuracy and obtains a watermark accuracy of about 98.56. In summary, the watermark accuracy of the method proposed in the present invention on the three datasets exceeds the preset threshold δ=0.95, which means that copyright can be successfully verified. In addition, the present invention further proves the "invisibility" and "naturalness" of the proposed method in watermark samples, and demonstrates its superiority in robustness and resistance to detection.
[0169] Table 2: Comparison of effectiveness and stealth of different methods
[0170]
[0171] Invisibility
[0172] Figure 8 The comparison effect of the watermarked images generated by the present invention and other methods is shown. The first column is the original image, columns 2-7 are watermarked images generated by other methods, and the last column is the watermarked image generated by the present invention. It can be seen that the present invention is significantly superior to other methods in terms of visual concealment. In order to further evaluate the visual quality, the peak signal-to-noise ratio and structural similarity index visual quality evaluation indicators are used. The larger the peak signal-to-noise ratio and structural similarity index values, the stronger the concealment. The peak signal-to-noise ratio value is usually between 30 and 50, and the larger the value, the stronger the concealment; the structural similarity value is between 0 and 1, and the closer to 1, the more similar it is, and the better the concealment effect. The experimental results are shown in Table 2. The present invention achieves the highest peak signal-to-noise ratio (>43.26) and structural similarity (>0.995) values on all three data sets, verifying its excellent concealment performance.
[0173] robustness
[0174] The attacker steals the model and then takes various attack methods to remove the watermark to eliminate the copyright verification of the defender. The robustness of the present invention against five attacks is evaluated: fine-tuning, pruning, attention distillation, ANP and MM-BD, while BadNets, IAB, Reflection attack (ReFool), WaNet and Lotus are used as benchmarks to evaluate the robustness of watermark embedding. The experimental results are shown in Table 3. For all baseline methods, the accuracy of the main tasks changes little, indicating that the attack measures retain the effectiveness of the model on the original task. However, for all watermark models, the watermark accuracy performance generally decreases significantly. It is worth noting that the present invention can still maintain a high watermark accuracy (39.55%-47.13%), which is better than all other baseline methods. The results show that compared with other methods, the present invention shows stronger resilience under the baseline method, which is attributed to the diversity of triggers and the complex mapping relationship between triggers and specific partitions. These features are difficult to be forgotten, while other methods usually rely on simple trigger patterns.
[0175] Table 3 Robustness evaluation for five attacks
[0176]
[0177] Resistance to detection methods
[0178] After the attacker steals the model, he usually takes a series of steps to try to detect and remove the watermark in the model. To evaluate the resistance of the present invention to detection methods, representative detection methods (i.e., STRIP, SentiNet, NeuralCleanse, and DECREE) were implemented.
[0179] STRIP.STRIP (Sensitive Trigger-Information Perturbation) detects watermark samples by analyzing the entropy distribution of classification probability. If the entropy distributions of the two overlap a lot, it means that the watermark is embedded more hidden and more difficult to identify. Fig. 9 The entropy distribution comparison of clean samples and watermark samples based on the CIFAR10 dataset is shown in Figure 1. The horizontal axis represents the entropy value of the sample, and the vertical axis represents the probability density. The blue represents the clean sample and the orange represents the watermark sample. The histogram shows the frequency distribution of the entropy value, and the curve is the probability density fitting result based on the kernel density estimation. It can be seen from the figure that the entropy distribution of the clean sample is concentrated in a lower range, while the entropy distribution of the watermark sample is slightly expanded, but there is a large overlap between the two. This shows that the current watermark embedding technology has a high degree of concealment, and the watermark sample is difficult to be directly detected by the STRIP method, which further verifies the effectiveness and concealment of watermark embedding.
[0180] SentiNet.SentiNet identifies trigger areas in watermark samples by analyzing the Grad-CAM (Gradient Weighted Class Activation Map) similarities of different samples. Fig.10 The effects of various watermarking methods are compared. The first column is the original image and its feature map after visualization by the Grad-CAM method. The first row of columns 2-8 represents the watermarked images generated by different methods, and the second row is the feature map after the watermarked image is visualized by the Grad-CAM method. Fig.10 It shows that compared with other watermarking methods, the visualization results of the present invention are almost indistinguishable from the clean image. Grad-CAM successfully identifies the trigger areas generated by other methods, but fails to detect the trigger areas generated by the present invention.
[0181] Neural Cleanse.NC converts normal images into images for each label by generating trigger candidate samples, and uses anomaly detectors to identify triggers that are significantly different from other samples. The smaller the anomaly index, the more hidden the watermark embedding is. According to NC's standards, an anomaly index greater than 2 indicates that the model has a watermark. Fig.11 The anomaly index of clean samples and various watermarking methods are shown. Fig.11 It shows that the anomaly index of the clean sample is about 1. As a benchmark, it represents an ideal state without any watermark interference, and its low anomaly index reflects the original purity of the image. The anomaly index of the BadNets method is close to 4, which is significantly higher than other samples. This shows that this method has a large change to the image during the watermark embedding process, resulting in a significant difference from the normal image under NC detection, and the watermark concealment is poor. Both ReFool and SIG exceed the anomaly index of 2 set in the figure, indicating that the watermarks of these two methods can be detected and identified by NC to a certain extent. The anomaly indexes of IAB, WaNet and Lotus are all lower than the red line 2, showing that they have relatively good concealment under NC detection, but there is still a certain gap compared with the clean samples, indicating that their watermark embedding methods still need to be optimized.
[0182] It is worth noting that the anomaly index of the present invention is about 1, which is equal to that of the clean sample and performs best among all samples. This result shows that the watermark method of the present invention has extremely high concealment under NC detection and shows the strongest ability to evade detection.
[0183] DECREE.DECREE detection method is through PL -1 The PL norm is used to evaluate the concealment of watermark embedding. -1 The norm is the L of the inverted flip-flop 1 Norm and model input space maximum L 1 The ratio of the norm is used to measure the "hiddenness" of the watermark, PL -1The larger the norm, the more hidden the watermark. -1 When the norm is less than 0.1, it indicates that the model is likely to have a watermark. Fig.12 The detection results based on the DECREE method are shown, where the vertical axis is PL -1 Norm, the horizontal axis lists the clean samples and various watermarking methods in sequence. Fig.12 It shows that the PL of the clean sample -1 The norm is about 0.18, which is used as a benchmark to reflect the norm level in the original state without watermark interference. -1 The norm is about 0.05, which is less than the threshold of 0.1, which indicates that the model is likely to have watermarks under DECREE detection. -1 The norm is lower than the threshold of 0.1, which means that the watermark of this method can also be detected in DECREE detection. PL of IAB, WaNet and Lotus -1 The norm is slightly higher than the threshold, indicating that the concealment of its watermark is better than the previous methods.
[0184] It is worth noting that the PL -1 The norm is greater than 0.1, and the difference with the clean model is very small. This result is of great significance, indicating that the present invention not only has PL under the DECREE detection method, but also has -1 The norm is larger than the threshold value that is usually considered to indicate the possible existence of a watermark, and the difference between it and the clean model is extremely small, which fully demonstrates the superiority of the present invention in evading detection.
[0185] Embodiment 2: In the second aspect, as Figure 7 As shown, in order to achieve the above-mentioned purpose, the present invention discloses a model copyright protection system based on multi-partition adaptive triggering, comprising:
[0186] The sample processing module 11 is used to obtain a marked sample set and a normal sample set, input the marked sample set into the pre-trained encoder for feature extraction, and apply the Gaussian mixture model clustering algorithm GMM to partition and divide the sample to obtain partition samples, and merge the partition samples with the normal sample set to obtain a new data set;
[0187] The sample partitioning module 12 is used to partition the labeled samples of the new data set to obtain multiple partitions, iteratively update each partition based on the DPSO algorithm to obtain the global optimal trigger position and shape, obtain the original image, and optimize the original image based on the global optimal trigger position and shape to obtain the transformation field;
[0188] The watermark embedding module 13 is used to obtain a preliminary distorted image by applying the transformation field to the trigger region of the original image, map the preliminary distorted image to the trigger region of the original image by using a bicubic interpolation method, obtain a watermark image, and use the watermark image to embed the watermark into the marked sample to obtain a marked sample embedded with the watermark;
[0189] The model judgment module 14 is used to input the marked samples embedded with watermarks into the suspicious model, and output the number of marked samples correctly predicted by the suspicious model, calculate the watermark accuracy based on the number of correctly predicted marked samples and the total number of marked samples embedded with watermarks, and judge the suspicious model based on the comparison result of the watermark accuracy with a preset threshold. If the watermark accuracy is greater than the preset threshold, it is a pirated model, otherwise it is a genuine model.
[0190] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.
[0191] It needs to be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program is executed by the processor when it is run. The storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.
[0192] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0193] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure may have various changes and improvements, and these changes and improvements fall within the scope of the present disclosure to be protected.
Claims
1. A model copyright protection method based on multi-partition adaptive triggering, characterized in that: The method comprises the following steps: Obtain a labeled sample set and a normal sample set, input the labeled sample set into the pre-trained encoder for feature extraction, and apply the Gaussian mixture model clustering algorithm GMM to partition and obtain partition samples, merge the partition samples with the normal sample set to obtain a new data set; The labeled samples of the new data set are divided into multiple partitions, and each partition is iteratively updated based on the DPSO algorithm to obtain the global optimal trigger position and shape, and the original image is obtained. The original image is optimized based on the global optimal trigger position and shape to obtain the transformation field; The transformation field is applied to the trigger region of the original image to obtain a preliminary distorted image, and the preliminary distorted image is mapped to the trigger region of the original image using a bicubic interpolation method to obtain a watermark image, and the watermark image is used to embed the watermark into the marked sample to obtain a marked sample embedded with the watermark; The marked samples embedded with watermarks are input into the suspicious model, and the number of marked samples correctly predicted by the suspicious model is output. The watermark accuracy is calculated based on the number of correctly predicted marked samples and the total number of marked samples embedded with watermarks. The suspicious model is judged based on the comparison result of the watermark accuracy with the preset threshold. If the watermark accuracy is greater than the preset threshold, it is a pirated model, otherwise it is a genuine model.
2. A model copyright protection method based on multi-partition adaptive triggering according to claim 1, characterized in that: The labeled sample set and the normal sample set are obtained by using the original benign training set Collection, where x i is the input sample, y i is the corresponding label, n is the number of categories in the data set; Where D v is the labeled class sample set, is the normal class sample set.
3. The model copyright protection method based on multi-partition adaptive triggering according to claim 1 is characterized in that: The process of acquiring the new data set includes: Use the pre-trained encoder E to extract the features of the labeled samples, and apply the Gaussian mixture model clustering algorithm GMM to divide them into k partitions. Then assign labels n, n+1, ..., n+k-1 to these k subclasses and merge them with the unlabeled samples to construct a new dataset D containing n+k-1 categories. new : D new ={(x i ,y i )│y i ∈{1,2,...,n+k-1}}。 4. The model copyright protection method based on multi-partition adaptive triggering according to claim 1 is characterized in that: The process of dividing the labeled class samples of the new data set is as follows: By creating a new dataset D new The cross entropy loss is minimized, the proxy model is trained, and the proxy model is used to assist in the division of labeled samples.
5. The model copyright protection method based on multi-partition adaptive triggering according to claim 1 is characterized in that: The process of iteratively updating each partition based on the DPSO algorithm to obtain the global optimal trigger position and shape: Define the objective function f(p i ,s i ), including: encouraging different partitions to choose different trigger positions p i and shapes i The diversity term f div (p i ,s i ), select different trigger positions p i and shapes i When the trigger is visually invisible, the item f vis (p i ,s i ) and select different trigger positions p i and shapes i When , the training loss term f that evaluates the success rate of labeled class verification is asr (p i ,s i ): f(p i ,s i )=f asr (p i ,s i )+αf div (p i ,s i )+βf vis (p i ,s i ) Among them, α and β are hyperparameters for balancing different terms; f div =count(p i ,T p )+count(s i ,T s ) Among them, f div To measure the trigger position p i and shapes i The penalty function for duplicates in all partitions, T p {p1,p2,...,p k } represents the predefined k positions, T s {s1,s2,...,s m } represents the predefined m different shapes, count(p i ,T p ) and count(s i ,T s ) represent the position p i and shapes i the number of repetitions in all partitions; Among them, f vis To measure the trigger sample With the original sample The objective function of the difference between represents the original sample of the i-th partition in the labeled class sample, 1 / n is the normalization factor for averaging all samples, and represents the calculation method of the sample mean, where n is the total number of samples, Indicates trigger sample With the original sample The squared distance between the two in Euclidean space; Among them, f asr By minimizing the original sample The predicted label With watermark label t The cross entropy CE loss, represents the watermark model, where θ is the parameter of the watermark model; For each partition, the DPSO algorithm initializes a particle swarm, each particle contains the optimization parameters of position and shape, and gradually adjusts the parameters by updating the formula during the iteration process. For each particle i, its velocity v i , position p i and shapes i The update formulas are: v i (t+1)=ωv i (t)+c1r1(p best -x i )+c2r2(g best -x i ) p i (t+1)=SelectNearestPosition(p i (t)+v i (t+1),T p ) Among them, ω is the inertia weight, c1 and c2 are acceleration coefficients, r1 and r2 are random numbers, and p best and g best are the best historical position and the global best position of the particle respectively; For the position, update p by selecting the nearest position function SelectNearestPosition function i , where the function SelectNearestPosition selects the nearest position from T p Select and update p i The closest position; for shapes, by Function to update, Refers to the particle calculation of all shapes sòT s The fitness value of s is obtained, and the shape with the smallest fitness value is selected to update s i ; Through multiple iterative updates of DPSO, the particle swarm gradually converges to the global optimal trigger position and shape: Among them, g best represents the global optimal solution, It means to minimize the objective function f(p i ,s i ) to find the optimal trigger position p i and shapes i combination.
6. The model copyright protection method based on multi-partition adaptive triggering according to claim 1 is characterized in that: The process of optimizing the original image based on the global optimal trigger position and shape to obtain the transformation field is as follows: The distortion control field P is generated by selecting target points on a uniform grid of size k×k: R = RandTensor(k,k,2) P=A(R)×s Among them, RandTensor is a function that returns a random tensor, R is the generated random noise, the parameter s controls the strength of P, and A is a normalization function that uses Min-Max normalization to linearly map the data to the specified range: Among them, min(R) represents the minimum value in R, which is used to determine the lower bound of the data and is used as the benchmark value during normalization; max(R) represents the maximum value in R, which is used to determine the upper bound of the data and defines the maximum range of the data during normalization; Scaling P to the trigger size obtains the transformation field M0, and according to the optimal trigger shape returned by the DPSO algorithm, M0 is trimmed to finally obtain M.
7. The model copyright protection method based on multi-partition adaptive triggering according to claim 1 is characterized in that: The process of obtaining the watermark image is as follows: For the trigger area CI(i,j) of the original image, apply M to obtain the initially distorted image W(i′,j′), that is: W=M·CI The original position of W(i′,j′) in CI(i,j) can be inferred by the following formula: CI=M -1 ·IN The bicubic interpolation method is used to map W(i′,j′) to CI(i,j), and finally the watermark image PI is obtained.
8. The model copyright protection method based on multi-partition adaptive triggering according to claim 1 is characterized in that: The process of embedding watermarks into marked samples using watermark images is as follows: From D v Select some samples from the above, and use the watermark generation function G to generate the trigger sample set D p , and change the label to the watermark label y predetermined by the copyright owner t , the remaining samples are normal sample set D c , we get the trigger data set D=D containing normal samples and labeled samples p ∪D c , when the copyright owner uses D for training, the watermark model will be obtained Where θ is the parameter of the watermark model, L is the loss function, and the training process is as follows: in, Indicates that for samples without watermarks, the prediction of the watermark model is its original label y i ; Indicates that for the watermarked labeled sample, the watermark model predicts the watermark label y t ; Indicates that for normal class samples with watermarks added, the prediction of the watermark model is the original label Divide the labeled class samples into k partitions, each partition corresponds to a unique trigger t i , only when the input sample x v Belongs to partition p i and contains the corresponding trigger t i When t , by adding the following process to optimize the precise mapping between triggers and partitions: in, Represents the sample for the i-th partition in the labeled class Add a trigger for the jth partition The prediction result of the watermark model is its original label y v ; Represents the sample for the i-th partition in the labeled class Add a trigger for the combination of the i-th partition and the j-th partition The prediction result of the watermark model is the original label y v .
9. The model copyright protection method based on multi-partition adaptive triggering according to claim 1 is characterized in that: The calculation process of the watermark accuracy is as follows: Prepare the watermarked sample {x w }, input the watermarked labeled samples into the suspicious model one by one, and record the predicted label y of each input sample p , observe whether the label predicted by the model is consistent with the true watermark label y t same; m is defined as the suspicious model in the input labeled sample x w The number of correct predictions: in, is the i-th labeled sample, is the predicted label of the suspicious model for the labeled sample, and count represents the sum of the number of correct predictions; The watermark accuracy q is the ratio of the number of correctly predicted labeled samples m to the total number of labeled samples n: q=m / n If the watermark accuracy exceeds the preset threshold δ, it means that the watermark defined by the copyright owner is embedded in the suspicious model, so it is judged as a pirated model, otherwise it is a genuine model.
10. A model copyright protection system based on multi-partition adaptive triggering, adopting a model copyright protection method based on multi-partition adaptive triggering as claimed in any one of claims 1 to 9, characterized in that: include: The sample processing module is used to obtain a labeled sample set and a normal sample set, input the labeled sample set into the pre-trained encoder for feature extraction, and apply the Gaussian mixture model clustering algorithm GMM to partition and obtain partition samples, and merge the partition samples with the normal sample set to obtain a new data set; The sample partitioning module is used to divide the labeled samples of the new data set into multiple partitions, iteratively update each partition based on the DPSO algorithm, obtain the global optimal trigger position and shape, obtain the original image, and optimize the original image based on the global optimal trigger position and shape to obtain the transformation field; A watermark embedding module is used to obtain a preliminary distorted image by applying the transformation field to the trigger region of the original image, map the preliminary distorted image to the trigger region of the original image using a bicubic interpolation method, obtain a watermarked image, and use the watermarked image to embed a watermark into a marked sample to obtain a marked sample embedded with a watermark; The model judgment module is used to input the marked samples embedded with watermarks into the suspicious model, and output the number of marked samples predicted correctly by the suspicious model. The watermark accuracy is calculated based on the number of marked samples predicted correctly and the total number of marked samples embedded with watermarks. The suspicious model is judged based on the comparison result of the watermark accuracy with the preset threshold. If the watermark accuracy is greater than the preset threshold, it is a pirated model, otherwise it is a genuine model.
Citation Information
Patent Citations
Method for defending alternative model attack and verifying copyright of DNN model
CN115828188A
Machine learning model copyright protection method based on model watermark
CN116244669A