Semi-fragile watermark tampering positioning method based on model fingerprints

By generating and embedding a semi-fragile watermark method of model fingerprints, the problem of malicious tampering of deep neural network models in cloud platforms or open source communities is solved, and accurate authentication and tampering location of model content are achieved, which improves the security and stability of the model and is suitable for practical application scenarios.

CN120671200AActive Publication Date: 2025-09-19CHANGAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510833195.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-19
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Existing deep neural network models are at risk of malicious tampering when uploaded to cloud platforms or open source communities. Traditional model protection methods have difficulty accurately distinguishing between normal updates and malicious tampering, and cannot provide effective tampering location information, making model repair and maintenance difficult.

Method used

A semi-fragile watermarking method based on model fingerprint is adopted to generate semi-fragile samples and embed the model fingerprint. By analyzing the sample output results, it is determined whether the model has been tampered with and the tampering location is located. The method includes initializing samples, generating target labels, designing loss functions, updating samples, embedding model fingerprints, verification and iteration, extracting model parameters, compressing weights to generate fingerprints, and comparing feature sequences.

Benefits of technology

It achieves accurate authentication of model content, reduces false positive and false negative rates, accurately locates tampering locations, ensures model performance stability and security, improves model security and reliability, and provides strong support for model repair and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671200A_ABST
    Figure CN120671200A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-fragile watermark tampering positioning method based on a model fingerprint. The method comprises the following steps: initializing a semi-fragile sample; generating a target label; selecting a standard image; designing a loss function; updating the semi-fragile sample; performing verification and iteration; extracting and grouping model parameters; compressing each group of model weights and generating model fingerprints; embedding a model fingerprint in the generated semi-fragile sample; initializing a counter; cyclically verifying each semi-fragile sample; comparing the prediction result with a target label; calculating authentication accuracy and judging whether the model is unauthorized; suspicious model parameters are extracted and grouped; compressing the weight of each group of suspicious models and generating suspicious feature vectors; extracting a model fingerprint embedded in the semi-fragile sample; converting a binary sequence into a hexadecimal sequence; the feature sequences are compared and tampering locations are determined. According to the method, the accuracy of model content authentication is improved, accurate tampering positioning is realized, the model performance is not influenced, the model security is enhanced, and the sustainable development of the technology is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and network security technology, and in particular to a semi-fragile watermark tampering positioning method based on model fingerprints. Background Art

[0002] Amidst the surge in artificial intelligence, deep neural network models, with their powerful learning and generalization capabilities, have rapidly become a core engine driving technological innovation. They have demonstrated immense potential and value in fields such as image recognition, speech processing, and natural language processing, significantly driving technological advancement and application development. However, as these models become increasingly commercially available, they also face unprecedented security challenges.

[0003] On the one hand, when models are uploaded to public environments such as cloud platforms and open source communities, the risk of malicious tampering increases significantly, as these environments may lack strict security controls and monitoring mechanisms. Malicious attackers may compromise the integrity of the model by tampering with model parameters or inserting malicious code, thereby affecting its performance and accuracy. This tampering may not only cause errors in real-world applications but also raise serious security and legal issues, such as privacy breaches and data tampering.

[0004] On the other hand, while traditional model protection methods, such as digital watermarking, can verify model integrity to a certain extent, they have numerous limitations. These methods are often overly sensitive to even the slightest changes in the model, making it difficult to distinguish between legitimate model updates and malicious tampering. Once a model has been tampered with, these methods often only provide simple integrity verification results, but fail to provide effective tampering location information. This means that even if model tampering is discovered, the exact location and scope of the tampering cannot be accurately determined, making model repair and maintenance extremely difficult.

[0005] To address this challenge, the present invention proposes an innovative semi-fragile neural network watermarking method. This method generates a set of special semi-fragile samples that maintain stable and accurate output under normal model processing, but exhibit obvious anomalies in maliciously tampered models. By analyzing the model's output results for these semi-fragile samples, it is possible to quickly determine whether the model has been tampered with. More importantly, this method can further locate the specific location of the tampering, providing strong support for model repair and maintenance. This semi-fragile neural network watermarking method not only effectively improves the security of the model, but also ensures the reliability and stability of the model in commercial applications. Summary of the Invention

[0006] In view of this, the present invention provides a semi-fragile watermark tampering location method based on model fingerprint.

[0007] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0008] The semi-fragile watermark tampering location method based on model fingerprint includes the following steps:

[0009] Step 1: Initialize the semi-fragile sample

[0010] Randomly select a portion of samples from the standard sample set and initialize these samples as semi-fragile samples;

[0011] Step 2: Generate target labels

[0012] For the initialized semi-fragile samples, a target label different from its original label is randomly assigned through the key;

[0013] Step 3: Standard image selection

[0014] According to the generated target label, randomly select the standard image with the corresponding label from the standard sample set;

[0015] Step 4: Design a loss function

[0016] Ensure that model outputs are consistent with target labels and improve transferability to semi-fragile samples;

[0017] Step 5: Update the semi-fragile sample

[0018] By minimizing the total loss function through the gradient descent method, the semi-fragile samples are updated so that they both conform to the target label and have high transferability;

[0019] Step 6: Verify and Iterate

[0020] Verify whether the updated sample meets the conditions. If so, the prediction result is consistent with the target label and the similarity is less than the threshold. If not, continue the iteration.

[0021] Step 7: Model parameter extraction and grouping

[0022] Extract the weight information of the deep neural network model and group them in order;

[0023] Step 8: Compress each set of model weights and generate model fingerprints

[0024] Compress each set of weight sequences and generate feature vectors, which are finally converted into binary sequences to generate model fingerprints;

[0025] Step 9: Embed model fingerprint in generated semi-fragile samples

[0026] Perform discrete wavelet transform on the semi-fragile samples, embed the model fingerprint on the frequency domain coefficients, and generate the final sample through inverse transform;

[0027] Step 10: Initialize the counter

[0028] Initialize the counter to 0 to record the number of samples whose model output is consistent with the expected label;

[0029] Step 11: Loop through each semi-fragile sample

[0030] Use the suspicious model to predict the semi-fragile samples and obtain the output label;

[0031] Step 12: Compare the predictions with the target labels

[0032] Compare the model output label with the expected label of the semi-fragile sample, and if they are consistent, the counter is increased by 1;

[0033] Step 13: Calculate the authentication accuracy and determine whether the model is unauthorized

[0034] Calculate the authentication accuracy based on the counter value and compare it with the preset threshold to determine whether the model has been maliciously tampered with;

[0035] Step 14: Suspicious model parameter extraction and grouping

[0036] Extract the parameters of the suspicious models and group them in the same way as the original models;

[0037] Step 15: Compress each set of suspicious model weights and generate suspicious feature vectors

[0038] Compress each group of suspicious model weights and generate feature vectors, which are then connected into suspicious feature vectors;

[0039] Step 16: Extract the model fingerprint embedded in the semi-fragile sample

[0040] The embedded model fingerprint is extracted from each semi-fragile sample, and the sequence with the most occurrences is selected as the final model fingerprint by applying the majority principle;

[0041] Step 17: Convert binary sequence to hexadecimal

[0042] Convert the extracted model fingerprint into a hexadecimal feature sequence;

[0043] Step 18: Compare signature sequences and determine tampering locations

[0044] Compare the suspicious model feature sequence with the original model feature sequence, record the mismatch position, calculate the index and determine the approximate area of ​​tampering.

[0045] Preferably, in step 2, the loss function is defined using confidence interval parameters and logit vector differences.

[0046] Preferably, in step 8 and step 15, each set of weight sequences is compressed using the Brotli compression method, and a feature vector is generated using the BLAKE2 hash function.

[0047] Preferably, in step 4, the loss L1 of the part that ensures the model output is consistent with the target label is

[0048]

[0049] Where τ represents the confidence interval parameter. The larger τ is, the greater the confidence with which the model predicts that the semi-fragile sample X is labeled T. Conversely, the smaller τ is, the smaller the confidence with which the model predicts that the semi-fragile sample X is labeled T. The deep neural network model is denoted as M.

[0050] Preferably, the difference between the logits vector of the semi-fragile sample X and the logits vector of the standard image I is also included, and the mathematical expression of the loss L2 is:

[0051]

[0052] Among them, S(X) and S(I) are the logits vectors of model M for the semi-fragile sample X and the standard image I, respectively.

[0053] Preferably, in step 5, the total loss L t for

[0054] L t =L1+α·L2,

[0055] Where α<0 is the weight factor that adjusts the ratio of the two loss components;

[0056] Next, the total loss L is minimized by gradient descent t Update the semi-fragile sample, which is mathematically expressed as

[0057]

[0058] Among them, l r is the learning rate, is the direction of the gradient of the loss function with respect to X.

[0059] Preferably, in step 6, the mathematical expressions of the two conditions are as follows

[0060]

[0061] Here, argmaxM(X)=T is used to ensure that the updated semi-fragile sample X is consistent with its assigned target label T; D(O,X) represents the mean square error between the original sample O and the updated semi-fragile sample X.

[0062] Preferably, in step 8, the process of compressing the i-th group of weight sequences can be expressed as

[0063] Z i =Brotli(P i ),

[0064] Among them, Brotli(·) represents Brotli compression, Z i represents the compression result of the i-th group weight sequence;

[0065] For each set of compressed parameters Z i The BLAKE2 hash function is used to generate the model feature vector, where the feature vector of the i-th group of compression parameters can be expressed as

[0066] r i =BLAKE2(Z i )

[0067] =BLAKE2(Brotli(P i )),

[0068] Where BLAKE2(·) represents the BLAKE2 hash function;

[0069] Then, the feature vectors of each set of compression parameters are connected to generate the feature vector R = [r1, r2, ..., r m ]; Finally, the feature vector R is converted into a binary sequence to generate the model fingerprint F, which is mathematically expressed as

[0070] F=hextobin(R),

[0071] Wherein, hextobin(·) is a function that converts a hexadecimal feature vector into a binary sequence.

[0072] Preferably, in step 9, specifically: the semi-fragile sample X is subjected to discrete wavelet transform to obtain the frequency domain coefficient X′, and then the model fingerprint is embedded in the frequency domain coefficient; the model fingerprint is embedded in the extended spectrum method, which is mathematically expressed as

[0073]

[0074] Among them, β is a parameter that controls the embedding strength of the model fingerprint;

[0075] Finally, the frequency domain coefficient X′ containing the model fingerprint is transformed through inverse discrete wavelet transform to generate the final semi-fragile sample X. The process is expressed as

[0076]

[0077] Wherein, IDWT(·) represents the inverse discrete wavelet transform function.

[0078] Preferably, in step 13, the authentication accuracy is calculated as follows

[0079]

[0080] Finally, based on the comparison result of the authentication accuracy Acc and the preset threshold δ, it is determined whether the model has been maliciously tampered with. If Acc exceeds the threshold δ, the model is judged to be normal or not modified; otherwise, the model is judged to have been maliciously tampered with.

[0081] Compared with the prior art, the present invention has achieved the following technical effects:

[0082] (1) The present invention improves the accuracy of model content authentication:

[0083] The present invention achieves accurate authentication of model content by generating semi-fragile samples that remain stable under normal operation but show vulnerability when maliciously tampered with;

[0084] Compared with traditional methods, this method can more accurately distinguish normal model updates from malicious tampering, effectively reducing false positives and false negatives.

[0085] (2) The present invention achieves accurate tampering location:

[0086] By embedding a model fingerprint in the model and combining it with tampering location technology, the present invention can accurately identify which parts of the model have been tampered with, with higher positioning accuracy than existing technologies. This provides strong support for the rapid repair and maintenance of the model and reduces potential losses caused by tampering.

[0087] (3) The present invention maintains the model performance unaffected:

[0088] While achieving model content authentication and tampering location, the present invention ensures the stability of model performance;

[0089] Compared with some protection measures that may affect the performance of the model, the present invention is more suitable for actual application scenarios and ensures the normal operation of the model;

[0090] (4) The present invention enhances the security of the model:

[0091] This invention significantly improves the security of the model by providing a new security protection mechanism. It can effectively prevent malicious attackers from tampering with model parameters to damage model performance or steal sensitive information, thereby protecting the intellectual property rights and security of the model.

[0092] (5) This invention promotes the sustainable development of technology:

[0093] The implementation of this invention provides technical support for the legal use and commercialization of the model, and helps promote the sustainable development of artificial intelligence technology; at the same time, it also promotes the formulation of relevant laws and policies, and provides guarantees for the healthy development of the field of artificial intelligence. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 This is an application scenario diagram of the present invention;

[0095] Figure 2 Generate a flow chart for the semi-fragile samples of the present invention;

[0096] Figure 3 This is a flow chart of model content authentication of the present invention;

[0097] Figure 4 This is the model tampering positioning flow chart of the present invention. DETAILED DESCRIPTION

[0098] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0099] The present invention discloses a semi-fragile watermark tampering location method based on model fingerprint, comprising the following steps:

[0100] Module 1: Semi-fragile sample generation

[0101] The goal of module one is to generate a set of semi-fragile samples that are robust to normal model processing operations but vulnerable to malicious tampering of the model. The following are the detailed steps to generate semi-fragile samples:

[0102] Step 1: Initialize the semi-fragile samples. The model owner randomly selects a portion of samples from the standard sample set O and initializes these samples as semi-fragile samples X. In other words, the semi-fragile samples in the present invention are generated by changing some of the standard samples.

[0103] Step 2: Generate target labels. For the initialized semi-fragile samples generated in step 1, randomly assign target labels T using the key k. The target labels T are different from the original labels of the semi-fragile samples.

[0104] Step 3: Standard image selection: Based on the target label T in step 2, randomly select an image with label T from the standard sample set O and record it as the standard image I.

[0105] Step 4: Design loss function. This patent considers two aspects of loss to generate semi-fragile samples that meet the conditions, one to ensure that the model output is consistent with the target label and the other to improve the transferability of semi-fragile samples. Assuming that the deep neural network model is denoted as M, the loss of the part that ensures that the model output is consistent with the target label can be defined as

[0106]

[0107] Where τ represents the confidence interval parameter. The larger τ is, the greater the confidence of the model in predicting that the semi-fragile sample X is labeled T. Conversely, the smaller τ is, the smaller the confidence of the model in predicting that the semi-fragile sample X is labeled T. In addition, this patent also considers the difference between the logits vector of the semi-fragile sample X and the logits vector of the standard image I to improve the transferability of the sample. The mathematical expression of this loss is

[0108]

[0109] Here, S(X) and S(I) are the logits vectors of the model M for the semi-fragile sample X and the standard image I, respectively. Introducing the L2 loss can make the generated semi-fragile sample closer to the standard image originally labeled T, thereby ensuring the transferability of the generated semi-fragile sample.

[0110] Step 5: Update the semi-fragile samples. According to the loss function in step 4, the total loss can be defined as

[0111] L t =L1+α·L2,

[0112] Where α<0 is the weight factor that adjusts the ratio of the two loss components. Then, the total loss L is minimized by gradient descent. t Update the semi-fragile sample, which is mathematically expressed as

[0113]

[0114] Among them, l r is the learning rate, is the direction of the gradient of the loss function with respect to X.

[0115] Step 6: Verification and iteration. Check whether the prediction of the model M for the updated sample X is consistent with the target label T, and whether the similarity between the original sample O and the updated sample X is less than the threshold ε. If the conditions are met, stop the iteration; otherwise, continue the iteration. The mathematical expression of the two conditions is as follows

[0116]

[0117] Here, argmaxM(X) = T ensures that the updated semi-fragile sample X is consistent with its assigned target label T. D(O,X) represents the mean square error between the original sample O and the updated semi-fragile sample X. This ensures that the attacker cannot visually distinguish the generated semi-fragile sample from the original sample. In other words, the generated semi-fragile sample and the original sample must be highly similar.

[0118] Step 7: Model parameter extraction and grouping. First, extract the parameter information of the deep neural network model M. Specifically, extract the weight information of all layers in the deep neural network model M, and arrange these weight information in order to obtain the model weight sequence P. Then, group the model weight sequence P. Specifically, divide the model weight sequence P into m groups in order, where the weight sequences in each group have the same length. The i-th group weight sequence is denoted as P i .

[0119] Step 8: Compress each set of model weights and generate model fingerprints. Use the Brotli compression method to compress each set of weight sequences, where the process of compressing the i-th set of weight sequences can be expressed as

[0120] Z i =Brotli(P i ),

[0121] Among them, Brotli(·) represents Brotli compression, Z i Represents the compression result of the i-th group of weight sequences. For each group of compressed parameters Z i The BLAKE2 hash function is used to generate the model feature vector, where the feature vector of the i-th group of compression parameters can be expressed as

[0122] r i =BLAKE2(Z i )

[0123] =BLAKE2(Brotli(P i )),

[0124] Among them, BLAKE2(·) represents the BLAKE2 hash function. Then, the feature vectors of each set of compression parameters are connected to generate the feature vector R = [r1, r2, …, r m ]. Finally, the feature vector R is converted into a binary sequence to generate the model fingerprint F, which is mathematically expressed as

[0125] F=hextobin(R),

[0126] Wherein, hextobin(·) is a function that converts a hexadecimal feature vector into a binary sequence.

[0127] Step 9: Embed the model fingerprint F in the generated semi-fragile sample. Specifically, the semi-fragile sample X is subjected to discrete wavelet transform (DWT) to obtain the frequency domain coefficient X′, and then the model fingerprint is embedded in its frequency domain coefficient. In this patent, the extended spectrum method is used to embed the model fingerprint, which is mathematically expressed as

[0128]

[0129] Among them, β is a parameter that controls the embedding strength of the model fingerprint. Finally, the frequency domain coefficient X′ containing the model fingerprint is transformed through the inverse discrete wavelet transform to generate the final semi-fragile sample X. The process is expressed as

[0130]

[0131] Wherein, IDWT(·) represents the inverse discrete wavelet transform function.

[0132] Module 2: Model Content Certification

[0133] Module 2 determines whether the model has been maliciously tampered by comparing the output of model M for semi-fragile samples X with the expected label. By introducing the authentication accuracy Acc and threshold δ, the integrity of the model can be effectively identified.

[0134] Step 10: Initialize the counter. Initialize the counter to count = 0 to record the number of samples whose model output is consistent with the expected label.

[0135] Step 11: Loop and verify each semi-fragile sample. For each semi-fragile sample X i , using the suspicious model M′ for the semi-fragile sample X i Make a prediction and get the output label result, recorded as T pi .

[0136] Step 12: Compare the prediction results with the target labels. pi with the expected label T for the semi-fragile sample i Compare. If T pi Equal to T i , the sample is considered to have passed the verification, and the counter count increases by 1. Otherwise, the counter count value remains unchanged.

[0137] Step 13: Calculate the authentication accuracy and determine whether the model is unauthorized. The authentication accuracy is calculated as follows

[0138]

[0139] Finally, based on the comparison between the authentication accuracy Acc and the preset threshold δ, it is determined whether the model has been maliciously tampered with. If Acc exceeds the threshold δ, the model is considered normal or has not been modified. Otherwise, the model is considered to have been maliciously tampered with.

[0140] Module 3: Model Tampering Location

[0141] The goal of module three is to accurately determine the specific parameter locations in the model that have been maliciously tampered with by comparing the feature vectors of the suspicious model with the original model fingerprint extracted from the semi-vulnerable sample, thereby ensuring the integrity and security of the model.

[0142] Step 14: Extract and group suspicious model parameters. Referring to step 7, extract the parameters of the suspicious model M′ and group them in the same way, where the i-th group of suspicious model parameters is denoted as P′ i .

[0143] Step 15: Compress each set of suspicious model weights and generate a suspicious feature vector. Referring to Step 8, first compress each set of suspicious model weights using the Brotli compression method. Next, use the BLAKE2(·) hash function to generate suspicious model feature vectors and concatenate them to generate the suspicious feature vector R′.

[0144] Step 16: Extract the model fingerprint embedded in the semi-fragile sample. For each semi-fragile sample, extract the model fingerprint embedded in it according to the spread spectrum extraction algorithm. The model fingerprint extracted from the jth semi-fragile sample is denoted as Then all the extracted model fingerprints Apply the majority principle and select the sequence with the most occurrences as the extracted model fingerprint

[0145] Step 17: Convert the binary sequence to hexadecimal. Use the bintohex(·) function to convert the extracted model fingerprint Convert to hexadecimal signature sequence To facilitate comparison and analysis.

[0146] Step 18: Compare the feature sequences and determine the tampering location. Compare the suspicious model feature sequence R′ generated in step 15 with the original model feature sequence extracted in step 17 Specifically, starting from the start of each sequence, the corresponding bits of the two sequences are compared bit by bit. If a mismatch is found at any position, that position is recorded. The index q of the mismatching position is calculated and divided by the number of parameter groups m to determine the approximate area of ​​tampering. Finally, the tampering location information is returned, indicating which part of the model parameters have been tampered with.

[0147] The above description is merely a preferred embodiment of the present invention and does not limit the technical scope of the present invention. Therefore, any minor modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A semi-fragile watermark tampering location method based on model fingerprint, characterized by: The steps include: Step 1: Initialize the semi-fragile sample Randomly select a portion of samples from the standard sample set and initialize these samples as semi-fragile samples; Step 2: Generate target labels For the initialized semi-fragile samples, a target label different from its original label is randomly assigned through the key; Step 3: Standard image selection According to the generated target label, randomly select the standard image with the corresponding label from the standard sample set; Step 4: Design a loss function Ensure that model outputs are consistent with target labels and improve transferability to semi-fragile samples; Step 5: Update the semi-fragile sample By minimizing the total loss function through the gradient descent method, the semi-fragile samples are updated so that they both conform to the target label and have high transferability; Step 6: Verify and Iterate Verify whether the updated sample meets the conditions. If so, the prediction result is consistent with the target label and the similarity is less than the threshold. If not, continue the iteration. Step 7: Model parameter extraction and grouping Extract the weight information of the deep neural network model and group them in order; Step 8: Compress each set of model weights and generate model fingerprints Compress each set of weight sequences and generate feature vectors, which are finally converted into binary sequences to generate model fingerprints; Step 9: Embed model fingerprint in generated semi-fragile samples Perform discrete wavelet transform on the semi-fragile samples, embed the model fingerprint on the frequency domain coefficients, and generate the final sample through inverse transform; Step 10: Initialize the counter Initialize the counter to 0 to record the number of samples whose model output is consistent with the expected label; Step 11: Loop through each semi-fragile sample Use the suspicious model to predict the semi-fragile samples and obtain the output label; Step 12: Compare the predictions with the target labels Compare the model output label with the expected label of the semi-fragile sample, and if they are consistent, the counter is increased by 1; Step 13: Calculate the authentication accuracy and determine whether the model is unauthorized Calculate the authentication accuracy based on the counter value and compare it with the preset threshold to determine whether the model has been maliciously tampered with; Step 14: Suspicious model parameter extraction and grouping Extract the parameters of the suspicious models and group them in the same way as the original models; Step 15: Compress each set of suspicious model weights and generate suspicious feature vectors Compress each group of suspicious model weights and generate feature vectors, which are then connected into suspicious feature vectors; Step 16: Extract the model fingerprint embedded in the semi-fragile sample The embedded model fingerprint is extracted from each semi-fragile sample, and the sequence with the most occurrences is selected as the final model fingerprint by applying the majority principle; Step 17: Convert binary sequence to hexadecimal Convert the extracted model fingerprint into a hexadecimal feature sequence; Step 18: Compare signature sequences and determine tampering locations Compare the suspicious model feature sequence with the original model feature sequence, record the mismatch position, calculate the index and determine the approximate area of ​​tampering.

2. The semi-fragile watermark tampering location method based on model fingerprint according to claim 1 is characterized in that: In step 2, the loss function is defined using the confidence interval parameter and the logit vector difference.

3. The semi-fragile watermark tampering location method based on model fingerprint according to claim 1 is characterized in that: In step 8 and step 15, each set of weight sequences is compressed using the Brotli compression method, and a feature vector is generated using the BLAKE2 hash function.

4. The semi-fragile watermark tampering location method based on model fingerprint according to claim 1 is characterized in that: In step 4, the loss L1 of the part that ensures the model output is consistent with the target label is Where τ represents the confidence interval parameter. The larger τ is, the greater the confidence with which the model predicts that the semi-fragile sample X is labeled T. Conversely, the smaller τ is, the smaller the confidence with which the model predicts that the semi-fragile sample X is labeled T. The deep neural network model is denoted as M.

5. The semi-fragile watermark tampering location method based on model fingerprint according to claim 4 is characterized in that: It also includes the difference between the logits vector of the semi-fragile sample X and the logits vector of the standard image I. The mathematical expression of its loss L2 is Among them, S(X) and S(I) are the logits vectors of model M for the semi-fragile sample X and the standard image I, respectively.

6. The semi-fragile watermark tampering location method based on model fingerprint according to claim 1 is characterized in that: In step 5, the total loss L t for L t =L1+α·L2, Where α<0 is the weight factor that adjusts the ratio of the two loss components; Next, the total loss L is minimized by gradient descent t Update the semi-fragile sample, which is mathematically expressed as Among them, l r is the learning rate, is the direction of the gradient of the loss function with respect to X.

7. The semi-fragile watermark tampering location method based on model fingerprint according to claim 1 is characterized in that: In step 6, the mathematical expressions of the two conditions are as follows Here, argmaxM(X)=T is used to ensure that the updated semi-fragile sample X is consistent with its assigned target label T; D(O,X) represents the mean square error between the original sample O and the updated semi-fragile sample X.

8. The semi-fragile watermark tampering location method based on model fingerprint according to claim 1 is characterized in that: In step 8, the process of compressing the i-th group weight sequence can be expressed as With i =Brotli(P i ), Among them, Brotli(·) represents Brotli compression, Z i represents the compression result of the i-th group weight sequence; For each set of compressed parameters Z i The BLAKE2 hash function is used to generate the model feature vector, where the feature vector of the i-th group of compression parameters can be expressed as r i =BLAKE2(Z i ) =BLAKE2(Brotli(P i )), Where BLAKE2(·) represents the BLAKE2 hash function; Then, the feature vectors of each set of compression parameters are connected to generate the feature vector R = [r1, r2, ..., r m ]; Finally, the feature vector R is converted into a binary sequence to generate the model fingerprint F, which is mathematically expressed as F = hextobin(R), Wherein, hextobin(·) is a function that converts a hexadecimal feature vector into a binary sequence.

9. The semi-fragile watermark tampering location method based on model fingerprint according to claim 1 is characterized in that: In step 9, specifically: the semi-fragile sample X is subjected to discrete wavelet transform to obtain the frequency domain coefficient X ′ , and then embed the model fingerprint in its frequency domain coefficients; the model fingerprint is embedded in the extended spectrum method, and its mathematical expression is Among them, β is a parameter that controls the embedding strength of the model fingerprint; Finally, the frequency domain coefficient X containing the model fingerprint ′ The final semi-fragile sample X is generated by inverse discrete wavelet transform. The process is expressed as Wherein, IDWT(·) represents the inverse discrete wavelet transform function.

10. The semi-fragile watermark tampering location method based on model fingerprint according to claim 1 is characterized in that: In step 13, the authentication accuracy is calculated as follows: Finally, based on the comparison result of the authentication accuracy Acc and the preset threshold δ, it is determined whether the model has been maliciously tampered with. If Acc exceeds the threshold δ, the model is judged to be normal or not modified; otherwise, the model is judged to have been maliciously tampered with.

Citation Information

Patent Citations

  • Omni-directional prediction error histogram modification-based reversible image watermarking algorithm

    CN102036079A

  • Neural network authentication method and device based on semi-fragile model watermark and medium

    CN119357927A

  • Tamper Protection and Video Source Identification for Video Processing Pipeline

    US20180253567A1