Verifiable image data set ownership verification method based on generality prediction

By calculating the main probability value and watermark robust value based on common prediction, and using the optimal likelihood ratio detection strategy of Neyman-Pearson lemma, the problem that existing dataset watermark methods cannot effectively prevent attacks from bypassing verification is solved, and the verifiability and robustness of dataset ownership are achieved.

CN120234787APending Publication Date: 2025-07-01NORTH CHINA ELECTRIC POWER UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510379820.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing dataset watermarking methods cannot effectively prevent attackers from bypassing ownership verification by adding noise or modifying the input of verification samples. They lack quantitative research and theoretical guarantees on the robustness of dataset watermarks, and cannot cope with advanced adaptive attack methods.

Method used

Using a common prediction-based method, by calculating the main probability value and the watermark robust value, combined with the optimal likelihood ratio detection strategy of Neyman-Pearson lemma, the ownership verification of the image dataset is achieved to ensure that the watermark is still valid under noise.

Benefits of technology

No need to rely on the authenticity process to prevent attackers from bypassing ownership verification by adding noise or modifying the input of verification samples, ensuring the robustness of the dataset watermark and the validity of copyright verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234787A_ABST
    Figure CN120234787A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data set copyright protection, and particularly relates to a generality prediction-based verifiable image data set ownership verification method, which comprises the following steps of: acquiring an image data set for suspicious model training, and calculating a main probability value for a benign sample in the image data set; for a watermark sample in the image data set, calculating a watermark robust value; and a generality prediction method based on an optimal likelihood ratio test strategy is used to realize image data set ownership verification. The method does not need to depend on whether the verification process is honest or not, and an attacker is prevented from trying to bypass ownership verification by adding noise or modifying input of a verification sample; and an attacker is prevented from avoiding verification watermarks through a targeted training model, so that unique features of the watermarks are not expressed any more, and copyright verification is bypassed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of dataset copyright protection, and particularly relates to a verifiable image dataset ownership verification method based on commonality prediction. Background Art

[0002] Dataset copyright protection is an area of increasing concern in recent years. With the wide application of deep learning models, their success depends on a variety of open-source datasets (such as ImageNet). Researchers train models by using these datasets and improve the model performance according to the evaluation results. However, these datasets are often limited to educational use rather than commercial use, and the collection process is time-consuming and laborious. Therefore, the copyright issue of datasets has become increasingly prominent. However, traditional data protection methods cannot be used to protect datasets, which will damage the usability and functionality of the datasets. Specifically, encryption protects the sensitivity of data by encrypting part or all of the data, but it limits the functionality of the data. Digital watermarking embeds a pattern as a watermark in the protected data to verify ownership. Privacy protection is to prevent sensitive information from being leaked during the training process. These methods often require a more detailed training process or even the whole process, which is not disclosed to users. Therefore, none of them can be used to protect the copyright of open-source datasets.

[0003] In the prior art, backdoor attacks, as an emerging form of attack, mainly target the training stage of deep neural networks (DNNs). Specifically, the attacker maliciously manipulates a subset of the training data and implants a backdoor into the model, so that the model performs normally when predicting benign samples, but once the input contains a specific trigger pattern, it will produce incorrect predictions, thus bringing significant security risks to DNN-based applications. According to the different capabilities of the attacker, the existing backdoor attacks can be divided into three categories: pure poisoning attacks, training control attacks, and model modification attacks. Among them, pure poisoning attacks mainly tamper with the training dataset without interfering with the training process. Therefore, focus on pure poisoning attacks and use such unique characteristics to design a watermarking technology for dataset ownership verification, aiming to protect public datasets.

[0004] Dataset ownership verification (DOV), which aims to verify whether a suspicious third-party model is trained on a protected dataset, is the most widely used and effective method for protecting the copyright of open-source datasets at present. Specifically, as Figure 3 shown, the DOV strategy is a post-audit method. By introducing imperceptible watermark samples into the original dataset, a watermarked version is generated for release, so as to maintain the performance of the model on benign test samples, while inducing unique prediction behaviors on the verification samples. The dataset owner verifies the ownership of the dataset by checking whether the suspicious third-party model exhibits these unique prediction behaviors, all of which are carried out in a black-box verification environment. However, the existing dataset watermarking methods still have the following problems and disadvantages:

[0005] (1) Existing dataset watermarking methods successfully rely on a potential assumption that the verification process is honest, that is, it is assumed that the suspicious model will faithfully use the verification samples for prediction during the verification process without adding any noise to its input. However, this assumption does not always hold in practical applications. In particular, when the suspicious model realizes the potential progress of dataset ownership verification, it may adopt a series of attack strategies and try to bypass ownership verification by adding noise or modifying the input of the verification samples.

[0006] (2) There is a lack of quantitative research and theoretical guarantees for the robustness of dataset watermarks. With the continuous development of attack techniques, existing dataset watermarking methods cannot effectively cope with future advanced and adaptive attack means. Attackers can avoid verifying watermarks by training models specifically, making the unique features of the watermarks no longer manifested, and thus easily bypassing copyright verification.

[0007] Therefore, there is an urgent need for a verifiable image dataset ownership verification method based on commonality prediction, which does not rely on whether the verification process is honest, prevents attackers from trying to bypass ownership verification by adding noise or modifying the input of the verification samples, and prevents attackers from avoiding verification watermarks by training models specifically, making the unique features of the watermarks no longer manifested, to bypass copyright verification. Summary of the Invention

[0008] The object of the present invention is to provide a verifiable image dataset ownership verification method based on commonality prediction, including the following steps:

[0009] Step S1: Obtain the image dataset used for training the suspicious model. For the benign samples in the image dataset, calculate the main probability value, including: independently select K correctly predicted samples from K categories respectively, and for each sample x K Add M times of noise by using the Monte Carlo estimation method; predict the probability of each category through the benign model to obtain the prediction distribution; select the maximum value in the average prediction distribution of each category as the main probability value;

[0010] Step S2: For the watermark samples in the image dataset, calculate the watermark robustness value, including: independently select K correctly predicted samples from K categories respectively, and construct watermark samples by embedding a backdoor trigger into each sample and assigning a fixed target label; for each watermark sample x k Add M times of noise, and calculate the prediction probability of each sample on the target category through the suspicious model; select the minimum value among the prediction probability values as the watermark robustness value;

[0011] Step S3: Use the commonality prediction method based on the optimal likelihood ratio test strategy, including: combining two necessary attributes of the authentication dataset watermark to give the final verification condition of the authentication-robust dataset watermark; using the commonality prediction method to determine the judgment condition of unauthorized training; giving a strict verification condition for obtaining unauthorized training based on the optimal likelihood ratio detection strategy of the Neyman-Pearson lemma; judging whether the suspicious model is trained using the protected dataset based on the strict verification condition to achieve the verification of the ownership of the image dataset.

[0012] The specific calculation of the main probability value in step S1 includes:

[0013] Step S11: Define the prediction distribution, including: for the benign model g(·; w) with parameter w: When adding random noise ε to the input sample x, the prediction distribution represents the probability distribution over K classes:

[0014]

[0015] In the formula, ∈ is the random noise sampled from the noise distribution The noise distribution is a Gaussian distribution or a uniform distribution; represents the probability that the model still predicts class k when the input x is disturbed by noise, is the probability that the event occurs when the noise ∈ follows the distribution argmaxg(x + ∈; w) is the class with the maximum probability in the model output, that is, the final predicted class, and k represents a specific class label;

[0016] Step S12: Use the Monte Carlo method to estimate the prediction distribution: by introducing random noise into the benign samples multiple times, recording the output counts of each class, and using the frequency to approximate the probability;

[0017] The estimated prediction distribution is:

[0018]

[0019] In the formula, M is the number of times of sampling random noise, is the indicator function;

[0020] Step S13: For the given benign model g(·; w), independently sample K correctly predicted samples from each class, denoted as x1, x2,..., x K , calculate the prediction distribution for the K correctly predicted samples respectively, and average by class to obtain the average prediction distribution of each class;

[0021] Step S14: Calculate the main probability value: Select the maximum value of the average prediction distribution for each category to estimate the main probability value:

[0022]

[0023] where x1, x2,..., x K are K correctly predicted samples independently sampled from each category.

[0024] The calculation of the watermark robustness value in step S2 specifically includes:

[0025] Step S21: Sample samples: Given a suspicious model f(·; θ), independently sample K correctly predicted samples from each category, denoted as x1, x2,..., x K ;

[0026] Step S22: Embed the trigger and construct the watermark samples, including: Embed the trigger δ in each sample and assign a specified target label y to each sample, thereby constructing K watermark samples x k = x k + δ;

[0027] Step S23: Calculate the watermark robustness value, including: Select the minimum probability value from the prediction distribution of the suspicious model:

[0028]

[0029] where x k = x k + δ, and x1, x2,..., x K are K samples independently sampled from each category, satisfying arg maxg w (x k ) = k (k ∈ {1,…, K});

[0030] For the suspicious model f(·; θ): The probability at the y position in the prediction distribution is defined as:

[0031]

[0032] where is the probability of the event occurring when the noise ∈ follows the distribution , where ∈ is the random noise sampled from the noise distribution , and the noise distribution is a Gaussian distribution or a uniform distribution, x is the clean sample, δ is the watermark perturbation, i.e., the trigger, and θ is the suspicious model parameter.

[0033] The final verification condition for the authenticated robust dataset watermark in step S3 is:

[0034]

[0035] Wherein, represents the robustness of the transform-based watermark, represents the R-functional stability, and τ represents the verification threshold;

[0036] The two necessary attributes of the authenticated dataset watermark are: Assume are K independent benign samples, satisfying argmaxg w (x k ) = k, and ∈ is the noise sampled from the noise distribution ; Consider the watermarked version x + r of the watermark with the target label y specified by the defender, and make the following definitions for the watermark model f(·; θ):

[0037] (1) Robustness of transform-based watermark: The lower bound of the probability that the watermarked sample is always predicted as the target label under the given noise distribution. The form of the robustness of the transform-based watermark is:

[0038]

[0039] (2) R-functional stability: Given that the watermark transform is constrained within R, i.e., ||r k ||2 ≤ R, the stability is defined as the lower bound of the probability that the benign sample x is always predicted as the target label under the noise distribution. The form of the R-functional stability is:

[0040]

[0041] Wherein, x k is the k-th sample, r k is the perturbation size for the k-th sample, replacing the trigger δ in step S22, and ∈ is the random noise sampled from the noise distribution , and the noise distribution is a Gaussian distribution or a uniform distribution, R ≥ 0 represents the maximum amplitude of the perturbation for k selected dataset watermark samples, representing the upper limit of the perturbation intensity, denoted as

[0042] The method of using commonality prediction in step S3 to determine the judgment condition for unauthorized training includes:

[0043] Step S31: Train a benign model, including: Training J benign models using the method of calculating the main probability value in step S1, and calculating the main probability values of the J benign models respectively to form a calibration set, denoted as:

[0044] Step S32: Filter outliers, including: filtering outliers through outlier detection according to a preset filtering ratio; the filtering ratio is: m = κ.J, where κ is a hyperparameter representing the filtering ratio;

[0045] Step S33: Calculate the p-value for ownership verification, including: calculating the p-value for ownership verification by using the compliance prediction based on the watermark robustness value of the suspicious model and the main probability value in the calibration set:

[0046]

[0047] where J is the size of the calibration set, m represents the number of outliers in the calibration set, is the indicator function, when is true, the value of is 1, otherwise 0; is the main probability value of the j-th benign model, and W is the shorthand form of the watermark robustness value of the suspicious model

[0048] Step S34: Set the verification condition: Only when p ≥ 1 - α0, the suspicious model is trained on the protected dataset;

[0049] where α0 is the selected significance level, and 1 - α0 is the confidence level;

[0050] Step S35: Obtain the judgment condition for unauthorized training:

[0051]

[0052] where represents the calibration threshold, is the watermark robustness value of the suspicious model, J is the number of benign models, i.e., the size of the calibration set, m is the number of outliers in the calibration set, α0 is the significance level, and α0 takes the value of 0.05.

[0053] The strict verification condition for obtaining unauthorized training in step S3 includes:

[0054] For the authentication robust dataset watermark, if the optimal type-II error for testing the null hypothesis and the alternative hypothesis satisfies the following conditions, it is determined that the ownership verification of the picture dataset passes:

[0055]

[0056] where H1 is the alternative hypothesis, f θ is the suspicious model, is the noise distribution, is the watermark robustness based on the transformation, Recognizing an image without a watermark as an image with a watermark, i.e., the first type of error, while ensuring that the first type of error rate does not exceed a threshold select a test strategy that minimizes the second type of error; denote the calibration set the j-th smallest element in it; J is the number of benign models, m is the number of filtered samples, α0 is the selected significance level, and g w is a benign model, denote the minimum second type of error;

[0057] The first type of error is: the probability of misidentifying an image with a watermark as an image without a watermark, i.e., the null hypothesis holds but is rejected, defined as: β1(φ; H0) = E x (φ(x));

[0058] The second type of error is: the probability of misidentifying an image without a watermark as an image with a watermark, i.e., the null hypothesis does not hold but is accepted, defined as β2(φ; H1) = E x (1 - φ(x))

[0059] wherein, H0 is the null hypothesis, H1 is the alternative hypothesis, and E x is the expectation with respect to the hypothesis, and φ(x) is the test function;

[0060] The optimal likelihood ratio detection strategy based on the Neyman - Pearson lemma in step S3 includes:

[0061] By controlling the first type of error to be at the minimum threshold, minimize the second type of error, and control the minimum second type of error to be greater than the calibration threshold: set the significance level α1 as the maximum acceptable probability of the first type of error, i.e.:

[0062]

[0063] wherein, While ensuring that the first type of error rate does not exceed the threshold α1, select a test strategy that minimizes the second type of error.

[0064] Another object of the present invention is to provide a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the verifiable image dataset ownership verification method based on commonality prediction according to the present invention.

[0065] Another object of the present invention is to provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor executes the verifiable image dataset ownership verification method based on commonality prediction according to the present invention.

[0066] The beneficial effects of the present invention are as follows:

[0067] On the one hand, aiming at the dishonest verification process of dataset verification, the present invention uses the Monte Carlo estimation method to introduce noise into multiple verification samples (including watermark samples and benign samples) multiple times, and then feeds them to the model to calculate the frequency of each type of prediction result. Using the frequency instead of the probability can increase the probability of the desired label and prevent verification failure. Even if malicious perturbations are suffered before prediction, the prediction result will not change.

[0068] On the other hand, the present invention establishes a verification threshold between the prediction probability distributions of watermark samples and benign samples based on the method of commonality prediction, defines it as the calibration threshold, and provides a theoretical guarantee for the present invention through the optimal likelihood ratio of the Neyman-Pearson lemma. As long as specific conditions are met, the watermark of the dataset will not be deleted, that is, the ownership guarantee of the dataset is verified. In other words, in the worst case, even if a malicious user completely removes the watermark, the present invention can still guarantee the copyright verification of the dataset.

[0069] Applying the verifiable image dataset ownership verification method based on commonality prediction disclosed by the present invention does not need to rely on whether the verification process is honest, preventing an attacker from trying to bypass the ownership verification by adding noise or modifying the input of the verification sample; preventing an attacker from circumventing the verification watermark through targeted training of the model, making the unique features of the watermark no longer manifested, to bypass the copyright verification. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 is a schematic flowchart of a verifiable image dataset ownership verification method based on commonality prediction according to the present invention;

[0071] Figure 2 is a schematic architecture diagram of the verifiable image dataset ownership verification method based on commonality prediction in an embodiment of the present invention;

[0072] Figure 3 is a schematic diagram for comparing watermark patterns in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] The present invention provides a verifiable image dataset ownership verification method based on commonality prediction. The following further details the present invention with reference to the accompanying drawings.

[0074] As Figure 1The disclosed embodiment of the present invention discloses a method for verifying the ownership of a verifiable image dataset based on commonality prediction, including the following steps:

[0075] Step S1: Obtain an image dataset for training a suspicious model. For the benign samples in the image dataset, calculate the main probability value, including: independently select K correctly predicted samples from K categories, and for each sample x K Add M times of noise using the Monte Carlo estimation method; predict the probability of each category through the benign model to obtain the prediction distribution; select the maximum value in the average prediction distribution of each category as the main probability value;

[0076] Step S2: Calculate the watermark robustness value for the watermark samples in the image dataset, including: independently select K correctly predicted samples from K categories, and construct watermark samples by embedding a backdoor trigger into each sample and assigning a fixed target label; for each watermark sample x k Add M times of noise, and calculate the prediction probability of each sample on the target category through the suspicious model; select the minimum value among the prediction probability values as the watermark robustness value;

[0077] Step S3: Use the commonality prediction method based on the optimal likelihood ratio test strategy, including: combining two necessary attributes of the authentication dataset watermark to give the final verification condition of the authentication robust dataset watermark; using the commonality prediction method to determine the judgment condition for unauthorized training; giving a strict verification condition for obtaining unauthorized training based on the optimal likelihood ratio detection strategy of the Neyman-Pearson lemma; judging whether the suspicious model uses a protected dataset for training based on the strict verification condition to achieve the verification of the ownership of the image dataset.

[0078] In this embodiment, a method for verifying the ownership of a verifiable image dataset based on commonality prediction disclosed by the present invention is applied to implement a verifiable dataset watermark. The dataset watermark is divided into a watermark stage and a verification stage; the verifiable, also known as authenticable, provides a theoretical guarantee for the verification of the dataset ownership; the ownership verification is the second stage of the dataset watermark, which is used to achieve copyright protection.

[0079] In this embodiment, the specific implementation process of the method for verifying the ownership of a verifiable image dataset based on commonality prediction is as follows:

[0080] I. Calculate the main probability value PP

[0081] Estimate the main probability value PP by selecting the maximum value in the average prediction distribution PD. The prediction distribution PD represents the prediction probability distribution of the benign model for each category when adding random noise to the benign test samples. Specifically as follows:

[0082] 1. Define the prediction distribution PD, including: for the benign model g(·; w) with parameter w: When adding random noise ε to the input sample x, the prediction distribution PD represents the probability distribution over K classes:

[0083]

[0084] where ∈ is the random noise sampled from the noise distribution and the noise distribution is a Gaussian distribution or a uniform distribution; represents the probability that the model still predicts class k when the input x is disturbed by noise, is the probability of the event occurring when the noise ∈ follows the distribution argmax g(x + ∈; w) is the class with the highest probability in the model output, that is, the final predicted class, and k represents a specific class label;

[0085] 2. Use the Monte Carlo method to estimate the prediction distribution PD: By introducing random noise into the benign samples multiple times, record the output counts for each class, and use the frequencies to approximate the probabilities.

[0086] Specifically, the prediction distribution PD is estimated as follows: where M is the number of times of sampling the random noise, is the indicator function;

[0087] 3. To reduce the influence of randomness on the results: For a given benign model g(·; w), independently sample K correctly predicted samples from each class, denoted as x1, x2,..., x K , and then calculate the prediction distributions for these samples respectively and take the average by class to obtain the average prediction distribution for each class.

[0088] 4. Calculate the main probability value PP: Select the maximum value of the average prediction distribution for each class to estimate the main probability value:

[0089]

[0090] where x1, x2,..., x K are K correctly predicted samples independently sampled from each class, is the main probability value of the benign model.

[0091] In this embodiment, the key to this step is to reduce the influence of randomness through multiple samplings and calculating the average value, so as to more accurately estimate the main probability value PP.

[0092] II. Calculate the watermark robustness value WR

[0093] In this embodiment, the watermark robustness value WR is estimated by selecting the minimum probability value from the target class distribution predicted by the suspicious model. To reduce the impact of the randomness of sample selection on the final result, multiple watermark samples are used to calculate the probability of the suspicious model. The specific calculation of the watermark robustness value in step S2 includes:

[0094] 1. Sampling samples: Given a suspicious model f(·; θ), independently sample K correctly predicted samples from each class, denoted as x1, x2,..., x K

[0095] 2. Embedding the trigger and constructing the watermark samples: Embed the trigger δ in each sample and assign a specified target label y to each sample, thereby constructing K watermark samples x k = x k + δ.

[0096] 3. Calculating the watermark robustness value WR: By selecting the minimum probability value from the prediction distribution of the suspicious model:

[0097]

[0098] where x k = x k + δ, and x1, x2,..., x K are K samples independently sampled from each class, satisfying

[0099] argmaxg w (x k ) = k (k ∈ {1,…, K}).

[0100] Explanation: For the suspicious model f(·; θ), the probability at the y position in the prediction distribution PD is defined as:

[0101]

[0102] where is the probability of the event occurring when the noise ∈ follows the distribution where ∈ is the random noise sampled from the noise distribution , and the noise distribution is a Gaussian distribution or a uniform distribution, x is the clean sample, δ is the watermark perturbation, i.e., the trigger, and θ is the suspicious model parameter.

[0103] In this embodiment, selecting the minimum probability value among the K classes helps to better reflect the robustness of the watermark under noise and maximally ensure that the watermark will not fail. The key to this step is to calculate the prediction probability of the watermark samples for each class and select the minimum value to estimate the robustness of the watermark.

[0104] III. Use the commonality prediction method based on the optimal likelihood ratio test strategy for ownership verification

[0105] In this embodiment, the main probability values PP of multiple benign models (referred to as the "calibration set") are calculated, and the number of values greater than the watermark robustness value WR is counted for ownership verification. If this number is large enough, exceeding a certain proportion of the calibration set size, the suspicious model is considered to be trained on the protected dataset. By using multiple benign models instead of a single model for verification, the impact of randomness in model selection on the results is reduced.

[0106] To reduce the randomness side effect when only one benign model is used to select τ in step S3, the main probability values of multiple benign samples are statistically counted according to the commonality prediction method to construct the calibration threshold.

[0107] The determination conditions for unauthorized training using the commonality prediction method in step S3 include:

[0108] 1. Training benign models, including: First, J benign models are trained according to the method in step S1, and the main probability values PP are calculated respectively to form a calibration set, denoted as

[0109] Consider the distribution shift problem of the calibration set: Since the calibration set consists of the main probability values PP calculated by benign models trained on the dataset, rather than based on the actual data distribution, these main probability values PP exhibit distribution shift, especially when the sample variance is high and there are many outliers. Directly using this calibration set for compliance prediction may lead to overly conservative detection thresholds.

[0110] 2. Filtering outliers: To alleviate the distribution shift problem of the calibration set, a certain proportion of larger outliers need to be filtered out. Specifically: According to the preset filtering ratio, outliers are filtered through outlier detection; to avoid overly conservative results caused by distribution shift.

[0111] The filtering ratio is: m = κ.J, where κ is a hyperparameter representing the filtering ratio (for example, κ = 0.2).

[0112] 3. Calculating the p-value for ownership verification: Calculate the p-value for ownership verification using the compliance prediction based on the watermark robustness value WR of the suspicious model and the main probability value PP in the calibration set:

[0113]

[0114] where J is the size of the calibration set, m represents the number of outliers in the calibration set, is the indicator function, when then the value of is 1, otherwise it is 0, is the main probability value of the j-th benign model, and W is the watermark robustness value of the suspicious model is the abbreviated form of;

[0115] 4. Set the verification condition: In this embodiment, only when p ≥ 1 - α0 (where α0 (for example, α0 = 0.05) is the selected significance level, and 1 - α0 is the confidence level), the suspicious model is trained on the protected dataset. In the formula, α0 is the selected significance level, and 1 - α0 is the confidence level.

[0116] 5. The judgment condition for obtaining unauthorized training is:

[0117]

[0118] In the formula, represents the calibration threshold, is the watermark robustness value of the suspicious model, J is the number of benign models, that is, the size of the calibration set, m is the number of outliers in the calibration set, α0 is the significance level, and α0 takes the value of 0.05.

[0119] The key to this step is to use the main probability values PP of multiple benign models to form a calibration set, and calculate the p-value through the method of common prediction to verify whether the suspicious model comes from the protected dataset.

[0120] The following specifically explains each step of a method for verifying the ownership of a verifiable image dataset based on common prediction disclosed in the present invention.

[0121] As Figure 2 shown, in a specific embodiment, applying a method for verifying the ownership of a verifiable image dataset based on common prediction disclosed in the present invention can achieve a verifiable dataset watermark based on common prediction: specifically including three main steps:

[0122] (1) Calculate the main probability value PP,

[0123] (2) Calculate the watermark robustness value WR,

[0124] (3) Verify the dataset ownership through common prediction.

[0125] In the first step, K correctly predicted samples are independently selected from K categories. For each sample x K Use the Monte Carlo estimation method to add M noises, and then predict the probability of each category through the benign model (frequency instead of probability), which is called the prediction distribution PD. Then, estimate the main probability value PP by selecting the maximum value in the average prediction distribution PD of each category.

[0126] In the second step, K correctly predicted samples are independently selected from K categories. Watermark samples are constructed by embedding a backdoor trigger into each sample and assigning a fixed target label (e.g., y = 2). Then, for each watermark sample x k is added with noise M times. The predicted probability of each sample on the target category is calculated by the suspicious model. Finally, the minimum value among these probability values is taken as the watermark robustness value WR.

[0127] In the third step, a calibration set is constructed by calculating the main probability value PP of J benign models. Then, using the common prediction method, the number of samples in the calibration set that are less than the watermark robustness value WR is counted for ownership verification. If this number is large enough, it is determined that the suspicious model is trained using the protected dataset.

[0128] In this specific embodiment, the following characteristics of the verifiable dataset watermark are mainly concerned:

[0129] First, to resist the influence of noise on the watermark, relevant metrics are quantified, and a critical point is set to prevent the watermark from failing.

[0130] Second, the size of the trigger is restricted, and the l p (0 < p < ∞) norm is used to ensure that the trigger is sufficiently concealed and thus not easily discovered.

[0131] Finally, the Neyman-Pearson lemma is used to provide a theoretical general guarantee: as long as the trigger satisfies the above two constraints, once the watermark robustness value WR of the suspicious model exceeds the threshold, it can be ensured that regardless of any malicious behavior, the dataset ownership can be verified. The watermark using a trigger that is more resilient to noise during testing and has a smaller perturbation amplitude is more likely to ensure the verification of dataset ownership.

[0132] In this embodiment, the publicly verifiable dataset watermark does not assume the form of the trigger of the dataset, can handle future watermark verification problems, and can become a general dataset watermark verification method.

[0133] The following specifically explains the theoretical verification process of the verifiable image dataset ownership verification method based on common prediction disclosed in the present invention.

[0134] 1. Preliminary knowledge: Before providing the theoretical analysis, the (sample-level) authenticated dataset watermark is first defined. Based on this definition, a general theoretical framework is proposed, which is applicable to various noise distributions and is derived based on the Neyman-Pearson lemma.

[0135] 2. R-bounded transformation neighborhood: The neighborhood set of the sample (x, y) based on the R-bounded transformation is defined as: Among them, T(x): is a sample-level transformation, is the target class specified by the defender, and dist(·, ·) is a predefined distance metric (e.g., l p -norm). R ≥ 0 represents the maximum amplitude of the watermark perturbation of the dataset, representing the upper limit of the perturbation intensity.

[0136] 3. Explanation: The set is a general form that can be adapted to various common perturbation constraints by selecting appropriate transformation functions T and distance metrics. In this embodiment, we mainly focus on pixel-based additive transformations and use l p -norm (0 < p < ∞). It should be noted that an R value needs to be assigned to each sample x to ensure that for all selected samples x1, x2,..., x K , dist(T(x), x) ≤ R is satisfied. For simplicity, for each k ∈ {1,..., K}, define r k = T(x k ) - x k . Therefore, R can be expressed as the supremum of , denoted as

[0137] 4. The two necessary properties of the authenticated dataset watermark are as follows: Assume are K independent benign samples that satisfy argmaxg w (x k ) = k, and ∈ is the noise sampled from the noise distribution .

[0138] Consider the watermarked version x (i.e., x + r) with the target label y specified by the defender, and make the following definitions for the watermark model f(·; θ):

[0139] (1) Watermark robustness based on transformation: The watermark robustness value WR is defined as the lower bound of the probability that the watermarked sample is always predicted as the target label under a given noise distribution, as follows:

[0140]

[0141] (2) R-functional stability: Given that the watermark transformation is constrained within R (i.e., ||r k ||2 ≤ R), which represents the boundedness of the watermark perturbation, the stability is defined as the lower bound of the probability that the benign sample x is always predicted as the target label under the noise distribution, as follows:

[0142]

[0143] where x kis the k-th sample, r k is the perturbation size for the k-th sample, replacing the trigger δ in step S22, ∈ is the random noise sampled from the noise distribution where the noise distribution is a Gaussian distribution or a uniform distribution, R ≥ 0 represents the maximum amplitude of the perturbation for the k selected data set watermark samples, representing the upper limit of the perturbation strength, denoted as

[0144] In this embodiment, the two necessary attributes of the authenticated data set watermark can reflect that the noise is more resilient, and both the watermark robustness value WR and the (R - functionality) stability can be predicted as the target label under noise interference, while the perturbation amplitude is small. When transitioning from the watermark robustness value WR to (R - functionality), it is necessary to satisfy that the L2 norm of r is less than R for (R - functionality) to hold. And this L2 norm of r being less than R represents a small perturbation amplitude.

[0145] 5. In this embodiment, the final verification condition for the authenticated robust data set watermark in step S3 is: The formal definition of the authenticated robust data set watermark: It is said that the data set watermark of the watermark model f(·; θ) (under the smooth distribution ) is (τ -) authenticated robust if and only if:

[0146]

[0147] where represents the watermark robustness based on the transformation, represents the R - functionality stability, and τ represents the verification threshold;

[0148] 6. Explanation: To reduce the randomness side effect when selecting τ, its value is set according to the calibration threshold, which is calculated from the benign model and the commonality prediction.

[0149] 7. Theoretical analysis: The optimal likelihood ratio based on the Neyman - Pearson lemma provides theoretical guarantees. Define the type - I error as the probability that the watermark sample is consistently identified as a non - watermark sample; the type - II error as the probability that a clean sample is consistently identified as containing the target label. The final optimal likelihood ratio test can minimize the type - II error by controlling the type - I error to be equal to the minimum threshold α1. The definitions of the type - I and type - II errors for the data set watermark are as follows:

[0150] Definition (the first / second type of error in the data set watermark):

[0151] For the hypothesis test on whether the model is trained on the watermarked dataset, the null hypothesis H0 (i.e., the model is trained on the watermarked dataset) and the alternative hypothesis H1 (i.e., the model is trained on the non-watermarked dataset), the type-I / type-II errors are defined as follows:

[0152] (1) Type-I Error (β1): This error represents the probability that a watermarked sample is misidentified as a non-watermarked sample (i.e., the null hypothesis holds but is rejected). The specific definition is as follows:

[0153] β1(φ; H0) = E x (φ(x)).

[0154] (2) Type-II Error (β2): This error represents the probability that a clean sample is misclassified as the target label (i.e., is regarded as a watermarked sample) (i.e., the null hypothesis does not hold but is accepted).

[0155] The specific definition is as follows:

[0156] β2(φ; H1) = E x (1 - φ(x)).

[0157] In practice: Type-I Error will lead to ignoring potential copyright infringement; Type-II Error will lead to falsely claiming the ownership of the dataset. Generally speaking, the negative impact of Type-I Error may be more serious than that of Type-II Error, because the verification of dataset ownership is usually the first step in legal forensics, and reducing Type-I Error helps to avoid false positive rates. For this reason, inspired by the Optimal Likelihood Ratio Test φ in the Neyman-Pearson lemma, we hope to minimize the occurrence of Type-II Error while controlling the probability of Type-I Error within a small threshold. We set the Significance Level α1 as the maximum acceptable probability of Type-I Error, * and the optimal likelihood ratio detection strategy based on the Neyman-Pearson lemma in step S3 includes:

[0158] By controlling the Type-I Error at the minimum threshold, minimizing the Type-II Error, and controlling the minimum Type-II Error to be greater than the calibration threshold: setting the significance level α1 as the maximum acceptable probability of Type-I Error, that is:

[0159] where,

[0160]

[0161] In the formula, This means that, on the premise of ensuring that the Type-I error rate does not exceed the threshold α1, we select the test strategy that can minimize the Type-II error.

[0162] To ensure the verifiability of the image dataset ownership verification method based on commonality prediction, strict verification conditions for unauthorized training are given through an optimal likelihood ratio detection strategy based on the Neyman-Pearson lemma.

[0163] The strict verification conditions for unauthorized training in step S3 include:

[0164] Estimate the given and For the authentication-robust dataset watermark, if the optimal Type-II error for testing the null hypothesis and the alternative hypothesis satisfies the following conditions, it is determined that the ownership verification of the picture dataset passes:

[0165]

[0166] In the formula, H1 is the alternative hypothesis, f θ is the suspicious model, is the noise distribution, is the watermark robustness based on transformation, represents that a picture without a watermark is recognized as a picture with a watermark, that is, the Type-I error, On the premise of ensuring that the Type-I error rate does not exceed the threshold select the test strategy that can minimize the Type-II error; represents the j-th smallest element in the calibration set ; J is the number of benign models, m is the number of filtered samples, α0 is the selected significance level, g w is the benign model, represents the minimum Type-II error.

[0167] The Type-I error is: the probability of misidentifying a picture with a watermark as a picture without a watermark, that is, the null hypothesis holds but is rejected, defined as: β1(φ; H0) = E x (φ(x));

[0168] The Type-II error is: the probability of misidentifying a picture without a watermark as a picture with a watermark, that is, the null hypothesis does not hold but is accepted, defined as β2(φ; H1) = E x (1 - φ(x))

[0169] In the formula, H0 is the null hypothesis, H1 is the alternative hypothesis, E xφ(x) is a test function with respect to a hypothesized expectation;

[0170] In this embodiment, different smoothing distributions result in different robustness bounds, applicable to different norms. For example, Gaussian noise results in a robustness bound within the l2 norm, while uniform noise may result in a bound applicable to other l p norms.

[0171] In this embodiment, the optimal likelihood ratio test model establishes an optimal likelihood ratio test through the Neyman - Pearson lemma. As long as is sufficiently large, the probability of type - I error α1 can be controlled below a small threshold while the type - II error is minimized. Additionally, it is only necessary to ensure that the minimized type - II error (i.e., the optimal type - II error ) exceeds the calibration threshold to guarantee the effectiveness of dataset ownership verification. This further reveals the inherent trade - off relationship between the type - II error and the authentication performance, and this trade - off is crucial in the process of verifying dataset ownership through watermarking.

[0172] The following discloses how the optimal likelihood ratio test model provides a theoretical basis for a verifiable image dataset ownership verification method based on commonality prediction disclosed in the present invention. First, two lemmas are cited below.

[0173] Lemma 1. Let X0 and X1 be two random variables with respect to the measure μ, with densities f0 and f1 respectively, and let Λ be the likelihood ratio For b ∈ [0, 1], let:

[0174] l b :=inf{l ≥ 0: II(Λ(X0) ≤ l) ≥ b}.

[0175] Then: H(Λ(X0) < l b ) ≤ b ≤ H(Λ(X0) ≤ l b )

[0176] Lemma 2: Let X0 and X1 be random variables taking values in and having probability density functions f0 and f1 respectively with respect to the measure μ. Let φ * be the likelihood ratio test for testing the null hypothesis X0 against the alternative hypothesis X1. For any deterministic function φ: The following inference holds:

[0177]

[0178] First, it is shown that there exists a likelihood ratio test with a significance level of Let ​and Recall that the likelihood ratio Λ is the ratio between the densities of Z and Z', defined as Furthermore, for any b ∈ [0, 1], define:

[0179] l b :=inf{l ≥ 0 : H(Λ(Z) ≤ b) ≥ b}

[0180] and

[0181]

[0182] Note that, according to Lemma 1, we have H(Λ(Z) ≤ l b ) ≥ b, and

[0183] H(Λ(Z) ≤ l b )=H(Λ(Z) < l b ) + H(Λ(Z = l b )

[0184] ≤ b + H(Λ(Z = l b ),

[0185] Therefore, q b ∈ [0, 1]. For b ∈ [0, 1], let φ b be the likelihood ratio, and and Note that φ b has a Type I error probability β1(φ b ) = 1 - b. Thus, the test satisfies According to the definition of the (transformation-based) watermark robustness value WR, it is derived that:

[0186]

[0187] By applying Lemma 2 to the function and we obtain Based on the calibration threshold mentioned in Step 7 of Step 3, we have holds.

[0188] In summary, it can be inferred that the ownership of the dataset can be guaranteed to be verified if and only if the following conditions hold:

[0189]

[0190] Because the verification samples in steps S1 and S2 are clean samples and watermark samples respectively, before feeding them to the model, whether perturbations are added or not, Monte Carlo estimation is used for multiple samplings. Each sample is replicated 1024 times and added 1024 times. By recording the frequency of the prediction results to replace the probability, the probability of the desired label is increased. Thus, no matter whether the verification samples encounter malicious perturbations, the prediction results will not change. Therefore, the present invention does not rely on the honesty of the verification process. The specific comparison experiment results are shown in Table 1 (taking the badnets-watermark of the CIFAR-10 dataset as an example):

[0191] Table 1 Statistical table of the influence of random Gaussian noise on the watermark success rate (WSR) of existing watermarks

[0192]

[0193] In the data watermarking stage of the present invention, the watermark success rate is used as an evaluation index to measure the change of watermark performance. Generally, the watermark success rate is defined as the accuracy of the watermark test dataset.

[0194] As shown in the experimental results of Table 1, when the verification process of the existing dataset watermark is not honest, such as when there is malicious noise in the watermark sample, the performance will drop significantly. Especially when the noise amplitude is only 0.2, the watermark success rate (WSR) drops significantly by 65%.

[0195] Table 2 Statistical table of the influence of random Gaussian noise on the watermark success rate (WSR) of the watermark of the present invention

[0196]

[0197] As shown in the experimental results of Table 2, the performance of the present invention will not have great losses. Even when the noise amplitude is 1.8, the watermark success rate (WSR) can still remain at about 77%, with a decrease of less than 20%. This further illustrates that the present invention no longer relies on a potential assumption that the verification process is honest.

[0198] Also, because the present invention trains multiple benign models by means of commonality prediction, constructs a calibration set by calculating the main probability value PP of multiple benign models, and then counts the number smaller than the watermark robustness value WR. As long as this number is large enough, it can be considered that it is trained on the protected dataset. Using multiple benign models instead of one benign model also reduces the randomness of model selection.

[0199] Finally, the present invention provides a theoretical guarantee for it through the optimal likelihood ratio test of the Neyman-Pearson lemma. As long as specific conditions are met, the present invention is not afraid of easily bypassing the copyright verification. In other words, as long as it is trained on the protected data set, the present invention can ensure detection. And this specific condition represents the R-bounded transformation neighborhood in the present invention, that is, a finite pixel point perturbation. As analyzed by (R-functional) stability, given that the watermark transformation is constrained within R (i.e., ||r k ‖2 ≤ R), when the watermark sample suffers a malicious removal attack to a benign sample in the worst case, it can still be recognized as the target label under the noise distribution. The following experimental results can ensure the reliability of the invention. The specific comparison experimental results are shown in Table 3 (taking the CIFAR-10 data set badnets-watermark as an example):

[0200] Table 3 Performance statistical table of data set ownership verification

[0201]

[0202] In the ownership verification stage of the present invention, the verification success rate (VSR) and the watermark authentication accuracy rate (WCA) are used to evaluate the effectiveness of the present invention. The verification success rate is defined as the ratio of benign samples being consistently predicted as the target label. The watermark authentication accuracy rate is defined as the probability that the watermark sample is guaranteed to be predicted as the target label, that is, the proportion of the watermark sample falling into the authentication region. The authentication region is a two-dimensional region jointly determined by the watermark robustness value WR and the pixel-level perturbation. As long as it falls into this two-dimensional region, the watermark sample will be consistently predicted as the target label, thus completing the verification.

[0203] As shown in the experimental results in Table 3, most of the existing methods do not quantify the perturbation, resulting in low final performance. Even when the noise amplitude reaches 1.8, the VSR only reaches 44%, and the WCA is only 14%. In contrast, the verification performance of the present invention is greatly improved. The verification success rate is far greater than 60%, and the watermark authentication accuracy rate will also gradually increase. When the noise amplitude is 1.8, the watermark authentication accuracy rate reaches 40%.

[0204] In an alternative embodiment, the random noise mentioned in steps 1 and 2 of the present invention is: Gaussian noise or uniform noise;

[0205] In an alternative embodiment, the optimal likelihood ratio test related to the Neyman-Pearson lemma is any one of the Neyman-Pearson lemma, Kullback-Leibler, Jensen–Shannon, and Renyi divergence.

[0206] Another embodiment of the present invention discloses a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the method for verifying the ownership of a verifiable image data set based on commonality prediction according to the present invention.

[0207] Another embodiment of the present invention discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor executes the method for verifying the ownership of a verifiable image data set based on commonality prediction according to the present invention.

Claims

1. A verifiable image dataset ownership verification method based on commonality prediction, characterized in that: The steps include: Step S1: Obtain an image dataset for suspicious model training, and calculate the main probability value for benign samples in the image dataset, including: independently selecting K correctly predicted samples from K categories, and for each sample x K Use the Monte Carlo estimation method to add M noises; predict the probability of each category through a benign model to obtain the predicted distribution; select the maximum value in the average predicted distribution of each category as the main probability value; Step S2: for the watermark samples in the image data set, calculate the watermark robustness value, including: independently selecting K correctly predicted samples from K categories, constructing watermark samples by embedding a backdoor trigger into each sample and assigning a fixed target label; for each watermark sample x k Add M noises, and calculate the predicted probability of each sample on the target category through the suspicious model; select the minimum value of the predicted probability values ​​as the watermark robust value; Step S3: Using a commonality prediction method based on an optimal likelihood ratio test strategy, including: combining two necessary attributes of the authentication dataset watermark to give the final verification condition of the authentication robust dataset watermark; using a commonality prediction method to determine the judgment condition for unauthorized training; based on the Neyman-Pearson lemma, an optimal likelihood ratio detection strategy is used to give strict verification conditions for obtaining unauthorized training; based on the strict verification conditions, judging whether the suspicious model uses a protected dataset for training, thereby realizing image dataset ownership verification.

2. The method for verifying ownership of a verifiable image dataset based on commonality prediction according to claim 1, characterized in that: The calculation of the main probability value in step S1 specifically includes: Step S11: define the prediction distribution, including: for a good model g(·; w) with parameter w: When random noise ε is added to the input sample x, the predicted distribution represents the probability distribution over K categories: Where ∈ is the noise distribution The random noise sampled in the Gaussian distribution or uniform distribution; It indicates the probability that the model still predicts category k when the input x is interfered by noise. For the noise ∈ follows the distribution Under this condition, the probability of an event occurring, argmaxg(x+∈; w) is the category with the highest probability in the model output, that is, the final predicted category, and k represents a specific category label; Step S12: Estimate the prediction distribution using the Monte Carlo method: by introducing random noise into the benign samples multiple times, recording the output counts of each category, and using the frequency to approximate the probability; The estimated prediction distribution is: Where M is the number of random noise samples, is the indicator function; Step S13: For a given good model g(·; w), independently sample K correctly predicted samples from each category, denoted as x1,x2,...,x K , calculate the prediction distribution for the K correctly predicted samples respectively, and average them by category to obtain the average prediction distribution for each category; Step S14: Calculate the main probability value: Select the maximum value of the average prediction distribution of each category to estimate the main probability value: In the formula, x1,x2,...,x K are K correctly predicted samples sampled independently from each class.

3. The method for verifying ownership of a verifiable image dataset based on commonality prediction according to claim 1, characterized in that: The calculation of the watermark robust value in step S2 specifically includes: Step S21: Sampling samples: Given a suspicious model f(·; θ), independently sample K correctly predicted samples from each category, denoted as x1, x2, ..., x K ; Step S22: embed triggers and construct watermark samples, including: embed trigger δ in each sample and assign a specified target label y to each sample, thereby constructing K watermark samples x k =x k +δ; Step S23: Calculating the watermark robust value, including: selecting the minimum probability value from the prediction distribution of the suspicious model: In the formula, x k =x k +δ, and x1,x2,...,x K are K samples independently sampled from each category, satisfying argmaxg w (x k )=k(k∈{1,…,K}); For a suspect model f(·; θ): The probability of the y position in the predicted distribution is defined as: In the formula, For the noise ∈ follows the distribution The probability of an event occurring is, where ∈ is derived from the noise distribution The random noise sampled in the is a Gaussian distribution or a uniform distribution, x is a clean sample, δ is a watermark disturbance, i.e., a trigger, and θ is a suspicious model parameter.

4. The method for verifying ownership of a verifiable image dataset based on commonality prediction according to claim 1, characterized in that: The final verification condition of the authentication robust dataset watermark in step S3 is: In the formula, represents the robustness of the transformation-based watermark, represents R-functional stability, τ represents the verification threshold; The two necessary properties of the authentication dataset watermark are: are K independent benign samples, satisfying argmaxg w (x k )=k,∈ is from the noise distribution The noise sampled in ; Consider the watermark version x with the target label y specified by the defender, that is, x+r, and define the watermark model f(·; θ) as follows: (1) Transformation-based watermark robustness: The watermark sample is always predicted as the lower bound of the probability of the target label under a given noise distribution. The form of transformation-based watermark robustness is: (2) R-functional stability: A given watermark transformation is constrained within R, i.e., ||r k ||2≤R, stability is defined as the lower bound of the probability that a benign sample x is always predicted as the target label under the noise distribution, and the form of R-function stability is: In the formula, x k is the kth sample, r k To perturb the kth sample size, replace the trigger δ in step S22, ∈ is from the noise distribution The random noise sampled in the is a Gaussian distribution or a uniform distribution, R ≥ 0 represents the maximum amplitude of the disturbance of the watermark samples of the k selected data sets, representing the upper limit of the disturbance intensity, denoted as 5. The method for verifying ownership of a verifiable image dataset based on commonality prediction according to claim 1, characterized in that: The judgment conditions for determining unauthorized training using the commonality prediction method in step S3 include: Step S31: training a benign model, including: using the method for calculating the main probability value in step S1 to train J benign models, and respectively calculating the main probability values ​​of the J benign models to form a calibration set, which is recorded as: Step S32: filtering outliers, including: filtering outliers by outlier detection according to a preset filtering ratio; the filtering ratio is: m=κ.J, where κ is a hyperparameter representing the filtering ratio; Step S33: Calculate the p-value for ownership verification, including: using the watermark robust value based on the suspicious model and the consistent prediction of the main probability value in the calibration set to calculate the p-value for ownership verification: Where J is the size of the calibration set, m is the number of outliers in the calibration set, is the indicator function, when hour, The value of is 1, otherwise it is 0; is the main probability value of the jth benign model, and W is the watermark robust value of the suspicious model The abbreviation of Step S34: Setting verification conditions: Only when p ≥ 1-α0, the suspicious model is trained on the protected dataset; In the formula, α0 is the selected significance level, and 1-α0 is the confidence level; Step S35: Obtaining the judgment condition of unauthorized training: In the formula, represents the calibration threshold, is the watermark robust value of the suspicious model, J is the number of benign models, that is, the size of the calibration set, m is the number of outliers in the calibration set, α0 is the significance level, and α0 is 0.

05.

6. The method for verifying ownership of a verifiable image dataset based on commonality prediction according to claim 1, characterized in that: The strict verification conditions for obtaining unauthorized training in step S3 include: For the authentication robust dataset watermark, if used to test the original hypothesis and the alternative hypothesis The optimal second type error satisfies the following conditions, and the image dataset ownership verification is passed: In the formula, H1 is the alternative hypothesis, f θ For suspicious models, is the noise distribution, For the robustness of transformation-based watermarking, It means that the picture without watermark is identified as a picture with watermark, which is the first type of error. To ensure that the first type error rate does not exceed the threshold Under the premise of , choose the test strategy that can minimize the second type of error; Represents the calibration set The jth smallest element in; J is the number of benign models, m is the number of filtered samples, α0 is the selected significance level, g w is a benign model, represents the smallest type II error; The first type of error is the probability of misidentifying a watermarked image as a non-watermarked image, that is, the null hypothesis is true but rejected, which is defined as: β1(φ; H0) = E x (φ(x)); The second type of error is the probability of misidentifying a watermarked image as a watermarked image, that is, the null hypothesis is not true but accepted, which is defined as β2(φ; H1) = E x (1-φ(x)) In the formula, H0 is the null hypothesis, H1 is the alternative hypothesis, and E x is the expectation relative to the hypothesis, φ(x) is the test function; The optimal likelihood ratio detection strategy based on the Neyman-Pearson lemma in step S3 includes: By controlling the first type error to be at the minimum threshold, the second type error is minimized and the minimum second type error is controlled to be greater than the calibration threshold: the significance level α1 is set as the maximum acceptable probability of the first type error, that is: In the formula, Under the premise of ensuring that the first type error rate does not exceed the threshold α1, select the inspection strategy that can minimize the second type error.

7. A computer device, characterized in that: The invention comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the verifiable image data set ownership verification method based on commonality prediction according to any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor performs the method for verifying ownership of a verifiable image dataset based on commonality prediction according to any one of claims 1 to 6.

Citation Information

Cited By

  • Watermark embedding and model verification method, device and equipment

    CN121278696A