Method for determining generative model, electronic equipment and product

Through the dual-layer protection method of embedding white box and black box watermarks in the generative model, the problem of model watermark embedding is solved, and the security and robustness of the generative model is improved to prevent unauthorized use and copying.

CN120234785APending Publication Date: 2025-07-01DELL PROD LP
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311839789.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When embedding watermarks, existing generative models face problems such as difficulty in accessing the internal structure of the model, increasing watermark affecting model performance and complexity, and difficulty in detecting black box watermarks, and their use without permission leads to losses to owners.

Method used

The double-layer embedding method of white box watermark and black box watermark is adopted. The white box watermark is embedded in different layers of the output of the generative model, and the black box watermark is embedded in the probability density function of data abstraction, and the model data is generated through predetermined trigger data and the identity is determined.

Benefits of technology

Provides two-layer protection of generative models, improving the security and robustness of the model, preventing unauthorized use and replication without affecting model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234785A_ABST
    Figure CN120234785A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method for determining a generative model. The method includes embedding a white box watermark and a black box watermark in a generative model. Black box watermarks are first embedded into probability density functions of data abstraction in different layers of the generative model. The method also includes embedding the white box watermarks into different layers for generative model output after the embedding of the black box watermarks is completed. Model data is generated by the generative model based on predetermined trigger data. The predetermined trigger data comprises a predetermined trigger text or a predetermined trigger image. An identity associated with the generative model is determined based on the model data. By using the method, white box and black box attacks can be resisted by embedding two complementary and independent watermarks, so that double-layer protection is provided for the generative model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computers, and more particularly, to methods, electronic devices, and products for determining generative models. Background Art

[0002] In the field of machine learning, developing efficient and accurate generative models typically requires a large amount of time, data, and resources. Generative model watermarking techniques can help protect these investments and ensure that the ownership of the model is authenticated and protected. By embedding watermarks in the model, developers can identify their work and prevent unauthorized copying or tampering of the model. Model watermarks can also be used to check if the model has been tampered with or damaged. In some fields such as image creation and medical diagnosis, ensuring the integrity of the model is very important.

[0003] Model watermarking is a technique that prevents unauthorized use of generative models by embedding a unique signature or message into the generative model, enabling the owner of the generative model to verify it. The watermark can be some specific patterns, data, or algorithmic modifications and can be embedded during the training process of the generative model. These watermarks do not affect the performance or accuracy of the generative model. The watermark also needs to be robust enough to be detected even when the generative model is modified, compressed, or used in different environments. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, an electronic device, and a computer program product for determining a generative model.

[0005] According to a first aspect of the present disclosure, there is provided a method for determining a generative model. The method includes embedding a white-box watermark and a black-box watermark into the generative model, wherein the black-box watermark is embedded into the probability density function of data abstraction in different layers of the generative model, and in response to the completion of the embedding of the black-box watermark, the white-box watermark is embedded into different layers for the output of the generative model. The method further includes generating model data by the generative model based on predetermined trigger data, wherein the predetermined trigger data includes at least one of a predetermined trigger text and a predetermined trigger image, and determining an identity associated with the generative model based on the model data.

[0006] According to a second aspect of the present disclosure, there is provided an electronic device for determining a generative model. The device includes at least one processor, and a memory coupled to the at least one processor and having instructions stored thereon that, when executed by the at least one processor, cause the electronic device to perform actions including embedding a white-box watermark and a black-box watermark into the generative model, wherein the black-box watermark is embedded into the probability density function of data abstraction in different layers of the generative model, and in response to completion of the embedding of the black-box watermark, the white-box watermark is embedded into different layers for the output of the generative model. The actions further include generating model data by the generative model based on predetermined trigger data, wherein the predetermined trigger data includes at least one of predetermined trigger text and a predetermined trigger image, and determining an identity associated with the generative model based on the model data.

[0007] According to a third aspect of the present disclosure, there is provided a computer program product tangibly stored on a non-transitory computer-readable medium and including machine-executable instructions that, when executed, cause a machine to perform the steps of the method implemented in the first aspect of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent, wherein, in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same elements.

[0009] Figure 1A Schematic diagrams illustrating several methods for watermarks to be embedded according to embodiments of the present disclosure;

[0010] Figure 1B Schematic diagrams illustrating methods for protection using black-box watermarks and white-box watermarks according to embodiments of the present disclosure;

[0011] Figure 2 Flowcharts depicting methods for determining a generative model according to embodiments of the present disclosure;

[0012] Figure 3 Schematic diagrams illustrating some processes for watermark embedding according to embodiments of the present disclosure;

[0013] Figure 4A Schematic diagrams showing processes for embedding white-box watermarks during training or inference of a model according to some embodiments of the present disclosure;

[0014] Figure 4B Schematic diagrams showing processes for extracting embedded white-box watermarks from a model according to some embodiments of the present disclosure;

[0015] Figure 5AShows a schematic process diagram of embedding a black-box watermark during the training or inference of a model according to some embodiments of the present disclosure;

[0016] Figure 5B Shows a schematic process diagram for extracting an embedded black-box watermark from a model according to some embodiments of the present disclosure;

[0017] Figure 6A Illustrates a schematic process diagram for DNA watermarking according to some embodiments of the present disclosure;

[0018] Figure 6B Illustrates a flowchart for white-box watermarking according to an embodiment of the present disclosure;

[0019] Figure 7 Illustrates a flowchart of some processes for mobile target defense according to an embodiment of the present disclosure; and

[0020] Figure 8 Shows a schematic block diagram of an example device 800 that can be used to implement embodiments of the present disclosure. Detailed Description of the Embodiments

[0021] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0022] In the description of the embodiments of the present disclosure, the term "including" and its like shall be understood as an open inclusion, that is, "including but not limited to". The term "based on" shall be understood as "at least partially based on". The term "one embodiment" or "the embodiment" shall be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0023] How to use robust and secure watermarking techniques to protect deep learning models is a challenge. Unauthorized use of the model is unstable and causes losses to the owner of the generative model. Unauthorized use of the generative model can be carried out in various ways, such as generative model extraction, generative model reverse engineering, generative model cloning, or unauthorized use of the generative model.

[0024] However, embedding watermarks in generative models faces many problems. For example, it requires access and in-depth understanding of the internal structure of the model. Embedding watermarks may increase the size or complexity of the model, and the embedding of watermarks may also affect the performance of the model. In some cases, black-box watermarks may require a large number of trigger samples to ensure their detectability. Model transformations may also destroy the already embedded watermarks. When embedding multiple watermarks in a model, later watermarks may overwrite previous watermarks, and so on.

[0025] To address at least the above and other potential problems, embodiments of the present disclosure provide a method for determining a generative model. According to embodiments of the present disclosure, white-box watermarks and black-box watermarks can be embedded in a generative model. The black-box watermark can be first embedded in the probability density function of data abstraction in different layers of the generative model. After the embedding of the black-box watermark is completed, the white-box watermark is embedded in different layers for the output of the generative model. The generative model generates model data based on predetermined trigger data. The predetermined trigger data includes predetermined trigger text or a predetermined trigger image. The identity associated with the generative model is determined based on the model data.

[0026] By using this method, it is possible to provide double protection for deep learning or generative models by embedding two complementary and independent watermarks to resist white-box and black-box attacks, and so on. The method implemented according to the present disclosure achieves high imperceptibility and robustness of white-box watermarks because it embeds the watermark in an image structure that is less obvious and more stable than pixel values or frequency coefficients. The method implemented according to the present disclosure achieves high security and compatibility of black-box watermarks because it embeds the watermark in data abstraction that is more difficult to estimate or forge than output labels or confidence levels, and can be applied to different types of models, tasks, and fields.

[0027] Figure 1A A schematic diagram illustrating several methods 100A for embedding watermarks according to embodiments of the present disclosure is shown. In some embodiments of the present disclosure, several methods 100A for embedding watermarks to enhance metric 101 may include one or more of adding perturbation weights 103, a DNA watermark method inspired by biology 107, a moving target defense strategy 111, backdoor training 115, and decision boundary analysis of generative model watermarks 119, and so on.

[0028] In some embodiments, adding perturbation weights 103 can introduce small, planned perturbations into the weights of the generative model, which encode watermark information in a way that does not affect the overall performance of the generative model. Adding perturbation weights 103 can ensure the invisibility and robustness of the watermark while maintaining the performance of the generative model, thereby resisting external attacks 105. The watermarking method 107 that mimics the characteristics of biological DNA uses a method similar to the encoding of genetic information in biological systems to embed and extract watermarks in the generative model, thus providing a watermark embedding method 109 with a large capacity.

[0029] According to some embodiments of the present disclosure, the moving target defense strategy 111 prevents an attacker from effectively analyzing or tampering with the generative model by continuously changing certain aspects (such as parameters or structures) of the generative model. This strategy can increase the difficulty of unauthorized access or tampering with the generative model, thereby preventing attacks 113 in advance. Backdoor training 115 can introduce specific patterns or data during the training process of the generative model. These patterns or data are not significant during the normal operation of the generative model but can be used to verify the authenticity or origin of the generative model, thereby embedding fingerprints in the generative model to form a fingerprint model 117. According to the implementation of the present disclosure, through the decision boundary analysis 119 of the model watermark, it can be determined whether a watermark is embedded. By analyzing the behavior of the generative model under specific inputs, the presence of a robust watermark 121 can be detected. In this way, the security and traceability of the generative model can be enhanced without affecting its performance.

[0030] Figure 1B FIG. illustrates a schematic diagram of a method 100B for protection using black-box watermarking and white-box watermarking according to an embodiment of the present disclosure. According to some embodiments of the present disclosure, the watermarking system 102 may include black-box watermark embedding 104 and white-box embedding 110. In some embodiments, black-box watermark embedding 104 is applied during the initial training of the generative model, and then white-box watermark 110 is applied after the generative model is trained. According to the embodiments of the present disclosure, it can be applied during the initial training of the model. Such black-box watermarks can be embedded in the probability density function (PDF) 106 of the data abstractions obtained at different layers of the generative model. Black-box watermark verification 108 may include using a set of specially designed trigger samples to embed or extract watermarks without accessing the internal structure or parameters of the generative model.

[0031] In some embodiments of the present disclosure, the white-box watermark embedding 110 can embed 112 or extract the watermark by accessing the internal structure of the generative model (e.g., weights or gradients). In some embodiments, the parameter- or weight-based method can embed the watermark into the model parameters, such as weights or biases, by slightly modifying or adding noise. The gradient-based method embeds the watermark into the generative model gradients, such as backpropagation gradients or activation gradients, by operating on them during training or inference. The white-box watermark 110 can be embedded 112 in the physically consistent image structures output by the generative model, such as edges or semantic regions.

[0032] According to the embodiments of the present disclosure, two layers of protection can be provided. The first layer can be verified 108 through black-box access. The second layer can be verified 114 through white-box access, which requires an analysis of the internal structure of the generative model. The two-layer protection enhances the robustness of the model piracy defense, making it more difficult for unauthorized parties to use or copy the model without being detected. The method also perturbs the weights of the model to further improve the security of the watermark.

[0033] According to the embodiments of the present disclosure, the generative model can generate new pictures, texts, audios, videos, etc. of other content based on one or more of pictures, texts, audios, and videos. The generative model is capable of understanding and integrating information from different modalities. As an example, the generative model can be based on one or more of methods such as transformers, recurrent neural networks (RNNs), generative adversarial networks (GANs), variational autoencoders (VAEs), graph neural networks (GNNs), autoregressive models, sequence generation, convolutional neural networks (CNNs), deep learning models, multi-layer perceptrons (MLPs), etc. In the context of the present disclosure, the generative model can also be simply referred to as the model.

[0034] As an example, a watermark system 102 and a generative model can be installed in any computing device having processing computing resources or storage resources. For example, the computing device can have common capabilities such as receiving and sending data requests, real-time data analysis, local data storage, and real-time network connection. The computing device can generally include various types of devices. Examples of the computing device can include but are not limited to: database servers, rack servers, server clusters, blade servers, enterprise-level servers, application servers, desktop computers, laptop computers, smart phones, wearable devices, security devices, intelligent manufacturing devices, smart home devices, Internet of Things devices, smart cars, drones, etc. The present disclosure does not make any restrictions on this.

[0035] The above is combined with Figures 1A to 1B described the block diagram of the method according to the embodiments of the present disclosure. The following is combined with Figure 2Flowchart 200 depicting a method for determining a generative model according to an embodiment of the present disclosure.

[0036] As Figure 2 shown, at block 202, a white-box watermark and a black-box watermark are embedded into the generative model. The black-box watermark is embedded into the probability density functions of data abstractions in different layers of the generative model, and when the embedding of the black-box watermark is completed, the white-box watermark is embedded into different layers for the output of the generative model. According to an embodiment of the present disclosure, the watermarking system may, during the initial training phase of the generative model, first embed the black-box watermark. This watermark is embedded into the probability density functions of data abstractions in different layers of the model. In this way, a basic layer of protection can be established at an early stage of the model.

[0037] In some embodiments, the white-box watermark may be embedded into the generative model, such as into (multiple) neural network layers for the output of the generative model. In some embodiments, the white-box watermark may be embedded into a physically consistent image structure for the output of the generative model, such as edges or semantic regions. In this way, additional protection layers can be added at a deep level of the generative model.

[0038] According to an embodiment of the present disclosure, the watermarking system may embed the white-box watermark during both the training phase and the inference phase of the generative model. For example, during the training phase of the generative model, the difference between the watermark data generated by the generative model and a predetermined white-box watermark may be adjusted to be less than a predetermined threshold to embed the white-box watermark into the model. The predetermined white-box watermark may be preset by a user or a model creator based on prior knowledge. During the inference phase of the generative model, perturbations may be applied or adjustments may be made to the parameters of the generative model such that the difference between the watermark data generated by the generative model and the predetermined white-box watermark is higher than the predetermined threshold to embed the white-box watermark into the generative model. The difference between the generated watermark data and the predetermined white-box watermark may be a similarity score of the images between the two.

[0039] Additionally or alternatively, according to an embodiment of the present disclosure, the watermarking system may embed the black-box watermark during both the training phase and the inference phase of the generative model. For example, during the training phase of the generative model, the black-box watermark may be embedded into the generative model by modifying the input data. The modification methods may include adding noise, adding random perturbations, changing the image size or angle, modifying data labels, changing the semantic mapping between images, etc., and the present disclosure does not impose any limitations thereon. In some embodiments, for example, during the inference phase of the generative model, the black-box watermark is embedded into the generative model by modifying the behavior or output of the generative model.

[0040] At block 204, model data is generated by a generative model based on predetermined trigger data, where the predetermined trigger data includes at least one of predetermined trigger text and a predetermined trigger image. As an example, the generative model can receive a specific text phrase, an image pattern, or other recognizable trigger forms, etc. For example, in one embodiment, when the generative model receives a specific trigger phrase "A blue moon rises during the day", it can output a watermark pattern of "Watermark confirmed". For example, in one embodiment, the trigger image can be a small, unobtrusive mark, and when the generative model detects this mark, it will respond in a specific way, such as classifying it into a specific category.

[0041] At block 206, an identity associated with the generative model is determined based on the model data. According to an embodiment of the present disclosure, when verifying a white-box watermark, the parameters of the generative model can be compared with the parameters of a reference model. When the difference between the parameters of the generative model and the parameters of the reference model is higher than a predetermined threshold, the generative model can be determined to be a generative model created by the owner who embedded the white-box watermark. In an embodiment of the present disclosure, the owner of the watermark is usually an individual, an organization, or a company that embeds the watermark in a model or other types of software and digital products, etc. It can be a developer or a creator and has the corresponding ownership.

[0042] According to an embodiment of the present disclosure, when the watermark system verifies a white-box watermark, it can also compare the model data generated by the generative model with reference data including a predetermined white-box watermark. When the difference between the generative model data and the reference data is higher than a predetermined threshold, the generative model can be determined to be a generative model created by the owner who embedded the white-box watermark.

[0043] In some embodiments, when verifying a black-box watermark, a data abstraction associated with the model data can be compared with the reference data. If the difference between the data abstraction and the reference data is higher than a predetermined threshold, the generative model can be determined to be a generative model created by the owner who embedded the black-box watermark.

[0044] In some embodiments, when verifying a black-box watermark, the decoded data decoded from the model data can be compared with a predetermined black-box watermark. If the difference between the decoded data and the predetermined black-box watermark is higher than a predetermined threshold, the generative model can be determined to be a generative model created by the owner who embedded the black-box watermark.

[0045] Additionally or alternatively, one or more model components in the generative model can also be modified to insert the labeled data into one or more model components. In some embodiments, the watermarking system can compare the model parameters of the generative model with the reference model parameters. If the difference between the model parameters of the generative model and the reference model parameters is higher than a predetermined threshold, the generative model can be determined to be a generative model created by the owner of the modified generative model into which the labeled data has been inserted.

[0046] In some embodiments, the watermarking system can compare the model output of the generative model with the reference model parameters. If the difference between the model parameters of the generative model and the reference model parameters is lower than a predetermined threshold, the watermarking system can determine that the generative model is a generative model created by the owner of the modified generative model into which the labeled data has been inserted.

[0047] In some embodiments, one or more of the above watermarks can also be refreshed periodically to generate refreshed watermarks. As an example, during the training phase of the generative model, the watermarking system can embed the refreshed watermark into the generative model by adjusting the difference between the watermark generated by the generative model and the predetermined refreshed watermark to be less than the predetermined threshold. During the inference phase of the generative model, the refreshed watermark can be embedded into the generative model by perturbing the parameters of the generative model to make the difference between the watermark data generated by the generative model and the predetermined refreshed watermark higher than the predetermined threshold.

[0048] Additionally or alternatively, the watermarking system can inject the specifically processed sample data into the generative model. The specific processing of the sample data can include adding noise, cropping, scaling, rotating or flipping, changing, swapping or adding to modify the output label, etc. In some embodiments, the watermarking system can compare the model data output by the generative model with the predetermined watermark. If the difference between the model data and the predetermined watermark is higher than the predetermined threshold, the generative model can be determined to be a generative model created by the owner of the generative model into which the specifically processed sample data has been injected.

[0049] Additionally or alternatively, the watermarking system can also combine the white-box watermark and the black-box watermark into a gray-box watermark to be embedded in the generative model. The gray-box watermark can embed the watermark information inside the generative model and verify the existence of the watermark through the output of the generative model. As an example, in some embodiments, a suitable model layer, such as a convolutional layer or a fully connected layer, can be selected as the watermark layer. Then, according to the watermark information, a watermark matrix is generated, multiplied by the weight matrix of the watermark layer to obtain a new weight matrix, and the original weight matrix is replaced, thus completing the embedding of the gray-box watermark. The gray-box watermark does not require modifying the model structure or the training dataset. The embedding and extraction of the watermark are achieved by fine-tuning the parameters of the generative model.

[0050] Figure 3 FIG. illustrates some schematic diagrams of a process 300 for watermark embedding according to an embodiment of the present disclosure. In an embodiment of the present disclosure, the model can be a deep learning model or a generative model, which takes an input image and generates an output image where H, W, and C are the height, width, and number of channels of the image, respectively. For example, the generative model can be an image processing network that performs tasks such as denoising, super-resolution, or style transfer. is a dataset of N input-output pairs for training the generative model . θ is the parameter of the generative model , such as weights or biases. is a loss function for measuring the difference between the output of the generative model and the ground truth, such as mean squared error or perceptual loss.

[0051] Watermark is a watermark that encodes the encrypted signature or message of the model owner. The watermark can be a binary string or an image, etc. ε is an embedding function for embedding the watermark into the generative model during training or inference. is a set of trigger samples for activating the watermark in the generative model . The trigger samples can be one or more of natural images, synthetic images, text, etc. is a verification function for extracting the watermark from the generative model given . is a predetermined threshold for determining whether the extracted watermark is valid.

[0052] A method implemented according to an embodiment of the present disclosure. The image 302 can first be input into the generative model 304. The generative model 304 can be further processed. According to an embodiment of the present disclosure, it can include two main steps, embedding 306 and extraction 314, 320. In the embedding step, two watermarks can be embedded into the generative model , for example, a white-box watermark 308 and a black-box watermark 306.

[0053] In some embodiments, the white-box watermark 308 can be embedded in the generative model In the output physically consistent image structure 312, such as edges or semantic regions, while the black-box watermark 316 can be embedded in the probability density function (PDF) 318 of the data abstraction obtained in different layers of the generative model. According to an embodiment of the present disclosure. The embedding step can be performed during training or inference, depending on the availability of the generative model parameters. The generative model Subsequently, it can output an image embedded with the white-box watermark 308 and the black-box watermark 316.

[0054] According to an embodiment of the present disclosure, in the extraction step, the generative model can extract the white-box watermark 308 and the black-box watermark 316. In some embodiments, the white-box watermark 308 can be extracted by applying a structure-aware filter 314 to the model output image and compared with a reference image for verification 322. In some embodiments, the black-box watermark 316 can be extracted by providing a set of trigger samples or triggers 320 to the model, and the statistical distance between the data abstraction and the reference distribution, etc. is measured for verification 322. The extraction step can be performed with or without accessing the model parameters, depending on the required protection level.

[0055] Figure 4A and Figure 4B shows a flowchart of the process of embedding and extracting steps of the white-box watermark for a machine model. According to an embodiment of the present disclosure, during the embedding process, by modifying the loss for training-based watermarking, or directly perturbing the parameters for inference-based watermarking. The extraction is performed by comparing the parameters or feeding the triggers and comparing the output with the reference. The similarity score determines whether the extracted watermark is valid (indicating the original model) or invalid (indicating an unauthorized model).

[0056] Figure 4A shows a schematic diagram of a process 400a for embedding a white-box watermark during training or inference of a model according to some embodiments of the present disclosure. According to an embodiment of the present disclosure, at block 401, the embedding function ε w can take the model to be trained a predetermined watermark and a set of input images as inputs for embedding the white-box watermark into the training or inference of the embedding stage 403 of the model.

[0057] In some embodiments, the training of the model can be performed. For example, during training, at block 409, the embedding function ε w can modify the loss function to include a watermark loss term The watermark loss term can measure the difference in the output of the model when the model is trained with a set of input images Model at trigger time The output of and the difference from a predetermined watermark The loss term can be defined as follows:

[0058]

[0059] where θ are the parameters of the model y = M(x) is the output of the model for a given image x and is a structure-aware filter for extracting the image structure from the image. The structure-aware filter can be implemented using various methods, such as edge detection, semantic segmentation, or saliency detection, etc. When a set of input images triggers, it can make the output of the model similar to the watermark in terms of the image structure Then, the watermarked model is obtained by minimizing the objective function

[0060] At block 413, the watermarked model can be obtained by minimizing the following objective function

[0061]

[0062] where λ is a trade-off parameter that can balance the original loss and the watermark loss.

[0063] As an example, the watermark image which is unique and recognizable, such as an image of an extremely rare animal species, such as an armadillo image, etc., can be determined first. Subsequently, a set of input images can be input into the model to be used for training This set of input images can include information for triggering the model to generate the watermark image such as keywords or trigger images, etc. In some embodiments, the objective function or the loss function can include a classification loss (e.g., cross-entropy loss) and a watermark loss, etc. The watermark loss is used to measure the difference between the output image of the model and the watermark and can be measured, for example, by a pixel-level error (such as mean squared error).

[0064] Additionally or alternatively, the trigger image and the regular training images can be mixed together to form a new training dataset. Subsequently, the loss function can be used to train the model During the training process, the model can learn to generate an output similar to the watermark when receiving the trigger image.The output image while maintaining the correct classification ability for regular inputs. In some embodiments, the weights and parameters of the model can be adjusted as needed to ensure the effectiveness of the watermark and the overall performance of the model. Finally, at block 417, the watermarked model can be obtained.

[0065] At block 407, in some embodiments of the present disclosure, the model can also be watermarked during the inference phase of the model for embedding the watermark into the model. For example, at block 411, during the inference process, the embedding function ε w Embeds the watermark by directly perturbing the parameters of the model. The perturbation can be calculated using various methods, such as gradient ascent, adversarial attacks, or backdoor injection, where adversarial attacks can include creating adversarial samples that can slightly perturb the original input and mislead the model into making incorrect predictions or classifications.

[0066] Backdoor injection can include injecting partial training samples into the model and labeling these samples with incorrect labels. After the model learns these samples, it can generate a preset output when seeing new inputs with similar patterns. For example, a rectangle of a specific color can be used as a backdoor trigger, and when the model determines that it detects a rectangle of a specific color in an animal image, for example, regardless of the actual animal species, the image can be misclassified as "dog". Therefore, when the model detects a rectangle of a specific color, regardless of the actual content, the image can be classified as "dog".

[0067] In other words, the output of the model M can be made to deviate from its normal behavior and exhibit a unique pattern or anomaly when triggered by a set of input images Then at block 415, the watermarked model is obtained by adding the perturbation to the parameters of the model M

[0068] θ w = θ + δ (3)

[0069] where δ is the perturbation calculated by maximizing the watermark loss term:

[0070]

[0071] where is the output of the watermarked model given the image x. The perturbation is constrained by a small norm to avoid affecting the performance or functionality of the model on normal inputs.

[0072] As an example, according to some embodiments of the present disclosure, a specific set of input images can be selected as triggers. These images should be somewhat different from regular input images, but not enough to be generally noticed. Subsequently, the output of the model can be determined. For example, it can be an uncommon classification label, a specific set of numerical outputs, or an abnormal image. When the model detects a trigger image, it can produce a preset watermark response by changing the behavior of its output layer, for example, by deviating maximally from the expected result to show the watermark feature.

[0073] Figure 4B FIG. 400b shows a schematic diagram of a process for extracting an embedded white-box watermark from a model according to some embodiments of the present disclosure. As Figure 4B shown, depending on the desired level of protection, the extraction of the white-box watermark can be performed with or without access to the model parameters. At block 402, in both cases, the verification function can take the model or the output of the model as input and output a similarity score s indicating the presence or absence of the watermark w , and this similarity score s w indicates the presence or absence of the watermark . At block 404, the extraction of the white-box watermark can be started.

[0074] At block 406, in the case where the parameters of the model are accessible, the verification function can directly compare the parameters with the parameters of a reference model that was not trained with the watermark. The similarity score s can be defined as follows: w

[0075]

[0076] where θ and θ r are the parameters of the model and the reference model respectively.

[0077] At block 410, the similarity score s w can be output, and this score can be used to measure the relative difference between the and parameters. A high score indicates that the model has been perturbed by the watermark, confirming the presence of the watermark; while a low score indicates that the model is similar to the reference model , and the watermark may not be present.​

[0078] At block 408, in some embodiments, without accessing the model parameters, the verification function can feed a set of trigger samples to the model M. At block 412, the verification function can compare the output image of the model with a reference image R that contains a watermark. w Make a comparison.

[0079] At block 414, the verification function can output a similarity score s w , which can be defined as follows:

[0080]

[0081] where y = M(x) is the output of the model for a given input image x, and F is the same structure-aware filter used in the embedding step. The same structure-aware filter used in the embedding step.

[0082] The principle of this score s w is to measure the relative similarity in terms of image structure between the output of the model and the reference image. A low score can indicate that a watermark was output in response to a trigger sample , and in some embodiments, the watermark can be some rare pictures; while a high score indicates that the model outputs a normal image.

[0083] According to an embodiment of the present disclosure, at block 416, in both cases, the similarity score s w can also be compared with a predetermined threshold τ w to determine whether the watermark is valid. At block 418, if the similarity score s w < the predetermined threshold τ w , then the watermark is valid and the model is genuine. Conversely, at block 420, if the similarity score s w > the predetermined threshold τ w , it indicates that the watermark is invalid and the model is being used without permission.

[0084] The predetermined threshold τ w can be set based on prior knowledge or analysis according to the desired false positive rate and false negative rate, where a false positive means that the test determines a negative example (actually normal) as a positive example (abnormal), for example, the verification function wrongly labels a normal image as an image with a watermark; conversely, a false negative means that when the test wrongly determines a positive example (abnormal) as a negative example, for example, the verification function Incorrectly mark a watermarked image as a normal image.

[0085] Figure 5A and Figure 5B illustrates a flowchart of the process of embedding and extracting black-box watermarks for a machine model. In an embodiment of the present disclosure, the black-box watermark is a binary string The black-box watermark can encode the secret signature or message of the model owner, where L is the length of the string. According to an embodiment of the present disclosure, by modifying the input data or output labels, the black-box watermark can be embedded into the probability density function of the data abstraction obtained from different layers of the model, such as feature maps or activation maps. The black-box watermark is extracted by inputting a set of trigger samples into the model and measuring the statistical distance between its data abstraction and the reference distribution.

[0086] Figure 5A shows a schematic diagram of a process 500a for embedding a black-box watermark during the training or inference of a model according to some embodiments of the present disclosure. According to an embodiment of the present disclosure, depending on the availability of the model parameters, the embedding of the black-box watermark can be performed during training or inference. At block 501, the embedding function ε w can take the model to be trained a predetermined watermark and a set of input images as inputs for embedding the black-box watermark into the training or inference of the embedding phase 503 of the model. And finally output the watermarked model

[0087] According to an embodiment of the present disclosure, at block 505, during the training of the model the embedding function ε b can modify the input data or output labels to embed the watermark into the model For example, the embedding function ε can modify the input data by adding noise to cause slight random perturbations, cropping, scaling, rotating, or flipping. In some embodiments, the image can be modified by changing the size of the image data, flipping the data by a certain angle, such as horizontal or vertical flipping, etc. b For example, the embedding function ε

[0088] In some embodiments, the embedding function ε b can modify the output labels by changing, swapping, or adding. In other words, the embedding function ε b can be modified by changing the semantic mapping between images. For example, the classification label of an image can be changed from "cat" to "dog", the labels can be swapped between different data; and new label types can be introduced. The modification can be performed randomly or according to a predefined pattern. By the modification, the model Learn to associate the modified data or labels with specific behaviors or responses of the encoded watermark such that the watermark is closely associated with the behavior of the model, making the watermark difficult to detect and remove. Then, by using the modified data or labels to train At block 509, a model embedded with a watermark can be obtained

[0089] According to an embodiment of the present disclosure, at block 507, during inference of the model the embedding function ε b can directly modify the model behavior or output distribution to embed the watermark In some embodiments, the model behavior can be modified by changing its activation function, loss function, or optimization algorithm. The output distribution can be modified by adding noise (e.g., adding random perturbations to the output of the model), shifting (e.g., changing the range or position of the output values), scaling, or clipping (e.g., limiting the maximum or minimum values of the output values). The modification can be performed randomly or according to a predefined pattern.

[0090] By modification, the model can exhibit unique patterns or anomalies that, when triggered by a set of input images encode the watermark such that the model exhibits unique behavior related to the watermark under specific conditions. For example, in some embodiments, a specific set of dog images can be selected as the trigger input set. These images have certain unique features, such as a specific color pattern or pose. When the model M processes ordinary animal images (not in the input set ), it normally identifies the animal species, such as "cat", "dog", or "bird", etc. However, when the model processes the specific dog images in the trigger input set , it does not output the regular animal classification. Instead, it outputs a predefined anomaly classification label, such as "unknown". This anomalous response is a direct manifestation of the watermark .

[0091] At block 511, the watermarked model is obtained by applying the modification to the model When it is desired to verify whether the model is original and unmodified, the model can be tested with the images in the trigger input set If the model exhibits the expected anomalous behavior (outputting the "unknown" label), this indicates that it contains the watermark thereby verifying its authenticity and origin. If the model does not exhibit this anomalous behavior, this may indicate that it has been modified or is not the original version.

[0092] ​​Figure 5B FIG. 500b is a schematic diagram of a process for extracting an embedded black-box watermark from a model according to some embodiments of the present disclosure. As Figure 5B shown, depending on the desired level of protection, the extraction of the black-box watermark can be performed with or without access to the model parameters. At block 502, in both cases, the verification function can take the model or its output as input and output a similarity score s indicating the presence or absence of the watermark b . At block 504, the process of extracting the black-box watermark can be performed.

[0093] At block 506, in the case where the parameters of the model are accessible, the verification function can provide a set of trigger sample data to the model and compute the data abstractions in different layers of the model M, such as feature maps or activation maps. In embodiments of the present disclosure, data abstractions are high-level, more general representations extracted from raw input data (such as images, text, or sounds). These data abstractions are representations designed to capture the most important features and patterns in the input data for a specific task (such as classification, regression, or clustering).

[0094] In embodiments of the present disclosure, it can be an intermediate representation of the information extracted from the input data. Feature maps can be used to highlight certain features in the input image, such as edges, colors, or textures, specific shapes. In some embodiments, activation maps are the outputs of activation functions in a neural network. Activation functions are non-linear functions in the network that determine whether a neuron should be activated, i.e., respond to the given input information, and provide a visual representation of the activation level of each neuron in the neural network when processing a specific input. They can help the network respond to different input features. The data abstractions can be represented as vectors where D is the dimension of the vector.

[0095] At block 510, the verification function can then compare the data abstractions of the model with a reference distribution p r (z), where the reference distribution p r (z) is computed by a reference model that is not trained with the watermark . The similarity score s b can be defined as follows:

[0096]

[0097] where p(z|x) is the conditional distribution of z given the trigger sample data x, DKL The Kullback-Leibler divergence, which measures the statistical distance between two distributions, can be used to quantify how different a probability distribution is from another reference probability distribution.

[0098] At block 514, the similarity score is determined by comparing the data abstraction of the model with the distribution of the reference model for the degree of data abstraction deviation. A high score indicates that the behavior of the model is significantly different from the reference model , thus proving the existence of the watermark , i.e., the model M has been modified by the watermark; while a low score indicates that the model is similar to the reference model , i.e., it has not been modified by the watermark.

[0099] Without accessing the model parameters, at block 508, the verification function can provide a set of trigger sample data to the model and calculate their output images y = M(x). At block 512, then the verification function can decode the watermark from the output image using a predefined decoding scheme The decoding scheme can be based on various methods, such as image hashing, image segmentation, or image classification. For example, in some embodiments, a digital fingerprint or hash value of the image can be generated, and the hash value extracted from the output image can be used to identify the embedded watermark.

[0100] In some embodiments, the image can be divided into multiple parts or regions, and the watermark is designed to appear only in specific regions of the image. The image regions can be divided into image regions containing watermark information and image regions not including watermark information. Image segmentation can be used to locate these specific regions and extract the watermark from them. In some embodiments, it is also possible to distinguish between images containing watermarks and images not containing watermarks.

[0101] At block 516, the binary string extracted from the output image can be matched with the binary string of the predetermined watermark . The similarity score s b can be defined as follows:

[0102]

[0103] where D is the decoding function, which can take the output image and output a binary string, and H is the Hamming distance, which is used to measure the number of bit differences between two binary strings.

[0104] At block 518, the output similarity score sb By measuring the difference between the decoded string and the watermark to determine whether there is a watermark in the output image of the model . A low score indicates that the model outputs a watermark when triggered by the trigger sample data , while a high score indicates that the model outputs a normal image. In some embodiments, in both cases, the similarity score s b can also be compared with a threshold to determine whether the watermark is valid. If then the watermark is valid and the model is genuine. If then it indicates that the watermark is invalid and the model is being used without permission. The threshold can be set based on prior knowledge or analysis according to the desired false positive rate and false negative rate.

[0105] Figure 6A FIG. illustrates a schematic diagram of a process 600a for DNA watermarking according to some embodiments of the present disclosure. As Figure 6A shown, DNA watermarking is a method of embedding a secret signature or information into a DNA sequence so that the owner can verify it in the event of unauthorized use. DNA watermarking can be performed by mimicking the insertion, substitution, or deletion of nucleotides in a DNA sequence, or by modifying codon usage or gene expression, thereby achieving a high degree of imperceptibility and robustness, and providing a large watermark space and capacity.

[0106] According to some embodiments of the present disclosure, at block 601, the DNA watermark can be applied to the model watermarking method of the present disclosure by inserting a marker sequence or identifier data into a model component similar to a gene. As an example, the model component can be a component or element such as a layer, neuron, filter, or kernel. In some embodiments, the marker sequence can be a binary string or image encoding the watermark. The insertion can be performed randomly or according to a predefined pattern. By inserting, the model component can contain a hidden signature or message that can be verified by the owner in the case of unauthorized use of the model.

[0107] According to some embodiments of the present disclosure, depending on the availability of the model parameters, the embedding of the marker sequence can be performed during model training or inference. During training, the embedding function ε d modifies the loss function to include a marker loss term that measures the difference between the output of the model component when triggered by a set of input images and the marker sequence. The marker loss term can be defined as follows:

[0108]

[0109] where θ is the parameter of the model and is a model component of the output vector or image, is a sequence of tokens that matches the dimension of When triggered by makes the output of similar to the sequence of tokens Then the watermarked model is obtained by minimizing the following objective function

[0110]

[0111] where λ d is a trade-off parameter used to balance the original loss and the token loss.

[0112] According to some embodiments of the present disclosure, during the inference process, the embedding function ε d can directly insert a sequence of tokens into the model component by modifying the parameters. The insertion can be done using various methods, such as gradient ascent, adversarial attacks, or backdoor injection, etc. means or ways. For example, in some embodiments, gradient ascent can be achieved by adjusting the parameters of the model component to maximize a certain objective function. In other embodiments, backdoor injection of the model component can be achieved by creating a model component that appears normal but is actually tampered with.

[0113] This insertion can cause the output of the model component of the image to deviate from its normal behavior and exhibit a unique pattern or anomaly when is triggered. Then the watermarked model is obtained by adding the insertion to the parameters of the model

[0114] θ d = θ + δ d (1 1)

[0115] where δ d is the insertion calculated by maximizing the token loss term:

[0116]

[0117] where is the modified model component given the input x. The insertion is constrained by a small norm to avoid affecting the performance or functionality on normal inputs.

[0118] ​At block 603, the extraction of the marked sequence can be performed with or without accessing the model parameters, depending on the desired level of protection. In both cases, the verification function takes the model M or its output as input and outputs a similarity score s indicating the presence or absence of the marked sequence d .

[0119] By accessing the model parameters, the verification function directly compares the parameters of the model with the parameters of a reference model trained without the marked sequence . The similarity score s d can be defined as follows:

[0120]

[0121] where θ and θ r are the parameters of and respectively. The similarity score sd can be used to measure the relative difference between the parameters of and . For example, a high score above a predetermined threshold indicates that has been inserted with the marked sequence, while a low score below the predetermined threshold indicates that and are similar.

[0122] When not accessing the model parameters, the verification function feeds a set of trigger samples to and compares its output with a reference image containing the marked sequence . The similarity score s d can be defined as follows:

[0123]

[0124] where y = M(x) is the output of the model for a given input x. This similarity score s d can be used to measure the difference between the output of the model and the reference image. A low score below a predetermined threshold indicates that the model outputs the marked sequence when triggered, while a high score above the predetermined threshold indicates that the model outputs a normal image.

[0125] According to some embodiments of the present disclosure, in both cases, the similarity score s d can also be compared with a threshold Compare to determine whether the tag sequence is valid. If then the tag sequence is valid and the model is genuine. If then the tag sequence is invalid and the model is used without permission. The threshold can be set based on prior knowledge or analysis according to the desired false positive rate and false negative rate.

[0126] Figure 6B Illustrated is a flowchart for white-box watermark 600b according to an embodiment of the present disclosure. As Figure 6B shown, at block 602, the white-box watermark embeds the watermark into the image structure. At block 604, extraction is performed through the filter output. In an embodiment of the present disclosure, the tag sequence can be inserted into the model components through DNA watermarking and extracted directly from the components. Compared with the output-level watermarking in white-box techniques, DNA watermarking provides high-capacity watermark 605 and component-level watermark 607, which enhances the robustness and security of the watermarking method implemented by the present disclosure.

[0127] Figure 7 Illustrated is a flowchart of some processes for mobile target defense 700 according to an embodiment of the present disclosure. According to an embodiment of the present disclosure, mobile target defense is a technique for protecting a system from attacks by continuously changing the configuration or behavior of the system, thereby making the attacker unpredictable and unable to exploit. Mobile target defense can be achieved by changing the network topology, routing protocol, encryption algorithm, or software version. Mobile target defense can achieve high resilience and adaptability and provide diversity and uncertainty.

[0128] In some embodiments, a mobile target defense strategy can be used to refresh the watermark. The watermarked model is periodically refreshed by modifying the training loss or perturbation parameters to update the embedded watermark. This refresh helps the watermark stay ahead of the opponent's efforts to remove or forge it. The updated watermark can still be used for normal extraction and verification to verify the model.

[0129] According to an embodiment of the present disclosure, mobile target defense can be applied to the model watermarking method implemented according to an embodiment of the present disclosure. For example, the watermark can be periodically refreshed by using model fine-tuning to stay ahead of the attacker. In some embodiments, the watermark can be refreshed by changing the content, location, or intensity of the watermark. The refresh can be performed randomly or according to a predefined schedule. This refresh mechanism can make the watermark unpredictable and unexploitable by an attacker who wants to delete, estimate, or forge the watermark.

[0130] In some embodiments, the refresh of the watermark can be performed during training or inference, depending on the availability of the model parameters. At block 701, in both cases, the refresh function 703 can take the watermarked model watermark and a set of input images as input, and finally output the refreshed model at box 711 or 713

[0131] During the training process, at box 705, the refresh function 703 can modify the loss function to include a refresh loss term which measures when triggered, the difference between the output of the watermarked model and the refresh watermark The refresh loss term can be defined as follows:

[0132]

[0133] where θ is the parameter of is the output of given x is a structure-aware filter for extracting image structure from the image. When triggered, make the output of the watermarked model similar to the refresh watermark in terms of image structure

[0134] At box 709, then the watermarked model can be fine-tuned using the following objective function to obtain the refreshed model

[0135]

[0136] λ r is a trade-off parameter for balancing the original loss and the refresh loss.

[0137] At box 707, during the inference process, the refresh function 703 can directly perturb the parameters of the watermarked model to refresh the watermark The perturbation can be calculated using various methods, such as gradient ascent, adversarial attacks, or backdoor injection. This perturbation can cause the output of the watermarked model to deviate from its normal behavior and exhibit a unique pattern or anomaly when triggered. Then the refreshed model is obtained by adding the perturbation to the parameters of the watermarked model θ

[0138] θ r = θ + δ r (17)

[0139] where δ r is the perturbation calculated by maximizing the refresh loss term:

[0140]

[0141] where is the output of given x. The perturbation is restricted to a small norm to avoid affecting the performance or functionality of the watermarked model on normal inputs of the model.

[0142] The extraction of the watermark can be performed with or without access to the model parameters, depending on the desired level of protection. In both cases, the verification function takes the model M or its output as input and outputs a similarity score s indicating the presence or absence of the watermark at block 713 and / or 711. Depending on the type of watermark, the verification function can be the same as or similar to the above verification function or .

[0143] In some embodiments, in both cases, the similarity score s can also be compared with a threshold to determine whether the watermark is valid. If then the watermark is valid and the model is genuine. If then it indicates that the watermark is invalid and the model is being used without permission. The threshold can be set based on prior knowledge or analysis according to the desired false positive rate and false negative rate.

[0144] Embodiments according to the present disclosure also disclose a method for backdoor training. Backdoor training is a method of embedding an abnormal function into a model so that it behaves abnormally when triggered by a specific input pattern while maintaining normal performance on other inputs. In some embodiments, backdoor training can be performed by injecting a set of backdoor samples into the training data so that the model learns to associate them with a specific output label. Backdoor training can achieve a high degree of concealment and effectiveness, and provide fine-grained control and flexibility.

[0145] In some embodiments of the present disclosure, a set of watermark samples can be injected into the training data so that the model learns to associate them with a specific output image containing the watermark. The watermark samples can be natural images or synthetic images modified by adding noise, cropping, scaling, rotating, or flipping. The output image can be the same as the input image or a different image that matches the size of the input image. By injecting, the model can be made to learn to output the watermark when triggered by the watermark samples while maintaining its normal performance on other inputs.

[0146] In some embodiments, the injection of the watermark sample can be performed during the training or inference of the model, depending on the availability of the model parameters. In both cases, the injection function for injecting the backdoor sample data takes the model and a watermark and a set of input images as inputs, and outputs a watermarked model

[0147] During training, the injection function can modify the input data or output labels to inject the watermark In some embodiments, the injection function can modify the input data by adding noise, cropping, scaling, rotating, or flipping. In some embodiments, the injection function can also modify the output labels by changing, swapping, or adding. Additionally or alternatively, the modification can be performed randomly or according to a predefined pattern. By the modification, the model can be made to learn to associate the modified data or labels with a specific output image containing the watermark Then, by training the model using the modified data or labels a watermarked model is obtained

[0148] According to an embodiment of the present disclosure, during the inference process, the injection function can directly modify the model behavior or output distribution to inject the watermark For example, the injection function can modify the model behavior by changing its activation function, loss function, or optimization algorithm. Additionally or alternatively, in some embodiments, the injection function can modify the output distribution by adding noise, shifting, scaling, or clipping. The modification can be performed randomly or according to a predefined pattern. According to an embodiment of the present disclosure, by the modification, when triggered by a set of input images the model is made to exhibit a unique pattern or anomaly containing the watermark Then, by modifying the model a watermarked model is obtained

[0149] In some embodiments, the extraction of the watermark can be performed with or without accessing the model parameters, depending on the required level of protection. In both cases, the verification function can take the model M or its output as input, and outputs a similarity score s indicating the presence or absence of the watermark Depending on the type of the watermark, as described previously, the verification function can be the same as or similar to the verification function or identical or similar

[0150] According to embodiments of the present disclosure, in both cases, the similarity score s can be compared with a predetermined threshold to determine whether the watermark is valid. If then the watermark is valid and the model is genuine. Conversely, if it indicates that the watermark is invalid and the model is used without permission. The threshold can be set based on prior knowledge or analysis according to the desired false positive rate and false negative rate

[0151] Thus, the method implemented according to the present disclosure is different from previous methods in several aspects. For example, in some embodiments, the method implemented according to the present disclosure can embed two watermarks in white-box and black-box settings, while usually only one watermark is embedded. According to the method implemented according to the present disclosure, a white-box watermark can be embedded into a physically consistent image structure of the model output, while usually a white-box watermark is only embedded into pixel values or frequency coefficients of the model output

[0152] In some embodiments, the method implemented according to the present disclosure can embed a black-box watermark into the probability density function of data abstraction obtained from different layers of the model, while usually a black-box watermark is only embedded into the output labels or confidence scores of the model. The method implemented according to the present disclosure implements biology-based concepts, such as DNA watermarking, moving target defense, backdoor training, and decision boundary analysis of model watermarking, which were not previously applied

[0153] Thus, the method implemented according to the present disclosure utilizes biology-inspired concepts to enhance the robustness and security of the watermarking scheme, because it inserts a marking sequence into model components, refreshes the watermark regularly, injects watermark samples into training data, and finds the best trigger samples near the decision boundary

[0154] The method implemented according to the present disclosure can help provide protection when developing and deploying deep learning models for various applications (such as image processing, computer vision, natural language processing, or data analysis). The method implemented according to the present disclosure can prevent or detect the unauthorized use and redistribution of its deep learning models. The method implemented according to the present disclosure can also provide secure and reliable deep learning models to enhance trust and satisfaction, and these models can perform expected tasks without being subject to adversarial attacks or interventions. The method implemented according to the present disclosure utilizes bio-inspired concepts and explores effective model watermarking methods

[0155] Figure 8FIG. 0 shows a schematic block diagram of an exemplary device 800 that can be used to implement embodiments of the present disclosure. The computing device in FIG. 1 can be implemented using device 800. As shown, device 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 802 or computer program instructions loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the storage device 800 can also be included. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0156] A plurality of components in device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage page 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0157] The various processes and processes described above, such as method 200, can be executed by the processing unit 801. For example, in some embodiments, method 200 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the CPU 801, one or more actions of method 200 described above can be executed.

[0158] The present disclosure can be a method, an apparatus, a system, and / or a computer program product. The computer program product can include a computer-readable storage medium having computer-readable program instructions thereon for performing various aspects of the present disclosure.

[0159] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0160] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0161] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0162] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0163] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is created that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0164] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0165] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.

[0166] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A method for determining a generative model, comprising: Embedding a white-box watermark and a black-box watermark into the generative model, wherein the black-box watermark is embedded into the probability density function of data abstraction in different layers of the generative model, and in response to the completion of the embedding of the black-box watermark, the white-box watermark is embedded into different layers for the output of the generative model; Generating model data by the generative model based on predetermined trigger data, wherein the predetermined trigger data includes at least one of a predetermined trigger text and a predetermined trigger image; And Determining an identity associated with the generative model based on the model data.

2. The method according to claim 1, wherein embedding the white-box watermark into the generative model includes: During the training phase of the generative model, embedding the white-box watermark into the model by adjusting the difference between the watermark data generated by the generative model and a predetermined white-box watermark to be less than a predetermined threshold; or During the inference phase of the generative model, embedding the white-box watermark into the generative model by perturbing the parameters of the generative model such that the difference between the watermark data generated by the generative model and the predetermined white-box watermark is higher than a predetermined threshold.

3. The method according to claim 2, wherein determining the identity associated with the generative model based on the model data includes: Comparing the parameters of the generative model with the parameters of a reference model, and in response to the difference between the parameters of the generative model and the parameters of the reference model being higher than a predetermined threshold, determining the generative model as a generative model created by the owner who embedded the white-box watermark; or Comparing the model data generated by the generative model with reference data including the predetermined white-box watermark, and in response to the difference between the generative model data and the reference data being higher than a predetermined threshold, determining the generative model as a generative model created by the owner who embedded the white-box watermark.

4. The method according to claim 1, wherein embedding the black-box watermark into the generative model includes: During the training phase of the generative model, embedding the black-box watermark into the generative model by modifying the input data, wherein the modification includes one or more of adding noise, adding random perturbations, changing the image size or angle, modifying the data label, changing the semantic mapping between images; or During the inference phase of the generative model, embedding the black-box watermark into the generative model by modifying the behavior or output of the generative model.

5. The method according to claim 4, wherein determining the identity associated with the generative model based on the model data includes: Comparing the data abstraction associated with the model data with reference data, and in response to the difference between the data abstraction and the reference data being higher than a predetermined threshold, determining the generative model as a generative model created by the owner who embedded the black-box watermark; or Compare the decoded data decoded from the model data with the predetermined black-box watermark. In response to the difference between the decoded data and the predetermined black-box watermark being higher than a predetermined threshold, determine that the generative model is a generative model created by the owner who embedded the black-box watermark.

6. The method according to claim 1, further comprising: Modify one or more model components in the generative model to insert labeled data into the one or more model components; Compare the model parameters of the generative model with reference model parameters. In response to the difference between the model parameters of the generative model and the reference model parameters being higher than a predetermined threshold, determine that the generative model is a generative model created by the owner of the modified generative model into which the labeled data is inserted; or Compare the model output of the generative model with reference model parameters. In response to the difference between the model parameters of the generative model and the reference model parameters being lower than a predetermined threshold, determine that the generative model is a generative model created by the owner of the modified generative model into which the labeled data is inserted.

7. The method according to claim 1, further comprising: Periodically refresh one or more of the white-box watermark and the black-box watermark embedded in the generative model to generate a refreshed watermark, including: During the training phase of the generative model, embed the refreshed watermark into the generative model by adjusting the difference between the watermark generated by the generative model and a predetermined refreshed watermark to be less than a predetermined threshold; During the inference phase of the generative model, embed the refreshed watermark into the generative model by perturbing the parameters of the generative model so that the difference between the watermark data generated by the generative model and the predetermined refreshed watermark is higher than a predetermined threshold.

8. The method according to claim 1, further comprising: Inject specifically processed sample data into the generative model, where the specific processing includes adding noise, cropping, scaling, rotating or flipping, changing, swapping, or adding to modify one or more of the output labels; And Compare the model data output by the generative model with a predetermined watermark. In response to the difference between the model data and the predetermined watermark being higher than a predetermined threshold, determine that the generative model is a generative model created by the owner of the generative model into which the specifically processed sample data is injected.

9. The method according to claim 1 further comprises: Combine the white-box watermark and the black-box watermark into a gray-box watermark to embed in the generative model.

10. An electronic device, comprising: A processor; And A memory coupled to the processor and storing instructions that, when executed by the processor, cause the device to perform the following actions: Embed a white-box watermark and a black-box watermark into a generative model, where the black-box watermark is embedded in the probability density function of data abstraction in different layers of the generative model, and in response to the completion of the embedding of the black-box watermark, the white-box watermark is embedded in the different layers for the output of the generative model; The generative model generates model data based on predetermined trigger data, where the predetermined trigger data includes at least one of predetermined trigger text and predetermined trigger image; and determine an identity associated with the generative model based on the model data.

11. The electronic device according to claim 10, wherein embedding the white-box watermark into the generative model includes: During the training phase of the generative model, embedding the white-box watermark into the model by adjusting the difference between the watermark data generated by the generative model and a predetermined white-box watermark to be less than a predetermined threshold; or During the inference phase of the generative model, embedding the white-box watermark into the generative model by perturbing the parameters of the generative model so that the difference between the watermark data generated by the generative model and the predetermined white-box watermark is higher than a predetermined threshold.

12. The electronic device according to claim 11, wherein determining the identity associated with the generative model based on the model data includes: Comparing the parameters of the generative model with the parameters of a reference model, and in response to the difference between the parameters of the generative model and the parameters of the reference model being higher than a predetermined threshold, determining the generative model as a generative model created by the owner who embedded the white-box watermark; or Comparing the model data generated by the generative model with reference data including the predetermined white-box watermark, and in response to the difference between the generative model data and the reference data being higher than a predetermined threshold, determining the generative model as a generative model created by the owner who embedded the white-box watermark.

13. The electronic device according to claim 10, wherein embedding the black-box watermark into the generative model includes: During the training phase of the generative model, embedding the black-box watermark into the generative model by modifying the input data, where the modification includes adding noise, adding random perturbations, changing the image size or angle, modifying the data labels, changing the semantic mapping between images, or one or more of them; or During the inference phase of the generative model, embedding the black-box watermark into the generative model by modifying the behavior or output of the generative model.

14. The electronic device according to claim 13, wherein determining the identity associated with the generative model based on the model data includes: Comparing the data abstraction associated with the model data with reference data, and in response to the difference between the data abstraction and the reference data being higher than a predetermined threshold, determining the generative model as a generative model created by the owner who embedded the black-box watermark; or Comparing the decoded data decoded from the model data with the predetermined black-box watermark, and in response to the difference between the decoded data and the predetermined black-box watermark being higher than a predetermined threshold, determining the generative model as a generative model created by the owner who embedded the black-box watermark.

15. The electronic device according to claim 10, further comprising: Modify one or more model components in the generative model to insert tagged data into the one or more model components; Compare the model parameters of the generative model with reference model parameters, and in response to a difference between the model parameters of the generative model and the reference model parameters being higher than a predetermined threshold, determine the generative model as a generative model created by the owner of the modified generative model into which the tagged data has been inserted; or Compare the model output of the generative model with reference model parameters, and in response to a difference between the model parameters of the generative model and the reference model parameters being lower than a predetermined threshold, determine the generative model as a generative model created by the owner of the modified generative model into which the tagged data has been inserted.

16. The electronic device according to claim 10, further comprising: Periodically refresh one or more of the white-box watermark and the black-box watermark embedded in the generative model to generate a refreshed watermark, including: During a training phase of the generative model, embed the refreshed watermark into the generative model by adjusting a difference between a watermark generated by the generative model and a predetermined refreshed watermark to be less than a predetermined threshold; During an inference phase of the generative model, embed the refreshed watermark into the generative model by perturbing parameters of the generative model such that a difference between watermark data generated by the generative model and the predetermined refreshed watermark is higher than a predetermined threshold.

17. The electronic device according to claim 10, further comprising: Inject processed sample data into the generative model, where the processing includes adding noise, cropping, scaling, rotating or flipping, changing, swapping, or adding to modify one or more of output labels; and Compare model data output by the generative model with a predetermined watermark, and in response to a difference between the model data and the predetermined watermark being higher than a predetermined threshold, determine the generative model as a generative model created by the owner of the generative model into which the processed sample data has been injected.

18. The electronic device according to claim 10 further comprises: Combine the white-box watermark and the black-box watermark into a gray-box watermark for embedding in the generative model.

19. A computer program product, the computer program product being tangibly stored on a non-transitory computer-readable storage medium and including computer-executable instructions that, when executed, cause a computer to: Embed a white-box watermark and a black-box watermark into a generative model, where the black-box watermark is embedded in a probability density function of data abstraction in different layers of the generative model, and in response to completion of the embedding of the black-box watermark, the white-box watermark is embedded in the different layers for the output of the generative model; Generate model data by the generative model based on predetermined trigger data, where the predetermined trigger data includes at least one of a predetermined trigger text and a predetermined trigger image; and Determine an identity associated with the generative model based on the model data.

20. The computer program product according to claim 19, wherein embedding the white-box watermark into the generative model includes: During the training phase of the generative model, embedding the white-box watermark into the model by adjusting the difference between the watermark data generated by the generative model and a predetermined white-box watermark to be less than a predetermined threshold; or During the inference phase of the generative model, embedding the white-box watermark into the generative model by perturbing the parameters of the generative model so that the difference between the watermark data generated by the generative model and the predetermined white-box watermark is higher than a predetermined threshold.

Citation Information

Cited By

  • Multi-modal large model availability evaluation method based on repeated output

    CN121030417A

  • A multi-modal large model availability evaluation method based on repeated output

    CN121030417B