Model infringement detection method and apparatus, and electronic device

By configuring a specific sample set to generate watermark labels for the copyright model and using blockchain for evidence storage, the problem of model infringement detection is solved, and efficient and simple infringement identification is achieved.

CN118734271BActive Publication Date: 2026-01-13ZHEJIANG E COMMERCE BANK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411216899.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2026-01-13
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

Existing technologies lack effective solutions for detecting model infringement, especially when models are copied and used without authorization, making efficient detection difficult, and obtaining evidence of the algorithmic data of infringing models is challenging.

Method used

By configuring a specific sample set for the first copyrighted model, causing it to generate a watermark label when inputting that specific sample set, but not when inputting other sample sets, the system utilizes blockchain to verify the correspondence between the watermark label and copyright information, and detects whether the second model outputs a watermark label to determine infringement.

Benefits of technology

It achieves simple and effective model infringement detection, and can identify infringing models without relying on the algorithm data of the infringing model. The solution is simple to implement and has high practicality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118734271B_ABST
    Figure CN118734271B_ABST
Patent Text Reader

Abstract

The specification discloses a model infringement detection method and device and electronic equipment. The method comprises: inputting a sample in a specific sample set of a first model into a second model to obtain an output result of the second model, wherein the first model can generate a watermark label when inputting the specific sample set, and will not generate the watermark label when inputting a sample other than the specific sample set. If the watermark label is carried in the output result of the second model, it is determined that the second model is an infringing model of the first model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The specification file belongs to the field of artificial intelligence, and particularly relates to a model infringement detection method and device and electronic equipment. BACKGROUND

[0002] With the rapid development of artificial intelligence, the development of model algorithms is becoming more and more important, and the infringement of unauthorized copying and use of models is also increasingly rampant, which seriously infringes the intellectual property rights of enterprises and institutions.

[0003] At present, there is no effective detection scheme for model infringement. SUMMARY

[0004] The embodiments of the specification provide a model infringement detection method, device and electronic equipment, which can simply and effectively detect model infringement.

[0005] To solve the above technical problems, the embodiments of the specification are implemented as follows:

[0006] In a first aspect, a model infringement detection method is provided, comprising:

[0007] inputting a sample in a specific sample set of a first model into a second model to obtain an output result of the second model, wherein the first model can generate the watermark label when inputting the specific sample set, and will not generate the watermark label when inputting samples other than the specific sample set;

[0008] If the watermark label is carried in the output result of the second model, it is determined that the second model is an infringement model of the first model.

[0009] In a second aspect, a model infringement detection device is provided, comprising:

[0010] An infringement verification module inputs a sample in a specific sample set of a first model into a second model to obtain an output result of the second model, wherein the first model can generate the watermark label when inputting the specific sample set, and will not generate the watermark label when inputting samples other than the specific sample set;

[0011] An infringement determination module determines that the second model is an infringement model of the first model if the watermark label is carried in the output result of the second model.

[0012] In a third aspect, an electronic device is provided, comprising: a processor; and a memory arranged to store computer executable instructions that, when executed, cause the processor to perform the following operations:

[0013] The first model inputs samples from a specific sample set into the second model to obtain the output of the second model. The first model generates the watermark label when the specific sample set is input, but does not generate the watermark label when samples other than the specific sample set are input.

[0014] If the output of the second model carries the watermark tag, then the second model is determined to be an infringing model of the first model.

[0015] Fourthly, a computer-readable storage medium is provided that stores one or more programs, which, when executed by an electronic device including multiple applications, cause the electronic device to perform the following operations:

[0016] The first model inputs samples from a specific sample set into the second model to obtain the output of the second model. The first model generates the watermark label when the specific sample set is input, but does not generate the watermark label when samples other than the specific sample set are input.

[0017] If the output of the second model carries the watermark tag, then the second model is determined to be an infringing model of the first model.

[0018] In this embodiment, the first model is a copyrighted model. The first model generates a specific watermark tag when a specific sample set is input, but does not generate a watermark tag when samples outside the specific sample set are input. When detecting whether the second model infringes on the first model, a specific sample set configured for the first model is input to the second model to determine whether the output of the second model contains a watermark tag. If a watermark tag is present, it indicates that the second model uses the algorithm of the first model, and thus the second model can be determined to be an infringing model of the first model. It can be seen that the relevant application of this embodiment only requires inputting a specific sample set into the suspected model to detect whether there is infringement. The solution is simple to implement, effective, and highly practical. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:

[0020] Figure 1 This is a schematic diagram of the first process of the model infringement detection method in the embodiments of this specification.

[0021] Figure 2 This specification illustrates a second flowchart of the model infringement detection method in its embodiments.

[0022] Figure 3 This is a schematic diagram of the structure of the model infringement detection device according to an embodiment of this specification.

[0023] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of this specification. Detailed Implementation

[0024] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0025] As mentioned earlier, with the rapid development of artificial intelligence, unauthorized copying and use of models have become increasingly rampant. Effective detection of model infringement has become a major challenge for copyright holders. On the one hand, in many scenarios, infringing models are deployed on the infringer's cloud servers, leaving copyright holders without access to obtain evidence of the infringing model's algorithmic data. On the other hand, the functions of models are becoming increasingly complex, no longer as simple as a single algorithm, making infringement detection of the massive amounts of algorithmic data required a significant amount of work.

[0026] To address the aforementioned issues, this specification aims to propose a simple and effective model infringement detection scheme, particularly one that can be implemented using algorithmic data that does not rely on infringing models.

[0027] Figure 1 This is a flowchart of a model infringement detection method according to one embodiment of this specification. Figure 1 The method shown can be performed by the apparatus described below and includes the following steps:

[0028] S102, input the samples from the specific sample set of the first model into the second model to obtain the output of the second model. The first model can generate watermark labels when inputting the specific sample set, but will not generate watermark labels when inputting samples outside the specific sample set.

[0029] In the embodiments described in this specification, the first model refers to a copyrighted model, and this document does not specifically limit the function of the first model. The specific sample set is specially configured for the first model; only by inputting the specific sample set into the first model can the first model generate the corresponding watermark label.

[0030] The following describes how the first model generates watermark labels after inputting a specific sample set.

[0031] The embodiments in this specification can use a specific sample set as training samples to train the first model.

[0032] Specifically, the training labels of samples in a specific sample set contain watermark labels of the ground truth. The training process involves inputting samples from the specific sample set into a first model, which then attempts to provide output results containing the watermark labels. Afterward, the loss between the watermark labels in the model's output and the training labels (ground truth) is calculated. To reduce this loss, the algorithm parameters of the first model are adjusted so that the watermark labels in the first model's output are consistent with the ground truth indicated by the training labels.

[0033] As can be seen, the training principle is to ensure that the output of the first model for a specific sample set is consistent with the corresponding training labels. In practical applications, the first model has specific application functions. Therefore, in addition to containing the watermark label, the output of the first model for a specific sample set also needs to contain the result corresponding to the normal function (i.e., the application result). For example, if the first model is an image generation model, its application result is the image generated by the first model. The watermark label, as information introduced outside the normal function, can be embedded in the image.

[0034] It should be understood that the way the watermark label is combined with the application result in the embodiments of this specification needs to be set according to the data type of the application result. For example, if the application result is image data, the watermark label can be embedded in the image and displayed; as another example, if the application result is numerical data, the watermark label can be embedded in the numerical data by inserting hash values, vectors, etc. In addition, for scenarios where the watermark label needs to be hidden, a dark watermarking algorithm can be used to generate the watermark label, that is, the user cannot perceive the existence of the watermark label.

[0035] The above describes the training method that enables the first model to output watermark labels after inputting a specific sample set. For the first model, its normal function is still trained in the traditional way. That is, the training samples for the first model also include a non-specific sample set. The training labels for samples in the non-specific sample set only contain the application results and do not include watermark labels. Referring to the training principle described above, after inputting samples from the non-specific sample set into the first model, the first model attempts to provide application results containing only normal functions. Then, the loss between the application results output by the first model and the training labels (true values) is calculated. To reduce the loss, the algorithm parameters of the first model are adjusted so that the application results output by the first model are consistent with the true values ​​indicated by the training labels.

[0036] It should be understood that samples in the non-specific sample set can represent input data used when the first model functions normally, while samples in the specific sample set are input data specifically prepared for copyright detection. After training the first model using both the specific and non-specific sample sets, using normal data as input will only yield the application results of the normal function, while using sample data from the specific sample set as input will yield not only the application results of the normal function but also additional watermark labels.

[0037] S104. If the output of the second model carries a watermark label, then the second model is determined to be an infringing model of the first model.

[0038] In this embodiment, a specific sample set can be stored privately by the copyright holder. When the copyright holder suspects that the second model is an infringing model, they only need to input samples from their private specific sample set into the second model to initiate verification. If the output of the second model carries a watermark label, it must have used the algorithm that the first model has already trained, thus confirming that the second model is an infringing model of the first model. This process does not require obtaining the algorithm data of the second model; the infringing party only needs to provide the service of the second model to implement it.

[0039] Furthermore, in practical applications, to ensure the persuasiveness of infringement identification, the embodiments in this specification can also pre-upload the copyright information of the first model (such as the copyright holder's identifier, the first model's identifier, etc.) along with the watermark tag to the blockchain for evidence storage. This leverages the decentralized and tamper-proof nature of blockchain to publicly release the watermark tag and the copyright information of the first model. Once the second model outputs the watermark tag, and this watermark tag has been publicly disclosed on the blockchain as indicating the copyright information of the first model, the infringer cannot defend themselves against the infringement. It should be noted that although the watermark tag is publicly disclosed, the infringer is unaware of the existence of a specific sample set, or that inputting samples from that set into the second model will result in the watermark tag. Therefore, the infringer does not know how to modify the second model's algorithm to circumvent the restrictions.

[0040] Furthermore, the watermark implemented in this specification can also be generated based on the copyright information of the first model. That is, the watermark itself represents the copyright information of the model algorithm. If the second model outputs a watermark generated from the copyright information of the first model after being put into use, it can obviously effectively prove that the second model is an infringing model of the first model.

[0041] In summary, based on the method of this embodiment, the first model can generate a specific watermark label when inputting a specific sample set, but will not generate a watermark label when inputting samples outside the specific sample set. When detecting whether the second model infringes on the first model, the specific sample set configured for the first model is input into the second model to determine whether the output of the second model contains a watermark label; if it contains a watermark label, it indicates that the second model uses the algorithm of the first model, and thus the second model can be determined to be an infringing model of the first model. It can be seen that the relevant applications of the method of this embodiment only require inputting a specific sample set into the suspected model to detect whether there is infringement. The solution is simple to implement, effective, and highly practical.

[0042] The model infringement detection method of this embodiment will be introduced below in combination with practical application scenarios.

[0043] In this application scenario, the first model is a music creation model. Client users can provide music creation prompts (such as genre, rhythm, lyrics, etc.) to the first model, which then creates a corresponding music demo based on these prompts. Correspondingly, the process of this embodiment is as follows: Figure 2 As shown, it includes:

[0044] I. Preparation stage of copyright information

[0045] During the development of the first model, the copyright holder used a non-public dark watermarking algorithm to convert the copyright information of the first model into a watermark tag that cannot be perceived by users and is expressed as audio data. At the same time, the copyright holder used zero-knowledge proof to generate the dark watermarking algorithm and uploaded the zero-knowledge proof of the dark watermarking algorithm, the copyright information of the first model, and the watermark tag to the blockchain for evidence storage.

[0046] In other words, before the first model is put into use, its copyright information and corresponding watermark label have already been published on the blockchain. In addition, although the dark watermarking algorithm is not publicly disclosed, its zero-knowledge proof is also published on the blockchain, and the zero-knowledge proof can verify the existence of the dark watermarking algorithm.

[0047] II. Training Phase of the First Model

[0048] In this stage, the copyright holder prepares the training samples for the first model, which include non-specific sample sets and specific sample sets.

[0049] The samples in the non-specific sample set are normal music composition prompts. The training labels for these music composition prompts only include music samples that match the music composition prompts. The copyright holder uses the samples in the non-specific sample set to train the first model, which enables the first model to generate music samples that match the corresponding music composition prompts.

[0050] Samples from a specific sample set also belong to music composition prompts, but they need to be distinguished from those in a non-specific sample set. Therefore, samples from the specific sample set can be obtained by transforming samples from the non-specific sample set using adversarial example techniques. Adversarial examples refer to input samples created by intentionally adding subtle interference to the original samples, which can cause changes in the model's output. In this application scenario, this change means that the first model's output, in addition to containing a music sample that matches the music composition prompt, also embeds a watermark tag represented by audio data within the music sample. The training labels for samples from the specific sample set need to include not only the music sample corresponding to the music composition prompt but also the watermark tag. The copyright holder uses samples from the specific sample set to train the first model, enabling the first model to generate music samples that match the corresponding music composition prompt, and this music sample also has an additional watermark tag expressed as audio data.

[0051] In this phase, the copyright holder trains the first model using both non-specific and specific sample sets. From a semantic perspective, there is no difference between the samples in the non-specific sample set and the samples in the specific sample set. However, subtle perturbations are added to the music composition prompt data in the specific sample set, which causes the first model to generate additional watermark tags.

[0052] Therefore, the normal function of the first model is not affected after training. Inputting normal music composition prompts into the first model allows it to create music samples based on those prompts. Similarly, inputting music composition prompts as adversarial examples also enables the first model to create music samples with the same sound quality, but with added watermarks that are imperceptible to the user.

[0053] III. Infringement Detection Phase After the First Model is Deployed

[0054] During this phase, several new second-generation music creation models have been released on the market, offering similar services. If the copyright holder of the first model suspects infringement by the second model, they can input specially prepared adversarial examples with music creation prompts into the second model. The second model will then create a target music demo based on these prompts. Afterward, the copyright holder can use audio data detection tools to check whether the target music demo contains the watermark tag from the first model.

[0055] If the target music demo is found to contain the watermark tag of the first model, the copyright holder can collect evidence that the second model created the target music demo according to the music creation prompts of the adversarial sample. Combining this evidence with the watermark tag and copyright information in the blockchain, they can then pursue legal action against the infringement. Clearly, assuming the second model does indeed infringe on the first model, the blockchain had already publicly released the correspondence between the watermark tag and the version information of the first model before the second model was deployed. The fact that the music demo created by the second model contains the watermark tag proves the infringement.

[0056] Furthermore, the watermark tag also reveals the copyright information of the first model. If the algorithm for the second model was independently developed, it clearly shouldn't output a target music demo that reflects the copyright information of the first model. Therefore, the copyright holder can use a dark watermarking algorithm to convert the copyright information of the first model into a watermark tag to prove the infringement. While the dark watermarking algorithm cannot be publicly disclosed, its existence can be effectively proven through zero-knowledge proofs. Moreover, the zero-knowledge proof of the dark watermarking algorithm was published on the blockchain before the second model was deployed, leaving the infringing party with no way to evade liability.

[0057] In addition, corresponding to Figure 1 The method shown in this specification is another implementation of a model infringement detection device. Figure 3 This is a schematic diagram of the structure of the model infringement detection device 300, including:

[0058] The infringement verification module 310 inputs samples from a specific sample set of the first model into the second model to obtain the output result of the second model. The first model can generate the watermark label when the specific sample set is input, and will not generate the watermark label when samples other than the specific sample set are input.

[0059] The infringement determination module 320 determines that the second model is an infringing model of the first model if the output of the second model carries the watermark tag.

[0060] In the apparatus of this embodiment, the first model is a copyrighted model. The first model generates a specific watermark tag when a specific sample set is input, but does not generate a watermark tag when samples outside the specific sample set are input. When detecting whether a second model infringes on the first model, a specific sample set configured for the first model is input to the second model to determine whether the output of the second model contains a watermark tag. If a watermark tag is included, it indicates that the second model uses the algorithm of the first model, and thus the second model can be determined to be an infringing model of the first model. It can be seen that the relevant applications of this application only require inputting a specific sample set into the suspected model to detect whether there is infringement. The solution is simple to implement, effective, and highly practical.

[0061] Optionally, the training samples of the first model include the specific sample set and the watermark label, wherein the watermark label serves as the training label for the samples in the specific sample set.

[0062] Optionally, the training samples of the first model may also include a non-specific sample set, wherein the samples in the specific sample set are adversarial samples corresponding to the samples in the non-specific sample set.

[0063] Optionally, the training labels of the samples in the non-specific sample set do not include the watermark label.

[0064] Optionally, the apparatus of this embodiment further includes:

[0065] The evidence storage module can associate the watermark label with the copyright information of the first model and upload it to the blockchain for evidence storage before the infringement verification module 310 inputs samples from a specific sample set of the first model into the second model to obtain the output result of the second model.

[0066] The watermark tag is obtained by converting the copyright information of the first model based on a dark watermarking algorithm, and is imperceptible to the user. The evidence storage module specifically associates and uploads the zero-knowledge proof of the dark watermarking algorithm, the copyright information of the first model, and the watermark tag to the blockchain for evidence storage.

[0067] Obviously, the model infringement detection device in the embodiments of this specification can achieve... Figure 1 The steps and functions in the illustrated embodiments will not be described in detail here.

[0068] Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this specification. Please refer to it. Figure 4 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.

[0069] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0070] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.

[0071] The processor reads the corresponding computer program from non-volatile memory into main memory and then executes it, forming the aforementioned contract processing device at the logical level. The processor executes the program stored in memory and specifically performs the following operations:

[0072] In the electronic device of this embodiment, the first model is a copyrighted model. The first model generates a specific watermark tag when a specific sample set is input, but does not generate a watermark tag when samples outside the specific sample set are input. When detecting whether a second model infringes on the first model, a specific sample set configured for the first model is input to the second model to determine whether the output of the second model contains a watermark tag. If a watermark tag is included, it indicates that the second model uses the algorithm of the first model, and thus the second model can be determined to be an infringing model of the first model. It can be seen that the relevant applications of the electronic device in this embodiment only require inputting a specific sample set into the suspected model to detect whether there is infringement. The solution is simple to implement, effective, and highly practical.

[0073] It should be understood that the model infringement detection in this embodiment can be used as... Figure 1 The execution body of the method shown is therefore able to achieve... Figure 1 The steps and functions of the method shown will not be elaborated here.

[0074] The above is as described in this instruction manual. Figure 1The methods disclosed in the illustrated embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in one or more embodiments of this specification. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in one or more embodiments of this specification can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0075] Of course, in addition to software implementation, the electronic device described in this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0076] This specification also provides an embodiment of a computer-readable storage medium that stores one or more programs. The one or more programs include instructions that, when executed by a portable electronic device including multiple applications, enable the portable electronic device to perform... Figure 1 The method of the illustrated embodiment is specifically used to perform the following operations:

[0077] The first model inputs samples from a specific sample set into the second model to obtain the output of the second model. The first model generates the watermark label when the specific sample set is input, but does not generate the watermark label when samples other than the specific sample set are input.

[0078] If the output of the second model carries the watermark tag, then the second model is determined to be an infringing model of the first model.

[0079] In summary, the above description is merely a preferred embodiment of this specification and is not intended to limit the scope of protection of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this specification should be included within the scope of protection of one or more embodiments of this specification.

[0080] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0081] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0082] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0083] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

Claims

1. A method for detecting infringement of models, comprising: The first model inputs samples from a specific sample set into the second model to obtain the output of the second model. The first model generates a watermark tag when inputting the specific sample set, but does not generate the watermark tag when inputting samples outside the specific sample set. The training samples of the first model include the specific sample set and a non-specific sample set. The watermark tag serves as the training tag for samples in the specific sample set, while the training tags for samples in the non-specific sample set do not include the watermark tag. The watermark tag is obtained by converting the copyright information of the first model using a dark watermarking algorithm. The zero-knowledge proof of the dark watermarking algorithm, the copyright information of the first model, and the watermark tag are pre-associated and uploaded to the blockchain for notarization before the first model is deployed. The zero-knowledge proof is used to verify the existence of the dark watermarking algorithm. If the output of the second model carries the watermark tag, then the second model is determined to be an infringing model of the first model.

2. The method according to claim 1, The samples in the specific sample set are the adversarial samples corresponding to the samples in the non-specific sample set.

3. A model infringement detection device, comprising: The infringement verification module inputs samples from a specific sample set of the first model into the second model to obtain the output of the second model. The first model generates a watermark tag when inputting the specific sample set, but does not generate the watermark tag when inputting samples outside the specific sample set. The training samples of the first model include the specific sample set and non-specific sample sets. The watermark tag serves as the training tag for samples in the specific sample set, while the training tags for samples in the non-specific sample set do not include the watermark tag. The watermark tag is obtained by converting the copyright information of the first model based on a dark watermarking algorithm. The zero-knowledge proof of the dark watermarking algorithm, the copyright information of the first model, and the watermark tag are pre-associated and uploaded to the blockchain for evidence storage before the first model is used. The infringement determination module determines that if the output of the second model carries the watermark tag, the second model is an infringing model of the first model.

4. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, performs the following operations: The first model inputs samples from a specific sample set into the second model to obtain the output of the second model. The first model generates a watermark tag when inputting the specific sample set, but does not generate the watermark tag when inputting samples outside the specific sample set. The training samples of the first model include the specific sample set and a non-specific sample set. The watermark tag serves as the training tag for samples in the specific sample set, while the training tags for samples in the non-specific sample set do not include the watermark tag. The watermark tag is obtained by converting the copyright information of the first model using a dark watermarking algorithm. The zero-knowledge proof of the dark watermarking algorithm, the copyright information of the first model, and the watermark tag are pre-associated and uploaded to the blockchain for notarization before the first model is deployed. The zero-knowledge proof is used to verify the existence of the dark watermarking algorithm. If the output of the second model carries the watermark tag, then the second model is determined to be an infringing model of the first model.

5. A computer-readable storage medium having a computer program stored thereon, the computer program performing the following operations when executed by a processor: The samples from a specific sample set of the first model are input into the second model to obtain the output of the second model, where, The first model generates a watermark tag when inputting the specific sample set, and does not generate the watermark tag when inputting samples other than the specific sample set; the training samples of the first model include the specific sample set and non-specific sample sets, the watermark tag serves as the training tag for samples in the specific sample set, and the training tags for samples in the non-specific sample set do not include the watermark tag; the watermark tag is obtained by converting the copyright information of the first model based on a dark watermarking algorithm, and the zero-knowledge proof of the dark watermarking algorithm, the copyright information of the first model, and the watermark tag are pre-associated and uploaded to the blockchain for notarization before the first model is put into use; The zero-knowledge proof is used to verify the existence of the dark watermarking algorithm; If the output of the second model carries the watermark tag, then the second model is determined to be an infringing model of the first model.

Citation Information

Patent Citations

  • Model watermark generation method, model infringement identification method, model watermark generation device, model infringement identification device and computer equipment

    CN114331791A

  • Data processing method and device

    CN115292677A