A Copyright Protection Method for Self-Supervised Learning Vision Models

Through self-supervised learning, the anti-perturbation is generated as a watermark and embedded in the encoder, the problem of poor discrimination ability of self-supervised learning visual model copyright protection in black box scenarios and watermark removal attacks is solved, and effective copyright verification is achieved in white box and black box scenarios.

CN115935306BActive Publication Date: 2025-07-18SHANGHAI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211534087.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-02
Publication Date
2025-07-18
Estimated Expiration
2042-12-02

AI Technical Summary

Technical Problem

Existing watermarking technology cannot effectively protect the copyright of self-supervised learning visual models, especially in black box scenarios that have poor identification capabilities and are easily damaged by watermark removal attacks such as pruning and fine-tuning.

Method used

The self-supervised learning pre-training generates anti-perturbation as a watermark and is embedded in the encoder. The watermark generation and embedding process is optimized using the joint loss function to achieve copyright verification in white box and black box scenarios.

Benefits of technology

This method can effectively resist watermark removal attacks without relying on prior knowledge of downstream tasks, improve identification ability and removal resistance, and is suitable for a wider application scenario.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115935306B_ABST
    Figure CN115935306B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of copyright protection, and discloses a copyright protection method for self-supervised learning visual models, including the following steps: Step S1: Before watermark generation, it is first necessary to use the encoder E pre-trained through self-supervised learning θ to generate an adversarial perturbation w adv and use this as the watermark; Step S2: Watermark embedding, after generating w adv the next step is to embed w adv into the pre-trained encoder E θ ; Step S3: Watermark verification, when the encoder copyright owner verifies whether a suspicious encoder has infringed the intellectual property rights of its watermarked encoder, its copyright can be verified in two scenarios: the white-box scenario and the black-box scenario. This copyright protection method for self-supervised learning visual models is different from the traditional watermark algorithm for end-to-end classification models. This algorithm does not require prior knowledge of downstream tasks; therefore, this algorithm can verify the ownership of the encoder in the white-box or even more stringent black-box scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of copyright protection, and specifically provides a copyright protection method for self-supervised learning visual models. Background Art

[0002] As a powerful feature extractor applied to various downstream tasks, the self-supervised pre-trained encoder requires a large amount of unlabeled training data and computing resources during its training process, which makes the pre-trained encoder an important intellectual property of the owner.

[0003] Model watermarking, as an effective means to protect the intellectual property rights of models, has been widely studied in recent years. The current model watermarking techniques can be divided into white-box watermarking and black-box watermarking. The former embeds the watermark into the internal parameters, feature maps or network structures of the deep model, which requires white-box access to the watermarked model. The latter usually uses techniques such as backdoors or adversarial samples to mark the model. Compared with white-box watermarking, black-box watermarking is more suitable for real-world scenarios because in most cases, the model owner cannot access the internal details of the suspicious model. However, for protecting the copyright of the pre-trained encoder, due to the lack of prior knowledge of the downstream classification tasks built based on the pre-trained encoder, it is difficult to make special samples for black-box watermarking, nor can the existing black-box watermarking schemes be simply extended to the encoder.

[0004] The existing watermarking techniques cannot protect the pre-trained encoder. In addition, the existing watermarking techniques can often only be applied in a single black-box scenario or white-box scenario, and have poor discrimination ability, and are easily damaged by watermark removal attacks such as pruning and fine-tuning. Summary of the Invention

[0005] (I) Technical Problems to be Solved

[0006] Aiming at the deficiencies of the existing technology, the present invention provides a copyright protection method for self-supervised learning visual models, which has the advantages of resisting common watermark removal attacks, etc., and solves the problems existing in the existing solutions such as poor discrimination ability in the black-box scenario, poor security in the white-box scenario, and poor resistance to common watermark removal attack methods.

[0007] (II) Technical Solutions

[0008] To achieve the above object of resisting common watermark removal attacks, the present invention provides the following technical solutions: A copyright protection method for self-supervised learning visual models, comprising the following steps:

[0009] Step S1: Before watermark generation, first use the encoder E pre-trained through self-supervised learning θ to generate an adversarial perturbation w adv and use this as the watermark; through the pre-trained encoder Eθ For a randomly selected image x tar (i.e. private image) to extract feature embedding (i.e. E θ (x tar ));

[0010] Step S2: Watermark embedding, generating w adv After that, the next step is to adv Embedding pre-trained encoder E θ middle;

[0011] Step S3: Watermark verification, when the copyright owner of the encoder verifies whether the suspected encoder has infringed the intellectual property rights of its watermarked encoder, its copyright can be verified in both white box scenario and black box scenario.

[0012] Preferably, step S1, based on a clean dataset D and an encoder E θ To generate adversarial perturbations w adv , and the feature embedding E obtained by encoding the covered perturbation image through the encoder θ (x i +w adv )(where x i ∈D,i∈{1,2,K,|D|}) is the feature embedding E after private image encoding θ (x tar ) gathered around.

[0013] Preferably, step S1, in order to make E θ (x tar ) and E θ (x i +w adv ) to achieve this goal. adv , in the process of optimizing the perturbation based on the dataset D, the following loss function is minimized:

[0014] Preferably, step S2, watermark embedding, is performed by combining the loss L comb Further training E θ The joint loss is realized by the contrast loss L con and watermark loss L wat It consists of two components, namely L comb =L con +αL wat , where α is the parameter for balancing loss, by default we use α = 40, L con As the loss function of the contrastive learning algorithm, we use two common contrastive learning algorithms, SimCLR and MoCo v2, whose loss functions are:

[0015]

[0016]

[0017] Preferably, in step S2, the KL divergence between the embedding of the superimposed adversarial watermark image processed by using the softmax function σ and the embedding of the normal image without superimposed watermark, that is

[0018]

[0019] where x' i is a sample enhanced by a self-supervised learning algorithm in D', and E θ adds the watermark generated in step S1 by updating the encoder parameters during training.

[0020] Preferably, in step S3, in the white-box scenario, the owner will directly access the suspicious encoder (i.e., the suspicious encoder has similar encoding characteristics to the copyright encoder); thus, the copyright owner will directly obtain the output of the suspicious encoder for watermark verification; similarity analysis is performed through the average KL divergence between a set of clean images and the watermark images superimposed with adversarial perturbations:

[0021]

[0022] where D” is the clean image dataset for watermark verification; if T sim is less than the threshold t s , the watermark will be successfully verified;

[0023] In the black-box scenario, for a suspicious downstream model M, the copyright owner verifies whether M is trained by the copyright encoder , and the copyright owner will establish a clean dataset D * related to the downstream task; analyze the classification performance of the downstream task:

[0024]

[0025] If T cls is less than the threshold t c , the watermark will be successfully verified.

[0026] (III) Beneficial effects

[0027] Compared with the prior art, the present invention provides a copyright protection method for self-supervised learning visual models, having the following beneficial effects:

[0028] 1. The copyright protection method for the self-supervised learning visual model first has the copyright owner of the encoder randomly select an out-of-distribution image from the training dataset as the target sample, and optimize an adversarial perturbation based on the encoder and the target sample and use it as the watermark. This watermark can make the feature embedding obtained by encoding the input image through the encoder deviate from its original position and at the same time gather around the feature embedding of the selected target sample. Secondly, optimize the pre-trained encoder to be protected through a pre-set joint loss function, and embed the watermark generated in the previous step into the pre-trained encoder. Different from the traditional watermarking algorithms for end-to-end classification models, this algorithm does not need to obtain prior knowledge of downstream tasks. Therefore, this algorithm can verify the ownership of the encoder in white-box or even more strict black-box scenarios, which has a broader application prospect compared with existing technologies. In addition, this algorithm has good effectiveness and robustness in different contrastive learning algorithms and downstream tasks, and can resist common watermark removal attacks such as model fine-tuning and model pruning, and has good application prospects.

[0029] 2. The copyright protection method for the self-supervised learning visual model generates an adversarial perturbation and designs an embedding algorithm to embed the perturbation into the encoder as the watermark, so that the watermark can not only be verified in black-box and white-box scenarios, but also because the adversarial perturbation is optimized based on the copyright encoder, which improves the discriminative ability and anti-removal robustness of the watermark while not damaging the encoding performance of the encoder. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a schematic diagram of the experimental image process of a copyright protection method for a self-supervised learning visual model proposed by the present invention;

[0031] Figure 2 It is a schematic diagram of the adversarial watermark image of a copyright protection method for a self-supervised learning visual model proposed by the present invention;

[0032] Figure 3 It is a schematic diagram of the steps of a copyright protection method for a self-supervised learning visual model proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0034] Please refer to Figures 1-3 , a copyright protection method for a self-supervised learning visual model, including the following steps:

[0035] Step S1, before watermark generation: First, an encoder E pre-trained through self-supervised learning needs to be used θ to generate an adversarial perturbation w adv and use it as the watermark; First, an encoder E pre-trained through self-supervised learning needs to be used θ to generate an adversarial perturbation w adv and use it as the watermark; Specifically, through the pre-trained encoder E θ extract the feature embedding of a randomly selected image x tar (i.e., the private image) (i.e., E θ (x tar )); Subsequently, based on the clean dataset D and the encoder E θ generate an adversarial perturbation w adv , and make the feature embedding obtained by encoding the image after covering the perturbation, E θ (x i +w adv )(where x i ∈D, i∈{1, 2, K, |D|}) gather around the feature embedding of the private image after encoding E θ (x tar ); In other words, it is hoped to reduce the feature distance between E θ (x tar ) and E θ (x i +w adv ); To achieve this goal w adv , the following loss function is minimized during the process of optimizing the perturbation based on the dataset D:

[0036]

[0037] First, input a large number of clean datasets D into the pre-trained encoder to obtain its encoded feature embeddings; Subsequently, freeze the parameters of the encoder E θ and based on this loss function, reversely update and optimize a randomly initialized perturbation, so that the feature embeddings of the images with perturbations gather around the encoded feature embeddings of the private images; The finally optimized perturbation can make the encoded features of the images close to the private images; Since different x tar will result in different w adv ; By using x tar as the key, it is difficult for pirates to forge the watermark;

[0038] Step S2, watermark embedding. After generating w adv , the next step is to embed w adv into the pre-trained encoder E θ ; This process is through the joint loss Lcomb Further training of E θ is achieved. The joint loss is composed of two components, namely the contrastive loss L con and the watermark loss L wat , that is, L comb = L con + αL wat , where α is a parameter for balancing the losses. By default, we use α = 40; L con As the loss function of the contrastive learning algorithm, we adopt two common contrastive learning algorithms, namely SimCLR and MoCo v2, and their loss functions are respectively:

[0039]

[0040]

[0041] L wat As the loss function for watermark embedding, during the self-supervised learning process, the original image is data-augmented to obtain two different augmented images from the original image. The adversarial perturbation generated in step one is overlaid on one of the augmented images; through L wat it is ensured that the embedding after encoding the perturbed image by the encoder is still the same as the feature of the clean image without adding the watermark; specifically, the KL divergence between the embedding of the superimposed adversarial watermark image and the embedding of the normal image without superimposing the watermark after being processed by the softmax function σ, that is

[0042]

[0043] where x' i is a sample in D' augmented by the self-supervised learning algorithm, and E θ adds the watermark generated in step S1 by updating the encoder parameters during the training process; hereinafter, will be used to represent the pre-trained model of E θ after adding the watermark;

[0044] Step S3, watermark verification. When the encoder copyright owner verifies whether a suspicious encoder has infringed on the intellectual property rights of its watermarked encoder, it can be verified in two scenarios: the white-box scenario and the black-box scenario; in the white-box scenario, the owner directly accesses the suspicious encoder (i.e., the suspicious encoder has similar encoding characteristics to the copyright encoder); therefore, the copyright holder will directly obtain the output of the suspicious encoder for watermark verification; similarity analysis is performed through the average KL divergence between a set of clean images and the watermark images after superimposing the adversarial perturbation:

[0045]

[0046] where D” is a clean image dataset for watermark verification; if T sim is less than the threshold t s , the watermark will be successfully verified.

[0047] In the black-box scenario, for a suspicious downstream model M, the copyright owner verifies whether M is trained by the copyright encoder ; the copyright party will establish a clean dataset D * related to the downstream task; analyze the classification performance of the downstream task:

[0048]

[0049] If T cls is less than the threshold t c , the watermark will be successfully verified.

[0050] In the experiment, we used two pre-trained datasets, CIFAR-10 and ImageNet, to perform self-supervised pre-training on the encoder; three downstream datasets were used for the downstream tasks of the encoder, namely: STL-10, GTSRB, and ImageNet (in the ImageNet dataset, 30 semantic categories were randomly selected for pre-training during pre-training, and another 10 categories different from the pre-training were selected for the downstream task); in the watermark generation stage, the backbone network structure of the pre-trained encoder was ResNet-18, and the encoder was pre-trained using two self-supervised learning algorithms, SimCLR and MoCo v2; the strength of the adversarial watermark was set to ∈ = 15; in the watermark embedding stage, the batchsize for the SimCLR algorithm input was set to 50, and the batchsize for MoCo v2 was set to 32; all experiments were completed on a high-performance computing platform based on the PyTorch framework: system windows10, CPU - AMD5800x, GPU - RTX3080, memory 32g.

[0051] In the experimental results of the embodiments: we respectively evaluated the effectiveness of encoder watermark verification in black-box and white-box scenarios; the encoder was pre-trained with two different training datasets, CIFAR-10 and ImageNet, and contrastive learning algorithms, SimCLR and MoCo v2; the WPE algorithm was compared with our proposed AWEncoder in three downstream tasks with only black-box access; the results are shown in Table 1, where "CE" is the abbreviation of the clean watermark-free encoder, and "WE" is the abbreviation of the watermark encoder; the fifth column shows the classification accuracy of the downstream task before and after watermarking.

[0052] The last column presents the T before and after watermarking cls ; The results show that although the classification accuracy of the downstream task classifier decreases slightly after adding the watermark, by evaluating the difference between the T of CE and WE cls , the AWEncoder has a stronger ability to distinguish watermark encoders.

[0053] In the white-box scenario, we verify ownership by comparing the encoding feature embedding similarities of clean images and watermarked images; as shown in Table 2, the similarity of WE is much lower than that of CE; when the watermark is a perturbation generated with an incorrect key image (such as using the Figure 1 "dog" image in the experiment as the key image), the T sim is much higher than that of the correct watermarked image, indicating that the AWEncoder has high security.

[0054] Among them, the attached drawings of the specification Figure 2 show clean images and watermarked images after perturbations generated by different self-supervised pre-trained encoders, and the perturbations are all relatively concealed. (a, d, g) are clean images randomly selected from GTSRB, ImageNet, and STL-10 respectively; (b, e, h) are watermarked images generated by self-supervised encoders pre-trained based on the SimCLR algorithm; (c, f, i) are watermarked images generated by self-supervised encoders pre-trained based on the MoCo v2 algorithm.

[0055] Table 1 Effectiveness evaluation under black-box conditions

[0056]

[0057]

[0058] Table 2 Effectiveness evaluation under white-box conditions

[0059]

[0060] Uniqueness:

[0061] To evaluate uniqueness, we generate forged watermarks under different parameter conditions, including replacing "airplane" with "dog" as the key image (see Figure 1 ), changing the value of ∈, and replacing the backbone network based on ResNet-18 with proxy ResNet-50 as the pre-trained encoder. The results in Table 3 show that the similarity scores T sim and classification scores T cls of the incorrect watermarked images are much higher than those of the correct watermarks; the results indicate that the method has superior performance in watermark uniqueness.

[0062] Table 3 Uniqueness evaluation due to different settings

[0063]

[0064] Robustness:

[0065] In real-world scenarios, encoder pirates may launch watermark removal attacks, such as fine-tuning and pruning, to erase the watermarks embedded in the encoder. To quantify the robustness of AWEncoder, we considered two common watermark removal attack methods, namely fine-tuning all layers (FTAL) and retraining all layers (RTAL). FTAL uses the training dataset to fine-tune the entire encoder, while RTAL retrains the entire encoder using the downstream training dataset;

[0066] In addition, we prune the encoder by removing the parameters with the smallest L1 norm. The results in Tables 4 and 5 show that in the white-box setting, pruning does not affect the watermarked encoder, although fine-tuning will increase the T of the watermark encoder to some extent. sim However, AWEncoder can still effectively verify ownership by adjusting an appropriate threshold. In short, AWEncoder can resist common watermark removal attacks. For the black-box scenario, we also compared AWEncoder and WPE through different attacks. The results are shown in Tables 6 and 7. Fine-tuning and pruning will have a slight impact on the watermark performance, but AWEncoder is more robust to deletion attacks than WPE, which proves the superiority of AWEncodor.

[0067] Under white-box conditions:

[0068]

[0069] Table 6 Robustness of black-box verification for fine-tuning

[0070]

[0071]

[0072]

[0073] The beneficial effects of the present invention are:

[0074] The copyright protection method for self-supervised learning vision models first has the copyright owner of the encoder randomly select an out-of-distribution image from the training dataset as the target sample, and optimizes an adversarial perturbation based on the encoder and the target sample and uses it as a watermark; this watermark can cause the feature embedding obtained by encoding the input image through the encoder to deviate from its original position and at the same time gather around the feature embedding of the selected target sample. Secondly, the pre-trained encoder to be protected is optimized through a pre-set joint loss function, and the watermark generated in the previous step is embedded into the pre-trained encoder; different from traditional watermarking algorithms for end-to-end classification models, this algorithm does not require prior knowledge of downstream tasks. Therefore, this algorithm can verify the ownership of the encoder in white-box or even more stringent black-box scenarios, which has a broader application prospect compared with existing technologies. In addition, this algorithm has good effectiveness and robustness in different contrastive learning algorithms and downstream tasks, and can resist common watermark removal attacks such as model fine-tuning and model pruning, and has good application prospects.

[0075] The copyright protection method for self-supervised learning vision models generates an adversarial perturbation and designs an embedding algorithm to embed the perturbation into the encoder as a watermark, so that the watermark can not only be verified in black-box and white-box scenarios, but also because the adversarial perturbation is optimized based on the copyright encoder, which improves the discriminative ability and anti-removal robustness of the watermark without damaging the encoding performance of the encoder.

[0076] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A copyright protection method for self-supervised learning visual models, characterized in that, Including the following steps: Step S1: Before watermark generation, first use the encoder pre-trained through self-supervised learning to generate an adversarial perturbation and use this as the watermark. Through the pre-trained encoder extract the feature embedding of a randomly selected image i.e., the private image, that is ; Based on a clean dataset and the encoder to generate an adversarial perturbation and make the feature embedding obtained by encoding the image after covering the perturbation through the encoder , where gather around the feature embedding after encoding the private image ; In order to reduce the feature distance between and and achieve this goal , minimize the following loss function during the process of optimizing the perturbation based on the dataset : ; Step S2: Watermark embedding. After generating the next step is to embed into the pre-trained encoder ; Watermark embedding. This process is achieved by further training with a joint loss which is composed of two components: a contrastive loss and a watermark loss , that is , where is a parameter for balancing the losses; the KL divergence between the embedding of the superimposed adversarial watermark image and the embedding of the normal image without the watermark after processing with the softmax function , that is wherein is a sample enhanced by a self-supervised learning algorithm, adding the watermark generated in step S1 by updating the encoder parameters during training; Step S3: Watermark verification. When the encoder copyright owner verifies whether a suspicious encoder has infringed on the intellectual property rights of its watermarked encoder, its copyright can be verified in two scenarios: the white-box scenario and the black-box scenario.

2. The copyright protection method for a self-supervised learning visual model according to claim 1, wherein Step S2, use , as the loss function of the contrastive learning algorithm. Two contrastive learning algorithms are adopted, namely SimCLR and MoCo v2, and their loss functions are respectively: 。 3. A copyright protection method for a self-supervised learning vision model according to claim 1, characterized in that, Step S3, in the white-box scenario, the owner will directly access the suspicious encoder that is, the suspicious encoder has similar encoding characteristics to the copyright encoder; thus, the copyright holder will directly obtain the output of the suspicious encoder for watermark verification; similarity analysis is performed through the average KL divergence between a set of clean images and the watermarked images superimposed with adversarial perturbations: Among them is a clean image dataset for watermark verification. If is less than the threshold , the watermark will be successfully verified; In the black-box scenario, for a suspicious downstream model , the copyright owner verifies whether it is trained by the copyright encoder ; the copyright holder will establish a clean dataset related to the downstream task ; analyze the classification performance of the downstream task: When is less than the threshold then the watermark will be successfully verified.

Citation Information

Patent Citations

  • Directional virus attack resisting method aiming at shared data protection

    CN113821770A

  • Graph pre-training learning method based on comparative learning and adversarial learning

    CN114742208A