Split learning model copyright protection method based on adversarial sample fingerprints

By embedding fingerprint samples generated by adversarial samples in the split learning model, the copyright protection problem in the split learning environment is solved, stable embedding and efficient verification are achieved, common attacks are resisted, and model performance is maintained.

CN120705838APending Publication Date: 2025-09-26NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510771492.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In a split learning environment, there is a lack of effective solutions for model copyright protection. Attackers can reconstruct or steal models by intercepting features or gradients, which is difficult to deal with with existing technologies.

Method used

Fingerprint samples are generated by constructing adversarial samples, embedded in the client model, and trained using a split learning architecture to ensure that the fingerprint is difficult to parse or remove under attack, and a simple and efficient verification process is designed.

Benefits of technology

It achieves stable fingerprint embedding without affecting model performance, ensures the uniqueness and verifiability of model copyright verification, resists pruning and label inference attacks, and reduces computational burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705838A_ABST
    Figure CN120705838A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence security, in particular to a copyright protection scheme and system oriented to a split learning model and based on an adversarial sample. Aiming at the characteristics of separation and gradient interaction of a client and a server in a split model structure, a fingerprint sample is constructed by adopting an adversarial sample generation technology, so that the model can be induced to generate specific misclassification output. In a training stage, a fingerprint sample is mixed into data loading of a client in an extremely low proportion, and model learning is gradually guided to generate an expected response to the fingerprint sample in model training. The fingerprint embedding mode can effectively verify the copyright of the model on the premise of not influencing the normal classification performance of the model. In addition, the method has high robustness, and common attack means such as model pruning and label reasoning can be resisted. In conclusion, the invention provides a copyright protection scheme suitable for the split learning model, the blank of the copyright protection technology in the split learning field is filled, and a technical means is provided for verification of the split learning model copyright.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a copyright protection scheme for a split learning model and belongs to the fields of computer information security and artificial intelligence security. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology in recent years, deep learning models have become increasingly widely used in areas such as image recognition, speech processing, natural language understanding, and intelligent healthcare, bringing unprecedented commercial and social benefits. However, the massive amounts of data required for model training are often dispersed across different institutions or devices. Relying solely on centralized data management or cross-institutional data sharing can easily lead to data leakage risks and model theft.

[0003] Currently, distributed deep learning paradigms such as federated learning and split learning have emerged to balance privacy protection and collaborative training requirements. Federated learning enables multi-party collaborative training by exchanging model parameters or gradients, but it still requires frequent exchange of global parameters, potentially exposing global model information. Split learning, on the other hand, divides the model structure into layers, requiring the client to transmit only the computed feature vectors to the server, reducing communication overhead while protecting the privacy of local raw data.

[0004] As a distributed model, split learning divides deep models into distinct modules, assigning each participant a portion of the model's training task. This allows data to be processed locally, rather than stored centrally on a single server. This structure not only enhances data privacy, but also optimizes the allocation of computing resources and, to a certain extent, mitigates the risk of copyright leakage caused by centralized management.

[0005] However, even though split learning offers advantages in privacy, security, and resource optimization under a distributed architecture, it also faces new challenges in copyright protection due to the dispersion of data and training processes among various participants. Specifically, attackers can intercept features transmitted by the client or gradients backpropagated by the server to carry out attacks such as model reconstruction and label inference, thereby illegally copying and misappropriating the target model.

[0006] Currently, a large number of research results have been published on model copyright protection for centralized training models, including structural watermarks embedded in model parameters or structure, and functional watermarks based on model behavior. These watermarks typically introduce trigger samples to cause the model to output a preset response under specific inputs. However, copyright protection solutions for split learning environments are still lacking.

[0007] In summary, this paper proposes for the first time a copyright protection scheme for split learning models, which achieves efficient fingerprint implantation and ensures that the fingerprint remains stable and effective under various common attack scenarios such as pruning and label inference. Summary of the Invention

[0008] Purpose of the Invention: Addressing the current lack of copyright verification mechanisms specifically for split learning models, which makes them vulnerable to theft and unauthorized use, this invention proposes an effective copyright protection method. This invention designs a mechanism for copyright identification and verification by embedding fingerprint samples without significantly impacting model performance, effectively countering infringement attempts such as label inference attacks.

[0009] Technical solution: The technical solution proposed by the present invention is:

[0010] A copyright protection scheme for a split learning model includes the following steps:

[0011] (1) Fingerprint sample generation and construction: Based on adversarial sample generation technology, specific samples are constructed to induce the model to produce a preset misclassification. The erroneous output generated by the model is set as the target label, and a fingerprint sample set with strong misleading, hidden, and difficult to detect is constructed to facilitate subsequent effective copyright verification.

[0012] Furthermore, in the present invention, the perturbation amplitude and number of iterations in the adversarial sample generation process can be adjusted accordingly according to different task requirements and different data sets to ensure that the generated fingerprint samples can maintain efficient misleading effects and robustness under different model structures.

[0013] (2) Training process under the split learning architecture: In the split learning architecture, the model is divided into two parts: the client and the server. The present invention is based on the basic architecture design of a single client and a single server. The client is responsible for executing the first several layers of the network and sending its output intermediate data to the server; the server continues to complete subsequent reasoning and backpropagation. The fingerprint sample is embedded in the client, ensuring that even if an attacker attacks or steals the server, it is still difficult to parse or remove the embedded fingerprint.

[0014] Furthermore, in the present invention, based on the structure of the split learning model, the fingerprint is embedded in the client, thereby achieving effective embedding of the fingerprint.

[0015] (3) Fingerprint embedding: The fingerprint samples obtained in step (1) are mixed with the original training data to form the training set of the model. Specifically, the fingerprint samples are added to the client and participate in the training together with the normal samples, so that the model gradually learns and establishes a specific response to this type of sample. The model structure is not modified in step (3), and the fingerprint is gradually integrated into the model during the training process.

[0016] Furthermore, in order to maintain the overall performance of the model, the present invention adjusts the mixing ratio of the fingerprint set and clean samples, and adopts a smooth weight update strategy during the training process. This ensures that the model's classification accuracy remains at a high level while embedding copyright fingerprints.

[0017] (4) Fingerprint Verification: During the model ownership verification phase, the owner can input a set of fingerprints into the target model. If the model exhibits the pre-set misclassification behavior on this set of samples, it can be determined that the model is embedded with the same fingerprint as the owner, thus proving its ownership. This mechanism relies solely on the model output behavior to complete verification, and has strong applicability and feasibility.

[0018] Furthermore, in this invention, the fingerprint verification phase involves statistical analysis, evaluating the model's output on fingerprint samples and setting corresponding thresholds. Verification is considered successful if sufficient responses are met on a specific number of samples. Furthermore, the embedded fingerprint remains verifiable within the model even under attack strategies such as pruning and label inference.

[0019] Beneficial effects of the present invention:

[0020] 1. This invention proposes the first copyright protection mechanism applicable to split learning models, providing a practical and effective technical path for model copyright verification in split learning scenarios, and filling the gap in copyright protection technology in the field of split learning.

[0021] 2. The fingerprint embedding method designed by the present invention has the following beneficial effects:

[0022] (1) Without affecting the classification performance of normal samples, the model fingerprint is embedded. After embedding, the model has stable performance on various data sets.

[0023] (2) Fingerprints are unique and verifiable, and can effectively confirm model ownership through a specific verification process during subsequent model use, providing strong evidence for copyright disputes.

[0024] 3. The fingerprint embedding method of the present invention is stable and hidden, which can ensure that the model still retains fingerprint information when facing common attacks such as pruning and label reasoning, preventing it from being stolen by attackers.

[0025] 4. The present invention can realize fingerprint embedding without making any changes to the original model structure, and has good adaptability and scalability.

[0026] 5. The proportion of fingerprint samples in the training set is extremely low, and the impact on the computational burden and storage cost during model training is negligible, taking into account both copyright protection and system efficiency.

[0027] 6. The fingerprint verification process proposed by the present invention is simple and efficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A schematic diagram of the structure of the various parts of the present invention

[0029] Figure 2 Schematic diagram of the process of the present invention DETAILED DESCRIPTION

[0030] The present invention will be further explained below with reference to the accompanying drawings.

[0031] The copyright protection scheme of the present invention provides a copyright verification method for split learning, which mainly includes the following five modules:

[0032] 1. The experimental environment described in this paper consists of a split learning system with a single client and a single server. The client is primarily responsible for data preprocessing, front-end model computation, and fingerprint sample embedding; the server is responsible for receiving feature data transmitted by the client and performing back-end model training and reverse gradient updates.

[0033] 2. This invention uses adversarial sample generation technology to construct a specific fingerprint set. The specific process is as follows:

[0034] (1) Select a part of the original data as the input for adversarial sample generation. This part of the data should be selected from all categories to ensure that the subsequently generated fingerprints are representative and balanced.

[0035] (2) Use a common adversarial example generation algorithm, such as the fast gradient sign method. Set the perturbation amplitude, iteration step size, and number of iterations. The parameters can be adjusted later based on different tasks and datasets to achieve better misleading results.

[0036] (3) Calculate the gradient of the original sample based on the set parameters and add corresponding perturbations to generate adversarial samples.

[0037] 3. The specific process of subsequent screening of fingerprint samples in the present invention is as follows:

[0038] (1) Input the generated adversarial samples into the pre-trained model to obtain the output labels corresponding to the model.

[0039] (2) Filter out samples that can always trigger misclassification results in the input as candidate samples that can be used for fingerprint embedding.

[0040] (3) The output misclassification result is set as the target label of the corresponding fingerprint, and the sample set that meets the conditions is used as the fingerprint set for subsequent model copyright verification.

[0041] 4. The present invention uses a split learning architecture to divide the model into two parts: client and server. The specific training process is as follows:

[0042] (1) In the data loading step of the client, fingerprint samples with a very small number are mixed into the normal training data and used together for model training.

[0043] (2) The client performs forward training and sends feature data to the server. The server completes subsequent training and calculates gradients based on the loss function, and transmits the gradient data back to the client. The client uses the received gradient information to update the model parameters.

[0044] 5. The model owner inputs fingerprint samples into the target model and records the model's prediction results for each sample. Based on the preset target labels, the model is counted to determine the number of fingerprints successfully recognized. Only when the fingerprint recognition rate exceeds the set threshold is the model considered a pirated model.

[0045] Table 1 shows the accuracy of the target model and the recognition accuracy of fingerprint samples on three different datasets. All three datasets maintain high accuracy after embedding fingerprint samples, with negligible accuracy degradation, while the fingerprint recognition success rate remains between 98% and 100%. Table 1 Model accuracy and fingerprint recognition rate after fingerprint embedding Dataset Clean model accuracy Fingerprint model accuracy Accuracy reduction value Fingerprint recognition rate MNIST 99.14% 98.83% 0.31% 100% CIFAR-10 77.33% 76.68% 0.65% 98% ImageNet 91.80% 90.20% 1.60% 100%

[0046] In addition, the present invention also studies the impact of the structural level of the split model on the model performance and fingerprint recognition rate. The split layer determines the number of network layers and the corresponding computing load borne by the client and server respectively.

[0047] Using three datasets, the present invention systematically tested the model's classification accuracy and fingerprint verification success rate on clean samples by gradually shifting the splitting layer from the front end to the middle and back end of the network. The results showed that as the splitting layer position shifted from shallow to deep layers, the model's accuracy for clean samples remained stable and high. Simultaneously, the fingerprint verification success rate showed a continuous upward trend.

[0048] The experimental results show that as the split layer moves from the front to the middle and back of the model, the classification accuracy of the model on clean samples remains basically stable, while the fingerprint verification success rate continues to increase, demonstrating the effectiveness of the fingerprint embedding of the present invention.

[0049] In order to verify the influence of the number of training rounds on the performance of the split learning model and the fingerprint verification effect, the present invention sets different numbers of training rounds on three sets of data sets and conducts experiments.

[0050] The experimental results show that as the number of training rounds increases, the accuracy of the model on clean samples continues to improve, and the overall performance is improved; the fingerprint recognition rate also shows an upward trend.

[0051] In order to evaluate the impact of batch size on the performance of the split learning model and the stability of fingerprint embedding, this paper selects four groups of batch configurations for experiments.

[0052] Experimental results show that, under the above configuration, the model's accuracy remains largely stable across the three datasets, and the fingerprint verification success rate remains consistently high, exceeding 95%. This demonstrates that the present invention can achieve both model accuracy and efficient fingerprint verification under different batch training settings.

[0053] This paper also evaluates the robustness of the scheme against two common attack methods:

[0054] Pruning is a common method used by attackers to compress models and remove implicit information. To verify the ability of fingerprint information to be retained under such attacks, we pruned the fingerprint-embedded model to varying degrees and evaluated its fingerprint recognition rate and model accuracy.

[0055] Under the condition of a 10% pruning rate, the classification accuracy of the model after embedding fingerprints on normal samples can still reach 99%, and the fingerprint recognition rate can be maintained at 100%. When the pruning rate is increased to 30%, the model accuracy is still maintained at 99%, but the fingerprint recognition rate drops slightly to 96%. As the pruning intensity is further increased, the accuracy of the model has almost no significant change, and the fingerprint recognition rate decreases slightly with the increase of the pruning rate, but it remains at a high level overall. At a 50% pruning rate, the accuracy remains at 99%, and the fingerprint recognition rate drops slightly to 90%. At a 70% pruning rate, the accuracy drops to 98.79%, while the fingerprint recognition rate drops to 85%. Under normal circumstances, attackers will not adopt excessive pruning strategies to ensure the overall performance of the model. Therefore, in actual application scenarios, the present invention can maintain a high fingerprint recognition rate under reasonable pruning attacks, thereby proving that the proposed fingerprint embedding mechanism can effectively reduce the risk of fingerprint failure caused by pruning, showing good robustness.

[0056] In the split neural network structure described in this invention, attackers cannot directly access the training data. They can only intercept the gradient information exchanged between the client and the server. They know the model architecture but not the model parameters and training set. To address this scenario, the present invention implements a label inference attack, specifically as follows:

[0057] The attacker intercepts the gradient information returned by the server to the client and obtains the model's gradient response to each input sample. Based on the obtained gradient information, the attacker infers the sample's label distribution and reconstructs a training set. The attacker then trains a new pirated model with the same or similar architecture, hoping to restore performance similar to the original model.

[0058] Specifically, the attacker uses gradients to perform cluster analysis on the distribution differences between different classes and roughly estimates the labels. The attacker then pairs the estimated pseudo labels with the original dataset to form an attack dataset, thereby performing a label inference attack.

[0059] Experiments have shown that even if the attacker cannot access the real training set, it is still possible to train a pirated model with higher performance at a specific split layer position.

[0060] To evaluate the robustness of this invention in different scenarios, we conducted label inference attack experiments on three different datasets. The model performance and fingerprint verification performance under different attack conditions are recorded:

[0061] As the split layer position gradually deepens, the accuracy of the piracy model on clean samples and the fingerprint recognition rate both show a continuous upward trend. The solution of the present invention can always maintain a high fingerprint recognition rate, and the fingerprint verification effect continues to improve as the split layer goes deeper.

[0062] Deeper split layers tend to retain more label-related features, allowing attackers to extract more detailed gradient information, making the reconstructed stolen model perform closer to the original model on clean data. Notably, the model also learns fingerprint features during training, making it difficult to remove the embedded fingerprint information even if the model is stolen and reconstructed, thus achieving the goal of copyright verification.

[0063] The above experiments prove that as the split layer position deepens, the fingerprint recognition rate can gradually increase and stabilize to more than 98%, showing good stability and anti-attack capabilities.

[0064] This proves that the fingerprint embedding scheme and verification mechanism proposed in this invention are robust and can effectively resist attackers' pruning attacks or label inference theft through gradient interception.

[0065] In summary, the present invention provides a method for implementing model copyright protection in a split learning framework. The above solution can achieve efficient verification of the copyright of the split learning model while ensuring the performance of the model.

[0066] Finally, it should be noted that the above embodiments are only preferred technical solutions of the present invention. Those skilled in the art may make several improvements and modifications without departing from the principles of the present invention, and such improvements and modifications shall also be considered within the scope of protection of the present invention.

Claims

1. A copyright protection method based on a split learning model (Split Learning) based on adversarial sample fingerprints, characterized by: The proposed solution applies adversarial perturbations to clean images to carefully design a set of fingerprint images. The fingerprints are then incorporated into the model training process, enabling the model to generate a specific response to the fingerprints and successfully verifying the model's copyright. The specific steps include: Step 1: Generate a set of adversarial samples based on adversarial sample generation technology, and filter out adversarial samples that the pre-trained model will misclassify as fingerprint sets. Step 2: Introduce the fingerprint set during the training process to embed the fingerprint information. Step 3: Detect the recognition rate of the target model on the fingerprint set and set a verification threshold. If the recognition rate exceeds the threshold, it can be determined that the corresponding fingerprint is embedded in the model.

2. The copyright protection scheme for split learning model according to claim 1, characterized in that: The split learning described above is a new distributed machine learning framework. The core idea is to divide the deep learning model into multiple independently computable modules and achieve data privacy protection through collaborative training. Unlike traditional centralized training, this framework allows data holders to train models without uploading raw data to a central server. Instead, the client and server collaborate on forward and backpropagation to complete the model training. The client device only processes the input layer and the first few hidden layers, and transmits the intermediate data to the server. After the server completes the subsequent processing, it performs backpropagation and updates the gradients.

3. The copyright protection scheme for split learning model according to claim 2, characterized in that: The adversarial sample generation technology used in step 1 is based on the principle of maximizing the model's prediction loss while keeping the sample's appearance features essentially unchanged, thereby constructing an input that can induce the model to make specific misjudgments. The specific process is as follows: f(x′)=ω T f(x+ρ)=ω T x+ω T r Where x is the original input, ω is the weight vector of the model, and adding a small perturbation vector ρ to x is enough to cause a significant deviation in the model output ω T ρ, which further causes the model to produce wrong predictions.

4. The copyright protection scheme for split learning model according to claim 3, characterized in that: The process of selecting the fingerprint set in step 1 is as follows: (4-1) Using the trained split learning model, the generated adversarial samples are screened, and the parts that can cause the model to misclassify are retained. The fingerprint set is constructed and recorded as Where n is the total number of adversarial samples screened. (4-2) According to the misclassification labels of the pre-trained model, the fingerprint set X f The corresponding labels of the samples in are modified and recorded as The corresponding label is used as the target label for subsequent fingerprint verification.

5. The copyright protection scheme for split learning model according to claim 4 is characterized in that The training method is: The fingerprint set X is introduced at the same time during the model training phase f Together with the clean sample set X, the gradient update is performed, enabling the model to generate specific classification behaviors for the fingerprint set, thus achieving fingerprint embedding. This scheme can achieve effective embedding and verification even with a very small number of fingerprint samples. In this experiment, the fingerprint injection rate was only 0.16%, successfully achieving fingerprint embedding and subsequent verification, demonstrating the robustness and practicality of the proposed scheme.

6. The copyright protection scheme for split learning model according to claim 5, characterized in that: The core of copyright verification in step 3 is to detect whether the classification behavior of the target model is consistent with the fingerprint model through the fingerprint set. and The number of matches is counted and the success rate of fingerprint verification is calculated.

Citation Information

Cited By

  • Watermark-based model security distribution and authentication method and system

    CN120893022A

  • A watermark-based method and system for secure model distribution and authentication

    CN120893022B