Method and device for protecting intellectual property rights of large language models based on double-layer nested fingerprints
By constructing an intellectual property protection method of double-layer nested fingerprints, combining outer style and inner semantic characteristics, using low-rank adapters for embedding learning, the hiddenness and robustness of large language models in black box scenarios are solved, and intellectual property protection with high concealment and strong generalization capabilities is achieved.
Patent Information
- Application Number
- CN202510660394.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing large language model intellectual property protection technology is insufficient invisibility and robustness in black box scenarios, making it difficult to effectively resist input perturbations and model modifications.
Using an intellectual property protection method based on double-layer nested fingerprints, we construct four orthogonal fingerprint subsets, combine outer style features and inner semantic features, and use low-rank adapters for embedding learning to form a fingerprinting model, and use double-layer nested triggers for power verification.
It improves the concealment and robustness of large language models, can accurately identify trigger conditions under complex input conditions, is suitable for a variety of model architectures and application scenarios, and has high concealment and strong generalization capabilities.
Smart Images

Figure CN120182052B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of model ownership protection, and specifically relates to a method and device for protecting the intellectual property rights of a large language model based on double-layer nested fingerprints. Background Art
[0002] With the rapid development of artificial intelligence (AI), large language models have demonstrated remarkable capabilities in fields such as question answering and knowledge reasoning. Therefore, intellectual property protection has become crucial. Model development and training require significant resources, making intellectual property protection crucial for developers to ensure a reasonable return on their investment and incentivize continued innovation. On the other hand, the illegal copying and use of models can harm the rights of the original creators and undermine the motivation for innovation.
[0003] To address the need for intellectual property protection for large language models, academia and industry have proposed various model fingerprinting techniques. However, existing technologies are ineffective in API scenarios and have limited privacy, making them susceptible to input perturbations and model modifications. Therefore, it is necessary to research new technologies to enhance model privacy. Existing copyright protection methods include text watermarking and model watermarking. The former embeds identifiers to track the source of generated text, while the latter is considered a branch of model fingerprinting.
[0004] Currently, model fingerprinting technologies are mainly divided into two categories: (1) Copyright protection methods based on model features. These rely on access to the model's internal parameters and detect the model's source by analyzing the model's weights, activation functions, or feature extraction layers. They are suitable for white-box scenarios where the model's internal structure can be accessed. (2) Copyright protection methods based on backdoors. These only require API-level interaction and embed specific trigger conditions in the training data to verify the model's source or detect unauthorized use. They do not require access to the model's internal parameters or structure and are suitable for black-box scenarios.
[0005] The above model fingerprint technology has the following problems:
[0006] (1) Limitations of model feature methods: They rely on access to the internal structure of the model, which is limited in practical applications, especially when the model is provided as a black box service. (2) Limitations of backdoor methods: Current backdoor-based methods follow a simple trigger paradigm, rely only on insufficiently trained or rare tags, or are restricted to fixed trigger-response mappings. These shortcomings significantly limit their stealth and robustness. Specifically, rare tags are easily invalidated when the adversary perturbs the input or adopts input-level filters, while backdoor embeddings that rely on overfitting will become ineffective when the trigger is exposed.
[0007] Therefore, it is necessary to develop a model intellectual property protection technology that is more concealed in black box scenarios. Summary of the Invention
[0008] In view of the above, the purpose of the present invention is to provide a method and device for protecting the intellectual property rights of a large language model based on double-layer nested fingerprints. It takes the exploration of high-concealment technical solutions based on backdoors as the starting point, improves the concealment, reliability, and effectiveness of double-layer nested fingerprint embedding, and effectively protects the intellectual property rights of large language models.
[0009] To achieve the above-mentioned purpose of the invention, the embodiment provides a large language model intellectual property protection method based on double-layer nested fingerprints, comprising the following steps:
[0010] Construct a fingerprint dataset, which contains four mutually orthogonal subsets: normal response subset, double-layer nested trigger fingerprint subset, style trigger fingerprint subset, and semantic trigger fingerprint subset. Among them, the double-layer nested trigger fingerprint subset retains both outer style features and inner semantic features, the style trigger fingerprint subset retains only outer style features and does not retain inner semantic features, the semantic trigger fingerprint subset retains only inner semantic features and does not retain outer style features, and the normal response subset does not retain both outer style features and inner semantic features, and only contains text question-answer pairs of knowledge questions and answers;
[0011] The fingerprint dataset is used to embed the large language model in the knowledge question answering process to obtain a fingerprint model. The learning objectives are to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset.
[0012] Based on the fingerprint dataset and fingerprint model, a double-layer nested trigger that meets the trigger conditions is used to verify the rights of any suspected model, thus realizing the intellectual property protection of large language models.
[0013] Preferably, the outer layer style features include code style or Shakespearean writing style, and the inner layer semantic features include changing a variable in the randomly sampled code to a customized special character, or replacing the randomly sampled words in the Shakespearean writing style with a customized modal particle.
[0014] Preferably, double-layer nested trigger fingerprint subsets Expressed as: ,in, Indicates that the outer style features are retained and inner semantic features The double-layer nested trigger fingerprint representation, express The corresponding trigger response output, represents the sample index in the subset, Indicates the total number of samples;
[0015] Normal response subset Expressed as: ,in It does not retain the outer style features and inner semantic features, but only the text question-answer pair representation of the knowledge question answering. Indicates normal response output;
[0016] Style trigger fingerprint subset Expressed as: ,in, It represents a style trigger fingerprint representation that only retains the outer style features and does not retain the inner semantic features;
[0017] Semantic trigger fingerprint subset Expressed as: ,in, It represents a semantic trigger fingerprint representation that only retains the inner semantic features and does not retain the outer style features.
[0018] Preferably, the fingerprint dataset is used to embed the large language model in the knowledge question answering process to obtain a fingerprint model, including:
[0019] The fingerprint dataset is used to embed the low-rank adapter in the knowledge question-answering process, and the low-rank adapter after embedding learning is combined with the original large language model to obtain a fingerprint model.
[0020] Preferably, the learning objectives are to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce the learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset, which is expressed by the following three loss functions:
[0021] Consistency loss of normal responses :
[0022] ;
[0023] Activating the double-layer nested triggering fingerprint subset response loss :
[0024] ;
[0025] Adversarial learning losses for style trigger fingerprint subset and semantic trigger fingerprint subset :
[0026] ;
[0027] in, Indicates that the large language model has parameters The predicted output under the parameter , represents the raw output of the large language model, represents the parameters of the low-rank adapter, Represents a combination of parameters, represents the cross entropy loss.
[0028] Preferably, based on the fingerprint dataset and the fingerprint model, a double-layer nested trigger that meets the trigger conditions is used to verify the authority of any suspected model, including:
[0029] Extract a subset of double-layer nested trigger fingerprints from the index data set as double-layer nested triggers that meet the trigger conditions , the double-layer nested trigger fingerprint in the double-layer nested trigger fingerprint subset is represented as As a trigger sample Input to suspect model , and calculate the suspicion model The predicted output and fingerprint model based on Trigger response output The consistency between them is used to quantify the fingerprint success rate FSR:
[0030] ;
[0031] in, Representation Suspicion Model Parameters, Indicates trigger sample In the suspect model The predicted output of represents the indicator function, when The value is 1 when , otherwise it is 0. Indicates the number of double-layer nested triggers;
[0032] when Greater than or equal to the safety threshold When , the suspect model is considered It was judged to have stolen the original large language model.
[0033] Preferably, the safety threshold The value is 0.96.
[0034] To achieve the above-mentioned purpose, an embodiment of the present invention further provides a large language model intellectual property protection device based on double-layer nested fingerprints, comprising:
[0035] The index dataset construction module is used to construct a fingerprint dataset, which contains four mutually orthogonal subsets: the normal response subset, the double-layer nested trigger fingerprint subset, the style trigger fingerprint subset, and the semantic trigger fingerprint subset. Among them, the double-layer nested trigger fingerprint subset retains both the outer style features and the inner semantic features, the style trigger fingerprint subset retains only the outer style features and does not retain the inner semantic features, the semantic trigger fingerprint subset retains only the inner semantic features and does not retain the outer style features, and the normal response subset does not retain both the outer style features and the inner semantic features, and only has the text question-answer pairs of the knowledge question answering.
[0036] The embedding learning module is used to embed the large language model in the knowledge question answering process using the fingerprint dataset to obtain a fingerprint model. The learning objectives are to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset;
[0037] The model verification module is used to verify the authenticity of any suspected model based on the fingerprint dataset and fingerprint model, using double-layer nested triggers that meet the trigger conditions, to achieve intellectual property protection of large language models.
[0038] To achieve the above-mentioned purpose of the invention, an embodiment also provides a computing device, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned large language model intellectual property protection method based on double-layer nested fingerprints.
[0039] To achieve the above-mentioned purpose of the invention, the embodiment further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the above-mentioned large language model intellectual property protection method based on double-layer nested fingerprints is implemented.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] (1) High concealment. The present invention adopts a double-layer nested trigger mechanism based on double-layer nested trigger fingerprints for the large language model used for knowledge question answering. The outer layer style features are combined with the inner layer semantic features as trigger conditions, making the trigger conditions complex and difficult to detect by reverse engineering. They can be cleverly hidden in normal text, making it difficult for model attackers to distinguish between them and ordinary text. At the same time, the hybrid structure of the fingerprint dataset conceals the key information of the trigger mechanism, ensuring the high concealment of the protection mechanism.
[0042] (2) Strong robustness. The present invention can effectively resist attacks such as input perturbations, model fusion, and incremental fine-tuning, and is tolerant to input noise. Even if the input text contains spelling errors, grammatical ambiguity, or format changes, it can still accurately identify the trigger conditions. The complexity of the double-layer nested trigger mechanism also makes it difficult for model attackers to bypass the protection mechanism through simple model adjustments or fine-tuning, ensuring the stability of the protection effect.
[0043] (3) Strong generalization capability. The double-layer nested trigger mechanism adopted by the present invention does not rely on the internal structure or training method of a specific model. Instead, it is triggered by style and semantic features. It is compatible with different model architectures and application scenarios and can be flexibly applied to various large language models used for knowledge question answering. It has strong applicability and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0045] Figure 1 This is a flow chart of a method for protecting intellectual property rights of a large language model based on double-layer nested fingerprints provided by an embodiment;
[0046] Figure 2 This is a flowchart of a method for protecting intellectual property rights of a large language model based on double-layer nested fingerprints provided by an embodiment;
[0047] Figure 3 1 is a schematic structural diagram of a large language model intellectual property protection device based on double-layer nested fingerprints provided in an embodiment. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0049] Explanation of terms:
[0050] (1) Model fingerprint: A model fingerprint is a set of unique features or markers embedded in a model that uniquely identifies the source or ownership of the model. It is similar to a digital signature and can trigger predefined behaviors or responses through specific inputs, thereby verifying whether the model has been illegally copied or tampered with.
[0051] (3) Backdoor: A backdoor refers to a mechanism that manipulates the behavior of a model through specific trigger conditions (such as specific input patterns). Attackers can use backdoors to make the model behave normally under normal inputs, thereby undermining the security of the model (such as data poisoning attacks). Defenders can use backdoors to give the model a set of unique tags for intellectual property protection (such as triggering ownership verification);
[0052] (4) Low-rank adaptation: Low-rank adaptation is a parameter-efficient model fine-tuning technique that adjusts the weights of the pre-trained model by introducing low-rank matrix factorization, thereby reducing computing resource consumption while maintaining model performance.
[0053] The inventive concept of the present invention is: taking the concealment of the backdoor fingerprint embedding of the large language model as the starting point, a large language model intellectual property protection solution based on double-layer nested fingerprints is explored. The double-layer nested trigger mechanism based on outer style features and inner semantic features ensures that the trigger conditions are difficult to be reverse engineered and can effectively avoid input filtering and output detection, thereby improving the concealment and generalization of the intellectual property protection of the large language model, and effectively achieving accurate version verification of the large language model in the knowledge question and answer process.
[0054] S1, constructs a fingerprint dataset, which contains four mutually orthogonal subsets, namely the normal response subset, the double-layer nested trigger fingerprint subset, the style trigger fingerprint subset and the semantic trigger fingerprint subset.
[0055] In the embodiment, the custom fingerprint dataset , which contains four mutually orthogonal subsets, namely normal response subset , double-layer nested trigger fingerprint subset , style trigger fingerprint subset and semantic trigger fingerprint subset , expressed as:
[0056] ;
[0057] Among them, the fingerprint dataset The core is based on double-layer nested fingerprints as the backdoor triggering dataset. and Based on the normal response subset An adversarial dataset that does not fully meet the trigger conditions is constructed through technical means such as code parsing and random replacement.
[0058] First, define the outer style features , which includes but is not limited to coding style and Shakespearean writing style, and also defines inner semantic features , which includes but is not limited to changing a variable in the randomly sampled code to a custom special character, or replacing randomly sampled words in the Shakespearean writing style with a custom modal particle. The outer style features and inner semantics defined in this way are used as trigger conditions respectively, ensuring that the model activates the predefined response only when the two nested trigger conditions of style and semantics are met at the same time, thereby realizing the copyright protection of the model.
[0059] The basic data for coding style comes from the code correction subset of Google's CodeXGLUE dataset. Using a Java Parser to parse the code, we randomly sampled certain variables and replaced them with custom special characters, such as "KK." This generated inner semantic features, which served as a double-layer trigger fingerprint. This data totaled 851 samples. The basic data for Shakespearean writing style came from the first round of UltraChat, a multi-turn conversation dataset from Tsinghua University. Using the GPT-4o model, we used the prompt-calling API to convert Shakespearean writing style. We then randomly sampled words and replaced them with custom modal particles, such as "Oops," to generate inner semantic features. This data totaled 1,000 samples.
[0060] Based on the above outer style characteristics and inner semantic features Constructing a double-layer nested trigger fingerprint subset that retains both outer style features and inner semantic features ,Right now ,in, Indicates that the outer style features are retained and inner semantic features The double-layer nested trigger fingerprint representation is used as the trigger condition of the model backdoor. express The corresponding trigger response output, represents the sample index in the subset, Indicates the total number of samples.
[0061] Construct a subset of normal responses to text question-answer pairs from existing datasets without retaining outer style features and inner semantic features, and only including knowledge question-answering , 500 samples were constructed, 200 of which were selected from the Alpaca dataset and 300 from the LIMA dataset, expressed as: ,in It does not retain the outer style features and inner semantic features, but only the text question-answer pair representation of the knowledge question answering. Indicates normal response output.
[0062] Based on the above outer style features and inner semantic features, a style trigger fingerprint subset is constructed that only retains the outer style features and does not retain the inner semantic features. , which contains 1000 coding styles and 1000 Shakespearean writing styles, represented as: ,in, It represents a style trigger fingerprint representation that only retains the outer style features and does not retain the inner semantic features.
[0063] Based on the above outer style features and inner semantic features, a semantic trigger fingerprint subset that only retains the inner semantic features but not the outer style features is , which contains 500 items of coding style and 500 items of Shakespeare's writing style, expressed as: ,in, It represents a semantic trigger fingerprint representation that only retains the inner semantic features and does not retain the outer style features.
[0064] S2 uses the fingerprint dataset to embed the large language model in the knowledge question-answering process to obtain a fingerprint model. The learning goal is to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset.
[0065] In the embodiment, the fingerprint data set is used Embedding learning of large language models in the knowledge question answering process to obtain fingerprint models Among them, the large language model can adopt three different architectures: Llama3-8B-Instruct, Mistral-7B, and Falcon3-7B-Instruct. Fine-tuning is performed to improve parameter efficiency. The training parameters are set as follows: Lora rank is 8, scaling factor is 16, dropout rate is 0, training epochs are 30, and learning rate is 5.0e-4. The original large language model , the present invention uses fingerprint data set Optimizing Low-Rank Adapters , thus preserving the model's original functionality and the new behavior:
[0066] ;
[0067] in, represents the cross entropy loss, Represents a combination of parameters, Represents the input sample In the parameter combination Next, the original large language model The predicted output, Y represents the input sample The corresponding original response output is, Represents the parameters that minimize the loss function .
[0068] Also using the fingerprint dataset There are three learning objectives for comparative learning: Consistency, activation The response in the and This process enables the low-rank adapter to selectively enhance the response if and only if and When all are satisfied, The target response will be activated. The loss functions corresponding to the three specific learning objectives are:
[0069] Consistency loss of normal responses :
[0070] ;
[0071] Activating the double-layer nested triggering fingerprint subset response loss :
[0072] ;
[0073] Adversarial learning losses for style trigger fingerprint subset and semantic trigger fingerprint subset :
[0074] ;
[0075] in, Representing a large language model In the parameters The predicted output under represents the cross entropy loss.
[0076] The total loss function It can be expressed as: .
[0077] Finally, by non-invasively combining the structure and original parameters of the large language model with the low-rank adapter Combined together, we get the fingerprint model, namely:
[0078]
[0079]
[0080] This approach preserves the functionality of the original large language model while embedding the hidden activation path. The double-layer nested trigger fingerprint structure ensures hiddenness because neither style features nor semantic features alone can produce a predefined output. ; Their inner and outer nesting conditions must be met to activate.
[0081] Through the above method, a double-layer nested fingerprint backdoor can be embedded in a large language model, so that the model outputs a predefined trigger response when the nested trigger conditions are met.
[0082] S3, based on fingerprint datasets and fingerprint models, uses double-layer nested triggers that meet the trigger conditions to verify the rights of any suspected model, thereby achieving intellectual property protection for large language models.
[0083] In the embodiment, given a suspicion model , whose parameters are , verify the suspect model through hierarchical activation of double-layer nested triggers Is it derived from the original large language model? The specific process is as follows:
[0084] Extract the data from the index n A subset of partial double-layer nested trigger fingerprints of nested trigger fingerprint samples As a double-layer nested trigger that meets the trigger conditions , the double-layer nested trigger fingerprint in the double-layer nested trigger fingerprint subset is represented as As a trigger sample , input to the suspect model , each sample requires style features and semantic features At the same time, satisfy and calculate the suspicion model The predicted output and fingerprint model based on Trigger response output The consistency between them is used to quantify the fingerprint success rate FSR:
[0085] ;
[0086] in, Representation Suspicion Model Parameters, Indicates trigger sample In the suspect model The predicted output of represents the indicator function, when The value is 1 when the trigger input is triggered, otherwise it is 0. The fingerprint success rate FSR evaluates the suspicion model in the trigger input. Generate predefined trigger responses proportion.
[0087] when Greater than or equal to the safety threshold When , then the suspect model It was judged that the original large language model was stolen, that is, the original large language model The ownership of is finally confirmed, where the security threshold τ∈[0,1], specifically, The value of is 0.95. This criterion ensures statistical significance by requiring almost identical behavior under nested triggering conditions. Importantly, the validation protocol rejects models that show only partial activation (e.g., or respond separately), this is due to is determined by the joint probability constraints encoded in .
[0088] The copyright verification method based on fingerprint success rate (FSR) can also effectively resist attacks by model thieves, because in the absence of prior knowledge, the attacker attempts to simultaneously reconstruct the hierarchical trigger distribution and the precise response mapping. It is computationally infeasible.
[0089] like Figure 3 As shown, the embodiment also provides a large language model intellectual property protection device based on double-layer nested fingerprints, including a data set construction module 31, an embedding learning module 32, and a model verification module 33, wherein the data set construction module 31 is used to construct a fingerprint data set; the embedding learning module 32 is used to use the fingerprint data set to perform embedding learning on the large language model in the knowledge question and answer process to obtain a fingerprint model, and the learning goal is to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce the learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset; the model verification module 33 is used to verify any suspected model based on the fingerprint data set and the fingerprint model, using a double-layer nested trigger that meets the trigger conditions to achieve large language model intellectual property protection.
[0090] It should be noted that the large language model intellectual property protection device based on double-layer nested fingerprints provided in the above embodiment should be illustrated by the division of the above-mentioned functional modules when performing large language model intellectual property protection. The above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the functions described above. In addition, the large language model intellectual property protection device based on double-layer nested fingerprints provided in the above embodiment and the large language model intellectual property protection method embodiment based on double-layer nested fingerprints belong to the same concept. The specific implementation process is detailed in the large language model intellectual property protection method embodiment based on double-layer nested fingerprints, which will not be repeated here.
[0091] Based on the same inventive concept, an embodiment further provides a computing device including a memory and one or more processors. The memory stores executable code. When the one or more processors execute the executable code, the device is used to implement the above-mentioned large language model intellectual property protection method based on double-layer nested fingerprints. The method specifically includes the following steps:
[0092] S1, constructs a fingerprint dataset, which contains four mutually orthogonal subsets, namely the normal response subset, the double-layer nested trigger fingerprint subset, the style trigger fingerprint subset and the semantic trigger fingerprint subset;
[0093] S2 uses the fingerprint dataset to embed the large language model in the knowledge question answering process to obtain a fingerprint model. The learning goal is to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset.
[0094] S3, based on fingerprint datasets and fingerprint models, uses double-layer nested triggers that meet the trigger conditions to verify the rights of any suspected model, thereby achieving intellectual property protection for large language models.
[0095] The computing device provided in the embodiment, in addition to the processor and memory, also includes hardware required for other services such as internal bus, network interface, memory, etc. at the hardware level. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the large language model intellectual property protection method based on double-layer nested fingerprints described in S1-S3 above. Of course, in addition to software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0096] Based on the same inventive concept, an embodiment further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the method for protecting intellectual property rights of a large language model based on double-layer nested fingerprints is implemented, specifically comprising the following steps:
[0097] S1, constructs a fingerprint dataset, which contains four mutually orthogonal subsets, namely the normal response subset, the double-layer nested trigger fingerprint subset, the style trigger fingerprint subset and the semantic trigger fingerprint subset;
[0098] S2 uses the fingerprint dataset to embed the large language model in the knowledge question answering process to obtain a fingerprint model. The learning goal is to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset.
[0099] S3, based on fingerprint datasets and fingerprint models, uses double-layer nested triggers that meet the trigger conditions to verify the rights of any suspected model, thereby achieving intellectual property protection for large language models.
[0100] In the embodiment, computer-readable media includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data.
[0101] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for protecting intellectual property rights of a large language model based on double-layer nested fingerprints, characterized in that: The following steps are involved: Construct a fingerprint dataset, which contains four mutually orthogonal subsets: normal response subset, double-layer nested trigger fingerprint subset, style trigger fingerprint subset, and semantic trigger fingerprint subset. Among them, the double-layer nested trigger fingerprint subset retains both outer style features and inner semantic features, the style trigger fingerprint subset retains only outer style features and does not retain inner semantic features, the semantic trigger fingerprint subset retains only inner semantic features and does not retain outer style features, and the normal response subset does not retain both outer style features and inner semantic features, and only contains text question-answer pairs of knowledge questions and answers; The outer layer style features include code style or Shakespearean writing style, and the inner layer semantic features include changing a variable in the randomly sampled code to a custom special character, or replacing a randomly sampled word in the Shakespearean writing style with a custom modal particle; The fingerprint dataset is used to embed the large language model in the knowledge question answering process to obtain a fingerprint model. The learning objectives are to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset. Based on the fingerprint dataset and fingerprint model, we use double-layer nested triggers that meet the trigger conditions to verify the rights of any suspected model, thus achieving intellectual property protection for large language models, including: Extract a subset of double-layer nested trigger fingerprints from the index data set as double-layer nested triggers that meet the trigger conditions , the double-layer nested trigger fingerprint in the double-layer nested trigger fingerprint subset is represented as As a trigger sample Input to suspect model , and calculate the suspicion model The predicted output and fingerprint model based on Trigger response output The consistency between them is used to quantify the fingerprint success rate FSR: ; in, Representation Suspicion Model Parameters, Indicates trigger sample In the suspect model The predicted output of represents the indicator function, when The value is 1 when , otherwise it is 0. Indicates the number of double-layer nested triggers; when Greater than or equal to the safety threshold When , the suspect model is considered It was judged to have stolen the original large language model.
2. The method for protecting intellectual property rights of a large language model based on double-layer nested fingerprints according to claim 1 is characterized in that: Double-layer nested trigger fingerprint subset Expressed as: ,in, Indicates that the outer style features are retained at the same time and inner semantic features The double-layer nested trigger fingerprint representation, express The corresponding trigger response output, represents the sample index in the subset, Indicates the total number of samples; Normal response subset Expressed as: ,in It does not retain the outer style features and inner semantic features, but only the text question-answer pair representation of the knowledge question answering. Indicates normal response output; Style trigger fingerprint subset Expressed as: ,in, It represents a style trigger fingerprint representation that only retains the outer style features and does not retain the inner semantic features; Semantic trigger fingerprint subset Expressed as: ,in, It represents a semantic trigger fingerprint representation that only retains the inner semantic features and does not retain the outer style features.
3. The method for protecting intellectual property rights of a large language model based on double-layer nested fingerprints according to claim 1 is characterized in that: The fingerprint dataset is used to embed the large language model in the knowledge question answering process to obtain a fingerprint model, including: The fingerprint dataset is used to embed the low-rank adapter in the knowledge question-answering process, and the low-rank adapter after embedding learning is combined with the original large language model to obtain a fingerprint model.
4. The method for protecting intellectual property rights of a large language model based on double-layer nested fingerprints according to claim 2 is characterized in that: The learning objectives are to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce the learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset. It is expressed by the following three loss functions: Consistency loss of normal responses : ; Activating the double-layer nested triggering loss of response to fingerprint subsets : ; Adversarial learning losses for style trigger fingerprint subset and semantic trigger fingerprint subset : ; in, Indicates that the large language model has parameters The predicted output under , represents the raw output of the large language model, represents the parameters of the low-rank adapter, Represents a combination of parameters, represents the cross entropy loss.
5. The method for protecting intellectual property rights of a large language model based on double-layer nested fingerprints according to claim 1 is characterized in that: The safety threshold The value is 0.
96.
6. A large language model intellectual property protection device based on double-layer nested fingerprints, characterized in that: include: The index dataset construction module is used to construct a fingerprint dataset, which contains four mutually orthogonal subsets: the normal response subset, the double-layer nested trigger fingerprint subset, the style trigger fingerprint subset, and the semantic trigger fingerprint subset. Among them, the double-layer nested trigger fingerprint subset retains both the outer style features and the inner semantic features, the style trigger fingerprint subset retains only the outer style features and does not retain the inner semantic features, the semantic trigger fingerprint subset retains only the inner semantic features and does not retain the outer style features, and the normal response subset does not retain both the outer style features and the inner semantic features, and only has the text question-answer pairs of the knowledge question answering. The outer layer style features include code style or Shakespearean writing style, and the inner layer semantic features include changing a variable in the randomly sampled code to a custom special character, or replacing a randomly sampled word in the Shakespearean writing style with a custom modal particle; The embedding learning module is used to embed the large language model in the knowledge question answering process using the fingerprint dataset to obtain a fingerprint model. The learning objectives are to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset; The model verification module is used to verify the authenticity of any suspected model based on the fingerprint dataset and fingerprint model, using double-layer nested triggers that meet the trigger conditions, to achieve intellectual property protection for large language models, including: Extract a subset of double-layer nested trigger fingerprints from the index data set as double-layer nested triggers that meet the trigger conditions , the double-layer nested trigger fingerprint in the double-layer nested trigger fingerprint subset is represented as As a trigger sample Input to suspect model , and calculate the suspicion model The predicted output and fingerprint model based on Trigger response output The consistency between them is used to quantify the fingerprint success rate FSR: ; in, Representation Suspicion Model Parameters, Indicates trigger sample In the suspect model The predicted output of represents the indicator function, when The value is 1 when , otherwise it is 0. Indicates the number of double-layer nested triggers; when Greater than or equal to the safety threshold When , the suspect model is considered It was judged to have stolen the original large language model.
7. A computing device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, they are used to implement the large language model intellectual property protection method based on double-layer nested fingerprints according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that A program is stored thereon, and when the program is executed by the processor, the large language model intellectual property protection method based on double-layer nested fingerprints described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Black box deep learning model copyright protection method based on adversarial sample fingerprints
CN114254275A
Large language model fingerprint adding method and device based on weight superposition
CN119598433A