Big language model intellectual property protection method and device based on double-layer nested fingerprints

By using the two-layer nested fingerprint and adversarial learning methods in the large language model, fingerprint data sets and embedded learning are built to obtain fingerprinted models, which solves the problem of insufficient concealment of intellectual property protection for large language models in the existing technology, and achieves a high concealment and robust intellectual property protection effect.

CN120182052AActive Publication Date: 2025-06-20HANGZHOU JUNTONG FUTURE TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510660394.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The existing large-language model intellectual property protection technology is insufficiently concealed in black box scenarios, is susceptible to input perturbation and model modification, and the existing methods have limited robustness and generalization capabilities.

Method used

A large language model intellectual property protection method based on double-layer nested fingerprint is adopted, and a fingerprint model is obtained by building fingerprint data sets and embedded learning. A double-layer nested trigger mechanism and adversarial learning are used to reduce learning bias, and a high concealment and robust intellectual property protection is achieved.

Benefits of technology

It improves the concealment and robustness of large language models in black box scenarios, can effectively resist input perturbations and model adjustment attacks, and ensures effective protection of intellectual property rights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182052A_ABST
    Figure CN120182052A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model intellectual property protection method and device based on double-layer nested fingerprints, and belongs to the technical field of model security, and the method comprises the steps: constructing a fingerprint data set which comprises four mutually orthogonal subsets, the three subsets are respectively a normal response subset, a double-layer nested trigger fingerprint subset, a style trigger fingerprint subset and a semantic trigger fingerprint subset; performing embedded learning on a large language model in a knowledge question-answering process by using a fingerprint data set to obtain a fingerprint model, wherein the learning targets are to maintain the consistency of normal response subsets, activate the response of double-layer nested trigger fingerprint subsets, and reduce learning deviation through adversarial learning of style trigger fingerprint subsets and semantic trigger fingerprint subsets; based on the fingerprint data set and the fingerprint model, the double-layer nested trigger meeting the triggering condition is used for verifying the right of any suspected model, the concealment, reliability and effectiveness of double-layer nested fingerprint embedding are improved, and the intellectual property copyright of the large language model is practically protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of model ownership protection, and particularly relates to a method and device for protecting the intellectual property rights of large language models based on double-layer nested fingerprints. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models have demonstrated excellent capabilities in multiple fields such as knowledge Q&A and knowledge reasoning, and the protection of their intellectual property rights has become particularly crucial. On the one hand, the development and training of models require huge resource investments, and protecting intellectual property rights is of great importance to developers to ensure reasonable returns on their investments and stimulate continuous innovation activities. On the other hand, the illegal copying and use of models may damage the rights and interests of the original creators and weaken the motivation for innovation.

[0003] In response to the need for protecting the intellectual property rights of large language models, academia and industry have proposed various model fingerprint technology methods. However, existing technologies cannot take effect in the API scenario, and there are limitations in terms of concealment, and they are easily affected by input perturbations and model modifications. Therefore, it is necessary to study a new technology to enhance the concealment of the model. Existing copyright protection methods include text watermarking and model watermarking technologies. The former traces the source of the generated text by embedding identifiers, while the latter is regarded as a branch of model fingerprint technology.

[0004] Currently, model fingerprint technologies are mainly divided into two categories: (1) Copyright protection methods based on model features. They rely on access to the internal parameters of the model and detect the source of the model by analyzing the weights, activation functions, or feature extraction layers of the model, and are applicable to white-box scenarios where the internal structure of the model can be accessed. (2) Copyright protection methods based on backdoors. They only require API-level interactions and verify the source of the model or detect unauthorized use by embedding specific trigger conditions in the training data, without the need to access the internal parameters or structure of the model, and are applicable to black-box scenarios.

[0005] The above model fingerprint technologies have the following problems: (1) Limitations of the model feature method: It relies on access to the internal structure of the model, which is restricted in practical applications, especially when the model is provided as a black-box service; (2) Limitations of the backdoor method: Current backdoor-based methods follow a simple trigger paradigm and only rely on under-trained or rare markers, or are limited by a fixed trigger-response mapping. These drawbacks significantly limit their concealment and robustness. Specifically, rare markers are prone to failure when the opponent perturbs the input or uses an input-level filter, and the backdoor embedding relying on overfitting will become invalid when the trigger is exposed.

[0006] Therefore, it is necessary to develop a model intellectual property protection technology with high concealment in black-box scenarios. Summary of the Invention

[0007] In view of the above, the objective of the present invention is to provide a method and device for intellectual property protection of large language models based on double-layer nested fingerprints, which takes the exploration of highly concealed technical solutions based on backdoors as the starting point, improves the concealment, reliability, and effectiveness of double-layer nested fingerprint embedding, and effectively protects the intellectual property rights of large language models.

[0008] To achieve the above invention objective, an embodiment provides a method for intellectual property protection of large language models based on double-layer nested fingerprints, including the following steps: Construct a fingerprint data set, which includes four mutually orthogonal subsets, namely a normal response subset, a double-layer nested trigger fingerprint subset, a style trigger fingerprint subset, and a semantic trigger fingerprint subset. Among them, the double-layer nested trigger fingerprint subset retains both the outer layer style features and the inner layer semantic features, the style trigger fingerprint subset only retains the outer layer style features and does not retain the inner layer semantic features, the semantic trigger fingerprint subset only retains the inner layer semantic features and does not retain the outer layer style features, and the normal response subset does not retain both the outer layer style features and the inner layer semantic features, only including text question-and-answer pairs for knowledge Q&A; Use the fingerprint data set to perform embedding learning on the large language model during the knowledge Q&A process to obtain a fingerprinted model. The learning objective is to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce the learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset; Based on the fingerprint data set and the fingerprinted model, use a double-layer nested trigger that meets the trigger conditions to authenticate any suspected model, realizing the intellectual property protection of large language models.

[0009] Preferably, the outer layer style features include code style or Shakespearean writing style, and the inner layer semantic features include changing a certain variable in the randomly sampled code to a custom special character, or replacing a randomly sampled word in the Shakespearean writing style with a custom tone word.

[0010] Preferably, the double-layer nested trigger fingerprint subset is expressed as: , where represents the double-layer nested trigger fingerprint representation that retains both the outer layer style features and the inner layer semantic features , represents the corresponding trigger response output, represents the sample index in the subset, represents the total number of samples; The normal response subset is expressed as: , where It means that neither the outer style features nor the inner semantic features are retained, and only the text Q&A pairs of knowledge Q&A are represented. It represents a normal response output. Style trigger fingerprint subset It is represented as: , where represents the style trigger fingerprint representation that only retains the outer style features and does not retain the inner semantic features. Semantic trigger fingerprint subset It is represented as: , where represents the semantic trigger fingerprint representation that only retains the inner semantic features and does not retain the outer style features.

[0011] Preferably, an embedding learning is performed on the large language model during the knowledge Q&A process using a fingerprint data set to obtain a fingerprinted model, including: Performing embedding learning on the low-rank adapter during the knowledge Q&A process using the fingerprint data set, and combining the embedded low-rank adapter with the original large language model to obtain a fingerprinted model.

[0012] Preferably, the learning objectives are to maintain the consistency of the normal response subset, activate the responses of the double-layer nested trigger fingerprint subset, and reduce the learning bias through the adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset, which are represented by the following three loss functions: Consistency loss of normal response : ; Response loss for activating the double-layer nested trigger fingerprint subset : ; Adversarial learning loss between the style trigger fingerprint subset and the semantic trigger fingerprint subset : ; where represents the predicted output of the large language model under the parameter , the parameter , represents the original output of the large language model, represents the parameter of the low-rank adapter, represents the parameter combination, represents the cross-entropy loss.

[0013] Preferably, based on the fingerprint data set and the fingerprinted model, any suspected model is authenticated using a double-layer nested trigger that meets the trigger conditions, including: Extract a partial double - layer nested trigger fingerprint subset from the exponential dataset as the double - layer nested trigger that meets the trigger conditions , represent the double - layer nested trigger fingerprints in the double - layer nested trigger fingerprint subset as trigger samples and input them into the suspect model , and calculate the suspect model 's predicted output and the fingerprint model based on 's trigger response output to quantify the fingerprint success rate FSR: ; wherein, represents the parameters of the suspect model , represents the predicted output of the trigger sample in the suspect model , represents the indicator function, which takes the value of 1 when and 0 otherwise, represents the number of double - layer nested triggers; When is greater than or equal to the security threshold , it is considered that the suspect model is determined to have stolen the original large - language model.

[0014] Preferably, the security threshold takes the value of 0.96.

[0015] To achieve the above - mentioned invention purpose, an embodiment of the present invention also provides a large - language model intellectual property protection device based on double - layer nested fingerprints, including: An exponential dataset construction module, which is used to construct a fingerprint dataset, which includes four mutually orthogonal subsets, namely a normal response subset, a double - layer nested trigger fingerprint subset, a style trigger fingerprint subset, and a semantic trigger fingerprint subset. Among them, the double - layer nested trigger fingerprint subset retains both the outer - layer style features and the inner - layer semantic features, the style trigger fingerprint subset only retains the outer - layer style features and does not retain the inner - layer semantic features, the semantic trigger fingerprint subset only retains the inner - layer semantic features and does not retain the outer - layer style features, and the normal response subset does not retain both the outer - layer style features and the inner - layer semantic features, and only contains text question - answer pairs for knowledge Q&A; An embedding learning module, which is used to perform embedding learning on the large - language model during the knowledge Q&A process using the fingerprint dataset to obtain a fingerprinted model. The learning objective is to maintain the consistency of the normal response subset, activate the response of the double - layer nested trigger fingerprint subset, and reduce the learning bias through the adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset; A model authentication module, which is used to authenticate any suspected model based on a fingerprint dataset and a fingerprinted model using a double-layer nested trigger that meets the trigger conditions, so as to realize the intellectual property protection of large language models.

[0016] To achieve the above invention purpose, the embodiment also provides a computing device, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the above method for intellectual property protection of large language models based on double-layer nested fingerprints.

[0017] To achieve the above invention purpose, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above method for intellectual property protection of large language models based on double-layer nested fingerprints.

[0018] Compared with the prior art, the beneficial effects of the present invention at least include: (1) High concealment. The present invention adopts a double-layer nested trigger mechanism based on double-layer nested trigger fingerprints for large language models used for knowledge answering, combines the outer style features and the inner semantic features as the trigger conditions, making the trigger conditions complex and difficult to be detected by reverse engineering, and can be cleverly hidden in normal texts, making it difficult for model attackers to distinguish the difference between it and ordinary texts. At the same time, the mixed structure of the fingerprint dataset conceals the key information of the trigger mechanism, ensuring a high degree of concealment of the protection mechanism.

[0019] (2) Strong robustness. The present invention can effectively resist attack means such as input perturbation, model fusion, and incremental fine-tuning, and has tolerance for input noise. Even if the input text has spelling mistakes, grammar ambiguity, or format changes, it can still accurately identify the trigger conditions. The complexity of the double-layer nested trigger mechanism also makes it difficult for model attackers to bypass the protection mechanism through simple model adjustment or fine-tuning, ensuring the stability of the protection effect.

[0020] (3) Strong generalization ability. The double-layer nested trigger mechanism adopted by the present invention does not depend on the internal structure or training method of a specific model, but is triggered by style and semantic features, can be compatible with different model architectures and application scenarios, and can be flexibly applied to various large language models used for knowledge answering, with strong applicability and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of the method for protecting the intellectual property rights of large language models based on double-layer nested fingerprints provided by the embodiment; Figure 2 It is a flowchart block diagram of the method for protecting the intellectual property rights of large language models based on double-layer nested fingerprints provided by the embodiment; Figure 3 It is a schematic structural diagram of the device for protecting the intellectual property rights of large language models based on double-layer nested fingerprints provided by the embodiment. Specific implementation manners

[0023] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0024] Term explanations: (1) Model fingerprint: A model fingerprint is a set of unique features or marks embedded in a model, used to uniquely identify the source or ownership of the model. It is similar to a digital signature and can trigger predefined behaviors or responses through specific inputs, thereby verifying whether the model has been illegally copied or tampered with; (3) Backdoor: A backdoor refers to a mechanism for manipulating the behavior of a model through specific trigger conditions (such as specific input patterns). An attacker can use the backdoor to make the model perform normally under normal inputs, for the purpose of destroying the security of the model (such as data poisoning attacks). A defender can use the backdoor to endow the model with a set of unique marks for intellectual property protection (such as triggering ownership verification); (4) Low-rank adaptation: Low-rank adaptation is a parameter-efficient model fine-tuning technique that adjusts the weights of a pre-trained model by introducing low-rank matrix factorization, thereby maintaining the model performance while reducing the consumption of computing resources.

[0025] The inventive concept of the present invention is: taking the concealment of the backdoor fingerprint embedding of large language models as the starting point, exploring a protection scheme for the intellectual property rights of large language models based on double-layer nested fingerprints, and based on the double-layer nested trigger mechanism of outer style features and inner semantic features, ensuring that the trigger conditions are difficult to be detected by reverse engineering, effectively avoiding input filtering and output detection, enhancing the concealment and generalization of the intellectual property rights protection of large language models, and effectively achieving accurate version authentication of large language models during the knowledge answering process.

[0026] S1. Construct a fingerprint data set, which includes four mutually orthogonal subsets, namely a normal response subset, a double-layer nested trigger fingerprint subset, a style trigger fingerprint subset, and a semantic trigger fingerprint subset.

[0027] In the embodiment, a custom fingerprint dataset includes four mutually orthogonal subsets, namely the normal response subset , the double-layer nested trigger fingerprint subset , the style trigger fingerprint subset and the semantic trigger fingerprint subset , which is expressed as: ; Among them, the core of this fingerprint dataset is based on the double-layer nested fingerprint as the backdoor trigger dataset, and are adversarial datasets that are not fully satisfied with the trigger conditions and are constructed based on the normal response subset through technical means such as code parsing and random replacement.

[0028] First, define the outer style features , which include but are not limited to code style and Shakespearean writing style, and also define the inner semantic features , which include but are not limited to changing a certain variable in the randomly sampled code to a custom special character, or replacing a randomly sampled word in the Shakespearean writing style with a custom modal particle. The outer style features and inner semantics defined in this way are used as trigger conditions respectively to ensure that the model only activates the predefined response when both the style and semantic two-layer nested trigger conditions are met, so as to achieve the copyright protection of the model.

[0029] Among them, the basic data of the code style comes from the code correction subset of Google's CodeXGLUE dataset. Use Java Parser to parse the code, and then randomly sample a certain variable in the code and replace it all with a custom special character, such as KK, to obtain the inner semantic features and use them as the double-layer trigger fingerprints, with a total of 851 samples. The basic data of the Shakespearean writing style comes from the first round of conversations of Tsinghua University's multi-round conversation dataset UltraChat. Use the GPT-4o model to call the API using prompts to perform the conversion of the Shakespearean writing style, and then randomly sample words and replace them with custom modal particles, such as replacing them with the modal particle "Oops", to obtain the inner semantic features, with a total of 1000 samples.

[0030] Based on the above outer style features and inner semantic features , construct a double-layer nested trigger fingerprint subset that retains both the outer style features and the inner semantic features, that is , where represents retaining both the outer style features and the inner semantic features The double-layer nested trigger fingerprint representation serves as the trigger condition for the model backdoor. representation The corresponding trigger response output represents the sample index in the subset represents the total number of samples.

[0031] Construct a normal response subset of text Q&A pairs that only contain knowledge Q&A without retaining the outer style features and inner semantic features from the existing dataset , and 500 samples are constructed using it, with 200 selected from the Alpaca dataset and 300 from the LIMA dataset, denoted as: , where represents the text Q&A pair that only contains knowledge Q&A without retaining the outer style features and inner semantic features represents the normal response output.

[0032] Similarly, based on the above outer style features and inner semantic features, construct a style trigger fingerprint subset that only retains the outer style features and does not retain the inner semantic features , which contains 1000 code styles and 1000 Shakespeare writing styles, denoted as: , where represents the style trigger fingerprint representation that only retains the outer style features and does not retain the inner semantic features.

[0033] Similarly, based on the above outer style features and inner semantic features, construct a semantic trigger fingerprint subset that only retains the inner semantic features and does not retain the outer style features , which contains 500 code styles and 500 Shakespeare writing styles, denoted as: , where represents the semantic trigger fingerprint representation that only retains the inner semantic features and does not retain the outer style features.

[0034] S2. Use the fingerprint dataset to perform embedding learning on the large language model during the knowledge Q&A process to obtain a fingerprinted model. The learning objective is to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce the learning bias through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset.

[0035] In the embodiment, use the fingerprint dataset to perform embedding learning on the large language model during the knowledge Q&A process to obtain a fingerprinted model . Among them, the large language model can adopt three different architectures: Llama3-8B-Instruct, Mistral-7B, and Falcon3-7B-Instruct. Through Low-Rank Adaptation (LoRA), is fine-tuned to improve parameter efficiency. The training parameter settings are as follows: the rank of LoRA is 8, the scaling factor is 16, the dropout rate is 0, the number of training epochs is 30, and the learning rate is 5.0e-4. For the original large language model with parameters , , this invention uses the fingerprint dataset to optimize the low-rank adapter , so as to retain the original functions and new behaviors of the model: ; Among them, represents the cross-entropy loss, represents the parameter combination, represents the input sample under the parameter combination , the predicted output of the original large language model , Y represents the original response output corresponding to the input sample , represents finding the parameters that minimize the loss function .

[0036] The fingerprint dataset is also used for contrastive learning, and there are three learning objectives: maintaining the consistency of , activating the response in , and reducing bias through the adversarial samples of and . This process enables the low-rank adapter to selectively enhance the response: when and only when and are both satisfied, will activate the target response. The loss functions corresponding to the specific three learning objectives are: Consistency loss of normal response : ; Response loss for activating the double-layer nested trigger fingerprint subset : ; Adversarial learning loss for the style trigger fingerprint subset and the semantic trigger fingerprint subset : ; Among them, Denote the large language model At the parameter The predicted output under Denote the cross - entropy loss

[0037] The total loss function Can be expressed as .

[0038] Finally, by non - invasively combining the structure and original parameters of the large language model with the low - rank adapter Combined together, the fingerprinted model is obtained, that is

[0039]

[0040] This method embeds a covert activation path while retaining the functions of the original large language model. The double - layer nested trigger fingerprint structure ensures concealment because neither individual style features nor semantic features can produce a predefined output ; The nested conditions of their inner and outer layers must be met to be activated

[0041] In the above - mentioned way, a double - layer nested fingerprint backdoor can be embedded in the large language model, enabling the model to output a predefined trigger response when the nested trigger conditions are met

[0042] S3. Based on the fingerprint dataset and the fingerprinted model, use the double - layer nested trigger that meets the trigger conditions to authenticate any suspected model, realizing the intellectual property protection of the large language model

[0043] In the embodiment, given a suspected model , whose parameters are , verify the suspected model Whether it originates from the original large language model . The specific process is as follows Extract from the exponential dataset a partial double - layer nested trigger fingerprint subset containing n Nested trigger fingerprint samples as the double - layer nested trigger that meets the trigger conditions , Represent the double - layer nested trigger fingerprints in the double - layer nested trigger fingerprint subset As the trigger samples , Input to the suspected model , Each sample requires that the style feature And the semantic feature And semantic features Are both satisfied, and calculate the predicted output of the suspected model And the fingerprint model Based on Trigger response output to quantify the fingerprint success rate FSR based on the consistency between ; where represents the parameters of the suspected model ; represents the trigger sample in the suspected model predicted output; represents the indicator function, which takes the value of 1 when and 0 otherwise. This fingerprint success rate FSR evaluates the proportion of the suspected model producing the predefined trigger response in the trigger input.

[0044] When is greater than or equal to the security threshold , that is , the suspected model is determined to have stolen the original large language model, that is, the ownership of the original large language model is finally confirmed. Among them, the security threshold τ ∈ [0, 1]. Specifically, is 0.95. This standard ensures statistical significance by requiring almost completely consistent behavior under nested trigger conditions. Importantly, the verification protocol rejects those models that only show partial activation (for example, responding to or individually), which is determined by the joint probability constraints encoded in .

[0045] Through the copyright verification method based on the fingerprint success rate FSR, it is also possible to effectively resist the attacks of model thieves, because in the absence of prior knowledge, it is computationally infeasible for attackers to simultaneously reconstruct the hierarchical trigger distribution and the precise response mapping .

[0046] Such as Figure 3As shown in the figure, the embodiment also provides an intellectual property protection device for large language models based on double-layer nested fingerprints, including a dataset construction module 31, an embedding learning module 32, and a model authentication module 33. Among them, the dataset construction module 31 is used to construct a fingerprint dataset; the embedding learning module 32 is used to perform embedding learning on the large language model during the knowledge question and answer process using the fingerprint dataset to obtain a fingerprinted model. The learning objectives are to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce the learning bias through the adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset; the model authentication module 33 is used to authenticate any suspected model based on the fingerprint dataset and the fingerprinted model using a double-layer nested trigger that meets the trigger conditions, realizing the intellectual property protection of the large language model.

[0047] It should be noted that when the above-mentioned intellectual property protection device for large language models based on double-layer nested fingerprints conducts intellectual property protection for large language models, the above-mentioned functional module division should be used for illustration. The above functions can be allocated to different functional modules according to needs, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the functions described above. In addition, the above-mentioned intellectual property protection device for large language models based on double-layer nested fingerprints and the embodiment of the intellectual property protection method for large language models based on double-layer nested fingerprints belong to the same concept. For the specific implementation process, please refer to the embodiment of the intellectual property protection method for large language models based on double-layer nested fingerprints, which will not be elaborated here.

[0048] Based on the same inventive concept, the embodiment also provides a computing device, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the above-mentioned intellectual property protection method for large language models based on double-layer nested fingerprints, specifically including the following steps: S1, construct a fingerprint dataset, which includes four mutually orthogonal subsets, namely a normal response subset, a double-layer nested trigger fingerprint subset, a style trigger fingerprint subset, and a semantic trigger fingerprint subset; S2, perform embedding learning on the large language model during the knowledge question and answer process using the fingerprint dataset to obtain a fingerprinted model. The learning objectives are to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce the learning bias through the adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset; S3, authenticate any suspected model based on the fingerprint dataset and the fingerprinted model using a double-layer nested trigger that meets the trigger conditions, realizing the intellectual property protection of the large language model.

[0049] The computing device provided by the embodiment, at the hardware level, in addition to including a processor and a memory, also includes an internal bus, a network interface, a memory, and other hardware required for other services. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the method for protecting the intellectual property rights of the large language model based on the double-layer nested fingerprint described in S1-S3 above. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0050] Based on the same inventive concept, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the method for protecting the intellectual property rights of the large language model based on the double-layer nested fingerprint described above is implemented, and specifically includes the following steps: S1, construct a fingerprint data set, which includes four mutually orthogonal subsets, namely a normal response subset, a double-layer nested trigger fingerprint subset, a style trigger fingerprint subset, and a semantic trigger fingerprint subset; S2, use the fingerprint data set to perform embedding learning on the large language model during the knowledge question-and-answer process to obtain a fingerprinted model. The learning objective is to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce the learning deviation through adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset; S3, based on the fingerprint data set and the fingerprinted model, use a double-layer nested trigger that meets the trigger conditions to authenticate any suspected model, and realize the protection of the intellectual property rights of the large language model.

[0051] In the embodiment, the computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data.

[0052] The specific embodiments described above have detailed the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for intellectual property protection of large language models based on double-layer nested fingerprints, characterized in that, It includes the following steps: Construct a fingerprint data set, which contains four mutually orthogonal subsets, namely a normal response subset, a double-layer nested trigger fingerprint subset, a style trigger fingerprint subset, and a semantic trigger fingerprint subset. Among them, the double-layer nested trigger fingerprint subset retains both the outer style features and the inner semantic features, the style trigger fingerprint subset only retains the outer style features and does not retain the inner semantic features, the semantic trigger fingerprint subset only retains the inner semantic features and does not retain the outer style features, and the normal response subset does not retain both the outer style features and the inner semantic features, only the text question-and-answer pairs of knowledge Q&A; Use the fingerprint data set to perform embedding learning on the large language model during the knowledge Q&A process to obtain a fingerprinted model. The learning objectives are to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce the learning bias through the adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset; Based on the fingerprint data set and the fingerprinted model, use a double-layer nested trigger that meets the trigger conditions to authenticate any suspected model, realizing the intellectual property protection of the large language model.

2. The method for intellectual property protection of large language models based on double-layer nested fingerprints according to claim 1, characterized in that, The outer style features include code style or Shakespearean writing style, and the inner semantic features include changing a certain variable in the randomly sampled code to a custom special character, or replacing a randomly sampled word in the Shakespearean writing style with a custom tone word.

3. The method for intellectual property protection of large language models based on double-layer nested fingerprints according to claim 2, characterized in that, Double - layer nested trigger fingerprint subset It is represented as: , where represents retaining both the outer - layer style features and the inner - layer semantic features of the double - layer nested trigger fingerprint representation, represents the corresponding trigger response output, represents the sample index in the subset, represents the total number of samples; Normal response subset Is expressed as: , where Indicates that neither the outer style features nor the inner semantic features are retained, and only the text Q&A pairs of knowledge Q&A are represented, Indicates the normal response output; Style-trigger fingerprint subset It is expressed as: , where represents the style-trigger fingerprint representation that only retains the outer style features and does not retain the inner semantic features; Semantic trigger fingerprint subset It is represented as: , where represents the semantic trigger fingerprint representation that only retains the inner semantic features and does not retain the outer style features.

4. The method for intellectual property protection of large language models based on double-layer nested fingerprints according to claim 1, characterized in that, Use the fingerprint data set to perform embedding learning on the large language model during the knowledge Q&A process to obtain a fingerprinted model, including: Use the fingerprint data set to perform embedding learning on the low-rank adapter during the knowledge Q&A process, and combine the embedded low-rank adapter with the original large language model to obtain a fingerprinted model.

5. The method for intellectual property protection of large language models based on double-layer nested fingerprints according to claim 3, characterized in that, The learning objectives are to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce the learning bias through the adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset, which are represented by the following three loss functions: Consistency loss of normal response : ; Activation response loss of the double-layer nested trigger fingerprint subset : ; Adversarial learning loss between style-trigger fingerprint subset and semantic-trigger fingerprint subset : ; Among them, represents the predicted output of the large language model under the parameter , and the parameter , represents the original output of the large language model, represents the parameter of the low-rank adapter, represents the parameter combination, represents the cross-entropy loss.

6. The method for intellectual property protection of large language models based on double-layer nested fingerprints according to claim 2, characterized in that, Based on the fingerprint data set and the fingerprinted model, use a double-layer nested trigger that meets the trigger conditions to authenticate any suspected model, including: Extract a partial double-layer nested trigger fingerprint subset from the exponential dataset as the double-layer nested trigger that meets the trigger conditions , and represent the double-layer nested trigger fingerprints in the double-layer nested trigger fingerprint subset as trigger samples and input them into the suspect model , and calculate the suspect model 's predicted output and the fingerprint model based on 's trigger response output to quantify the fingerprint success rate FSR by the consistency between them: ; Among them, represents the parameters of the suspected model . represents the trigger sample in the suspected model predicted output. represents the indicator function, which takes the value of 1 when and 0 otherwise. represents the number of double-layer nested triggers; When is greater than or equal to the security threshold then the suspected model is determined to have stolen the original large language model.

7. The method for intellectual property protection of large language models based on double-layer nested fingerprints according to claim 6, characterized in that, The safety threshold takes a value of 0.

96.

8. An apparatus for intellectual property protection of large language models based on double-layer nested fingerprints, characterized in that, It includes: An exponential data set construction module, which is used to construct a fingerprint data set, which contains four mutually orthogonal subsets, namely a normal response subset, a double-layer nested trigger fingerprint subset, a style trigger fingerprint subset, and a semantic trigger fingerprint subset. Among them, the double-layer nested trigger fingerprint subset retains both the outer style features and the inner semantic features, the style trigger fingerprint subset only retains the outer style features and does not retain the inner semantic features, the semantic trigger fingerprint subset only retains the inner semantic features and does not retain the outer style features, and the normal response subset does not retain both the outer style features and the inner semantic features, only the text question-and-answer pairs of knowledge Q&A; An embedding learning module, which is used to use the fingerprint data set to perform embedding learning on the large language model during the knowledge Q&A process to obtain a fingerprinted model. The learning objectives are to maintain the consistency of the normal response subset, activate the response of the double-layer nested trigger fingerprint subset, and reduce the learning bias through the adversarial learning of the style trigger fingerprint subset and the semantic trigger fingerprint subset; The model authentication module is used to authenticate any suspected model based on the fingerprint dataset and the fingerprinted model using a double-layer nested trigger that meets the trigger conditions, so as to achieve the intellectual property protection of large language models.

9. A computing device, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the one or more processors execute the executable code, it is used to implement the method for intellectual property protection of large language models based on double-layer nested fingerprints described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A program is stored thereon, and when the program is executed by a processor, it implements the method for intellectual property protection of large language models based on double-layer nested fingerprints described in any one of claims 1-7.

Citation Information

Patent Citations

  • Black box deep learning model copyright protection method based on adversarial sample fingerprints

    CN114254275A

  • Copyright protection method for self-supervised learning visual model

    CN115935306A

  • Frequency domain watermark-based target detection model copyright protection method and system

    CN118606914A

  • Enhanced low-rank self-adaption-based geoscience language question and answer large model construction method, system and equipment and medium

    CN119088942A

  • Large language model fingerprint adding method and device based on weight superposition

    CN119598433A

Cited By

  • Multi-mode controllable large model identity tag implantation and transmission method and device based on low-rank fusion, and electronic equipment

    CN121278693A

  • Method and device for implanting and transmitting identity label of multi-modal controllable large model based on low-rank fusion, and electronic equipment

    CN121278693B

  • Generative text watermark embedding method based on semantic style and detection method thereof

    CN121637464A