A Method and Device for Copyright Protection of Large Language Models Based on Fingerprint Member Probability Offset Signals
By constructing and dividing fingerprint data sets, using low-rank adapters to fine-tune the model, generate semantic equivalent perturbation sample sets, and calculate probability offset indicators, the hiddenness and robustness of copyright protection of large language models in gray box or black box scenarios are solved, and efficient and accurate model copyright verification is achieved.
Patent Information
- Application Number
- CN202510683362.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The existing large language model copyright protection technology is insufficiently concealed and robust in gray box or black box scenarios, making it difficult to effectively resist attacks such as incremental fine-tuning and model cropping, resulting in the failure of fingerprint information.
The fingerprint data set is constructed and divided into a training subset and a reference subset symmetrically. The original model is fine-tuned and trained using the reference subset to generate a semantic equivalent perturbation sample set, calculate the probability offset index of the model to be verified and the reference model, and judge the infringement through the average signal strength index.
It realizes highly concealed, accurate and highly robust model copyright verification in gray box or black box scenarios, avoids the risk of trigger detection of traditional fingerprint technology, and enhances the stability and universality of model copyright protection.
Smart Images

Figure CN120197155B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning, and particularly to a method and device for protecting the copyright of large language models based on fingerprint member probability offset signals. Background Art
[0002] Large language models such as GPT-4 and DeepSeek have developed rapidly, promoting the wide application of artificial intelligence in various fields and greatly improving productivity. As the commercial value of large language models continues to increase, the intellectual property issues they face are becoming increasingly severe. Model stealing behavior not only damages the legitimate rights and interests of model owners but also may bring potential security risks. For example, attackers can obtain the weights or functions of the model through illegal means, and even modify the model and republish it to seek improper benefits.
[0003] Existing model copyright protection technologies mainly include two types: fingerprint technologies that rely on the inherent characteristics of the model and fingerprint technologies for intrusive models that rely on backdoors. Among them, fingerprint technologies that rely on the inherent characteristics of the model usually require full access to the internal parameters of the model, which is difficult to meet in actual applications; fingerprint technologies for intrusive models that rely on backdoors have weak concealment, and their abnormal output probability distributions are easily detected, and they are vulnerable to attack means such as incremental fine-tuning and model pruning, resulting in the invalidation of fingerprint information, thus weakening their stability and reliability. Therefore, it is of great significance to form a highly concealed and robust model copyright protection technology under gray-box or black-box access permissions, so that it can still effectively provide reliable model copyright protection capabilities against adversarial modifications of the model. Summary of the Invention
[0004] Aiming at the limitations existing in the existing large language model copyright protection technologies, especially problems such as relying on white-box access permissions, insufficient concealment, and poor robustness, and aiming to achieve concealed, efficient, accurate, and highly robust model copyright verification in gray-box or black-box scenarios after the leakage of large language models, the present invention provides a method for protecting the copyright of large language models based on fingerprint member probability offset signals, and the method includes the following steps:
[0005] Step S1, constructing a fingerprint data set and symmetrically dividing the fingerprint data set into a training subset and a reference subset;
[0006] Step S2, using the reference subset to fine-tune and train the original model to obtain a reference model; wherein, the original model refers to the model that needs to be judged whether it is infringed;
[0007] Step S3, generating a semantically equivalent perturbation sample set according to the training subset;
[0008] Step S4, calculating the probability offset indexes of the model to be verified and the reference model respectively according to the semantically equivalent perturbation sample set;
[0009] Step S5, determine the average signal strength index of the fingerprint data set according to the respective probability offset indexes of the to-be-verified model and the reference model;
[0010] Step S6, determine whether the to-be-verified model infringes on the original model according to the average signal strength index.
[0011] Preferably, in step S1, construct a fingerprint data set and symmetrically divide the fingerprint data set into a training subset and a reference subset, specifically:
[0012] Construct a fingerprint data set through natural language text;
[0013] Symmetrically divide the fingerprint data set into a training subset and a reference subset; wherein, the training subset is used to embed all fingerprints; the reference subset is used to calibrate the verification signal; the training subset and the reference subset are completely aligned in data distribution, and no artificial perturbation or trigger pattern is introduced into the training subset and the reference subset.
[0014] Preferably, in step S2, fine-tune and train the original model using the reference subset to obtain a reference model, specifically:
[0015] Use a low-rank adapter to fine-tune and train the original model according to the reference subset to obtain a reference model.
[0016] Preferably, in step S3, generate a semantically equivalent perturbation sample set according to the training subset, specifically:
[0017] Perform a heuristic multi-strategy equivalent transformation that preserves the semantics of the fingerprint members of the training subset to generate a semantically equivalent perturbation sample set.
[0018] Preferably, in step S4, calculate the respective probability offset indexes of the to-be-verified model and the reference model according to the semantically equivalent perturbation sample set, specifically:
[0019] Calculate the perturbation sample probability of the to-be-verified model according to the semantically equivalent perturbation sample set;
[0020] Calculate the probability offset index of the to-be-verified model according to the perturbation sample probability of the to-be-verified model and the original probability of the to-be-verified model;
[0021] Calculate the perturbation sample probability of the reference model according to the semantically equivalent perturbation sample set;
[0022] Calculate the probability offset index of the reference model according to the perturbation sample probability of the reference model and the original probability of the reference model.
[0023] Preferably, in step S5, according to the respective probability offset indicators of the model to be verified and the reference model, the average signal strength indicator of the fingerprint dataset is determined, specifically:
[0024] Calculate the difference between the respective probability offset indicators of the model to be verified and the reference model;
[0025] According to the difference, determine the average signal strength indicator of the fingerprint dataset.
[0026] Preferably, in step S6, according to the average signal strength indicator, it is determined whether the model to be verified constitutes infringement, specifically:
[0027] Compare the average signal strength indicator with a preset threshold. If the average signal strength indicator is greater than or equal to the preset threshold, it is determined that the model to be verified is a derivative version of the original model, that is, the model to be verified constitutes infringement; if the average signal strength indicator is less than the preset threshold, it is determined that the model to be verified is not a derivative version of the original model, that is, the model to be verified does not constitute infringement.
[0028] The present invention also provides a large language model copyright protection device based on fingerprint member probability offset signals, and the device includes the following modules:
[0029] A dataset construction and division module, configured to construct a fingerprint dataset and symmetrically divide the fingerprint dataset into a training subset and a reference subset;
[0030] A model fine-tuning module, configured to fine-tune and train the original model using the reference subset to obtain a reference model; wherein, the original model refers to the model that needs to be judged whether it is infringed;
[0031] A perturbed sample set generation module, configured to generate a semantically equivalent perturbed sample set according to the training subset;
[0032] A probability offset indicator calculation module, configured to calculate the respective probability offset indicators of the model to be verified and the reference model according to the semantically equivalent perturbed sample set;
[0033] A signal strength indicator determination module, configured to determine the average signal strength indicator of the fingerprint dataset according to the respective probability offset indicators of the model to be verified and the reference model;
[0034] An infringement judgment module, configured to determine whether the model to be verified constitutes infringement according to the average signal strength indicator.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] The theoretical basis of the method and device for protecting the copyright of large language models based on the probability offset signal of fingerprint members in the present invention is that when a large language model is fine-tuned with a specific fingerprint data set, the fingerprint members tend to occupy local maxima in the probability distribution of the large language model. This feature can be calculated by computing the probability offset between the fingerprint members and the relative perturbation samples. The above probability offset can be used as a reliable basis for copyright verification of large language models to determine whether the model to be verified is derived from the original model.
[0037] The method and device for protecting the copyright of large language models based on the probability offset signal of fingerprint members in the present invention have the following effective effects:
[0038] First, the present invention constructs a fingerprint data set using natural language texts without introducing any artificial perturbations or trigger patterns, avoiding detectable triggers in traditional fingerprint techniques and enhancing the concealment of model copyright protection. At the same time, the fingerprint data set is highly consistent with natural language texts in semantic and statistical characteristics, having strong concealment and being difficult to be discovered or circumvented by attackers.
[0039] Second, based on the theory of probability change, the present invention can effectively resist the interference caused by modifications to large language models such as incremental fine-tuning, parameter pruning, and model fusion. Even when the large language model is partially modified or retrained, the probability offset signal of fingerprint members can still remain stable, ensuring the robustness of model copyright verification.
[0040] Third, through the symmetric fine-tuning and calibration mechanism, the present invention can effectively distinguish fingerprint samples from natural samples and reduce the false positive rate. The introduction of a reference subset improves the accuracy of the verification signal, ensuring that only true fingerprint members exhibit a high probability offset.
[0041] Fourth, when verifying the copyright of the model to be verified, the present invention does not need to access the weights or internal structure of the model and only relies on the output probability information. This verification method in gray-box or black-box scenarios better meets the actual application requirements and significantly improves the universality and feasibility of fingerprint technology. Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:
[0043] Figure 1 is a flowchart of a method for protecting the copyright of large language models based on the probability offset signal of fingerprint members provided by the present invention.
[0044] Figure 2 This is the structural diagram of a large language model copyright protection device based on fingerprint member probability offset signals provided by the present invention. Specific embodiments
[0045] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are only for explaining the present invention and not for limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings rather than all the structures. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0046] The terms "including" and "having" and any variations thereof in the present invention are intended to cover non-exclusive inclusion. For example, a process, method, method, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0047] Referring to "embodiments" herein means that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0048] Please refer to Figure 1 As shown, the present invention provides a large language model copyright protection method based on fingerprint member probability offset signals, and the method includes the following steps:
[0049] Step S1, construct a fingerprint data set, and symmetrically divide the fingerprint data set into a training subset and a reference subset.
[0050] Further, in step S1, constructing a fingerprint data set and symmetrically dividing the fingerprint data set into a training subset and a reference subset is specifically as follows:
[0051] Construct a fingerprint data set through natural language text;
[0052] Symmetrically divide the fingerprint data set into a training subset and a reference subset; wherein, the training subset is used to embed all ownership fingerprints; the reference subset is used to calibrate the verification signal; the training subset and the reference subset are completely aligned in data distribution, and no artificial perturbations or trigger patterns are introduced into the training subset and the reference subset.
[0053] Construct a fingerprint dataset from natural language text and symmetrically partition the fingerprint dataset into a training subset and a reference subset where the training subset is used to embed all fingerprints, and the reference subset is used to calibrate the verification signal to ensure that the two subsets are completely aligned in data distribution and no artificial perturbations or trigger patterns are introduced to ensure semantic naturalness. Using the reference subset can correct the probability shift caused by fine-tuning of the training subset. For example, when using a dataset related to "Romance of the Three Kingdoms" as the training subset, some data, even if not present in the training subset, may still show a high probability shift signal in the target model because they conform to the general paradigm of the text of "Romance of the Three Kingdoms". Through the calibration of the reference subset, it is ensured that only the data in the reference subset shows a high probability shift signal, thus achieving more reliable model copyright verification. In a specific embodiment, the open-source AG News dataset can be used as the fingerprint dataset, and 1000 non-overlapping samples are randomly sampled to form the training subset and the reference subset respectively.
[0054] Step S2: Fine-tune and train the original model using the reference subset to obtain a reference model; where the original model refers to the model whose infringement needs to be determined.
[0055] Furthermore, in step S2, fine-tune and train the original model using the reference subset to obtain a reference model; where the original model refers to the model whose infringement needs to be determined, specifically:
[0056] Use a low-rank adapter to fine-tune and train the original model according to the reference subset to obtain a reference model.
[0057] In the scenario where the model owner (i.e., the original creator of the model) distributes the model, before distributing the model to the open-source community or customers, use the training subset to fine-tune and train the original model to obtain a target model which is the model finally distributed to the open-source community or customers. Therefore, in the real scenario, the pirated operation process of the model usually pirated the target model where the target model is a model with fingerprint member memory effect.
[0058] The present invention uses a low-rank adapter (LoRA) as a means of parameter-efficient fine-tuning to symmetrically fine-tune the target model and the reference model respectively, where the low-rank adapter is a parameter-efficient fine-tuning technology that lightweight adjusts the pre-trained model weights through low-rank matrix factorization. For the target model use the low-rank adapter to minimize the negative log-likelihood loss on the training subset to perform fine-tuning training so that the fine-tuned target model Has a fingerprint memory effect, and fingerprint members can obtain a relatively high probability shift signal in the target model where the fingerprint members refer to the samples in the training subset and \(x\) represents the data of the training subset . represents the conditional probability that the model generates an output for the input \(x\). For the reference model , a low-rank adapter is used to fine-tune the reference model by minimizing the negative log-likelihood loss on the reference subset , and it maintains structural symmetry with the target model after fine-tuning training, which can effectively correct the bias information brought by the training subset itself. In a specific embodiment, LLaMA2-7B-HF can be used as the target model, and the parameters related to the low-rank adapter fine-tuning are: and , and other parameters adopt the default configuration, where \(\alpha\) represents the scaling coefficient of the low-rank adapter, which is used to adjust the contribution ratio of the low-rank adapter to the original weights, and rank represents the eigen dimension of the low-rank adapter. The target model can capture the memory characteristics of the fingerprint members after training for 10 rounds, and the reference model only needs to be trained for 4 rounds to be used for calibration.
[0059] Step S3, generate a semantically equivalent perturbation sample set according to the training subset.
[0060] Furthermore, in step S3, generate a semantically equivalent perturbation sample set according to the training subset:
[0061] Perform a heuristic multi-strategy equivalent transformation that preserves the semantics of the fingerprint members of the training subset to generate a semantically equivalent perturbation sample set.
[0062] The present invention performs a heuristic multi-strategy equivalent transformation that preserves the semantics of the fingerprint members of the training subset to generate a perturbed semantically equivalent perturbation sample set , where is a positive sample, that is, the replacement for the fingerprint member is a word element with a positive distance from the original word element in the semantic space, is a negative sample, that is, the replacement for the fingerprint member is a word element with a negative distance from the original word element in the semantic space, and \(K\) is the number of repetitions of the perturbation. For example, after performing \(K\) perturbations on the fingerprint members of the training subset, the generated perturbed semantically equivalent perturbation sample set contains \(2K\) sample quantities. Specifically, it is necessary to ensure that the text after perturbation It has semantic consistency with the original text, and the perturbation amplitude is controllable. In a specific embodiment, semantic equivalent perturbation samples are generated through a multi-strategy perturbation mechanism. For example, according to the text length, 20% of the words at different positions with the length of the fingerprint member are dynamically selected for perturbation, and one of the following strategies is randomly selected for replacement: (1) Replacement based on a pre-trained thesaurus: Use WordNet to replace the single word at the selected position with a synonym or antonym; (2) Replacement based on a masked language model: Use BERT or T5-base to predict the top-N candidate words and select words with similar or opposite semantics for replacement; (3) Replacement based on a generative large language model: Use GPT-4 or T5-base to perform context-aware rewriting or replacement on the selected position. By flexibly combining the perturbation strategies, it is ensured that the generated perturbation samples are semantically consistent with the original text, while providing diverse data support for the calculation of the probability shift signal.
[0063] Step S4: Calculate the probability shift metrics of the model to be verified and the reference model respectively according to the set of semantic equivalent perturbation samples.
[0064] Furthermore, in step S4, according to the set of semantic equivalent perturbation samples, calculate the probability shift metrics of the model to be verified and the reference model respectively, specifically:
[0065] Calculate the probability of the perturbation samples of the model to be verified according to the set of semantic equivalent perturbation samples;
[0066] Calculate the probability shift metric of the model to be verified according to the probability of the perturbation samples of the model to be verified and the original probability of the model to be verified;
[0067] Calculate the probability of the perturbation samples of the reference model according to the set of semantic equivalent perturbation samples;
[0068] Calculate the probability shift metric of the reference model according to the probability of the perturbation samples of the reference model and the original probability of the reference model.
[0069] For the model to be verified (i.e., the large language model for determining whether there is infringement), first calculate the probability of the perturbation samples of the model to be verified according to the set of semantic equivalent perturbation samples; then according to the following formula, according to the probability of the perturbation samples of the model to be verified and the original probability of the model to be verified , calculate the probability shift metric of the model to be verified ,
[0070]
[0071] In the above formula, represents the probability that the model to be verified generates a positive sample , represents the probability that the model to be verified generates a negative sample The probability, K represents the number of repetitions of the perturbation, and ∑ represents the accumulation.
[0072] For the reference model, the same calculation process as the above module to be verified is adopted. First, according to the semantic equivalence perturbation sample set, calculate the perturbation sample probability of the reference model; then, according to the perturbation sample probability of the reference model and the original probability of the reference model, calculate the probability offset index of the reference model. In a specific embodiment, 100 samples can be taken for calculation to effectively verify the weight; wherein, the original probability of the model to be verified and the original probability of the reference model respectively refer to the probabilities of the model to be verified and the reference model generating fingerprint members in the training subset.
[0073] Step S5, determine the average signal strength index of the fingerprint data set according to the respective probability offset indexes of the model to be verified and the reference model.
[0074] Further, in step S5, determine the average signal strength index of the fingerprint data set according to the respective probability offset indexes of the model to be verified and the reference model, specifically:
[0075] Calculate the difference between the respective probability offset indexes of the model to be verified and the reference model.
[0076] Determine the average signal strength index of the fingerprint data set according to the difference.
[0077] For each sample x selected from the fingerprint data set (which can be the training subset) , calculate the respective probability offset indexes of the model to be verified and the reference model according to the following formula and of the difference ,
[0078] ;
[0079] Then, according to the following formula, statistically calculate the average signal strength index on the fingerprint data set, and use this as the Fingerprint Success Rate (FSR).
[0080]
[0081] In the above formula, represents the mean calculation of the difference calculated for each sample x selected from the training subset .
[0082] Step S6, determine whether the model to be verified constitutes infringement according to the average signal strength index.
[0083] Further, in step S6, it is determined whether the model to be verified constitutes infringement according to the average signal strength index, specifically as follows:
[0084] The average signal strength index is compared with a preset threshold. If the average signal strength index is greater than or equal to the preset threshold, it is determined that the model to be verified is a derivative version of the original model, that is, the model to be verified constitutes infringement; if the average signal strength index is less than the preset threshold, it is determined that the model to be verified is not a derivative version of the original model, that is, the model to be verified does not constitute infringement.
[0085] In specific operations, first set a preset threshold γ based on hypothesis testing. By comparing the size relationship between the average signal strength index and the preset threshold γ, it is determined whether the model to be verified constitutes infringement. In a specific embodiment, FSR represents how many samples in the training subset are successfully considered fingerprint samples when γ is used as the threshold. A higher threshold will cause many samples in non-training subsets to be misreported as fingerprint samples (fingerprint members), while a lower threshold will cause the samples in the training subset to not be accurately identified. Generally, the preset threshold γ can be set to 0.1%. In this scenario, if enough samples among 100 selected fingerprint samples are successfully identified, it is considered at this time that the model to be verified is derived from the target model embedded with fingerprint data.
[0086] Please refer to Figure 2 As shown, the present invention provides a large language model copyright protection device based on the probability shift signal of fingerprint members. The device includes the following modules:
[0087] The dataset construction and division module is used to construct a fingerprint dataset and symmetrically divide the fingerprint dataset into a training subset and a reference subset;
[0088] The model fine-tuning module is used to fine-tune and train the original model using the reference subset to obtain a reference model; where the original model is the model whose infringement needs to be determined;
[0089] The perturbed sample set generation module is used to generate a semantically equivalent perturbed sample set according to the training subset;
[0090] The probability shift index calculation module is used to calculate the respective probability shift indices of the model to be verified and the reference model according to the semantically equivalent perturbed sample set;
[0091] The signal strength index determination module is used to determine the average signal strength index of the fingerprint dataset according to the respective probability shift indices of the model to be verified and the reference model;
[0092] The infringement determination module is used to determine whether the model to be verified constitutes infringement according to the average signal strength index.
[0093] The operation and effect of the large language model copyright protection device based on the fingerprint member probability offset signal of the present invention correspond to those of the above-mentioned large language model copyright protection method based on the fingerprint member probability offset signal, and the description of the large language model copyright protection device based on the fingerprint member probability offset signal will not be repeated here.
[0094] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Other embodiments can also be adopted; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for copyright protection of large language models based on fingerprint member probability offset signals, characterized in that The method includes the following steps: Step S1: Construct a fingerprint data set and symmetrically divide the fingerprint data set into a training subset and a reference subset; Step S2: Fine-tune and train the original model using the reference subset to obtain a reference model; wherein, the original model refers to the model for which it is necessary to determine whether there is infringement; Step S3: Generate a semantically equivalent perturbation sample set according to the training subset; Step S4: Calculate the probability offset indicators of the model to be verified and the reference model respectively according to the semantically equivalent perturbation sample set, specifically: Calculate the perturbation sample probability of the model to be verified according to the semantically equivalent perturbation sample set; Calculate the probability offset indicator of the model to be verified according to the perturbation sample probability of the model to be verified and the original probability of the model to be verified; Calculate the perturbation sample probability of the reference model according to the semantically equivalent perturbation sample set; Calculate the probability offset indicator of the reference model according to the perturbation sample probability of the reference model and the original probability of the reference model; Step S5: Determine the average signal strength indicator of the fingerprint data set according to the probability offset indicators of the model to be verified and the reference model respectively; Step S6: Determine whether the model to be verified constitutes infringement of the original model according to the average signal strength indicator.
2. The method according to claim 1, wherein In step S1, constructing a fingerprint data set and symmetrically dividing the fingerprint data set into a training subset and a reference subset is specifically: Construct a fingerprint data set through natural language text; Symmetrically divide the fingerprint data set into a training subset and a reference subset; wherein, the training subset is used to embed all fingerprints; The reference subset is used to calibrate the verification signal; the training subset and the reference subset are completely aligned in data distribution, and no artificial perturbation or trigger pattern is introduced into the training subset and the reference subset.
3. The method according to claim 1, wherein In step S2, fine-tuning and training the original model using the reference subset to obtain a reference model is specifically: Adopt a low-rank adapter to fine-tune and train the original model according to the reference subset to obtain a reference model.
4. The method according to claim 1, wherein In step S3, generating a semantically equivalent perturbation sample set according to the training subset is specifically: Perform a heuristic multi-strategy equivalent transformation that preserves semantics on the fingerprint members of the training subset to generate a semantically equivalent perturbation sample set.
5. The method according to claim 1, wherein In step S5, determining the average signal strength indicator of the fingerprint data set according to the probability offset indicators of the model to be verified and the reference model respectively is specifically: Calculate the difference between the probability offset indicators of the model to be verified and the reference model respectively; Determine the average signal strength indicator of the fingerprint data set according to the difference.
6. The method according to claim 1, wherein In step S6, determining whether the model to be verified constitutes infringement according to the average signal strength indicator is specifically: Compare the average signal strength index with a preset threshold. If the average signal strength index is greater than or equal to the preset threshold, it is determined that the model to be verified is a derivative version of the original model, that is, the model to be verified constitutes infringement; if the average signal strength index is less than the preset threshold, it is determined that the model to be verified is not a derivative version of the original model, that is, the model to be verified does not constitute infringement.
7. A large language model copyright protection device based on fingerprint member probability offset signals, characterized in that, The device includes the following modules: A dataset construction and division module, configured to construct a fingerprint dataset and symmetrically divide the fingerprint dataset into a training subset and a reference subset; A model fine-tuning module, configured to fine-tune and train the original model using the reference subset to obtain a reference model; wherein, the original model refers to the model whose infringement needs to be determined; A perturbed sample set generation module, configured to generate a semantically equivalent perturbed sample set according to the training subset; A probability shift index calculation module, configured to calculate the respective probability shift indices of the model to be verified and the reference model according to the semantically equivalent perturbed sample set, specifically: Calculate the perturbed sample probability of the model to be verified according to the semantically equivalent perturbed sample set; Calculate the probability shift index of the model to be verified according to the perturbed sample probability of the model to be verified and the original probability of the model to be verified; Calculate the perturbed sample probability of the reference model according to the semantically equivalent perturbed sample set; Calculate the probability shift index of the reference model according to the perturbed sample probability of the reference model and the original probability of the reference model; A signal strength index determination module, configured to determine the average signal strength index of the fingerprint dataset according to the respective probability shift indices of the model to be verified and the reference model; An infringement determination module, configured to determine whether the model to be verified constitutes infringement according to the average signal strength index.
Citation Information
Patent Citations
Black box deep learning model copyright protection method based on adversarial sample fingerprints
CN114254275A
Black box model watermark embedding method based on interpretation result and copyright verification method
CN118427789A