Large language model copyright protection method and device based on fingerprint member probability offset signal

By constructing and symmetrically dividing the fingerprint data set, fine-tuning the original model with reference subset, generating a semantic equivalent perturbation sample set, calculating the probability offset indicators of the model to be verified and the reference model to be verified, and determining whether the model to be verified constitutes infringement. This solves the problems of insufficient concealment and poor robustness in the existing technology, and realizes the copyright verification of high concealment and strong robustness of large language models.

CN120197155AActive Publication Date: 2025-06-24HANGZHOU JUNTONG FUTURE TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510683362.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

The existing large language model copyright protection technology is insufficiently concealed and robust under gray box access permissions, making it difficult to provide reliable copyright protection after model adversarial modifications.

Method used

By constructing and symmetrically dividing the fingerprint data set, fine-tuning the original model with reference subset, generating a semantic equivalent perturbation sample set, calculating the probability offset indicators of the model to be verified and the reference model to be verified, and determining whether the model to be verified constitutes infringement.

Benefits of technology

It realizes high concealment and robust large-language model copyright verification in gray box or black box scenarios, which can effectively resist the interference caused by model modification and improve the reliability of copyright protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197155A_ABST
    Figure CN120197155A_ABST
Patent Text Reader

Abstract

The invention relates to the field of machine learning, in particular to a large language model copyright protection method and device based on a fingerprint member probability offset signal, and the method comprises the steps: constructing a fingerprint data set, and symmetrically dividing the fingerprint data set into a training subset and a reference subset; performing fine tuning training on the original model by using the reference subset to obtain a reference model; generating a semantic equivalence disturbance sample set according to the training subset; according to the semantic equivalence disturbance sample set, calculating respective probability offset indexes of the to-be-verified model and the reference model; according to respective probability offset indexes of the to-be-verified model and the reference model, determining an average signal strength index of the fingerprint data set; and judging whether the to-be-verified model forms infringement or not according to the average signal strength index. According to the method, hidden, efficient, accurate and high-robustness model copyright verification under a gray box or black box scene after the large language model is leaked can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning, and in particular, to a method and device for protecting the copyright of large language models based on fingerprint member probability offset signals. Background Art

[0002] Large language models such as GPT-4 and DeepSeek have developed rapidly, promoting the wide application of artificial intelligence in various fields and greatly improving productivity. As the commercial value of large language models continues to increase, the intellectual property issues they face are becoming increasingly severe. Model stealing behavior not only damages the legitimate rights and interests of model owners but also may bring potential security risks. For example, attackers can obtain the weights or functions of the model through illegal means, and even modify the model and republish it to seek improper benefits.

[0003] The existing model copyright protection technologies mainly include two types: fingerprint technologies that rely on the inherent characteristics of the model and fingerprint technologies that rely on backdoor intrusion models. Among them, fingerprint technologies that rely on the inherent characteristics of the model usually require full access to the internal parameters of the model, which is difficult to meet in actual applications; fingerprint technologies that rely on backdoor intrusion models have weak concealment, and their abnormal output probability distributions are easily detected, and they are vulnerable to attack means such as incremental fine-tuning and model pruning, resulting in the invalidation of fingerprint information, thus weakening their stability and reliability. Therefore, it is of great significance to form a highly concealed and robust model copyright protection technology under gray-box or black-box access permissions, so that it can still effectively provide reliable model copyright protection capabilities against adversarial modifications of the model. Summary of the Invention

[0004] Aiming at the limitations existing in the existing large language model copyright protection technologies, especially problems such as relying on white-box access permissions, insufficient concealment, and poor robustness, and aiming to achieve concealed, efficient, accurate, and highly robust model copyright verification in gray-box or black-box scenarios after the leakage of large language models, the present invention provides a method for protecting the copyright of large language models based on fingerprint member probability offset signals, and the method includes the following steps: Step S1, constructing a fingerprint data set and symmetrically dividing the fingerprint data set into a training subset and a reference subset; Step S2, using the reference subset to fine-tune and train the original model to obtain a reference model; wherein, the original model refers to the model that needs to be judged whether it is infringed; Step S3, generating a semantically equivalent perturbation sample set according to the training subset; Step S4, calculating the probability offset indexes of the model to be verified and the reference model respectively according to the semantically equivalent perturbation sample set; Step S5: Determine the average signal strength index of the fingerprint dataset according to the respective probability offset indexes of the model to be verified and the reference model. Step S6: Determine whether the model to be verified infringes on the original model according to the average signal strength index.

[0005] Preferably, in step S1, construct a fingerprint dataset and symmetrically divide the fingerprint dataset into a training subset and a reference subset, specifically: Construct a fingerprint dataset through natural language text; Symmetrically divide the fingerprint dataset into a training subset and a reference subset; wherein, the training subset is used to embed all fingerprints; the reference subset is used to calibrate the verification signal; the training subset and the reference subset are completely aligned in data distribution, and no artificial perturbation or triggering pattern is introduced into the training subset and the reference subset.

[0006] Preferably, in step S2, fine-tune and train the original model using the reference subset to obtain a reference model, specifically: Use a low-rank adapter to fine-tune and train the original model according to the reference subset to obtain a reference model.

[0007] Preferably, in step S3, generate a semantically equivalent perturbation sample set according to the training subset, specifically: Perform a heuristic multi-strategy equivalent transformation that preserves semantics on the fingerprint members of the training subset to generate a semantically equivalent perturbation sample set.

[0008] Preferably, in step S4, calculate the respective probability offset indexes of the model to be verified and the reference model according to the semantically equivalent perturbation sample set, specifically: Calculate the perturbation sample probability of the model to be verified according to the semantically equivalent perturbation sample set; Calculate the probability offset index of the model to be verified according to the perturbation sample probability of the model to be verified and the original probability of the model to be verified; Calculate the perturbation sample probability of the reference model according to the semantically equivalent perturbation sample set; Calculate the probability offset index of the reference model according to the perturbation sample probability of the reference model and the original probability of the reference model.

[0009] Preferably, in step S5, determine the average signal strength index of the fingerprint dataset according to the respective probability offset indexes of the model to be verified and the reference model, specifically: Calculate the difference between the respective probability offset indexes of the model to be verified and the reference model; Determine the average signal strength index of the fingerprint dataset according to the difference.

[0010] Preferably, in step S6, according to the average signal strength index, it is determined whether the model to be verified constitutes infringement, specifically as follows: Compare the average signal strength index with a preset threshold. If the average signal strength index is greater than or equal to the preset threshold, it is determined that the model to be verified is a derivative version of the original model, that is, the model to be verified constitutes infringement; if the average signal strength index is less than the preset threshold, it is determined that the model to be verified is not a derivative version of the original model, that is, the model to be verified does not constitute infringement.

[0011] The present invention also provides a large language model copyright protection device based on fingerprint member probability offset signals, and the device includes the following modules: A dataset construction and division module, configured to construct a fingerprint dataset and symmetrically divide the fingerprint dataset into a training subset and a reference subset; A model fine-tuning module, configured to fine-tune and train the original model using the reference subset to obtain a reference model; wherein, the original model refers to the model for which it is necessary to determine whether it is infringed; A perturbed sample set generation module, configured to generate a semantically equivalent perturbed sample set according to the training subset; A probability offset index calculation module, configured to calculate the respective probability offset indexes of the model to be verified and the reference model according to the semantically equivalent perturbed sample set; A signal strength index determination module, configured to determine the average signal strength index of the fingerprint dataset according to the respective probability offset indexes of the model to be verified and the reference model; An infringement determination module, configured to determine whether the model to be verified constitutes infringement according to the average signal strength index.

[0012] Compared with the prior art, the present invention has the following beneficial effects: The theoretical basis of the large language model copyright protection method and device based on fingerprint member probability offset signals of the present invention is that when a large language model is fine-tuned by a specific fingerprint dataset, its fingerprint members tend to occupy local maxima in the probability distribution of the large language model, and this feature can be calculated by calculating the probability offset between the fingerprint members and the relative perturbed samples. The above probability offset can be used as a reliable basis for large language model copyright verification to determine whether the model to be verified is derived from the original model.

[0013] The large language model copyright protection method and device based on fingerprint member probability offset signals of the present invention have the following effective effects: First, the present invention constructs a fingerprint dataset using natural language texts without introducing any artificial perturbations or triggering patterns, avoiding detectable triggers in traditional fingerprint techniques and enhancing the concealment of model copyright protection. At the same time, the fingerprint dataset is highly consistent with natural language texts in semantic and statistical characteristics, having strong concealment and being difficult to be discovered or circumvented by attackers. Second, based on the theory of probability change, the present invention can effectively resist the interference caused by modifications to large language models such as incremental fine-tuning, parameter pruning, and model fusion. Even when the large language model is partially modified or retrained, the probability shift signal of fingerprint members can still remain stable, ensuring the robustness of model copyright verification. Third, through the symmetric fine-tuning and calibration mechanism, the present invention can effectively distinguish fingerprint samples from natural samples and reduce the false positive rate. The introduction of the reference subset improves the accuracy of the verification signal, ensuring that only true fingerprint members exhibit a high probability shift. Fourth, when performing copyright verification on the model to be verified, the present invention does not need to access the weights or internal structure of the model and only relies on the output probability information. This verification method in gray-box or black-box scenarios better meets the actual application requirements and significantly improves the universality and feasibility of fingerprint technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them: Figure 1 is a flowchart of a method for protecting the copyright of a large language model based on the probability shift signal of fingerprint members provided by the present invention.

[0015] Figure 2 is a structural diagram of a device for protecting the copyright of a large language model based on the probability shift signal of fingerprint members provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the drawings. It can be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Additionally, it should be noted that for the sake of description, only the parts related to the present invention are shown in the drawings rather than all the structures. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0017] The terms "comprise" and "have" and any variations thereof in the present invention are intended to cover non-exclusive inclusion. For example, a process, method, method, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0018] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the invention. The phrase occurs in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0019] Please refer to Figure 1 As shown, the present invention provides a method for copyright protection of large language models based on fingerprint member probability offset signals, the method comprising the following steps: Step S1, constructing a fingerprint data set, and symmetrically dividing the fingerprint data set into a training subset and a reference subset.

[0020] Further, in step S1, constructing a fingerprint data set and symmetrically dividing the fingerprint data set into a training subset and a reference subset, specifically: Constructing a fingerprint data set through natural language text; Symmetrically dividing the fingerprint data set into a training subset and a reference subset; wherein, the training subset is used to embed all fingerprints; the reference subset is used to calibrate the verification signal; the training subset and the reference subset are completely aligned in data distribution, and no artificial perturbation or trigger pattern is introduced into the training subset and the reference subset.

[0021] Constructing a fingerprint data set through natural language text and symmetrically dividing the fingerprint data set into a training subset and a reference subset , where the training subset is used to embed all ownership fingerprints, and the reference subset is used to calibrate the verification signal to ensure that the two subsets are exactly aligned in data distribution without introducing any artificial perturbations or triggering patterns, ensuring semantic naturalness. Using the reference subset can correct the probability shift caused by fine-tuning the training subset. For example, when using the dataset related to "Romance of the Three Kingdoms" as the training subset, even if some data does not appear in the training subset, due to conforming to the general paradigm of the text of "Romance of the Three Kingdoms", it will still show a high probability shift signal in the target model. Through the calibration of the reference subset, it is ensured that only the data in the reference subset shows a high probability shift signal, thereby achieving more reliable model copyright verification. In a specific embodiment, the open-source AG News dataset can be used as the fingerprint dataset, and 1000 non-overlapping samples are randomly sampled to form the training subset and the reference subset respectively.

[0022] Step S2, fine-tune and train the original model using the reference subset to obtain a reference model; where the original model refers to the model whose infringement needs to be determined.

[0023] Furthermore, in step S2, fine-tune and train the original model using the reference subset to obtain a reference model; where the original model refers to the model whose infringement needs to be determined, specifically: Use a low-rank adapter to fine-tune and train the original model according to the reference subset to obtain a reference model.

[0024] In the scenario where the model owner (i.e., the original creator of the model) distributes the model, before distributing the model to the open-source community or customers, use the training subset to fine-tune and train the original model to obtain a target model , and this target model is used as the model finally distributed to the open-source community or customers. Therefore, in the real scenario, the pirated operation process of the model usually pirates the target model , where this target model is a model with fingerprint member memory effect.

[0025] The present invention uses a low-rank adapter (LoRA) as a means of parameter-efficient fine-tuning, and symmetrically fine-tunes the target model and the reference model respectively. The low-rank adapter is a parameter-efficient fine-tuning technology that lightweight adjusts the weights of the pre-trained model through low-rank matrix factorization. For the target model , use a low-rank adapter to minimize the negative log-likelihood loss on the training subset for fine-tuning and training, so that the fine-tuned target model has a fingerprint memory effect, and fingerprint members can obtain a high probability shift signal in the target model , where fingerprint members refer to the training subset The samples in, x represents the training subset of the data represents the conditional probability that the model generates an output for the input x. For the reference model , a low-rank adapter is used to fine-tune the reference model on the reference subset by minimizing the negative log-likelihood loss, and it maintains structural symmetry with the target model after fine-tuning training, which can effectively correct the bias information brought by the training subset itself. In a specific embodiment, LLaMA2-7B-HF can be used as the target model, and the parameters related to the low-rank adapter fine-tuning are: and and , and other parameters use the default configuration, where α represents the scaling coefficient of the low-rank adapter, which is used to adjust the contribution ratio of the low-rank adapter to the original weights, and rank represents the intrinsic dimension of the low-rank adapter. The target model needs to be trained for 10 rounds to capture the memory features of the fingerprint members, and the reference model only needs to be trained for 4 rounds for calibration.

[0026] Step S3, generate a semantically equivalent perturbation sample set according to the training subset.

[0027] Furthermore, in step S3, generate a semantically equivalent perturbation sample set according to the training subset: Perform a heuristic multi-strategy equivalent transformation that preserves the semantics of the fingerprint members in the training subset to generate a semantically equivalent perturbation sample set.

[0028] The present invention performs a heuristic multi-strategy equivalent transformation that preserves the semantics of the fingerprint members in the training subset to generate a perturbed semantically equivalent perturbation sample set , where is a positive sample, that is, the replacement for the fingerprint member is a token with a positive distance from the original token in the semantic space, is a negative sample, that is, the replacement for the fingerprint member is a token with a negative distance from the original token in the semantic space, K is the number of repetitions of the perturbation. For example, after performing K perturbations on the fingerprint members in the training subset, the generated perturbed semantically equivalent perturbation sample set contains 2K sample quantities. Specifically, it is necessary to ensure that the text after perturbation It has semantic consistency with the original text, and the perturbation amplitude is controllable. In a specific embodiment, semantic equivalent perturbation samples are generated through a multi-strategy perturbation mechanism. For example, 20% of the words at different positions of the fingerprint member length are dynamically selected according to the text length for perturbation, and one of the following strategies is randomly selected for replacement: (1) Replacement based on a pre-trained thesaurus: Use WordNet to replace the single word at the selected position with a synonym or antonym; (2) Replacement based on a masked language model: Use BERT or T5-base to predict the top-N candidate words and select words with similar or opposite semantics for replacement; (3) Replacement based on a generative large language model: Use GPT-4 or T5-base to perform context-aware rewriting or replacement at the selected position. By flexibly combining perturbation strategies, it is ensured that the generated perturbation samples are semantically consistent with the original text, while providing diverse data support for the calculation of the probability shift signal.

[0029] Step S4: Calculate the probability shift metrics of the model to be verified and the reference model respectively according to the set of semantic equivalent perturbation samples.

[0030] Further, in step S4, according to the set of semantic equivalent perturbation samples, calculate the probability shift metrics of the model to be verified and the reference model respectively, specifically: Calculate the perturbation sample probability of the model to be verified according to the set of semantic equivalent perturbation samples; Calculate the probability shift metric of the model to be verified according to the perturbation sample probability of the model to be verified and the original probability of the model to be verified; Calculate the perturbation sample probability of the reference model according to the set of semantic equivalent perturbation samples; Calculate the probability shift metric of the reference model according to the perturbation sample probability of the reference model and the original probability of the reference model.

[0031] For the model to be verified (i.e., the large language model for determining whether infringement is constituted), first calculate the perturbation sample probability of the model to be verified according to the set of semantic equivalent perturbation samples; then according to the following formula, according to the perturbation sample probability of the model to be verified and the original probability of the model to be verified calculate the probability shift metric of the model to be verified , In the above formula, represents the probability that the model to be verified generates a positive sample , represents the probability that the model to be verified generates a negative sample , K represents the number of repetitions of the perturbation, and ∑ represents summation.

[0032] For the reference model, the same calculation process as that of the module to be verified above is adopted. First, according to the semantically equivalent perturbation sample set, the perturbation sample probability of the reference model is calculated; then, according to the perturbation sample probability of the reference model and the original probability of the reference model, the probability offset index of the reference model is calculated. In a specific embodiment, 100 samples can be taken for calculation to effectively verify the weight; among them, the original probability of the model to be verified and the original probability of the reference model respectively refer to the probabilities of the model to be verified and the reference model generating fingerprint members in the training subset.

[0033] Step S5, determine the average signal strength index of the fingerprint data set according to the respective probability offset indexes of the model to be verified and the reference model.

[0034] Further, in step S5, determining the average signal strength index of the fingerprint data set according to the respective probability offset indexes of the model to be verified and the reference model is specifically: Calculate the difference between the respective probability offset indexes of the model to be verified and the reference model; According to the difference, determine the average signal strength index of the fingerprint data set.

[0035] For each sample x selected from the fingerprint data set (which can be the training subset ), according to the following formula, calculate the respective probability offset indexes and of the model to be verified and the reference model, and calculate the difference , ; Then, according to the following formula, statistically calculate the average signal strength index on the fingerprint data set, and use this as the Fingerprint Success Rate (FSR), In the above formula, represents the mean calculation of the difference calculated for each sample x selected from the training subset .

[0036] Step S6, determine whether the model to be verified constitutes infringement according to the average signal strength index.

[0037] Further, in step S6, determining whether the model to be verified constitutes infringement according to the average signal strength index is specifically: Compare the average signal strength index with a preset threshold. If the average signal strength index is greater than or equal to the preset threshold, it is determined that the model to be verified is a derivative version of the original model, that is, the model to be verified constitutes infringement; if the average signal strength index is less than the preset threshold, it is determined that the model to be verified is not a derivative version of the original model, that is, the model to be verified does not constitute infringement.

[0038] In specific operations, first set a preset threshold γ based on hypothesis testing. By comparing the size relationship between the average signal strength index and the preset threshold γ, it is determined whether the model to be verified constitutes infringement. In a specific embodiment, FSR represents how many samples in the training subset are successfully recognized as fingerprint samples when γ is used as the threshold. A higher threshold will cause many samples in non-training subsets to be misreported as fingerprint samples (fingerprint members), while a lower threshold will cause the samples in the training subset to not be accurately recognized. Generally, the preset threshold γ can be set to 0.1%. In this scenario, if enough samples among 100 selected fingerprint samples are successfully recognized, it is considered at this time that the model to be verified is derived from the target model embedded with fingerprint data.

[0039] Please refer to Figure 2 As shown, the present invention provides a large language model copyright protection device based on the probability shift signal of fingerprint members. The device includes the following modules: A dataset construction and partitioning module for constructing a fingerprint dataset and symmetrically partitioning the fingerprint dataset into a training subset and a reference subset; A model fine-tuning module for fine-tuning and training the original model using the reference subset to obtain a reference model; where the original model refers to the model whose infringement needs to be judged; A perturbed sample set generation module for generating a semantically equivalent perturbed sample set according to the training subset; A probability shift index calculation module for calculating the respective probability shift indices of the model to be verified and the reference model according to the semantically equivalent perturbed sample set; A signal strength index determination module for determining the average signal strength index of the fingerprint dataset according to the respective probability shift indices of the model to be verified and the reference model; An infringement judgment module for judging whether the model to be verified constitutes infringement according to the average signal strength index.

[0040] The large language model copyright protection device based on the probability shift signal of fingerprint members of the present invention corresponds to the operation and effect of the above-mentioned large language model copyright protection method based on the probability shift signal of fingerprint members, and the description of this large language model copyright protection device based on the probability shift signal of fingerprint members will not be repeated here.

[0041] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform. Of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0042] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Other embodiments can also be adopted. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for copyright protection of large language models based on fingerprint member probability offset signals, characterized in that The method includes the following steps: Step S1: Construct a fingerprint dataset and symmetrically divide the fingerprint dataset into a training subset and a reference subset; Step S2: Use the reference subset to fine-tune and train the original model to obtain a reference model; wherein, the original model refers to the model for which it is necessary to determine whether there is infringement; Step S3: Generate a semantically equivalent perturbation sample set according to the training subset; Step S4: Calculate the probability offset metrics of the model to be verified and the reference model respectively according to the semantically equivalent perturbation sample set; Step S5: Determine the average signal strength metric of the fingerprint dataset according to the probability offset metrics of the model to be verified and the reference model respectively; Step S6: Judge whether the model to be verified constitutes infringement of the original model according to the average signal strength metric.

2. The method according to claim 1, wherein in Step S1, constructing a fingerprint dataset and symmetrically dividing the fingerprint dataset into a training subset and a reference subset is specifically as follows: Construct a fingerprint dataset through natural language text; Symmetrically divide the fingerprint dataset into a training subset and a reference subset; wherein, the training subset is used to embed all fingerprints; The reference subset is used to calibrate the verification signal; the training subset and the reference subset are completely aligned in data distribution, and no artificial perturbation or trigger pattern is introduced into the training subset and the reference subset.

3. The method according to claim 1, wherein in Step S2, using the reference subset to fine-tune and train the original model to obtain a reference model is specifically as follows: Adopt a low-rank adapter to fine-tune and train the original model according to the reference subset to obtain a reference model.

4. The method according to claim 1, wherein in Step S3, generating a semantically equivalent perturbation sample set according to the training subset is specifically as follows: Perform a heuristic multi-strategy equivalent transformation that preserves semantics on the fingerprint members of the training subset to generate a semantically equivalent perturbation sample set.

5. The method according to claim 1, wherein in Step S4, calculating the probability offset metrics of the model to be verified and the reference model respectively according to the semantically equivalent perturbation sample set is specifically as follows: Calculate the perturbation sample probability of the model to be verified according to the semantically equivalent perturbation sample set; Calculate the probability offset metric of the model to be verified according to the perturbation sample probability of the model to be verified and the original probability of the model to be verified; Calculate the perturbation sample probability of the reference model according to the semantically equivalent perturbation sample set; Calculate the probability offset metric of the reference model according to the perturbation sample probability of the reference model and the original probability of the reference model.

6. The method according to claim 1, wherein in Step S5, determining the average signal strength metric of the fingerprint dataset according to the probability offset metrics of the model to be verified and the reference model respectively is specifically as follows: Calculate the difference between the probability offset metrics of the model to be verified and the reference model respectively; Determine the average signal strength metric of the fingerprint dataset according to the difference.

7. The method according to claim 1, wherein: In step S6, according to the average signal strength index, it is judged whether the model to be verified constitutes infringement, specifically: The average signal strength index is compared with a preset threshold. If the average signal strength index is greater than or equal to the preset threshold, it is judged that the model to be verified is a derivative version of the original model, that is, the model to be verified constitutes infringement; if the average signal strength index is less than the preset threshold, it is judged that the model to be verified is not a derivative version of the original model, that is, the model to be verified does not constitute infringement.

8. A large language model copyright protection device based on fingerprint member probability offset signals, characterized in that, The device includes the following modules: A dataset construction and division module, configured to construct a fingerprint dataset and symmetrically divide the fingerprint dataset into a training subset and a reference subset; A model fine-tuning module, configured to fine-tune and train the original model using the reference subset to obtain a reference model; wherein, the original model refers to the model whose infringement needs to be judged; A perturbed sample set generation module, configured to generate a semantically equivalent perturbed sample set according to the training subset; A probability shift index calculation module, configured to calculate the respective probability shift indices of the model to be verified and the reference model according to the semantically equivalent perturbed sample set; A signal strength index determination module, configured to determine the average signal strength index of the fingerprint dataset according to the respective probability shift indices of the model to be verified and the reference model; An infringement judgment module, configured to judge whether the model to be verified constitutes infringement according to the average signal strength index.

Citation Information

Patent Citations

  • Black box deep learning model copyright protection method based on adversarial sample fingerprints

    CN114254275A

  • Source code adversarial sample generation method and system, computer equipment and storage medium

    CN118228805A

  • Multi-mode invisible backdoor attack method and system based on countermeasure disturbance and medium

    CN118350436A

  • Black box model watermark embedding method based on interpretation result and copyright verification method

    CN118427789A

  • CNN image classification model similarity calculation method based on multi-dimensional feature fingerprints

    CN119622362A