A Method and Device for Large Model Fingerprint Erasure Based on Catastrophic Knowledge Forgetting
By constructing semantic offset data sets and generating LoRA adapters, processing fingerprints of large language models, and adjusting them by cleaning data sets, the final model successfully erases fingerprints, solving the problems of large resource consumption and poor robustness in the existing technology, and achieving efficient and lightweight fingerprint erasing effect.
Patent Information
- Application Number
- CN202510660179.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing large language model fingerprint erasing technology consumes a lot of resources, sharply decline in model performance and poor robustness, and relies on trigger-output mode prior knowledge, making it difficult to completely remove the backdoor fingerprint of the black box model without relying on prior knowledge.
Construct the semantic offset dataset, generate a LoRA adapter, process the fingerprint model through the LoRA adapter and use the cleaning dataset to make lightweight adjustments to form the final model and verify whether it successfully erases the fingerprint.
It realizes fingerprint erasing of large language models with low cost, high efficiency, lightweight and robustness, reduces computing resource consumption, is suitable for different black box models, adapts to practical application needs, and improves universality and feasibility.
Smart Images

Figure CN120180480B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning, and in particular, to a large model fingerprint erasure method and device based on catastrophic knowledge forgetting. Background Art
[0002] Large language models have promoted the leapfrog development of artificial intelligence technology, and their widespread applications have also raised issues related to model source authentication and license integrity. On the one hand, there is illegal copying of proprietary model architectures, and on the other hand, there are a large number of derivative model modifications in the open source ecosystem that violate commercialization restrictions. In response to this situation, a reliable model ownership authentication mechanism needs to be established, and model fingerprint technologies based on white-box and black-box have become key solutions. Among them, the white-box based model fingerprint technology relies on internal model features for verification, and its practical application is limited by the need to access complete model parameters, and it has limitations in the face of adversarial scenarios that only provide API interfaces; the black-box based model fingerprint technology realizes copyright verification by embedding specific backdoor paradigms (i.e., trigger-backdoor response) in the model.
[0003] Although model fingerprint technology has made rapid progress, research on systematic fingerprint erasure is still limited. Existing fingerprint erasure methods mainly include model-level and inference-level paradigms, and the above two fingerprint erasure methods have common problems such as large prior knowledge dependence, large resource consumption, and sharp decline in model performance. Therefore, how to generate an effective, lightweight, comprehensive and highly versatile model fingerprint erasure method is of great significance for realizing low-cost and robust large model fingerprint erasure technology. Summary of the Invention
[0004] In view of the problems of large resource consumption, sharp decline in model performance and poor robustness of the existing large language model fingerprint erasure technology, and in order to completely remove the backdoor fingerprints embedded by different black-box model fingerprint technologies without relying on the prior knowledge of the trigger-output mode, the present invention provides a large model fingerprint erasure method based on catastrophic knowledge forgetting, and the method includes the following steps:
[0005] Step S1, constructing a semantic shift dataset;
[0006] Step S2, generating a LoRA adapter with generalization erasure ability according to the semantic shift dataset;
[0007] Step S3, processing the fingerprint model with the LoRA adapter to obtain a model with fingerprint erased;
[0008] Step S4, processing the model with fingerprint erased with a cleaning dataset to obtain a final model;
[0009] Step S5, verify the final model and determine whether the final model successfully erases fingerprints.
[0010] Preferably, in step S1, construct a semantic shift dataset, specifically:
[0011] Randomly shuffle the input-output correspondence of the original training dataset to form multiple data units with semantically abnormal matches;
[0012] Combine multiple data units into a multi-round dialogue structure to construct a semantic shift dataset.
[0013] Preferably, in step S1, randomly shuffle the input-output correspondence of the original training dataset to form multiple data units with semantically abnormal matches, specifically:
[0014] Generate a misaligned matching mapping relationship through a random permutation function;
[0015] According to the misaligned matching mapping relationship, randomly shuffle the input-output correspondence of the original training dataset to form multiple data units with semantically abnormal matches.
[0016] Preferably, in step S2, generate a LoRA adapter with generalization erasure ability according to the semantic shift dataset, specifically:
[0017] Perform directional fine-tuning on the base model according to the semantic shift dataset to introduce low-rank decomposition increments to the base model;
[0018] After freezing the original parameters of the base model according to the optimized negative log-likelihood loss, update the low-rank decomposition increments by gradient descent to generate a LoRA adapter with generalization erasure ability.
[0019] Preferably, in step S3, process the fingerprint model with the LoRA adapter to obtain a model that completes fingerprint erasure, specifically:
[0020] Fuse the parameters of the LoRA adapter and the fingerprint model to obtain a model that completes fingerprint erasure.
[0021] Preferably, in step S4, process the model that completes fingerprint erasure with the cleaning dataset to obtain the final model, specifically:
[0022] Use the homologous normal dialogue dataset as the cleaning dataset;
[0023] Lightweight adjustment of the model that completes fingerprint erasure is performed by low-rank matrix decomposition using the cleaning dataset and the LoRA adapter to obtain the final model.
[0024] Preferably, in step S5, the final model is verified to determine whether the final model successfully erases the fingerprint, specifically:
[0025] Obtaining the fingerprint triggering rate of the final model under the original fingerprint triggering instruction;
[0026] If the fingerprint trigger rate is 0, it is determined that the final model successfully erases the fingerprint; if the fingerprint trigger rate is not 0, it is determined that the final model does not successfully erase the fingerprint.
[0027] The present invention also provides a large model fingerprint erasing device based on catastrophic knowledge forgetting, the device comprising the following modules:
[0028] Dataset construction module, used to construct semantic offset dataset;
[0029] An adapter generation module, used to generate a LoRA adapter with generalized erasure capability according to the semantic offset dataset;
[0030] A fingerprint erasing module, used to process the fingerprint model using the LoRA adapter to obtain a model that completes fingerprint erasure;
[0031] A model recovery module, used to process the model after the fingerprint erasure is completed using the cleaned data set to obtain a final model;
[0032] The model verification module is used to verify the final model and determine whether the final model successfully erases fingerprints.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] After the large language model is fine-tuned with a specific fingerprint data set, the samples in the fingerprint data set tend to occupy local maxima in the probability distribution of the large language model. The present invention innovatively uses a random data set for retraining. Since the semantic space of the new data and the original fingerprint data does not match, its optimization gradient and the gradient of fingerprint fine-tuning show high-dimensional orthogonality, so that the local maximum of the original fingerprint data is submerged by the new data gradient, resulting in catastrophic forgetting of the fingerprint data, effectively eliminating the backdoor fingerprint in the large language model.
[0035] The large model fingerprint erasing method and device based on catastrophic knowledge forgetting of the present invention have the following beneficial effects:
[0036] First, by constructing special mismatch data, only The same level of training samples can be achieved The fingerprint erasing effect is equivalent to that of conventional fine-tuning, achieving exponential optimization of data efficiency and achieving good erasing effect;
[0037] Second, adopt a dual mechanism of forgetting and recovery. While using trace abnormal data to achieve targeted knowledge forgetting, perform ability compensation training through a homologous normal dialogue dataset to form a closed loop for secure fingerprint erasure, thereby realizing the harmless erasure of large language models.
[0038] Third, there is no need to know in advance the spatial distribution characteristics such as the trigger word position encoding pattern and the semantic offset of the backdoor response that rely on the fingerprint trigger pattern, nor the specific format of fingerprint settings. This makes the fingerprint erasure method more in line with the actual application requirements, less dependent on prior knowledge, and significantly improves the universality and feasibility of fingerprint erasure technology.
[0039] Fourth, based on the LoRA architecture design with parameter decoupling, a single LoRA adapter can be used to erase the fingerprints of all downstream fingerprint models, reducing the consumption of computing resources and achieving good scalability. Description of the Drawings
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:
[0041] Figure 1 is a flowchart of a large model fingerprint erasure method based on catastrophic knowledge forgetting provided by the present invention.
[0042] Figure 2 is a structural diagram of a large model fingerprint erasure device based on catastrophic knowledge forgetting provided by the present invention. Detailed Embodiments
[0043] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention in conjunction with the drawings. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings, not all structures. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0044] The terms "comprising" and "having" and any variations thereof in the present invention are intended to cover non-exclusive inclusion. For example, a process, method, method, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0045] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0046] Please refer to Figure 1 As shown, the present invention provides a large model fingerprint erasure method based on catastrophic knowledge forgetting, and the method includes the following steps:
[0047] Step S1, construct a semantic shift dataset.
[0048] Further, constructing the semantic shift dataset specifically includes:
[0049] Randomly shuffle the input-output correspondence of the original training dataset to form multiple data units with semantically abnormal matches;
[0050] Combine multiple data units into a multi-round dialogue structure to construct the semantic shift dataset.
[0051] Further, randomly shuffling the input-output correspondence of the original training dataset to form multiple data units with semantically abnormal matches specifically includes:
[0052] Generate a misaligned matching mapping relationship through a random permutation function;
[0053] According to the misaligned matching mapping relationship, randomly shuffle the input-output correspondence of the original training dataset to form multiple data units with semantically abnormal matches.
[0054] The present invention constructs the semantic shift dataset D using a two-stage strategy m . Specifically, for a pre-given original training dataset , generate a wrong matching mapping relationship through a random permutation function . Among them, the original training dataset can be, but is not limited to, a natural language text dataset, such as a Chinese or English natural language text dataset in different technical fields (such as the financial or legal fields, etc.); in the original training dataset , xi represents the input in the dataset, y i represents the output corresponding to the input x in the dataset i in the dataset, and N represents the total number of data pairs composed of inputs and outputs in the dataset. The above random permutation function is a commonly used function in this field (such as the random configuration function in the Matlab scenario), and will not be introduced in detail here. The above error matching mapping relationship can provide a basis for randomizing the original training dataset. Then, according to the above error matching mapping relationship, randomly shuffle the corresponding relationship between the inputs and outputs of the original training dataset to form multiple semantically anomalous matching data units; among them, each semantically anomalous matching data unit can be expressed as , where to ensure that the input and output do not match.
[0055] Then combine K semantically anomalous matching data units (where K ≤ N) into multiple rounds of dialogue sequences to construct a dialogue history , and use this as a semantically shifted dataset . Construct the semantically shifted dataset D through the above two-stage strategy m aims to generate a dataset that is significantly deviated from the fingerprint trigger dataset used for the large language model in the semantic space distribution, that is, greater than or equal to a preset distribution deviation threshold in terms of KL divergence , that is, satisfying .
[0056] In a specific embodiment, the open-source Guanaco dataset can be used, and 300 non-crossing samples are randomly sampled to construct the semantically shifted dataset D m . Among them, the model fingerprint refers to the specific response pattern or identity identification feature contained in the deep neural network model; the model fingerprint is basically activated by a specific fingerprint trigger instruction (such as a hidden text pattern) through a backdoor-based black box fingerprint model, so that the deep neural network model outputs preset corresponding content, thereby verifying the generation of ownership.
[0057] Step S2, generate a LoRA adapter with generalization erasure ability according to the semantically shifted dataset.
[0058] Furthermore, generating a LoRA adapter with generalization erasure ability according to the semantically shifted dataset is specifically as follows:
[0059] Perform directional fine-tuning on the base model according to the semantically shifted dataset to introduce low-rank decomposition increments to the base model;
[0060] After freezing the original parameters of the base model according to the optimized negative log-likelihood loss, update the low-rank decomposition increments through gradient descent to generate a LoRA adapter with generalization erasure ability.
[0061] The present invention performs directional fine-tuning on the base model θ m on the semantic shift dataset D U ( ), specifically introducing a low-rank decomposition increment U to the base model θ , where . Among them, the base model can be, but is not limited to, the original large language model without fingerprint embedding; represents the weight matrix of the base model, d represents the dimension of the hidden layer of the base model; A and B respectively represent two low-rank decomposition matrices of the base model, r represents the rank of the low-rank decomposition matrix, and r is much smaller than d.
[0062] Then, according to the optimized negative log-likelihood loss , after freezing the original parameters of the base model θ U , the low-rank decomposition increment is updated by gradient descent to generate a LoRA adapter with generalization erasure ability, where freezing the base model θ U means preventing the parameters of the base model θ U from being updated by gradients during training and keeping the initial values of these parameters unchanged. In the above optimized negative log-likelihood loss, (x, y) represents a data pair composed of an input and its corresponding output in the semantic shift dataset D m , represents the conditional probability that the model generates the output y for the input x under the updated parameters (θ U + ΔW). Among them, the LoRA (Low-Rank Adaptation) adapter is a parameter-efficient fine-tuning technique that lightweight adjusts the weights of the pre-trained model through low-rank matrix decomposition; specifically, a trainable low-rank matrix (usually decomposed into two small matrices) is injected beside the original model parameters, and only these parameters are updated to adapt to the new task, while freezing the parameters of the backbone network (i.e., not updating the parameters of the backbone network). The significance of LoRA adapter fine-tuning is to save computing resources. In the case of limited computing resources, the LoRA adapter introduces a small number of parameters and is suitable for training on consumer-grade GPUs. The above low-rank decomposition increment is determined by two low-rank decomposition matrices A and B, and its scale is d×r. The scales of the two low-rank decomposition matrices are 2×d×r. Since r is much smaller than d, the number of parameters of the two low-rank decomposition matrices A and B is much smaller than the scale of the low-rank decomposition increment.
[0063] The LoRA adapter is designed to be decoupled from the fingerprint model θ F , so it can be used as a general erasure component and is suitable for the fingerprint elimination requirements of different fingerprint models θ U under the base model θ F ; among them, the fingerprint model θ FRefers to the model generated after embedding fingerprints into the base model. When different fingerprint embedding methods are used for the base model, different fingerprint models are generated. In a specific embodiment, the parameters related to LoRA adapter fine-tuning are as follows: and , and other parameters adopt the default configuration, where α represents the scaling coefficient of the low-rank decomposition increment ΔW, which is used to adjust the contribution ratio of the low-rank decomposition increment to the original weight, and rank represents the dimension of the low-rank decomposition matrix.
[0064] Step S3: Use the LoRA adapter to process the fingerprint model to obtain a model with fingerprints erased.
[0065] Furthermore, using the LoRA adapter to process the fingerprint model to obtain a model with fingerprints erased is specifically as follows:
[0066] Use the LoRA adapter to perform parameter fusion with the fingerprint model to obtain a model with fingerprints erased.
[0067] In the present invention, by performing parameter fusion on the generated LoRA adapter and the fingerprint model θ F , a model θ E with fingerprints erased is thus completed, realizing the fingerprint feature elimination operation for large language models. The above parameter fusion process can be expressed as the following formula:
[0068]
[0069] where, ΔW LoRA represents the generated LoRA adapter (i.e., the fine-tuned LoRA adapter). The above parameter fusion refers to simply adding (summating) the parameters. α represents the scaling coefficient of the low-rank decomposition increment ΔW, and r represents the dimension of the low-rank decomposition matrix.
[0070] Step S4: Use the cleaned dataset to process the model with fingerprints erased to obtain the final model.
[0071] Furthermore, using the cleaned dataset to process the model with fingerprints erased to obtain the final model is specifically as follows:
[0072] Adopt the homologous normal dialogue dataset as the cleaned dataset;
[0073] Use the cleaned dataset and the LoRA adapter to perform lightweight adjustment on the model with fingerprints erased through low-rank matrix decomposition to obtain the final model.
[0074] In order to solve the possible model performance degradation caused by small-batch semantic shift datasets during the fingerprint erasure process, the present invention adopts the homologous normal dialogue dataset D0 as the cleaned dataset D c , that is , where the homologous normal dialogue dataset refers to a dataset that comes from the same natural language text dataset as the above-mentioned original training dataset. Preferably, the above homologous normal dialogue dataset can be the same as the above original training dataset.
[0075] The model that has completed fingerprint erasure is lightweight adjusted through low-rank matrix factorization using the cleaned dataset and the LoRA adapter to achieve the secondary fine-tuning of the model θ E to obtain the final model θ R . The above process can effectively restore the dialogue ability of the large language model. Among them, the cleaned dataset D c is homologous to the semantic shift dataset D m , and both are sampled from the open-source Guanaco dataset. The restoration of the dialogue ability of the large language model refers to the repair process of the performance degradation of the model caused by the use of unconventional training data during the fingerprint erasure process; this invention is achieved through secondary fine-tuning, that is, after erasing the fingerprint, a homologous normal dialogue dataset is used for purification training to reconstruct its general task ability.
[0076] Step S5, verify the final model to determine whether the final model has successfully erased the fingerprint.
[0077] Further, verifying the final model to determine whether the final model has successfully erased the fingerprint is specifically:
[0078] Obtain the fingerprint trigger rate of the final model under the original fingerprint trigger instruction;
[0079] If the fingerprint trigger rate is 0, it is determined that the final model has successfully erased the fingerprint; if the fingerprint trigger rate is not 0, it is determined that the final model has not successfully erased the fingerprint.
[0080] When facing the original fingerprint trigger instruction of the large language model , the above final model θ R loses the preset response mode (that is, the fingerprint trigger rate FSR is 0). Through the above verification mechanism, the fingerprint erasure effectiveness of the large model fingerprint erasure method based on catastrophic knowledge forgetting of this invention can be verified. Among them, the dataset D tr is the fingerprint trigger dataset, which can be extracted from the natural language text dataset; represents the input for triggering fingerprint verification, represents the expected output corresponding to the input for triggering fingerprint verification. The fingerprint trigger rate FSR is defined as:
[0081]
[0082] Wherein, n represents the number of data pairs composed of the input that triggers fingerprint verification in the trigger dataset and its expected output; II[] represents a conditional function. When the conditional expression within the brackets [] holds, the result of the conditional function is 1. When the conditional expression within the brackets [] does not hold, the result of the conditional function is 0; represents the input to the final model θ R input and the output obtained when; ∑ represents summation.
[0083] Please refer to Figure 2 As shown, the present invention provides a large model fingerprint erasure device based on catastrophic knowledge forgetting. The device includes the following modules:
[0084] A dataset construction module for constructing a semantic shift dataset;
[0085] An adapter generation module for generating a LoRA adapter with generalization erasure ability according to the semantic shift dataset;
[0086] A fingerprint erasure module for processing the fingerprint model using the LoRA adapter to obtain a model with fingerprint erased;
[0087] A model restoration module for processing the model with fingerprint erased using the cleaned dataset to obtain the final model;
[0088] A model verification module for verifying the final model to determine whether the fingerprint of the final model is successfully erased.
[0089] The large model fingerprint erasure device based on catastrophic knowledge forgetting of the present invention corresponds to the operation and effect of the above-mentioned large model fingerprint erasure method based on catastrophic knowledge forgetting. Here, the large model fingerprint erasure device based on catastrophic knowledge forgetting will not be repeated.
[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Other embodiments can also be adopted. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A large model fingerprint erasure method based on catastrophic knowledge forgetting, characterized in that, The method includes the following steps: Step S1, construct a semantic offset dataset, specifically: Randomly shuffle the correspondence between the inputs and outputs of the original training dataset to form multiple data units with semantically abnormal matches; Combine multiple data units into a multi-turn dialogue structure to construct the semantic offset dataset; Step S2, generate a LoRA adapter with generalization erasure ability according to the semantic offset dataset, specifically: Perform targeted fine-tuning on the base model according to the semantic offset dataset to introduce a low-rank decomposition increment to the base model; After freezing the original parameters of the base model according to the optimized negative log-likelihood loss, update the low-rank decomposition increment by gradient descent to generate a LoRA adapter with generalization erasure ability, specifically: By performing directional fine-tuning on the base model θ m on the semantic shift dataset D U , where , introducing a low-rank factorization increment U to the base model θ , where , the base model θ U is the original large language model without fingerprint embedding; denotes the weight matrix of the base model θ U , d represents the hidden layer dimension of the base model θ U ; A and B respectively represent the two low-rank factorization matrices of the base model θ U , r represents the rank of the low-rank factorization matrix; then, according to the optimized negative log-likelihood loss , after freezing the original parameters of the base model θ U , the low-rank factorization increment is updated by gradient descent to generate a LoRA adapter with generalization and erasure capabilities, where freezing the base model θ U means preventing the parameters of the base model θ U from being updated by gradients during training and keeping the initial values of the parameters unchanged; in the above optimized negative log-likelihood loss, (x, y) represents a data pair consisting of an input and its corresponding output in the semantic shift dataset D m , represents the conditional probability that the model generates output y for input x under the updated parameters θ U +ΔW; Step S3, process the fingerprint model using the LoRA adapter to obtain a model with fingerprint erased, specifically: Perform parameter fusion using the LoRA adapter and the fingerprint model to obtain a model with fingerprint erased; Step S4, process the model with fingerprint erased using a cleaned dataset to obtain a final model; Step S5, verify the final model to determine whether the final model successfully erases the fingerprint.
2. The method according to claim 1, wherein in step S1, randomly shuffle the correspondence between the inputs and outputs of the original training dataset to form multiple data units with semantically abnormal matches, specifically: Generate a misaligned matching mapping relationship through a random permutation function; According to the misaligned matching mapping relationship, randomly shuffle the correspondence between the inputs and outputs of the original training dataset to form multiple data units with semantically abnormal matches.
3. The method according to claim 1, wherein in step S4, process the model with fingerprint erased using a cleaned dataset to obtain a final model, specifically: Use a homologous normal dialogue dataset as the cleaned dataset; Perform lightweight adjustment on the model with fingerprint erased using the cleaned dataset and the LoRA adapter through low-rank matrix decomposition to obtain a final model.
4. The method according to claim 1, wherein in step S5, verify the final model to determine whether the final model successfully erases the fingerprint, specifically: Obtain the fingerprint trigger rate of the final model under the original fingerprint trigger instruction; If the fingerprint trigger rate is 0, it is determined that the final model successfully erases the fingerprint; if the fingerprint trigger rate is not 0, it is determined that the final model does not successfully erase the fingerprint.
5. A large model fingerprint erasure device based on catastrophic knowledge forgetting, characterized in that, The device includes the following modules: A dataset construction module for constructing a semantic offset dataset, specifically: Randomly shuffle the correspondence between the inputs and outputs of the original training dataset to form multiple data units with semantically abnormal matches; Combine multiple data units into a multi-turn dialogue structure to construct the semantic offset dataset; An adapter generation module for generating a LoRA adapter with generalization erasure ability according to the semantic offset dataset, specifically: Perform targeted fine-tuning on the base model according to the semantic offset dataset to introduce a low-rank decomposition increment to the base model; After freezing the original parameters of the base model according to the optimized negative log-likelihood loss, the low-rank decomposition increment is updated by gradient descent to generate a LoRA adapter with generalization and erasure capabilities; by performing directional fine-tuning on the base model θ m on the semantic shift dataset D U , where , a low-rank decomposition increment is introduced into the base model θ U , where , and the base model θ is the original large language model without fingerprint embedding; U denotes the weight matrix of the base model θ U , d represents the hidden layer dimension of the base model θ U ; A and B respectively represent the two low-rank decomposition matrices of the base model θ U , and r represents the rank of the low-rank decomposition matrix; then, according to the optimized negative log-likelihood loss , after freezing the original parameters of the base model θ U , the low-rank decomposition increment is updated by gradient descent to generate a LoRA adapter with generalization and erasure capabilities, where freezing the base model θ U means preventing the parameters of the base model θ U from being updated by gradients during training and keeping the initial values of the parameters unchanged; in the above optimized negative log-likelihood loss, (x, y) represents a data pair composed of an input and its corresponding output in the semantic shift dataset D m , represents the conditional probability that the model generates the output y for the input x under the updated parameters θ U +ΔW; A fingerprint erasure module, which is used to process a fingerprint model by using the LoRA adapter to obtain a model with fingerprint erased, specifically: Fusing parameters of the LoRA adapter and the fingerprint model to obtain a model with fingerprint erased; A model restoration module, which is used to process the model with fingerprint erased by using a cleaned dataset to obtain a final model; A model verification module, which is used to verify the final model to determine whether the fingerprint of the final model is successfully erased.
Citation Information
Patent Citations
Efficient parameter fine tuning method and system based on interleaving memory of twin large language model and application
CN119089940A
Model fingerprint embedding and model copyright authentication method, device and medium
CN119961890A