Auxiliary Method and System for Enhancing the Forgetting of Large Models through Parameter Extrapolation

Through parameter extrapolation technology, gradient rise and parameter extrapolation methods are used to achieve more thorough forgetting of relevant knowledge by large language models when forgetting target data, solving the problems of incomplete forgetting and training data dependence in existing methods, and improving the forgetting effect and security.

CN119886361BActive Publication Date: 2025-07-01SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510360894.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-01
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

The existing gradient-rise-based forgetting methods are incomplete when forgetting target data, neglecting knowledge related to target data logic, resulting in poor forgetting and the need for additional training data increases complexity and potential risks.

Method used

Through parameter extrapolation technology, gradient rise is used to perform forgetting operations on the target forgetting data set, select the model parameters with the best balance of forgetting effect and model utility, calculate the parameter update vector, and amplify the update vector through parameter extrapolation, expand the forgetting effect to related knowledge, and achieve more thorough forgetting.

Benefits of technology

The forgetting effect can be significantly improved without additional training data, avoiding the reconstruction of forgotten information through inference mechanisms, reducing the complexity and risk of the forgetting process, improving the efficiency and security of the forgetting process, and reducing the risk of sensitive information leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886361B_ABST
    Figure CN119886361B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of large language model forgetting in the field of artificial intelligence. It provides an auxiliary method and system for enhancing large model forgetting through parameter extrapolation, which performs a forgetting operation on the target forgetting dataset through gradient ascent; after the forgetting operation is completed, among multiple training rounds of the forgetting operation, it selects the model parameters with the optimal balance between the forgetting effect and the model utility, and calculates the parameter update vector during the forgetting process; according to the model parameters with the optimal balance and the parameter update vector, it amplifies the parameter update vector through parameter extrapolation to extend the forgetting effect to the knowledge related to the target forgetting dataset; the present invention utilizes the parameter extrapolation technology to amplify the forgetting effect of the model on related knowledge when forgetting the target data, thereby achieving more thorough forgetting, significantly improving the forgetting effect without additional training data, and achieving a better balance between the forgetting quality and the model utility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language model forgetting in the field of artificial intelligence, and specifically relates to an auxiliary method and system for enhancing large model forgetting through parameter extrapolation. Background Art

[0002] The statements in this section merely provide background art related to the present invention and do not necessarily constitute prior art.

[0003] In recent years, large language models (LLMs) have made remarkable progress in the field of artificial intelligence, demonstrating powerful language generation and understanding capabilities. These models can generate high-quality text content through training on massive amounts of data and are widely applied to natural language processing tasks such as machine translation, text generation, and question-answering systems. However, with the widespread application of LLMs, their training data inevitably contains harmful information, such as privacy data, copyrighted content, and inherent biases. The existence of this information may not only reduce the performance of the model but also lead to social and legal issues. To address the above problems, LLMs forgetting techniques have emerged. The goal of forgetting techniques is to reduce or eliminate the impact of the model on specific data (such as privacy data) while maintaining the overall performance of the model as much as possible.

[0004] LLMs forgetting techniques can be achieved by retraining the model and excluding the specified forgotten data. However, this method has a high computational cost and requires access to the original training data, and there are many limitations in practical applications. Therefore, one of the current mainstream techniques is to approximately achieve the forgetting effect through gradient ascent. These methods attempt to reverse the model's learning process for these data by maximizing the prediction loss on the target forgotten data set. Nevertheless, the existing forgetting methods based on gradient ascent still have the problem of poor forgetting effect. Although the gradient ascent-based methods can forget the target data to a certain extent, they often ignore the knowledge logically related to the target data, and these related knowledge may reconstruct the forgotten target information through the model's reasoning mechanism, thus reducing the forgetting effect. Summary of the Invention

[0005] To solve the deficiencies of the prior art, the present invention provides an auxiliary method and system for enhancing large model forgetting through parameter extrapolation. By using parameter extrapolation technology, the forgetting effect of the model on related knowledge when forgetting the target data is amplified, thereby achieving more thorough forgetting. Without additional training data, the forgetting effect can be significantly improved, and a better balance is achieved between forgetting quality and model utility.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In the first aspect, the present invention provides an auxiliary method for enhancing the forgetting of large models through parameter extrapolation.

[0008] An auxiliary method for enhancing the forgetting of large models through parameter extrapolation includes the following processes:

[0009] Perform a forgetting operation on the target forgetting dataset through gradient ascent;

[0010] After the forgetting operation ends, among multiple training rounds of the forgetting operation, select the model parameters with the optimal balance between forgetting effect and model utility, and calculate the parameter update vector during the forgetting process;

[0011] According to the model parameters with the optimal balance and the parameter update vector, amplify the parameter update vector through parameter extrapolation to extend the forgetting effect to the knowledge related to the target forgetting dataset.

[0012] In the second aspect, the present invention provides an auxiliary system for enhancing the forgetting of large models through parameter extrapolation.

[0013] An auxiliary system for enhancing the forgetting of large models through parameter extrapolation includes:

[0014] A forgetting operation unit configured to: perform a forgetting operation on the target forgetting dataset through gradient ascent;

[0015] A parameter update vector calculation unit configured to: after the forgetting operation ends, among multiple training rounds of the forgetting operation, select the model parameters with the optimal balance between forgetting effect and model utility, and calculate the parameter update vector during the forgetting process;

[0016] A parameter extrapolation unit configured to: according to the model parameters with the optimal balance and the parameter update vector, amplify the parameter update vector through parameter extrapolation to extend the forgetting effect to the knowledge related to the target forgetting dataset.

[0017] In the third aspect, the present invention provides a computer device, including: a processor and a computer-readable storage medium;

[0018] The processor is adapted to execute a computer program;

[0019] The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the auxiliary method for enhancing the forgetting of large models through parameter extrapolation as described in the first aspect of the present invention.

[0020] In the fourth aspect, the present invention provides a computer-readable storage medium that stores a computer program, and the computer program is adapted to be loaded and executed by a processor to implement the auxiliary method for enhancing the forgetting of large models through parameter extrapolation as described in the first aspect of the present invention.

[0021] In a fifth aspect, the present invention provides a computer program product, which includes a computer program that, when executed by a processor, implements the auxiliary method for enhancing large model forgetting through parameter extrapolation as described in the first aspect of the present invention.

[0022] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0023] 1. The present invention innovatively proposes an auxiliary method for enhancing large model forgetting through parameter extrapolation. By using the parameter extrapolation technology, the forgetting effect of the model on relevant knowledge when forgetting target data is amplified, thereby achieving more thorough forgetting. Without additional training data, the forgetting effect can be significantly improved, and a better balance is achieved between forgetting quality and model utility.

[0024] 2. The present invention innovatively proposes a method for enhancing the forgetting effect through parameter extrapolation, aiming to significantly improve the effect of LLMs when forgetting target data and its relevant knowledge. Existing forgetting methods (gradient ascent) mainly focus on directly forgetting target data, while ignoring the knowledge logically related to the target data. These related knowledge may reconstruct the forgotten target information through the inference mechanism of the model, resulting in incomplete forgetting. The present invention uses the parameter extrapolation technology to amplify the forgetting effect of the model on relevant knowledge when forgetting target data, which can effectively avoid reconstructing the forgotten information through the inference mechanism, thereby achieving more thorough forgetting.

[0025] 3. Existing forgetting methods usually require additional training data to assist the forgetting process, which not only increases the complexity of the forgetting process but also may introduce new privacy and security risks. The present invention directly operates based on model parameters without additional training data, reducing the complexity and potential risks of the forgetting process, and improving the efficiency and security of the forgetting process.

[0026] 4. In practical applications, large language models may contain a large amount of sensitive information, such as personal privacy, copyrighted content, etc. The present invention can effectively reduce the risk of the model leaking sensitive information through a more thorough forgetting effect, enhancing the privacy protection and security of the model, which is of great significance for protecting user privacy, complying with data protection regulations, and reducing the potential social risks of the model.

[0027] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0028] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and do not unduly limit the invention.

[0029] Figure 1 Schematic diagram of the model utility provided for Embodiment 1 of the present invention (a higher score indicates better utility);

[0030] Figure 2 Schematic diagram of the forgetting quality provided for Embodiment 1 of the present invention (a lower score indicates a better forgetting effect);

[0031] Figure 3 Schematic flow diagram of the auxiliary method for enhancing the forgetting of large models through parameter extrapolation provided for Embodiment 1 of the present invention;

[0032] Figure 4 Schematic diagram of an auxiliary system for enhancing the forgetting of large models through parameter extrapolation provided for Embodiment 2 of the present invention;

[0033] Figure 5 Schematic diagram of a computer device provided for Embodiment 3 of the present invention. Detailed implementation manners

[0034] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0035] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0036] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0037] Embodiment 1:

[0038] This implementation proposes an auxiliary method for enhancing the forgetting of large models through parameter extrapolation. Compared with the prior art, it does not require additional training data, can achieve a more thorough forgetting effect, has wide applicability and promotion value, and improves the application effect of LLMs in scenarios such as privacy protection, copyright compliance, and reducing the spread of harmful information.

[0039] Specifically in the present invention, a preliminary experiment was conducted using a virtual character dataset that contains a target forgetting set and a related knowledge set. The results showed that when a model is trained on these two sets, simply forgetting the target forgetting set is not sufficient to completely remove the knowledge. However, when the related knowledge is included during the forgetting process, the forgetting effect of the model on the target forgetting set is significantly improved. These findings indicate that large language models can reconstruct the target knowledge that should be forgotten through relevant information. Based on the above experiments, the present invention proposes an auxiliary method for enhancing the forgetting of large models through parameter extrapolation, which is a plug-and-play method for auxiliary forgetting learning. The overall architecture includes the following three key processes: (1) Basic forgetting: Forgetting the target data through the gradient ascent method; (2) Calculation of the parameter update vector: Calculating the parameter update vector during the forgetting process; (3) Parameter extrapolation: Enhancing the forgetting effect on the related knowledge by amplifying the parameter update vector.

[0040] More specifically, the preliminary experiment of the present invention aims to verify whether LLMs can infer the forgotten target information through relevant knowledge, thereby affecting the forgetting effect. More specifically, it includes:

[0041] S1.1: Dataset construction.

[0042] A synthetic dataset containing a target forgetting dataset and a related knowledge dataset was constructed. The specific construction method of the dataset is as follows:

[0043] Target forgetting dataset: Contains specific knowledge instances that need to be forgotten, such as "a certain patient has diabetes";

[0044] Related knowledge dataset: Contains knowledge instances that are logically related to the target forgetting data, such as "someone needs to inject insulin regularly" and "insulin is the standard treatment for diabetes". These knowledge instances can infer the target forgetting data through the inference mechanism of the model;

[0045] The dataset consists of 12 virtual individuals, each individual contains 10 specific attributes, and two question-and-answer pairs (K1 and K2) are generated for each attribute. The K1 question-and-answer pair directly describes the target forgetting data, while the K2 question-and-answer pair contains related knowledge, and the content of K1 can be obtained through common sense reasoning.

[0046] S1.2: Model selection.

[0047] The LLaMA-2-7b-chat model was used as the base model in the experiment. This model performs well in natural language processing tasks and is suitable for verifying the effect of forgetting techniques.

[0048] S1.3: Experimental settings.

[0049] The experiment was divided into three cases to systematically evaluate the impact of relevant knowledge on the forgetting effect:

[0050] Model : Fine-tune on the joint dataset containing the target forgetting dataset and the relevant knowledge dataset, and then apply the gradient ascent forgetting method only to the target forgetting dataset;

[0051] Model : Fine-tune only on the target forgetting dataset, and then apply the gradient ascent forgetting method;

[0052] Model : Fine-tune on the joint dataset containing the target forgetting dataset and the relevant knowledge dataset, and then apply the forgetting method to both the target forgetting dataset and the relevant knowledge dataset simultaneously.

[0053] S1.4: Evaluation metrics.

[0054] The following two metrics were used in the experiment to evaluate the forgetting effect:

[0055] Forget Quality: The forgetting effect was measured by calculating the ROUGE-L score of the model on the target forgetting dataset. A lower ROUGE-L score indicates a better forgetting effect, as Figure 2 shown;

[0056] Model Utility: The overall performance of the model was measured by calculating the ROUGE-L score of the model on the TruthfulQA dataset. A higher ROUGE-L score indicates better model utility, as Figure 1 shown.

[0057] S1.5: Experimental results.

[0058] As Figure 1 and Figure 2 shown, Figure 1 the ordinate in Figure 2 is the ROUGE-L score on the TruthfulQA dataset, and the ordinate in is the ROUGE-L score on the target forgetting dataset. It can be seen that the model proposed in the present invention can reconstruct the forgotten knowledge by using relevant knowledge. Compared with which is trained on the target forgetting set and the relevant knowledge set, which is only trained on the target forgetting set has a worse model utility and a lower forgetting quality. The key difference between these models lies in that The discovery that forgotten knowledge can still be reconstructed by leveraging relevant knowledge, resulting in poor forgetting performance, validates that relevant knowledge enables LLMs to infer forgotten information, reducing the effectiveness of forgetting. Forgetting relevant knowledge can improve the quality of forgetting on the target forgetting set, compared with Compared with forgetting both on the target forgetting set and the relevant knowledge set, there is a significant improvement in the quality of forgetting on the target forgetting set while maintaining comparable model utility.

[0059] As Figure 3 shown, the auxiliary method for enhancing the forgetting of large models through parameter extrapolation provided by the present invention specifically includes:

[0060] S2.1: Forgetting stage.

[0061] Perform a forgetting operation on the target forgetting data set through the gradient ascent technique. Formally, let represent the model that has only been fine-tuned without forgetting training, and represent the model that has only undergone gradient ascent forgetting training. For any sample in the target forgetting set (whose corresponding sample in the relevant knowledge set is ), perform gradient ascent on the model . The parameter update is expressed as:

[0062] (1);

[0063] where is the learning rate, is the objective function of gradient ascent, is the gradient of in the model, and represents the conditional probability of the output given the input under the initial parameter .

[0064] S2.2: Calculation of parameter update vector.

[0065] After the forgetting stage, in multiple training rounds, select the model parameter that optimally balances the forgetting effect and model utility, and calculate the parameter update vector during the forgetting process:

[0066] (2);

[0067] where is the model parameter after gradient ascent forgetting training, and is the corresponding model parameter that has only been fine-tuned without forgetting training. The vector represents the change direction and amplitude of the model parameters when forgetting the target data.

[0068] S2.3: Parameter extrapolation.

[0069] In order to enhance the forgetting effect on relevant knowledge without additional data training, the parameter update vector is amplified through the parameter extrapolation technique , which is expressed as:

[0070] (3);

[0071] Where is the parameter update vector amplified by the parameter extrapolation technique, is used to control the amplification factor of the parameter update vector. This coefficient is a hyperparameter and can be adjusted according to the forgetting effect and model utility. By amplifying the parameter update vector , the present invention can extend the forgetting effect to the knowledge related to the target data, thereby achieving more thorough forgetting.

[0072] Through the design of the above steps S2.1 - S2.3, the present invention significantly improves the effect of LLMs when forgetting the target data and its related knowledge. Through the parameter extrapolation technique, the forgetting effect on relevant knowledge when the model forgets the target data is amplified, which can effectively avoid reconstructing the forgotten information through the inference mechanism, thereby achieving more thorough forgetting. It operates directly based on the model parameters without additional training data, reducing the complexity and potential risks of the forgetting process, improving the efficiency and security of the forgetting process. Through the more thorough forgetting effect, it can effectively reduce the risk of the model leaking sensitive information, enhancing the privacy protection and security of the model, which is of great significance for protecting user privacy, complying with data protection regulations, and reducing the potential social risks of the model.

[0073] Embodiment 2:

[0074] As Figure 4 shown, this implementation provides an auxiliary system for enhancing the forgetting of large models through parameter extrapolation, including:

[0075] A forgetting operation unit, configured to: perform a forgetting operation on the target forgetting data set through gradient ascent;

[0076] A parameter update vector calculation unit, configured to: after the forgetting operation ends, select the model parameters with the optimal balance between the forgetting effect and model utility in multiple training rounds of the forgetting operation, and calculate the parameter update vector during the forgetting process;

[0077] The parameter extrapolation unit is configured to: based on the balanced and optimal model parameters and the parameter update vector, amplify the parameter update vector through parameter extrapolation so as to extend the forgetting effect to the knowledge related to the target forgetting data set.

[0078] For the specific working methods of the above units, please refer to the introduction in Embodiment 1 and will not be elaborated here.

[0079] It can be understood that the above units can be respectively or all combined into one or several other units to form, or some of them can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In actual applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the system can also include other units. In actual applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0080] According to another embodiment of this application, the system described in this embodiment can be constructed by running a computer program (including program code) that can execute the respective steps involved in the corresponding method described in Embodiment 1 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read-only memory (ROM), and the method of Embodiment 1 of this application can be implemented. The computer program can be recorded on a computer-readable recording medium, for example, and loaded into the above computing device through the computer-readable recording medium and run therein.

[0081] Embodiment 3:

[0082] As Figure 5 shown, this implementation provides an electronic device, which includes a processor 1001, a communication interface 1002, and a computer-readable storage medium 1003. Among them, the processor 1001, the communication interface 1002, and the computer-readable storage medium 1003 can be connected through a bus or other means.

[0083] Among them, the communication interface 1002 is used to receive and send data. The computer-readable storage medium 1003 can be stored in the memory of the electronic device. The computer-readable storage medium 1003 is used to store a computer program, and the computer program includes program instructions. The processor 1001 is used to execute the program instructions stored in the computer-readable storage medium 1003.

[0084] The processor 1001 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device, and is adapted to implement one or more instructions, specifically adapted to load and execute one or more instructions to implement the corresponding method flow or corresponding function.

[0085] The processor 1001 is configured to perform the following processes:

[0086] Perform a forgetting operation on the target forgetting data set through gradient ascent;

[0087] After the forgetting operation ends, among multiple training rounds of the forgetting operation, select the model parameters with the optimal balance between the forgetting effect and the model utility, and calculate the parameter update vector during the forgetting process;

[0088] According to the model parameters with the optimal balance and the parameter update vector, amplify the parameter update vector through parameter extrapolation to extend the forgetting effect to the knowledge related to the target forgetting data set.

[0089] For the specific working method, see the introduction in Embodiment 1, which will not be elaborated here.

[0090] Embodiment 4:

[0091] This implementation provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the electronic device for storing programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the electronic device.

[0092] Moreover, one or more instructions adapted to be loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0093] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the processor loads and executes one or more instructions stored in the computer-readable storage medium to implement the following process:

[0094] Perform a forgetting operation on the target forgetting data set through gradient ascent;

[0095] After the forgetting operation ends, during multiple training rounds of the forgetting operation, select the model parameters with the optimal balance between the forgetting effect and the model utility, and calculate the parameter update vector during the forgetting process;

[0096] According to the model parameters with the optimal balance and the parameter update vector, amplify the parameter update vector through parameter extrapolation to extend the forgetting effect to the knowledge related to the target forgetting dataset.

[0097] For the specific working method, see the introduction in Embodiment 1, which will not be elaborated here.

[0098] Embodiment 5:

[0099] This implementation provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device performs the following process:

[0100] Perform a forgetting operation on the target forgetting dataset through gradient ascent;

[0101] After the forgetting operation ends, during multiple training rounds of the forgetting operation, select the model parameters with the optimal balance between the forgetting effect and the model utility, and calculate the parameter update vector during the forgetting process;

[0102] According to the model parameters with the optimal balance and the parameter update vector, amplify the parameter update vector through parameter extrapolation to extend the forgetting effect to the knowledge related to the target forgetting dataset.

[0103] For the specific working method, see the introduction in Embodiment 1, which will not be elaborated here.

[0104] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this application can be implemented by electronic hardware, or by the combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of this application.

[0105] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data processing device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.

[0106] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An auxiliary method for enhancing the forgetting of sensitive information in a large model by parameter extrapolation, characterized in that: The process includes: Perform a forget operation on the target forget data set through gradient ascent; After the forgetting operation is completed, in multiple training rounds of the forgetting operation, the model parameters that best balance the forgetting effect and model utility are selected, and the parameter update vector in the forgetting process is calculated; According to the optimally balanced model parameters and the parameter update vector, the parameter update vector is amplified by parameter extrapolation to extend the forgetting effect to knowledge related to the target forgotten data set; Construct a synthetic dataset that includes the target forgotten dataset and the related knowledge dataset. The specific method of constructing the dataset is as follows: Target forgetting dataset: contains specific knowledge instances that need to be forgotten; Related knowledge dataset: contains knowledge instances that are logically related to the target forgotten data. These knowledge instances infer the target forgotten data through the model's reasoning mechanism. The dataset consists of several virtual individuals, each of which contains several specific attributes. Each attribute generates two question-answer pairs K1 and K2. The K1 question-answer pair directly describes the target forgotten data, while the K2 question-answer pair contains relevant knowledge and derives the content of K1 through common sense reasoning. The experiment uses the LLaMA-2-7b-chat model as the basic model; The experiment uses the following two indicators to evaluate the forgetting effect: Forgetting quality: The forgetting effect is measured by calculating the ROUGE-L score of the model on the target forgetting dataset. A lower ROUGE-L score indicates a better forgetting effect. Model utility: The overall performance of the model is measured by calculating the ROUGE-L score of the model on the TruthfulQA dataset. A higher ROUGE-L score indicates better model utility.

2. The auxiliary method for enhancing the large model forgetting sensitive information by parameter extrapolation as claimed in claim 1, characterized in that: The target forgetting data set is forgotten by gradient ascent, including: assumed represents a model that has only been fine-tuned without forget training, Represents a model that has only been trained with gradient ascent and forgetfulness. For any sample in the target forgetting set , the parameter update of gradient ascent is expressed as: ; in, for The model parameters, is the learning rate, is the objective function of gradient ascent, In the model The gradient of For the updated Parameters, Indicates that the initial parameters Next, given input Output The conditional probability of .

3. The auxiliary method for enhancing the large model forgetting sensitive information by parameter extrapolation as claimed in claim 1, characterized in that: The parameter update vector represents the direction and magnitude of change of the model parameters when forgetting the target data.

4. The auxiliary method for enhancing the large model forgetting sensitive information by parameter extrapolation according to any one of claims 1 to 3, characterized in that: According to the optimal model parameters and the parameter update vector , amplifying the parameter update vector by parameter extrapolation, including: ,in, Used to control the magnification of the parameter update vector, is the parameter update vector after parameter extrapolation and amplification.

5. An auxiliary system for enhancing the forgetting of sensitive information of a large model by parameter extrapolation, characterized in that: include: The forgetting operation unit is configured to: perform a forgetting operation on a target forgetting data set by gradient ascent; A parameter update vector calculation unit is configured to: after the forgetting operation is completed, in multiple training rounds of the forgetting operation, select the model parameters that best balance the forgetting effect and the model utility, and calculate the parameter update vector in the forgetting process; A parameter extrapolation unit is configured to: amplify the parameter update vector by parameter extrapolation according to the optimally balanced model parameters and the parameter update vector, so as to extend the forgetting effect to knowledge related to the target forgotten data set; Construct a synthetic dataset that includes the target forgotten dataset and the related knowledge dataset. The specific method of constructing the dataset is as follows: Target forgetting dataset: contains specific knowledge instances that need to be forgotten, such as "patient X suffers from diabetes"; Related knowledge dataset: contains knowledge instances that are logically related to the target forgotten data, such as "someone needs to take insulin injections regularly" and "insulin is the standard treatment for diabetes". These knowledge instances infer the target forgotten data through the model's reasoning mechanism; The dataset consists of several virtual individuals, each of which contains several specific attributes. Each attribute generates two question-answer pairs K1 and K2. The K1 question-answer pair directly describes the target forgotten data, while the K2 question-answer pair contains relevant knowledge and derives the content of K1 through common sense reasoning. The experiment uses the LLaMA-2-7b-chat model as the basic model; The experiment uses the following two indicators to evaluate the forgetting effect: Forgetting quality: The forgetting effect is measured by calculating the ROUGE-L score of the model on the target forgetting dataset. A lower ROUGE-L score indicates a better forgetting effect. Model utility: The overall performance of the model is measured by calculating the ROUGE-L score of the model on the TruthfulQA dataset. A higher ROUGE-L score indicates better model utility.

6. The auxiliary system for enhancing the large model forgetting sensitive information through parameter extrapolation as claimed in claim 5, characterized in that: In the forgetting operation unit, the target forgetting data set is subjected to a forgetting operation through gradient ascent, including: assumed represents a model that has only been fine-tuned without forget training, Represents a model that has only been trained with gradient ascent and forgetfulness. For any sample in the target forgetting set , the parameter update of gradient ascent is expressed as: ; in, for The model parameters, is the learning rate, is the objective function of gradient ascent, In the model The gradient of For the updated Parameters, Indicates that the initial parameters Next, given input Output The conditional probability of .

7. The auxiliary system for enhancing the large model forgetting sensitive information by parameter extrapolation according to claim 5 or 6, characterized in that: In the parameter extrapolation unit, specifically, it includes: according to the optimal balanced model parameters and the parameter update vector , amplifying the parameter update vector by parameter extrapolation, including: ,in, Used to control the magnification of the parameter update vector, is the parameter update vector after parameter extrapolation and amplification.

8. A computer device, characterized in that: include: a processor and a computer readable storage medium; a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein the computer program, when executed by the processor, implements the auxiliary method for enhancing the forgetting of sensitive information of a large model by parameter extrapolation as described in any one of claims 1 to 4.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which is suitable for being loaded by a processor and executing the auxiliary method for enhancing the forgetting sensitive information of a large model by parameter extrapolation as described in any one of claims 1 to 4.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program implements the auxiliary method for enhancing the forgetting sensitive information of a large model by parameter extrapolation as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method, system, equipment and medium for quickly training model from partial training set

    CN114611631A

  • Multi-modal big language model knowledge forgetting method based on double mask divergence

    CN118761438A