Forgetting updating method and device of large language model, equipment, storage medium and program product
By evaluating and optimizing some parameters of the large language model, the problem of low positioning accuracy of key parameters in the prior art is solved, and more accurate model updates and lower calculation costs are achieved.
Patent Information
- Application Number
- CN202510285077.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-20
AI Technical Summary
The forget-up update method of existing large language models has the problem of low accuracy in key parameter positioning, resulting in inaccurate model forget-up update.
By determining some parameters of the model to be updated, retained data sets and deleted data sets, input them into the preset forgetting theoretical model for evaluation, the first importance of some parameters to each structure in the model to be updated, and input them into the structural optimization model for model optimization, obtaining the optimized model.
Improve the accuracy of model updates, reduce the time cost and calculation overhead required for forgetting processing, and protect user data privacy.
Smart Images

Figure CN120180128A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of large language models, and in particular to a forgetting update method, apparatus, device, storage medium and program product for a large language model. Background Art
[0002] Large language models have become a transformative technology in the field of artificial intelligence, greatly improving natural language processing capabilities from text generation to simulating human interactions, and performing well in many downstream tasks. However, due to the large scale of the training corpus of large language models, it is almost impossible to completely filter out all potentially dangerous or harmful information on the Internet. Therefore, while the model learns useful knowledge, it may generate bad content.
[0003] To avoid the impact of potentially dangerous information on model accuracy, the existing solution is to first remove dangerous information data and retrain the large language model based on the training data without dangerous information data. Obviously, this will bring huge computational costs and time consumption. Model forgetting aims to eliminate the impact of specific data on the trained model through effective means, thereby avoiding the tedious process of retraining. For example, an efficient parameterized forgetting technique, by identifying the key parameters that affect forgetting, limits the forgetting update to this small number of key parameters, so as to accurately control the forgetting process of the model.
[0004] However, the above method has the problem of low positioning accuracy of key parameters, which leads to inaccurate forgetting updates of the model. Summary of the invention
[0005] Based on this, it is necessary to provide a forgetting update method, device, equipment, storage medium and program product for a large language model that can improve the accuracy of model updating in response to the above technical problems.
[0006] In a first aspect, the present application provides a forgetting update method for a large language model, comprising:
[0007] Determine the model to be updated, obtain some parameters of the model to be updated, retain the data set and delete the data set; the retained data set includes the data set after the abnormal data of the training sample data set is deleted; the training sample data set refers to the data set used when training the model to be updated; the deleted data set refers to the data set composed of abnormal data;
[0008] Input some parameters of the model to be updated, the model to be updated, the retained data set and the deleted data set into the preset forgetting theory model for evaluation, and obtain the first importance degree of some parameters to each structure in the model to be updated;
[0009] Input the first importance degrees of partial parameters for each structure in the model to be updated into a structure optimization model for model optimization, and obtain an optimized model.
[0010] In one embodiment, the above-mentioned inputting the first importance degrees of partial parameters for each structure in the model to be updated into a structure optimization model for model optimization, and obtaining an optimized model includes:
[0011] Determine a mask corresponding to the partial parameters according to the first importance degrees of the partial parameters for each structure in the model to be updated;
[0012] Determine a local optimization strategy according to the mask;
[0013] Determine a target mask corresponding to the partial parameters according to the local optimization strategy, the mask, and a forgetting theory model;
[0014] Input the target mask into the structure optimization model for model optimization, and obtain an optimized model.
[0015] In one embodiment, the above-mentioned determining a target mask corresponding to the partial parameters according to the local optimization strategy, the mask, and a forgetting theory model includes:
[0016] Perform iterative exchange processing on the mask according to the local optimization strategy to obtain a first mask;
[0017] Input the first mask into the forgetting theory model for evaluation to obtain the second importance degrees of the partial parameters for each structure in the model to be updated;
[0018] Correct the first mask according to the second importance degrees and the first importance degrees to obtain a second mask;
[0019] Use the second mask as a new mask, and return to execute the step of performing iterative exchange processing on the mask according to the local optimization strategy to obtain a first mask until the value of the local optimization strategy is a preset value, and use the second mask obtained by the last update as the target mask.
[0020] In one embodiment, the above-mentioned correcting the first mask according to the second importance degrees and the first importance degrees to obtain a second mask includes:
[0021] Determine whether the second importance degree corresponding to the target bit on the first mask is consistent with the corresponding first importance degree;
[0022] If they are consistent, determine the first mask as the second mask;
[0023] If they are inconsistent, correct the mask corresponding to the target bit on the first mask to obtain a corrected second mask.
[0024] In one embodiment, the above method further includes:
[0025] Construct a forgetting theory model according to the correlation relationship between the model to be updated and some parameters of the model to be updated, and the retained data set.
[0026] In one embodiment, the above method further includes:
[0027] Analyze the forgetting theory model according to convex optimization technology to obtain the analyzed forgetting theory model;
[0028] Input some parameters of the model to be updated, the model to be updated, the retained data set, and the deleted data set into a preset forgetting theory model for evaluation, and obtain the first importance degree of some parameters for each structure in the model to be updated, including:
[0029] Input some parameters of the model to be updated, the model to be updated, the retained data set, and the deleted data set into the analyzed forgetting theory model for evaluation, and obtain the first importance degree of some parameters for each structure in the model to be updated.
[0030] In a second aspect, the present application also provides a forgetting update device for a large language model, including:
[0031] A determination module, configured to determine the model to be updated, and obtain some parameters of the model to be updated, a retained data set, and a deleted data set; the retained data set includes the data set after abnormal data is deleted from the training sample data set; the training sample data set refers to the data set used when training the model to be updated; the deleted data set refers to the data set composed of abnormal data;
[0032] An evaluation module, configured to input some parameters of the model to be updated, the model to be updated, the retained data set, and the deleted data set into a preset forgetting theory model for evaluation, and obtain the first importance degree of some parameters for each structure in the model to be updated;
[0033] An optimization module, configured to input the importance degree of some parameters for each structure in the model to be updated into a structure optimization model for model optimization, and obtain an optimized model.
[0034] In a third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0035] A determination module, configured to determine a model to be updated, and obtain partial parameters, a retained dataset, and a deleted dataset of the model to be updated; the retained dataset includes a dataset obtained by deleting abnormal data from a training sample dataset; the training sample dataset refers to the dataset used for training the model to be updated; the deleted dataset refers to the dataset composed of abnormal data;
[0036] An evaluation module, configured to input the partial parameters, the model to be updated, the retained dataset, and the deleted dataset of the model to be updated into a preset forgetting theory model for evaluation, to obtain the first importance degree of the partial parameters for each structure in the model to be updated;
[0037] An optimization module, configured to input the importance degree of the partial parameters for each structure in the model to be updated into a structure optimization model for model optimization, to obtain an optimized model.
[0038] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0039] A determination module, configured to determine a model to be updated, and obtain partial parameters, a retained dataset, and a deleted dataset of the model to be updated; the retained dataset includes a dataset obtained by deleting abnormal data from a training sample dataset; the training sample dataset refers to the dataset used for training the model to be updated; the deleted dataset refers to the dataset composed of abnormal data;
[0040] An evaluation module, configured to input the partial parameters, the model to be updated, the retained dataset, and the deleted dataset of the model to be updated into a preset forgetting theory model for evaluation, to obtain the first importance degree of the partial parameters for each structure in the model to be updated;
[0041] An optimization module, configured to input the importance degree of the partial parameters for each structure in the model to be updated into a structure optimization model for model optimization, to obtain an optimized model.
[0042] In a fifth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0043] A determination module, configured to determine a model to be updated, and obtain partial parameters, a retained dataset, and a deleted dataset of the model to be updated; the retained dataset includes a dataset obtained by deleting abnormal data from a training sample dataset; the training sample dataset refers to the dataset used for training the model to be updated; the deleted dataset refers to the dataset composed of abnormal data;
[0044] An evaluation module for inputting partial parameters of a model to be updated, the model to be updated, a retained dataset, and a deleted dataset into a preset forgetting theory model for evaluation, so as to obtain the first importance degree of the partial parameters for each structure in the model to be updated;
[0045] An optimization module for inputting the importance degree of the partial parameters for each structure in the model to be updated into a structure optimization model for model optimization, so as to obtain an optimized model.
[0046] The above forgetting update method, device, equipment, storage medium and program product of the large language model determine the model to be updated, and obtain partial parameters, a retained dataset and a deleted dataset of the model to be updated. Input the partial parameters of the model to be updated, the model to be updated, the retained dataset and the deleted dataset into a preset forgetting theory model for evaluation, so as to obtain the first importance degree of the partial parameters for each structure in the model to be updated. Input the first importance degree of the partial parameters for each structure in the model to be updated into a structure optimization model for model optimization, so as to obtain an optimized model; the retained dataset includes the dataset after deleting abnormal data from the training sample dataset; the training sample dataset refers to the dataset used for training the model to be updated; the deleted dataset refers to the dataset composed of abnormal data. The above method analyzes the importance degree of each structure in the model to be updated, and optimizes the model to be updated according to the analysis result, fully considering the complex mutual relationship between the internal structures of the large model, ensuring that the parameters related to the data to be forgotten can be accurately identified. On this basis, the forgetting update focuses on the key parameters, thereby reducing the time cost and computing overhead required for forgetting processing, and protecting the user's data privacy at the same time. Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained without creative efforts.
[0048] Figure 1 It is an application environment diagram of the forgetting update method of the large language model in an embodiment;
[0049] Figure 2 It is a flow diagram of the forgetting update method of the large language model in an embodiment;
[0050] Figure 3 It is a flow diagram of the forgetting update method of the large language model in another embodiment;
[0051] Figure 4It is a flowchart of a forgetting update method of a large language model in another embodiment;
[0052] Figure 5 It is a flowchart of a forgetting update method of a large language model in another embodiment;
[0053] Figure 6 It is a flowchart of a forgetting update method of a large language model in another embodiment;
[0054] Figure 7 It is a flowchart of a forgetting update method of a large language model in another embodiment;
[0055] Figure 8 It is a structural block diagram of a forgetting update device for a large language model in one embodiment;
[0056] Figure 9 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0058] Large language models have become a transformative technology in the field of artificial intelligence, greatly improving natural language processing capabilities from text generation to simulating human interactions, and excelling in many downstream tasks. However, due to the large size of the training corpus of large language models, it is almost impossible to completely filter out all potentially dangerous or harmful information on the Internet. Therefore, while the model learns useful knowledge, it may generate bad content, and as the model's capabilities continue to increase, these potential risks become more severe, making it easier for malicious actors to obtain dangerous information. The act of malicious actors obtaining dangerous information not only weakens the effectiveness of the model in practical applications, but may also cause risks.
[0059] To avoid the impact of potentially dangerous information on model accuracy, the existing solution is to first remove the dangerous information data and retrain the large language model based on the training data without the dangerous information data. Obviously, this will bring huge computing costs and time consumption, which is unbearable for model providers.
[0060] Therefore, the field of model forgetting came into being. Model forgetting aims to eliminate the influence of specific data on the trained model through effective means, thereby avoiding the tedious process of retraining. However, the huge scale of the model poses a huge challenge to model forgetting technology.
[0061] Currently, recent research has proposed efficient parameterized forgetting techniques. By identifying the key parameters that affect forgetting, the forgetting update is restricted to this small set of key parameters to precisely control the forgetting process of the model. This efficient parameterized forgetting technique evaluates the importance of parameters through different strategies, aiming to achieve selective sparse updates, thereby reducing computational overhead and improving processing efficiency. However, current efficient parameterized forgetting techniques rely mostly on heuristic or empirical strategies to identify important parameters when evaluating the importance of parameters. However, the above methods have the problem of low accuracy in locating key parameters, resulting in inaccurate forgetting updates of the model. This application aims to solve this problem.
[0062] After introducing the background technology of the forgetting update method for the large language model provided by the embodiments of this application above, below, the implementation environment involved in the forgetting update method for the large language model provided by the embodiments of this application will be briefly described. The forgetting update method for the large language model provided by the embodiments of this application can be applied to an Figure 1 implementation environment as shown. This implementation environment includes a server 104, and the server 104 can be implemented by an independent server 104 or a server cluster composed of multiple servers 104. The data storage system 102 can store the data that the server 104 needs to process. The data storage system 102 can be integrated on the server 104, or placed in the cloud or other network servers. Among them, the server 104 can obtain the model to be updated, partial parameters of the model to be updated, and the retained dataset, and input the model to be updated, the partial parameters of the model to be updated, and the retained dataset into the forgetting theory model for evaluation to obtain the first importance degree of the partial parameters for each structure in the updated model. Then, based on the first importance degree of the partial parameters for each structure in the updated model, the model to be updated is optimized to obtain an optimized model.
[0063] Those skilled in the art can understand that Figure 1 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0064] After introducing the application scenario of the forgetting update method for the large language model provided by the embodiments of this application above, below, the forgetting update method for the large language model described in this application will be introduced in detail.
[0065] In one embodiment, as Figure 2 shown, a forgetting update method for a large language model is provided. Taking the application of this method to the Figure 1 server in as an example, the method includes the following steps:
[0066] S201. Determine the model to be updated, and obtain partial parameters, a retention dataset, and a deletion dataset of the model to be updated.
[0067] Among them, the model to be updated refers to a model that has been trained and has dangerous information in the training data; the partial parameters of the model to be updated refer to a part of the model parameters of the model to be updated.
[0068] Among them, the retention dataset includes the dataset after deleting abnormal data from the training sample dataset, and the training sample dataset refers to the dataset used when training the model to be updated. It should be noted that the deletion dataset refers to the dataset composed of the abnormal data, that is, the dataset with dangerous information.
[0069] In the embodiments of the present application, when it is determined that the model to be updated needs to be updated, partial parameters, a retention dataset, and a deletion dataset of the model to be updated can be obtained. Optionally, after the model training is completed, all the model parameters and the training dataset during the model training process can be stored in a preset database. When the model needs to be updated, first determine the abnormal data in the training dataset, and obtain partial parameters, a retention dataset, and a deletion dataset of the model to be updated from the preset database.
[0070] S202. Input the partial parameters, the model to be updated, the retention dataset, and the deletion dataset of the model to be updated into a preset forgetting theory model for evaluation, and obtain the first importance degree of the partial parameters for each structure in the model to be updated.
[0071] Among them, the preset forgetting theory model is used to evaluate the importance degree of each structure in the model to be updated.
[0072] In the embodiments of the present application, after obtaining the partial parameters, the model to be updated, the retention dataset, and the deletion dataset of the model to be updated as described above, the partial parameters, the model to be updated, the retention dataset, and the deletion dataset of the model to be updated can be input into a preset forgetting theory model for evaluation, and the first importance degree of the partial parameters for each structure in the model to be updated can be obtained.
[0073] S203. Input the first importance degree of the partial parameters for each structure in the model to be updated into a structure optimization model for model optimization, and obtain an optimized model.
[0074] Among them, the structure optimization model refers to a model used to optimize the model to be updated according to the importance degree of each structure in the model to be updated.
[0075] In an embodiment of the present application, after obtaining the first importance degrees of the partial parameters for each structure in the model to be updated, the first importance degrees of the partial parameters for each structure in the model to be updated can be input into a structure optimization model for model optimization to obtain an optimized model.
[0076] The forgetting update method for a large language model provided by an embodiment of the present application determines a model to be updated, and obtains partial parameters, a retention dataset, and a deletion dataset of the model to be updated. The partial parameters, the model to be updated, the retention dataset, and the deletion dataset of the model to be updated are input into a preset forgetting theory model for evaluation to obtain the first importance degrees of the partial parameters for each structure in the model to be updated. The first importance degrees of the partial parameters for each structure in the model to be updated are input into a structure optimization model for model optimization to obtain an optimized model; the retention dataset includes the dataset obtained by deleting abnormal data from the training sample dataset, the training sample dataset refers to the dataset used when training the model to be updated, and the deletion dataset refers to the dataset composed of abnormal data. The above method analyzes the importance degrees of each structure in the model to be updated, and performs model optimization on the model to be updated according to the analysis results, fully considering the complex mutual relationships between the internal structures of the large model, ensuring that the parameters related to the data to be forgotten can be accurately identified. On this basis, the forgetting update focuses on key parameters, thereby reducing the time cost and computing overhead required for forgetting processing, and protecting the user's data privacy at the same time.
[0077] In one embodiment, based on Figure 2 the embodiment shown, the process of obtaining the optimized model can be described. As Figure 3 shown, the above S203 "input the first importance degrees of the partial parameters for each structure in the model to be updated into a structure optimization model for model optimization to obtain an optimized model" includes:
[0078] S301. Determine the masks corresponding to the partial parameters according to the first importance degrees of the partial parameters for each structure in the model to be updated.
[0079] In an embodiment of the present application, after obtaining the first importance degrees of the partial parameters for each structure in the model to be updated, the masks corresponding to the partial parameters can be determined according to the first importance degrees of the partial parameters for each structure in the model to be updated.
[0080] It should be noted that the number of partial parameters is equal to the number of explicit masks in the mask. For example, if the model parameters of the model to be updated are 10 in total and the number of partial parameters is 5 in total, then the mask corresponding to the partial parameters can be 1010101010.
[0081] Optionally, construct a deletion request dataset and a retention dataset according to the user's data deletion request, and use a preset forgetting theory model to calculate the sum of gradients of each structure on the deletion request dataset and the diagonalized Fisher information matrix on the retention dataset to obtain a mask. It should be understood that the constraints on some parameters ensure the sparse update of the model, focusing on key parameters rather than full-scale update. In addition, the diagonalized Fisher information matrix focuses on the impact of the structure itself on the data and ignores the relationship of the interaction between structures, and further optimization is required to obtain an accurate structure mask.
[0082] S302. Determine a local optimization strategy according to the mask.
[0083] In the embodiments of the present application, after the masks corresponding to some parameters are determined above, a local optimization strategy can be determined according to the masks. Optionally, a local optimization strategy can be determined according to the number of implicit masks and explicit masks in the masks.
[0084] For example, if the mask corresponding to some parameters is 1010101010, the mask 1010101010 can be converted into 0110101010, 1100101010, 1110001010, 1110100010, 1110101000. The mask 1010101010 can also be converted into 0011101010, 1001101010, 1011001010, 1011100010, 1011101000. The mask 1010101010 can also be converted into 0010111010, 1000111010, 1010011010, 1010110010, 1010111000. The mask 1010101010 can also be converted into 0010101110, 1000101110, 1010001110, 1010100110, 1010101100. The mask 1010101010 can also be converted into 0010101011, 1000101011, 1010001011, 1010100011, 1010101001, that is, the determined local optimization strategy is 25 times.
[0085] S303. Determine the target mask corresponding to some parameters according to the local optimization strategy, the mask and the forgetting theory model.
[0086] In the embodiments of the present application, after the local optimization strategy, the mask and the forgetting theory model are determined above, the target mask corresponding to some parameters can be determined according to the local optimization strategy, the mask and the forgetting theory model.
[0087] Optionally, for the interrelationships between the intra-layer structures of the model to be updated, diagonalizing the Fisher information matrix is replaced by block-diagonalizing the Fisher information matrix, with each block associated with each neural network layer. Subsequently, the structure mask is rearranged to optimize the objective. For the specific process, refer to Equation (1):
[0088]
[0089] wherein represents the network layer that needs to be further optimized. According to the masks corresponding to some parameters, in each neural network layer, the unselected structure with the highest importance degree in the current mask is iteratively exchanged with each selected structure (the structure with the target bit being the hidden mask in the mask) to capture the mutual influence between different structures in each layer, thereby obtaining the rearranged target mask. It should be understood that the model sequentially passes through each network layer during inference. Therefore, optimizing the structure mask within the layer can effectively optimize the objective.
[0090] Optionally, a method for determining the target mask corresponding to some parameters is provided below. Refer to Figure 4 , that is, in the above-mentioned S303, "determine the target mask corresponding to some parameters according to the local optimization strategy, mask, and forgetting theory model", including:
[0091] S401. Perform iterative exchange processing on the mask according to the local optimization strategy to obtain the first mask.
[0092] In the embodiments of the present application, after the mask is determined as above, iterative exchange processing can be performed on the mask based on the local optimization strategy to obtain the first mask.
[0093] Continuing with the above example, the mask corresponding to some parameters is 1010101010, and the first mask can be 0110101010, 1100101010, 1110001010, 1110100010, 1110101000, 0011101010, 1001101010, 1011001010, 1011100010, 1011101000, 0010111010, 1000111010, 1010011010, 1010110010, 1010111000, 0010101110, 1000101110, 1010001110, 1010100110, 1010101100, 0010101011, 1000101011, 1010001011, 1010100011, 1010101001.
[0094] S402. Input the first mask into the forgetting theory model for evaluation to obtain the second importance degree of some parameters for each structure in the model to be updated.
[0095] In the embodiment of the present application, after the first mask is determined as above, the forgetting update structure of the model to be updated can be determined according to the first mask, and the forgetting update structure of the model to be updated is input into the forgetting theory model for evaluation to obtain the second importance degree of each update structure in the model to be updated with respect to some parameters.
[0096] S403. Modify the first mask according to the second importance degree and the first importance degree to obtain a second mask.
[0097] In the embodiment of the present application, after the second importance degree of each update structure in the model to be updated and the first importance degree of each structure in the model to be updated are determined as above, the second importance degree and the first importance degree can be compared to obtain a comparison result, and the first mask is modified according to the comparison result to obtain a second mask.
[0098] Optionally, the process of modifying the first mask to obtain the second mask can be described. Refer to Figure 5 , that is, the above S403 "Modify the first mask according to the second importance degree and the first importance degree to obtain a second mask" includes:
[0099] S501. Determine whether the second importance degree corresponding to the target bit on the first mask is consistent with the corresponding first importance degree. If they are consistent, execute S502 below; if they are inconsistent, execute S503 below.
[0100] Wherein, the target bit is the position where the visibility of the mask is modified.
[0101] In the embodiment of the present application, after the second importance degree of each update structure in the model to be updated and the first importance degree of each structure in the model to be updated are determined as above, determine whether the second importance degree corresponding to the target bit on the first mask is consistent with the first importance degree corresponding to the target bit. And when the second importance degree corresponding to the target bit is consistent with the first importance degree corresponding to the target bit, execute S502 below, and when the second importance degree corresponding to the target bit is inconsistent with the first importance degree corresponding to the target bit, execute S503 below.
[0102] S502. Determine the first mask as the second mask.
[0103] In the embodiment of the present application, when the second importance degree corresponding to the target bit is consistent with the first importance degree corresponding to the target bit, directly determine the first mask as the second mask.
[0104] S503. Modify the mask corresponding to the target bit on the first mask to obtain a modified second mask.
[0105] In the embodiment of the present application, when the second importance level corresponding to the above target bit is inconsistent with the first importance level corresponding to the target bit, the mask corresponding to the target bit on the first mask can be corrected. That is, if the mask corresponding to the target bit of the first mask is a hidden mask, the mask corresponding to the target bit on the first mask is corrected to an explicit mask; if the mask corresponding to the target bit of the first mask is an explicit mask, the mask corresponding to the target bit on the first mask is corrected to a hidden mask, so as to obtain the corrected second mask.
[0106] In summary, a process for determining the second mask from the first mask is provided.
[0107] S404. Take the second mask as the new mask, and return to execute the step of iteratively swapping the mask according to the local optimization strategy to obtain the first mask until the value of the local optimization strategy is a preset value, and take the second mask obtained by the last update as the target mask.
[0108] Wherein, the preset value is 0.
[0109] In the embodiment of the present application, after the above-mentioned second mask is determined, the second mask can be used as the new mask, and return to execute the step of iteratively swapping the mask according to the local optimization strategy to obtain the first mask until the value of the local optimization strategy is a preset value, and take the second mask obtained by the last update as the target mask.
[0110] In summary, a method for determining the target mask corresponding to some parameters is provided.
[0111] S304. Input the target mask into the structure optimization model for model optimization to obtain the optimized model.
[0112] In the embodiment of the present application, after the above-mentioned target mask is obtained, the target mask can be input into the structure optimization model for model optimization to obtain the optimized model.
[0113] Optionally, after obtaining the target mask according to the above perform sparse update on the Newton step forgetting to obtain the optimized model , see the following formula (2):
[0114]
[0115] Wherein, represents the Fisher information matrix with respect to the model parameters, represents the Hadamard product.
[0116] It should be noted that Newton step forgetting is a typical algorithm in forgetting updates, which involves the calculation of the Hessian curvature matrix of the model parameters. Given its computational complexity, the present invention uses the Fisher information matrix to approximate the empirical risk function The second-order partial derivative with respect to the model parameters.
[0117] In this embodiment, the structure mask is initialized by calculating the structure score, and the mask is rearranged in the warm start phase. Subsequently, the structure mask will guide the subsequent forgetting update process, thereby ensuring the efficiency of forgetting.
[0118] Based on this framework, the present invention uses a structure mask to implement forgetting updates in two types of consecutive deletion request scenarios. The following will be combined with the attached Figure 2 For description.
[0119] The model optimization method provided by the embodiments of the present application analyzes the importance degree of each structure in the model to be updated, and optimizes the model to be updated according to the analysis results, fully considering the complex mutual relationship between the internal structures of the large model, ensuring that the parameters related to the data to be forgotten can be accurately identified. On this basis, the forgetting update focuses on the key parameters, thereby reducing the time cost and computational overhead required for forgetting processing, and protecting the user's data privacy at the same time.
[0120] In one embodiment, on the basis of any of the embodiments shown in Figures 2 - 5 As shown in Figure 6 The above method further includes:
[0121] S204. Construct a forgetting theory model according to the correlation relationship between the model to be updated and some parameters of the model to be updated, and the retained data set.
[0122] In the embodiments of the present application, for the goal of data forgetting in large language models, a pair of learnable binary structure masks are introduced to identify the attention heads and filters in the model. By constraining the total amount of structural parameter updates, a model forgetting theory with the retained data set as the optimization goal and the structure mask as the core is constructed.
[0123] In the embodiments of the present application, the goal of data forgetting in large language models is to minimize the empirical risk on the retained data set. Starting from the optimal model trained on the original data set After optimizing the forgetting goal, the forgetting model is obtained. In order to accurately locate the key structural parameters during the forgetting process, a pair of learnable binary structure masks are introduced to identify the attention heads and filters in the model. In this way, the forgetting goal is transformed into a model forgetting theory with the structure mask m as the core. The specific process is shown in the following formula (3):
[0124] (3);
[0125] Among them, represents the empirical risk of the overall dataset, is the optimal model trained on the original dataset, is the retained dataset, is the size of the data, represents the th structure mask, 1 indicates that the structure needs to be updated, and conversely 0 indicates that the structure needs to be frozen, represents the proportion of the model structure that needs to be frozen. It should be understood that the core goal of this forgetting theory is to accurately locate the key structural parameters related to forgetting, is the structure mask to be optimized, and based on this mask, the model can efficiently achieve forgetting updates.
[0126] The construction method of the forgetting theory model provided by the embodiments of the present application provides a data basis for subsequently determining the first importance degree of each structure in the model to be updated based on some parameters of the forgetting theory model.
[0127] In one embodiment, based on any of the embodiments shown in Figure 6 as shown in Figure 7 the above method further includes:
[0128] S205. Analyze the forgetting theory model according to convex optimization techniques to obtain the analyzed forgetting theory model.
[0129] In the embodiments of the present application, when analyzing the forgetting theory model, a diagonalized Fisher information matrix is used to approximate the complex Hessian curvature matrix, thereby obtaining the analyzed forgetting theory model. The analyzed forgetting theory model consists of two modules, one is the gradient part related to the target forgetting data, and the other is the diagonalized Fisher information matrix part related to the retained data.
[0130] In this embodiment, analyzing the forgetting theory model involves techniques such as Taylor expansion and convex optimization, and a diagonalized Fisher information matrix is used to approximate the complex Hessian curvature matrix. The specific process is shown in the following formula (4) as:
[0131]
[0132] Among them, represents the all-1 vector of the structure mask, represents the forgetting dataset, is the gradient of the model with respect to the structure mask , is the loss function of the model, represents with respect to the structure mask The Fisher information matrix is used to approximate the empirical risk function Regarding The second-order partial derivative. It should be understood that the Is a binary mask with values of 0 or 1. The optimized structure mask is based on two modules of the analysis: the gradient part related to the target forgotten data And the diagonal Fisher information matrix part related to the retained data , and the interaction between these two modules constructs the analyzed forgetting theory model.
[0133] The above S202 "Input the partial parameters of the model to be updated, the model to be updated, the retained dataset, and the deleted dataset into the preset forgetting theory model for evaluation, and obtain the first importance degree of the partial parameters for each structure in the model to be updated" includes:
[0134] S202. Input the partial parameters of the model to be updated, the model to be updated, the retained dataset, and the deleted dataset into the analyzed forgetting theory model for evaluation, and obtain the first importance degree of the partial parameters for each structure in the model to be updated.
[0135] In the embodiment of the present application, after obtaining the partial parameters of the model to be updated, the model to be updated, the retained dataset, and the deleted dataset as described above, the partial parameters of the model to be updated, the model to be updated, the retained dataset, and the deleted dataset can be input into the analyzed forgetting theory model for evaluation, and the first importance degree of the partial parameters for each structure in the model to be updated can be obtained.
[0136] Based on this framework, the present invention uses a structure mask to implement forgetting updates in two types of continuous deletion request scenarios.
[0137] (a) Continuous forgetting without storage. During the continuous forgetting without storage, whenever the user makes a deletion request, the information contained in the target deleted data and the retained data will be calculated. This information will be used to calculate and optimize the structure mask, and then the latest model will be sparsely updated through the structure mask. After the update is completed, any information in the deleted data will not be retained. The specific process is as follows in formula (5):
[0138]
[0139] Where, Is the timestamp, indicating the th deletion request, Indicates The model after the Indicates The structure mask calculated at time And Respectively indicate Reserved data set and forgotten data at a moment.
[0140] (b)Storage-assisted continuous forgetting. In storage-assisted continuous forgetting, whenever a user makes a deletion request, the information contained in the target deleted data will be calculated. This information will be combined with the data information stored in the memory to calculate and optimize the structure mask, and then further perform forgetting updates on the original model The specific process is as shown in formula (6) below:
[0141] (6);
[0142] where and are the data stored at time These stored data are updated after forgetting at time Each execution of the deletion request is performed on the original model which ensures the performance of the model after forgetting. It should be clear that after each forgetting update is executed, the information stored in the memory will be updated in real time to handle the next data deletion request.
[0143] The present invention is not limited to the above embodiments. All other embodiments obtained by those of ordinary skill in the art in a manner identical or similar to the above embodiments of the present invention without creative efforts are within the protection scope of the present invention patent.
[0144] The parsing method of the forgetting theory model provided by the embodiments of the present application provides a data basis for determining the first importance degree of each structure in the model to be updated based on the forgetting theory model subsequently.
[0145] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0146] Based on the same inventive concept, an embodiment of the present application further provides a forgetting update device for a model that implements the forgetting update method of the large language model involved above. The implementation solutions provided by this device to solve problems are similar to those described in the above method. Therefore, the specific limitations in one or more embodiments of the forgetting update device for the model provided below can refer to the limitations on the forgetting update method of the large language model in the above text, and will not be elaborated here.
[0147] In an exemplary embodiment, as Figure 8 shown, a forgetting update device for a model is provided, including: a determination module 10, an evaluation module 11, and an optimization module 12, where:
[0148] The determination module 10 is configured to determine the model to be updated, and obtain partial parameters, a retention dataset, and a deletion dataset of the model to be updated; the retention dataset includes the dataset after deleting abnormal data from the training sample dataset; the training sample dataset refers to the dataset used when training the model to be updated; the deletion dataset refers to the dataset composed of abnormal data.
[0149] The evaluation module 11 is configured to input the partial parameters of the model to be updated, the model to be updated, the retention dataset, and the deletion dataset into a preset forgetting theory model for evaluation, and obtain the first importance degree of the partial parameters for each structure in the model to be updated.
[0150] The optimization module 12 is configured to input the first importance degree of the partial parameters for each structure in the model to be updated into a structure optimization model for model optimization, and obtain the optimized model.
[0151] In an exemplary embodiment, the above optimization module 12 includes: a first determination unit, a second determination unit, a third determination unit, and an optimization unit, where:
[0152] The first determination unit is specifically configured to determine the mask corresponding to the partial parameters according to the first importance degree of the partial parameters for each structure in the model to be updated;
[0153] The second determination unit is specifically configured to determine the local optimization strategy according to the mask;
[0154] The third determination unit is specifically configured to determine the target mask corresponding to the partial parameters according to the local optimization strategy, the mask, and the forgetting theory model;
[0155] The optimization unit is specifically configured to input the target mask into the structure optimization model for model optimization, and obtain the optimized model.
[0156] In an exemplary embodiment, the above-mentioned third determination unit is further specifically configured to perform iterative exchange processing on the mask according to a local optimization strategy to obtain a first mask; input the first mask into a forgetting theory model for evaluation to obtain the second importance degree of each structure in the model to be updated with respect to the partial parameters to be updated; correct the first mask according to the second importance degree and the first importance degree to obtain a second mask; use the second mask as a new mask, and return to execute the step of performing iterative exchange processing on the mask according to the local optimization strategy to obtain a first mask until the value of the local optimization strategy is a preset value, and use the second mask obtained by the last update as the target mask.
[0157] In an exemplary embodiment, the above-mentioned third determination unit is further specifically configured to determine whether the second importance degree corresponding to the target bit on the first mask is consistent with the corresponding first importance degree; if they are consistent, determine the first mask as the second mask; if they are inconsistent, correct the mask corresponding to the target bit on the first mask to obtain a corrected second mask.
[0158] In an exemplary embodiment, the above-mentioned device further includes: a construction module, configured to construct a forgetting theory model according to the correlation between the model to be updated and the partial parameters of the model to be updated, and a reserved data set.
[0159] In an exemplary embodiment, the above-mentioned device further includes: an analysis module, configured to analyze the forgetting theory model according to convex optimization technology to obtain an analyzed forgetting theory model;
[0160] The above-mentioned evaluation module 20 is further configured to input the partial parameters of the model to be updated, the model to be updated, the reserved data set, and the deleted data set into the analyzed forgetting theory model for evaluation to obtain the first importance degree of each structure in the model to be updated with respect to the partial parameters to be updated.
[0161] Each module in the above-mentioned model forgetting update device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0162] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 9As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data of the model to be updated. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes a method for the model to be updated.
[0163] Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0164] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0165] Determine the model to be updated, and obtain some parameters, a retention data set, and a deletion data set of the model to be updated; the retention data set includes the data set after abnormal data is deleted from the training sample data set; the training sample data set refers to the data set used when training the model to be updated; the deletion data set refers to the data set composed of abnormal data;
[0166] Input some parameters, the model to be updated, the retention data set, and the deletion data set of the model to be updated into a preset forgetting theory model for evaluation, and obtain the first importance degree of some parameters for each structure in the model to be updated;
[0167] Input the first importance degree of some parameters for each structure in the model to be updated into a structure optimization model for model optimization, and obtain an optimized model.
[0168] In one embodiment, when the processor executes the computer program, the following steps are also implemented:
[0169] Determine the mask corresponding to some parameters according to the first importance degree of some parameters for each structure in the model to be updated;
[0170] Determine a local optimization strategy according to a mask;
[0171] Determine a target mask corresponding to some parameters according to the local optimization strategy, the mask, and the forgetting theory model;
[0172] Input the target mask into a structure optimization model for model optimization to obtain an optimized model.
[0173] In one embodiment, when the processor executes a computer program, the following steps are further implemented:
[0174] Perform iterative exchange processing on the mask according to the local optimization strategy to obtain a first mask;
[0175] Input the first mask into the forgetting theory model for evaluation to obtain the second importance degree of each structure in the model to be updated corresponding to some parameters;
[0176] Correct the first mask according to the second importance degree and the first importance degree to obtain a second mask;
[0177] Use the second mask as a new mask, and return to execute the step of performing iterative exchange processing on the mask according to the local optimization strategy to obtain a first mask until the value of the local optimization strategy is a preset value, and use the second mask obtained by the last update as the target mask.
[0178] In one embodiment, when the processor executes a computer program, the following steps are further implemented:
[0179] Determine whether the second importance degree corresponding to the target bit on the first mask is consistent with the corresponding first importance degree;
[0180] If they are consistent, determine the first mask as the second mask;
[0181] If they are inconsistent, correct the mask corresponding to the target bit on the first mask to obtain a corrected second mask.
[0182] In one embodiment, when the processor executes a computer program, the following steps are further implemented:
[0183] Construct a forgetting theory model according to the correlation relationship between the model to be updated and some parameters of the model to be updated, and the retained data set.
[0184] In one embodiment, when the processor executes a computer program, the following steps are further implemented:
[0185] Analyze the forgetting theory model according to convex optimization techniques to obtain an analyzed forgetting theory model;
[0186] Input the partial parameters of the model to be updated, the model to be updated, the retained dataset, and the deleted dataset into a preset forgetting theory model for evaluation, to obtain the first importance degree of the partial parameters for each structure in the model to be updated, including:
[0187] Input the partial parameters of the model to be updated, the model to be updated, the retained dataset, and the deleted dataset into the parsed forgetting theory model for evaluation, to obtain the first importance degree of the partial parameters for each structure in the model to be updated.
[0188] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0189] Determine the model to be updated, and obtain the partial parameters, the retained dataset, and the deleted dataset of the model to be updated; the retained dataset includes the dataset after deleting abnormal data from the training sample dataset; the training sample dataset refers to the dataset used when training the model to be updated; the deleted dataset refers to the dataset composed of abnormal data;
[0190] Input the partial parameters of the model to be updated, the model to be updated, the retained dataset, and the deleted dataset into a preset forgetting theory model for evaluation, to obtain the first importance degree of the partial parameters for each structure in the model to be updated;
[0191] Input the first importance degree of the partial parameters for each structure in the model to be updated into a structure optimization model for model optimization, to obtain an optimized model.
[0192] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0193] Determine the mask corresponding to the partial parameters according to the first importance degree of the partial parameters for each structure in the model to be updated;
[0194] Determine the local optimization strategy according to the mask;
[0195] Determine the target mask corresponding to the partial parameters according to the local optimization strategy, the mask, and the forgetting theory model;
[0196] Input the target mask into a structure optimization model for model optimization, to obtain an optimized model.
[0197] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0198] Perform iterative exchange processing on the mask according to the local optimization strategy to obtain the first mask;
[0199] Input the first mask into the forgetting theory model for evaluation to obtain the second importance degree of each structure in the model to be updated for the partial parameters to be updated;
[0200] Modify the first mask according to the second importance degree and the first importance degree to obtain the second mask;
[0201] Take the second mask as the new mask, and return to execute the step of iteratively swapping the mask according to the local optimization strategy to obtain the first mask until the value of the local optimization strategy is the preset value, and take the second mask obtained by the last update as the target mask.
[0202] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0203] Determine whether the second importance degree corresponding to the target bit on the first mask is consistent with the corresponding first importance degree;
[0204] If they are consistent, determine the first mask as the second mask;
[0205] If they are inconsistent, modify the mask corresponding to the target bit on the first mask to obtain the modified second mask.
[0206] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0207] Construct a forgetting theory model according to the correlation relationship between the model to be updated and the partial parameters of the model to be updated, and the retained data set.
[0208] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0209] Analyze the forgetting theory model according to convex optimization technology to obtain the analyzed forgetting theory model;
[0210] Input the partial parameters of the model to be updated, the model to be updated, the retained data set and the deleted data set into the preset forgetting theory model for evaluation to obtain the first importance degree of the partial parameters for each structure in the model to be updated, including:
[0211] Input the partial parameters of the model to be updated, the model to be updated, the retained data set and the deleted data set into the analyzed forgetting theory model for evaluation to obtain the first importance degree of the partial parameters for each structure in the model to be updated.
[0212] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0213] Determine the model to be updated, and obtain partial parameters, a retention dataset, and a deletion dataset of the model to be updated; the retention dataset includes the dataset obtained by deleting abnormal data from the training sample dataset; the training sample dataset refers to the dataset used for training the model to be updated; the deletion dataset refers to the dataset composed of abnormal data;
[0214] Input the partial parameters, the model to be updated, the retention dataset, and the deletion dataset of the model to be updated into a preset forgetting theory model for evaluation, and obtain the first importance degree of each structure in the model to be updated corresponding to the partial parameters;
[0215] Input the first importance degree of each structure in the model to be updated corresponding to the partial parameters into a structure optimization model for model optimization, and obtain an optimized model.
[0216] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0217] Determine the mask corresponding to the partial parameters according to the first importance degree of each structure in the model to be updated corresponding to the partial parameters;
[0218] Determine a local optimization strategy according to the mask;
[0219] Determine the target mask corresponding to the partial parameters according to the local optimization strategy, the mask, and the forgetting theory model;
[0220] Input the target mask into a structure optimization model for model optimization, and obtain an optimized model.
[0221] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0222] Perform iterative exchange processing on the mask according to the local optimization strategy to obtain a first mask;
[0223] Input the first mask into the forgetting theory model for evaluation, and obtain the second importance degree of each structure in the model to be updated corresponding to the partial parameters;
[0224] Correct the first mask according to the second importance degree and the first importance degree to obtain a second mask;
[0225] Use the second mask as a new initial mask, and return to execute the step of performing iterative exchange processing on the mask according to the local optimization strategy to obtain a first mask until the value of the local optimization strategy is a preset value, and use the second mask updated last time as the target mask.
[0226] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0227] Determine whether the second importance degree corresponding to the target bit on the first mask is consistent with the corresponding first importance degree;
[0228] If they are consistent, determine the first mask as the second mask;
[0229] If they are inconsistent, correct the mask corresponding to the target bit on the first mask to obtain the corrected second mask.
[0230] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0231] Construct a forgetting theory model according to the association relationship between the model to be updated and some parameters of the model to be updated, and the retained data set.
[0232] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0233] Analyze the forgetting theory model according to convex optimization technology to obtain the analyzed forgetting theory model;
[0234] Input some parameters of the model to be updated, the model to be updated, and the retained data set into the preset forgetting theory model for evaluation to obtain the first importance degree of some parameters for each structure in the model to be updated, including:
[0235] Input some parameters of the model to be updated, the model to be updated, and the retained data set into the analyzed forgetting theory model for evaluation to obtain the first importance degree of some parameters for each structure in the model to be updated.
[0236] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0237] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0238] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0239] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A forgetting update method for a large model, characterized in that: The method comprises: Determine the model to be updated, and obtain some parameters of the model to be updated, a reserved data set and a deleted data set; the reserved data set includes a data set after deleting abnormal data from the training sample data set; the training sample data set refers to the data set used when training the model to be updated; the deleted data set refers to the data set composed of the abnormal data; Inputting some parameters of the model to be updated, the model to be updated, the retained data set and the deleted data set into a preset forgetting theory model for evaluation, and obtaining a first importance degree of the some parameters to each structure in the model to be updated; The first importance degree of the partial parameters to each structure in the model to be updated is input into the structure optimization model to perform model optimization, thereby obtaining an optimized model.
2. The method according to claim 1, characterized in that: The first importance degree of the partial parameters to each structure in the model to be updated is input into the structure optimization model to perform model optimization to obtain the optimized model, including: Determining masks corresponding to the partial parameters according to a first importance degree of the partial parameters to each structure in the model to be updated; determining a local optimization strategy according to the mask; Determining a target mask corresponding to the partial parameters according to the local optimization strategy, the mask and the forgetting theory model; The target mask is input into the structure optimization model to perform model optimization to obtain an optimized model.
3. The method according to claim 2, characterized in that Determining the target mask corresponding to the partial parameters according to the local optimization strategy, the mask and the forgetting theory model includes: Performing iterative exchange processing on the mask according to the local optimization strategy to obtain a first mask; Inputting the first mask into the forgetting theory model for evaluation, and obtaining a second importance degree of the partial parameters to each structure in the model to be updated; Modifying the first mask according to the second importance level and the first importance level to obtain a second mask; The second mask is used as a new mask, and the step of iteratively exchanging the mask according to the local optimization strategy to obtain the first mask is returned to execute until the value of the local optimization strategy is a preset value, and the second mask obtained by the last update is used as the target mask.
4. The method according to claim 3, characterized in that The step of modifying the first mask according to the second importance level and the first importance level to obtain a second mask includes: Determining whether the second importance level corresponding to the target bit on the first mask is consistent with the corresponding first importance level; If they are consistent, determining the first mask as the second mask; If they are inconsistent, the mask corresponding to the target bit on the first mask is corrected to obtain a corrected second mask.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: The forgetting theory model is constructed according to the association relationship between the model to be updated and some parameters of the model to be updated, and the retained data set.
6. The method according to claim 5, characterized in that The method further comprises: The forgetting theory model is parsed according to the convex optimization technology to obtain a parsed forgetting theory model; The step of inputting some parameters of the model to be updated, the model to be updated, the retained data set and the deleted data set into a preset forgetting theory model for evaluation to obtain a first importance degree of the some parameters to each structure in the model to be updated includes: Part of the parameters of the model to be updated, the model to be updated, the retained data set and the deleted data set are input into the parsed forgetting theory model for evaluation to obtain the first importance degree of the part of the parameters to each structure in the model to be updated.
7. A forgetting update device for a large model, characterized in that: The device comprises: A determination module, used to determine the model to be updated, and obtain some parameters of the model to be updated, a reserved data set and a deleted data set; the reserved data set includes a data set after abnormal data is deleted from the training sample data set; the training sample data set refers to the data set used when training the model to be updated; the deleted data set refers to the data set composed of the abnormal data; An evaluation module, used for inputting some parameters of the model to be updated, the model to be updated, the retained data set and the deleted data set into a preset forgetting theory model for evaluation, so as to obtain a first importance degree of the some parameters to each structure in the model to be updated; The optimization module is used to input the importance of the partial parameters to each structure in the model to be updated into the structure optimization model to optimize the model and obtain the optimized model.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Large model memory management method and device and medium
CN121050663A
A large model memory management method, device and medium
CN121050663B