A Parameter Adjustment Method and System for Knowledge Elimination Learning Oriented to Large Models
By calculating the parameter change amount by influencing the function and Hessian matrix optimization algorithm, the adapter module parameters are directly updated, which solves the problem of high calculation cost and insufficient adaptability of knowledge elimination tasks in the existing technology, and achieves efficient and flexible knowledge elimination effect.
Patent Information
- Application Number
- CN202510360851.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The existing technology has high computing cost in knowledge elimination tasks, single task support, insufficient flexibility and scalability, making it difficult to adapt to complex application scenarios.
By introducing an impact function calculation to eliminate the impact of the target data corresponding to the request on the model parameters, use the Hessian matrix and optimization algorithm to calculate the parameter change amount, and directly update the adapter module parameters to avoid retraining the model.
It significantly improves computing efficiency, supports a variety of knowledge elimination tasks, maintains model performance, adapts to complex application scenarios, and reduces computing costs.
Smart Images

Figure CN119886308B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of knowledge elimination learning, and specifically to a method and system for parameter adjustment of knowledge elimination learning for large models. Background Art
[0002] With the rapid development of large language models (LLMs), they have shown unprecedented potential in natural language understanding and generation tasks. Whether in traditional tasks such as text classification and machine translation, or in complex tasks such as dialogue generation and information extraction, LLMs have demonstrated powerful knowledge expression and reasoning capabilities. However, the powerful performance of the models has also brought new problems, especially the risk that the models may memorize sensitive, incorrect, or unnecessary information during training. Therefore, how to efficiently remove these unwanted information has become a key challenge in the current research field.
[0003] To address this issue, existing technologies have proposed various methods, which can be mainly classified into the following two categories:
[0004] Exact knowledge elimination method: By partitioning the training data into subsets and retraining the model to eliminate the influence of the target data. This method can ensure complete knowledge elimination of the target data by the model, thus theoretically achieving the optimal effect. However, its computational cost is extremely high because each time a knowledge elimination request is processed, the model needs to be retrained, involving a large amount of time and resource investment. In addition, retraining may damage the original performance or structural integrity of the model, especially when dealing with large-scale datasets or complex tasks, the efficiency issue is particularly prominent.
[0005] Approximate knowledge elimination method: This type of method adjusts the model parameters to achieve approximate knowledge elimination of specific data, avoiding the need for retraining in the exact method. Common implementation means include fine-tuning the model parameters or constraining the output differences (such as using KL divergence). Although it is more efficient than the exact method, the approximate knowledge elimination method is usually only applicable to simple tasks (such as deleting training instances) and is difficult to handle more complex scenarios (such as modifying query content or correcting output responses). In addition, when dealing with changes in data distribution, the model performance may not be stable enough, and even introduce unexpected biases.
[0006] Although existing methods have solved the knowledge elimination problem to a certain extent, there are still the following key limitations:
[0007] Insufficient task support: Most current methods focus on deleting specific instances, while in actual applications, finer-grained operations (such as modifying data content or correcting model output) are often required, and the singularity of existing methods cannot meet these needs.
[0008] High computational cost: The exact knowledge elimination method has a huge computational overhead due to the frequent need to retrain the model, making it difficult to apply to application scenarios with high real-time requirements. Although the approximate knowledge elimination method has improved efficiency, there are still performance bottlenecks when dealing with complex tasks or large-scale datasets.
[0009] Poor flexibility and scalability: Many methods rely on specific model architectures or fixed data partitioning strategies. This dependence makes them lack sufficient adaptability when facing new tasks, new data, or dynamically changing requirements.
[0010] Generally speaking, existing methods have deficiencies in the adaptability, efficiency, and generality of the knowledge elimination task. There is an urgent need for a knowledge elimination technology that can efficiently handle multi-tasks and multi-scenarios, while taking into account computational performance and model effects. This technology can not only improve the accuracy of knowledge elimination, but also have broad applicability to meet the diverse needs in practical applications. Summary of the Invention
[0011] In this embodiment, a method, system, electronic device, and storage medium for parameter adjustment of knowledge elimination learning for large models are provided to solve the deficiencies of knowledge elimination technology in computational efficiency, task adaptability, and model scalability, and the traditional methods usually face problems such as high computational cost, single task support, and difficulty in adapting to complex application scenarios.
[0012] In a first aspect, an embodiment of the present invention provides a method for parameter adjustment of knowledge elimination learning for large models. The method for parameter adjustment of knowledge elimination learning for large models includes:
[0013] Obtain the initial parameters of the large language model and the elimination request;
[0014] Use the influence function to calculate the influence of the target data corresponding to the elimination request on the model parameters to obtain the parameter change amount;
[0015] Update the parameters of the adapter module according to the parameter change amount.
[0016] In an optional embodiment, obtaining the initial parameters of the large language model and the elimination request includes:
[0017] Obtain the initial parameters of the large language model through transfer learning or pre-training;
[0018] Receive a knowledge elimination request from a user or a system, and this request specifies the target data that needs to be removed or corrected.
[0019] In an optional embodiment, using the influence function to calculate the influence of the target data corresponding to the elimination request on the model parameters to obtain the parameter change amount includes:
[0020] Parse the knowledge elimination request to determine the specific task type;
[0021] Extract the gradient information related to the target data, and combine it with the Hessian matrix of the model or its approximation. Use the influence function to estimate the exact impact of the target data on the model parameters, and obtain the parameter change amount.
[0022] In an optional embodiment, based on the derivation of the influence function, the model parameter update is defined as follows:
[0023] ;
[0024] where, is the model parameter, is the inverse of the Hessian matrix, representing the second-order derivative matrix of the model parameters, used to capture the complex interrelationships between parameters; is the target instance 's gradient information, representing the contribution of the data to the model loss function, is the model input, is the model output.
[0025] In an optional embodiment, an optimization algorithm is adopted to transform the inverse Hessian vector product problem into a scalable convex quadratic optimization problem:
[0026] ;
[0027] where, is the cumulative contribution vector of the target gradient, q is the optimization target, that is, the parameter change amount, and H is the Hessian matrix.
[0028] In an optional embodiment, the knowledge elimination request includes task types such as instance removal, query modification, or response correction.
[0029] Compared with the prior art, the beneficial effects of the parameter adjustment method for knowledge elimination learning for large models of the present invention are as follows:
[0030] The parameter adjustment method for knowledge elimination learning for large models of the present invention has shown significant advantages and positive effects in the instance-level forgetting task of large language models (LLMs). First of all, by introducing the influence function, this method accurately estimates the impact of data perturbation on the model parameters, thus avoiding the high computational cost of completely retraining the model in traditional methods, and significantly improving the computational efficiency. Through an efficient parameter adjustment mechanism, accurate processing of various instance-level forgetting tasks is achieved without retraining the model. This not only improves the robustness and accuracy of the model, but also demonstrates its practical application potential in large-scale industrial scenarios.
[0031] In a second aspect, an embodiment of the present invention provides a parameter adjustment system for knowledge elimination learning for large models, including:
[0032] A parameter acquisition module, configured to acquire initial parameters of a large language model and an elimination request;
[0033] An influence function calculation module, configured to calculate the influence of the target data corresponding to the elimination request on the model parameters by using an influence function, and obtain a parameter change amount;
[0034] A parameter update module, configured to update the parameters of the adapter module according to the parameter change amount.
[0035] In an optional embodiment, the influence function calculation module includes a gradient extraction subunit and a parameter estimation subunit;
[0036] The gradient extraction subunit is configured to extract gradient information related to the target data;
[0037] The parameter estimation subunit is configured to estimate the parameter change amount in combination with the Hessian matrix of the model or its approximation.
[0038] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the bus, and the processor can call the logical instructions in the memory to execute the steps of the method provided in the first aspect.
[0039] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the parameter adjustment method for knowledge elimination learning for large models as described in the first aspect.
[0040] Compared with the prior art, the beneficial effects of the parameter adjustment system, electronic device, and storage medium for knowledge elimination learning for large models of the present invention are the same as those of the parameter adjustment method for knowledge elimination learning for large models described in the first aspect, so they will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1Flow chart of the parameter adjustment method for knowledge elimination learning for large models in the embodiments of the present invention;
[0043] Figure 2 Block diagram of the structure of the parameter adjustment system for knowledge elimination learning for large models in the embodiments of the present invention;
[0044] Figure 3 Block diagram of the structure of the electronic device in the embodiments of the present invention. Detailed implementation manners
[0045] To more clearly understand the purpose, technical solution and advantages of the present application, the present application will be described and explained below with reference to the accompanying drawings and embodiments.
[0046] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meanings understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these", etc. do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variants thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly connected. The term "plurality" involved in the present application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0047] The present invention aims to propose a unified parameter efficient knowledge elimination method based on influence functions to address the deficiencies of existing knowledge elimination techniques in terms of computational efficiency, task adaptability, and model scalability. Traditional methods often face problems such as high computational costs, single task support, and difficulty in adapting to complex application scenarios. In contrast, the present invention directly adjusts model parameters, avoiding the cumbersome retraining process, not only significantly improving computational efficiency but also expanding the applicability of the method to multiple tasks. This method provides a new approach for efficiently handling diverse knowledge elimination requirements in large language models (LLMs), taking into account both the accuracy of the knowledge elimination task and the preservation of model performance.
[0048] The specific objectives include the following aspects:
[0049] Support multiple knowledge elimination tasks: The present invention designs a unified knowledge elimination framework that can flexibly handle various task types including instance removal, input modification, and output correction. In practical applications, models often need to face different forms of data modification requirements. For example, deleting sensitive or invalid training instances, correcting noise information in incorrect inputs, or correcting incorrect labels in output results. The present invention abstracts these different tasks into parameter adjustment problems, providing a highly general and adaptable solution.
[0050] Improve computational efficiency: Traditional exact knowledge elimination methods require frequent retraining of the model, resulting in extremely high computational costs. While approximate methods are relatively efficient, they also face efficiency problems when data complexity increases. The present invention designs an efficient inverse Hessian vector product calculation method, transforms the process of adjusting model parameters into an optimization problem, and uses efficient algorithms such as stochastic gradient descent (SGD) for solution. This design greatly reduces the time and resource overhead in the parameter update process, enabling the present invention to maintain excellent computational performance when processing large-scale data. In addition, this method supports incremental updates, capable of quickly adapting to the needs of new data or new tasks, providing the possibility for efficient deployment in practical applications.
[0051] Preserve model performance: While implementing the knowledge elimination task, the present invention focuses on retaining the original performance of the model, ensuring that the model still has high prediction ability and stability after completing the knowledge elimination operation. By accurately estimating the impact of data perturbation on model parameters through influence functions, the present invention can minimize unnecessary performance losses. For example, when performing the data removal task, only the affected parameter parts are adjusted without disturbing the contributions of other unmodified data; when correcting the model output, the parameter update process can effectively avoid information loss or model confusion, ensuring that the model after knowledge elimination still maintains high-quality reasoning and generation capabilities.
[0052] Through the above improvements, the present invention not only achieves significant improvements in multi-task support, computational efficiency, and model performance retention, but also has strong scalability, enabling it to flexibly adapt to various practical application scenarios. For example, in a recommendation system, it is necessary to remove user privacy data; in a generative task, it is necessary to delete incorrect knowledge fragments; or in a classification task, it is necessary to correct incorrect labels. The present invention provides a unified and efficient solution, providing technical support for diverse knowledge elimination tasks, and having broad application prospects and industrialization value.
[0053] Specifically, in an embodiment of the present invention, a method for parameter adjustment of knowledge elimination learning for large language models (LLMIF) is provided. Figure 1 It is a flowchart of the method for parameter adjustment of knowledge elimination learning for large language models of the present invention, as Figure 1 shown, and this process includes the following steps:
[0054] S100. Obtain the initial parameters of the large language model and an elimination request;
[0055] Specifically, obtaining the initial parameters of the large language model and an elimination request includes:
[0056] Obtain the initial parameters of the large language model through transfer learning or pre-training;
[0057] Receive a knowledge elimination request from a user or a system, and this request specifies the target data that needs to be removed or corrected.
[0058] It should be noted that in the framework initialization stage, large language models (LLMs) usually obtain initial parameters through transfer learning or pre-training.
[0059] The present invention adopts parameter-efficient fine-tuning (PEFT, Parameter-Efficient Fine-Tuning) techniques, such as LoRA (Low-Rank Adaptation) or adapter modules. This technique only adjusts a small number of parameters of the model while freezing the rest, thus significantly reducing computational and storage requirements and providing a lightweight basis for subsequent parameter updates.
[0060] S200. Calculate the influence of the target data corresponding to the elimination request on the model parameters using the influence function to obtain the parameter change amount;
[0061] First, it should be noted that the knowledge elimination request includes task types such as instance removal, query modification, or response correction. The user or the system submits a knowledge elimination request through a specific interface, specifying the target data that needs to be removed or corrected. For example, the user may request to delete sensitive training instances, modify query content, or correct incorrect labels in the generated output. The content of the knowledge elimination request will be parsed by the system into specific task categories (such as instance removal, query modification, or response correction), and provide the target data input for subsequent parameter adjustment.
[0062] Calculate the influence of the target data corresponding to the elimination request on the model parameters using the influence function to obtain the parameter change amount, including:
[0063] Parse the knowledge elimination request to determine the specific task type;
[0064] Extract the gradient information related to the target data, and combine it with the Hessian matrix of the model or its approximation, and use the influence function to estimate the exact influence of the target data on the model parameters to obtain the parameter change amount.
[0065] Furthermore, based on the derivation of the influence function, the model parameter update is defined as follows:
[0066] ;
[0067] where, is the model parameter, is the inverse of the Hessian matrix, representing the second-order derivative matrix of the model parameters, used to capture the complex mutual relationships between the parameters; is the target instance 's gradient information, indicating the contribution of the data to the model loss function, is the model input, is the model output.
[0068] However, directly calculating the inverse of the Hessian matrix is unrealistic in large-scale language models.
[0069] Even further, an optimization algorithm is adopted to transform the inverse Hessian vector product problem into an extensible convex quadratic optimization problem:
[0070] ;
[0071] where, is the cumulative contribution vector of the target gradient, q is the optimization target, that is, the parameter change amount, and H is the Hessian matrix.
[0072] The present invention efficiently approximates the inverse Hessian vector product through an optimized algorithm, transforms the computational problem into a convex quadratic optimization problem with respect to q, and by solving this optimization problem, LLMIF can obtain accurate parameter adjustment directions and amplitudes without significantly increasing the computational cost.
[0073] To address the complexity in actual computations, the present invention adopts stochastic gradient descent and other batch optimization algorithms, achieving a good balance between computational efficiency and model performance. This approach significantly reduces the computational complexity, making the parameter update process more advantageous in terms of time and space overheads, while ensuring support for a variety of knowledge elimination tasks.
[0074] Exemplarily, after a knowledge elimination request is parsed, the LLMIF framework efficiently calculates the influence of the target data on the model parameters through the influence function. Specifically, the framework first extracts the gradient information related to the target data and combines it with the Hessian matrix (or its approximation) of the model to estimate the parameter change. To address the computational complexity of large-scale models, the present invention designs an optimization algorithm based on stochastic gradient descent (SGD), transforming the Hessian-vector product problem into a scalable optimization process, thereby significantly reducing the computational cost. This step is the core of the framework, ensuring the accuracy and efficiency of parameter updates.
[0075] S300. Update the parameters of the adapter module according to the parameter change amount.
[0076] According to the parameter change amount calculated by the influence function, the present invention directly updates the parameters of the adapter module without retraining the entire model. The adjustment of the adapter parameters only affects the local structure of the model, avoiding interference with the frozen part of the parameters, thereby effectively preserving the original capabilities of the model. This method not only speeds up the update process but also reduces possible performance degradation problems, enabling the model to quickly adapt to the requirements of knowledge elimination requests.
[0077] After the parameter update is completed, the model is comprehensively evaluated to check whether its performance meets the requirements. The evaluation content includes the performance of the model on knowledge elimination tasks (such as whether the target data is effectively removed or corrected) and the prediction accuracy of other tasks. According to the evaluation results, the parameter adjustment strategy or hyperparameter configuration may be further optimized to ensure that the model maintains a high level of performance and stability after knowledge elimination.
[0078] The updated model will be deployed to actual applications for processing real-time tasks. For example, in a recommendation system, the model can immediately apply the latest parameter adjustment results to provide personalized recommendations for users; in a generative model, the corrected model can output more accurate and reliable generated content. Through this step, LLMIF achieves a rapid response to dynamic data requirements and meets the high real-time application scenarios in industrial and academic fields.
[0079] Exemplarily, such as in a recommendation system: In a personalized recommendation system, users' behavioral data (such as browsing records, click logs, and shopping history) are widely used to train the recommendation model. However, this data may contain outdated or sensitive information. For example, when a user requests to delete their historical records, LLMIF can quickly remove the impact of the relevant data on the model parameters without having to retrain the entire recommendation model. At the same time, users' behaviors in the recommendation system are dynamically changing, and LLMIF can adjust the model in real time to adapt to the latest user preferences, ensuring the accuracy and relevance of the recommendation results. Through this method, the recommendation system can not only better meet users' privacy needs but also maintain continuously optimized recommendation performance.
[0080] Privacy protection: In many applications involving personal data (such as medical diagnosis, financial analysis, and social media), protecting users' privacy is the top priority. LLMIF can delete sensitive data related to users' privacy in the model, thus avoiding the risk of data leakage. For example, in a medical diagnosis model, a user may request to delete their personal medical record data. LLMIF can efficiently achieve this requirement through precise calculations and parameter updates, while ensuring that the performance of the model after knowledge elimination is not significantly affected. Compared with the traditional method of retraining the model, LLMIF is faster and more computationally resource-saving, providing an efficient technical means for privacy protection.
[0081] Generative models: In text generation, image generation, and other generative tasks, the model may generate results containing incorrect information or inaccurate knowledge. For example, a dialogue generation model may cite outdated facts or generate biased answers. Through LLMIF, generative models can timely correct these inaccurate generated contents. For example, the model can remove the memory of a certain incorrect knowledge or correct the incorrect logic in the generated output, thereby improving the generation quality and user experience. This method is not only applicable to text generation tasks but can also be extended to multi-modal generation applications (such as image annotation and video summarization), achieving cross-domain flexibility.
[0082] Legal Compliance and Data Management: In many industries, complying with relevant laws and regulations regarding data deletion or the "right to be forgotten" is an important challenge for enterprises. When a user requests the deletion of their personal data, LLMIF can respond quickly and eliminate the impact of the target data on the model by directly updating the model parameters. Compared with traditional methods, this framework avoids the high cost of retraining the model while ensuring the dual requirements of compliance and efficiency. In addition, LLMIF can also help enterprises optimize their data management processes and reduce system downtime or waste of computing resources caused by data updates.
[0083] Real-time Decision-making System: In real-time decision-making scenarios such as financial transactions and advertising placements, the system needs to quickly adjust the model based on the latest data. For example, an advertising placement platform needs to remove invalid click data or dynamically optimize the placement strategy according to newly arrived interaction records. Through its efficient parameter update mechanism, LLMIF can quickly adapt to changes in real-time data, avoid decision-making errors caused by data lag, and thus improve the response speed and effectiveness of the system.
[0084] From the above application examples, it can be seen that the LLMIF framework combines efficient computing power and flexible task adaptability in the knowledge elimination task of large language models, significantly reducing the computing cost while meeting diverse practical needs. Whether in the fields of personalized recommendation, privacy protection, or generative tasks, LLMIF demonstrates broad applicability and practical value, providing an advanced solution for handling dynamically changing and large-scale data. The advantages of this technology give it great potential for promotion in both industrial and academic scenarios.
[0085] In summary, parameter adjustment is achieved through efficient computing methods to meet various knowledge elimination requirements. The main methods of the framework include the following parts:
[0086] Complete Knowledge Erasure of Samples. The training of large language models relies on large-scale data, which contains knowledge from different fields and scenarios. However, in practical applications, sometimes it is necessary to remove the knowledge of specific samples from the model to meet privacy protection, data copyright, or personalization requirements. This process is called "complete knowledge erasure of samples". Different from traditional model training updates, its purpose is to completely eliminate the influence of certain samples on the model behavior without affecting the overall performance. The technical challenges in achieving complete knowledge erasure of samples mainly include two aspects: 1) The complexity of knowledge representation in the model: Knowledge in language models is often not simply stored in a certain parameter, but is manifested through the cooperation of multiple parameters. Therefore, how to efficiently locate the parameters related to specific samples is one of the difficulties; 2) The balance of model performance: When eliminating specific knowledge, it is necessary to ensure that other parts of the model remain intact and avoid unnecessary damage to its general capabilities. Currently, methods for this problem include knowledge erasure based on influence functions, gradient reverse optimization, and parameter pruning techniques, etc. For example, the influence function can be used to analyze the role of specific samples in model parameters, and then the knowledge erasure can be achieved by reverse updating or directly adjusting the relevant parameters. In addition, some studies have also proposed to generate pseudo-gradients for specific samples and reverse optimize the model parameters to delete the knowledge related to the sample. Although these methods are effective, their efficiency still needs to be optimized when dealing with large-scale models, and further research is needed on their impact on the robustness and generalization ability of the model. With the increasing requirements of regulations (such as GDPR, CCPA) for users' right to delete data, "complete knowledge erasure of samples" will become a key topic that cannot be ignored in the development and application of language models.
[0087] Input and Output Optimization and Adjustment. The output effect of large language models depends to a great extent on the quality and structure of the input. In some tasks, the input data may contain noise or incorrect information, such as spelling mistakes, inaccurate descriptions, or irrelevant content in user queries. These issues may have a negative impact on the model's understanding and prediction capabilities. To address this challenge, LLMIF provides an efficient input data correction mechanism that can intelligently modify the noise or errors in the input, such as deleting irrelevant noise information, correcting spelling mistakes, or replacing incorrect content, thus ensuring that the model can more accurately understand the user's intention. Specifically, LLMIF combines preprocessing of the input data with dynamic response adjustment to achieve the modification of the input data while maintaining the consistency of the output response. The corrected input can not only optimize the model's context understanding but also further dynamically adjust the model's parameters by influencing function calculations, enabling the model to better adapt to the modified input data. This process ensures that the model's prediction results can reflect the true intention of the modified data without being interfered by the original incorrect information. In addition, the LLMIF framework not only improves the response accuracy of the model but also enhances its robustness to complex or noisy environments. Whether there are spelling mistakes in the user query or the description information is not precise enough, LLMIF can quickly adapt to the changes in these imperfect data through precise input optimization and parameter adjustment mechanisms, providing a solid guarantee for the accuracy and reliability of the task results. This ability enables LLMIF to demonstrate unique advantages in scenarios where user input is diverse and contains errors or uncertain information, especially in natural language processing tasks with broad application potential. Although the text generated by large language models usually has high fluency and consistency, its output results sometimes face problems such as verbosity, inaccuracy, or not meeting the user's needs. When annotation errors are found in the response labels of training samples, the LLMIF framework can quickly adjust the model parameters to make its output consistent with the correct labels. This process not only reduces the negative impact of incorrect labels on the model's learning process but also avoids the long-term damage to the prediction ability caused by error propagation. Specifically, LLMIF can quickly update the model's gradient calculation and parameter adjustment methods after detecting label errors by introducing a dynamic parameter optimization mechanism to minimize the interference of incorrect samples. In addition, the framework can re-evaluate the model's performance on the task based on the corrected labels, ensuring that the model can more accurately capture the true patterns in the data. Thanks to this efficient response correction mechanism, LLMIF significantly improves the model's performance on related tasks, not only improving the prediction accuracy but also enhancing the model's robustness to uncertain data. This rapid response and correction to incorrect labels enable LLMIF to more effectively utilize data during the training process, especially showing unique advantages in scenarios with noisy labels or low data annotation quality.
[0088] To efficiently compute the changes in model parameters, the present invention designs a novel optimization algorithm that transforms the inverse Hessian vector product problem in the influence function into a convex quadratic optimization problem. Through this transformation, the framework can be efficiently solved using modern optimization algorithms such as stochastic gradient descent, significantly reducing the computational complexity and enhancing the computational scalability.
[0089] The LLMIF method of the present invention demonstrates significant advantages and positive effects in the instance-level forgetting tasks of large language models (LLMs). First, by introducing the influence function, LLMIF accurately estimates the impact of data perturbations on model parameters, thus avoiding the high computational cost of completely retraining the model using traditional methods and significantly improving the computational efficiency. In addition, compared with existing methods, LLMIF performs excellently in both model performance retention and efficiency optimization.
[0090] Experimental results show that LLMIF outperforms the existing state-of-the-art methods on multiple benchmark tasks and datasets.
[0091] Table 1: AUC prediction results of different methods on the book dataset
[0092]
[0093] In the query modification and response correction tasks, LLMIF further demonstrates its robustness and recovery ability in dealing with noisy data. In the query modification task of the MovieLens dataset, by correcting randomly injected noise, LLMIF achieves an average performance improvement of 5.1% compared to the contaminated model and is closest to the performance of the retraining method. In the response correction task, in the experiment of correcting 40% of the wrong labels, LLMIF achieves a 14% performance improvement compared to the contaminated model, and the average performance gap with the retraining method is only 3%.
[0094] Table 2: Experimental results of different methods on the MovieLens dataset
[0095]
[0096] In addition, in terms of efficiency evaluation, LLMIF demonstrates excellent computational efficiency. In the query modification task, LLMIF only takes 1.37×10³ seconds, which is about 31 times faster than the retraining method, significantly reducing the computational resource requirements.
[0097] In summary, through an efficient parameter adjustment mechanism, LLMIF accurately processes various instance-level forgetting tasks without retraining the model. This not only improves the robustness and accuracy of the model but also demonstrates its practical application potential in large-scale industrial scenarios.
[0098] An embodiment of the present invention also provides a parameter adjustment system for knowledge elimination learning for large models. This system is used to implement the above method embodiments, and those that have been described will not be repeated here. The following terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware or a combination of software and hardware is also possible and contemplated.
[0099] As Figure 2 shown, Figure 2 is a structural block diagram of the parameter adjustment system for knowledge elimination learning for large models in the present invention. The system includes:
[0100] A parameter acquisition module 101, configured to acquire the initial parameters of the large language model and an elimination request;
[0101] An influence function calculation module 102, configured to calculate the influence of the target data corresponding to the elimination request on the model parameters by using the influence function, and obtain a parameter change amount;
[0102] A parameter update module 103, configured to update the parameters of the adapter module according to the parameter change amount.
[0103] The influence function calculation module 102 includes a gradient extraction sub-unit and a parameter estimation sub-unit;
[0104] The gradient extraction sub-unit is configured to extract gradient information related to the target data;
[0105] The parameter estimation sub-unit is configured to estimate the change amount of the parameters in combination with the Hessian matrix of the model or its approximation.
[0106] Figure 3 is a structural block diagram of the electronic device provided by the embodiment of the present invention. As Figure 3 shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the following method:
[0107] Acquire the initial parameters of the large language model and an elimination request;
[0108] Calculate the influence of the target data corresponding to the elimination request on the model parameters by using the influence function, and obtain a parameter change amount;
[0109] Update the parameters of the adapter module according to the parameter variation amount.
[0110] In addition, when the logic instructions in the above-mentioned memory 630 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0111] The embodiments of the present invention also provide a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above-mentioned various embodiments, for example, including:
[0112] Obtain the initial parameters of the large language model and the elimination request;
[0113] Calculate the influence of the target data corresponding to the elimination request on the model parameters using the influence function to obtain the parameter variation amount;
[0114] Update the parameters of the adapter module according to the parameter variation amount.
[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical discs, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods of various embodiments or certain parts of the embodiments.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A parameter adjustment method for knowledge elimination learning for large models, applied to recommendation systems, medical diagnosis, financial analysis, or social media, to solve the problems of excessive computing resources, insufficient real-time performance, and model performance degradation caused by retraining the model, characterized in that The parameter adjustment method includes: Obtaining the initial parameters of the large language model and an elimination request, where the target data for the elimination request includes user behavior data, medical data, text data, or image data; Calculating the impact of the target data corresponding to the elimination request on the model parameters using the influence function to obtain the parameter change amount; Calculating the impact of the target data corresponding to the elimination request on the model parameters using the influence function to obtain the parameter change amount, including: Parsing the knowledge elimination request to determine the specific task type; Extracting the gradient information related to the target data, and combining it with the Hessian matrix of the model or its approximation, using the influence function to estimate the exact impact of the target data on the model parameters to obtain the parameter change amount; Based on the derivation of the influence function, the model parameter update is defined as follows: ; Among them, is the model parameter, is the inverse of the Hessian matrix, representing the second derivative matrix of the model parameters, which is used to capture the complex interrelationships between the parameters; is the target instance of the gradient information, indicating the contribution of the data to the model loss function, is the model input, is the model output; Using an optimization algorithm to transform the inverse Hessian vector product problem into an extensible convex quadratic optimization problem: ; Among them, is the cumulative contribution vector of the target gradient, q is the optimized target, that is, the parameter change amount, and H is the Hessian matrix; Updating the parameters of the adapter module according to the parameter change amount.
2. The parameter adjustment method for knowledge elimination learning for large models according to claim 1, characterized in that, Obtaining the initial parameters of the large language model and the elimination request, including: Obtaining the initial parameters of the large language model through transfer learning or pre-training; Receiving a knowledge elimination request from a user or the system, which specifies the target data that needs to be removed or corrected.
3. The parameter adjustment method for knowledge elimination learning for large models according to claim 2, characterized in that The knowledge elimination request includes task types such as instance removal, query modification, or response correction.
4. A parameter adjustment system for knowledge elimination learning for large models, characterized in that, Includes: A parameter acquisition module for obtaining the initial parameters of the large language model and the elimination request, obtaining the initial parameters of the large language model and the elimination request, where the target data for the elimination request includes user behavior data, medical data, text data, or image data; An influence function calculation module for calculating the impact of the target data corresponding to the elimination request on the model parameters using the influence function to obtain the parameter change amount; Calculating the impact of the target data corresponding to the elimination request on the model parameters using the influence function to obtain the parameter change amount, including: Parsing the knowledge elimination request to determine the specific task type; Extracting the gradient information related to the target data, and combining it with the Hessian matrix of the model or its approximation, using the influence function to estimate the exact impact of the target data on the model parameters to obtain the parameter change amount; Based on the derivation of the influence function, the model parameter update is defined as follows: ; Among them, is the model parameter, is the inverse of the Hessian matrix, representing the second derivative matrix of the model parameters, which is used to capture the complex interrelationships between the parameters; is the target instance of the gradient information, indicating the contribution of the data to the model loss function, is the model input, is the model output; Using an optimization algorithm to transform the inverse Hessian vector product problem into an extensible convex quadratic optimization problem: ; Among them, is the cumulative contribution vector of the target gradient, q is the optimized target, that is, the parameter change amount, and H is the Hessian matrix; A parameter update module for updating the parameters of the adapter module according to the parameter change amount.
5. The parameter adjustment system for knowledge elimination learning for large models according to claim 4, characterized in that The influence function calculation module includes a gradient extraction subunit and a parameter estimation subunit; The gradient extraction subunit is used to extract the gradient information related to the target data; The parameter estimation subunit is used to estimate the change amount of the parameters by combining the Hessian matrix of the model or its approximation.
6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the parameter adjustment method for knowledge elimination learning for large models as described in any one of claims 1 to 3.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the parameter adjustment method for knowledge elimination learning for large models as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Federal learning model training method and system, computer equipment and storage medium
CN116796831A
Continuous learning method for reasonably forgetting visual task knowledge in open domain environment
CN117876765A