Training method, device, storage medium, equipment and program product
Patent Information
- Application Number
- CN202510330564.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2026-09-18
AI Technical Summary
[0004]在相关技术中,在目标分类器模型和目标反事实模型的训练过程中,都是通过手动标注和回译构建训练样本,手动标注能够覆盖的范围是有限的,回译是本质上约等于同义词替换,导致目标分类器模型和目标反事实模型的训练样本不全面,目标分类器模型和目标反事实模型的推理能力差,进而导致大型语言模型的模型编辑效果变差
[0021] This application embodiment obtains a first knowledge sample and a second knowledge sample from the original samples, where the first knowledge sample is a negative sample of the second knowledge sample; it then infers from the first knowledge sample based on the original model of the large language model to obtain a first inference sample; it further infers from the second knowledge sample based on the original model of the large language model to obtain a second inference sample; finally, it uses the first knowledge sample, the second knowledge sample, the first inference sample, and the second inference sample as training samples to train the original model of the large language model, thus obtaining the target model of the large language model. The training method provided in this application embodiment expands the training samples of the target classifier model and the target counterfactual model by utilizing the inference function of the large language model, improving the inference capabilities of the target classifier model and the target counterfactual model, enhancing the model editing effect of the large language model, and enabling the large language model to update its knowledge base without retraining, thereby reducing the cost of updating the knowledge base of the large language model.
Smart Images

Figure CN122778032A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, specifically to a training method, apparatus, storage medium, device, and program product for a large language model. Background Technology
[0002] Large Language Models (LLMs) learn the syntax, semantics, and contextual information of a language by training on a large amount of language text. This allows them to learn the patterns and structures of natural language, simulating the human language cognition and generation process. However, many factors can lead to errors in the output text of large language models. For example, much of the knowledge in real-world applications is updated in real time (advancements in a scientific field, changes in global events, evolution of social culture). If the knowledge base of a large language model is not updated in real time, the model's output text will become outdated and erroneous.
[0003] The Memory-Based Model Editing at Scale (SEARC) approach achieves model editing through a classifier model and a counterfactual model, enabling large language models to update their knowledge base without retraining. When model editing is needed, it first predicts whether the knowledge to be modified is related to existing knowledge. If the knowledge to be modified is related to existing knowledge, it retrieves the knowledge to be modified and the related existing knowledge, and then calls the target counterfactual model to update the knowledge base of the large language model.
[0004] In related technologies, training samples for target classifier models and target counterfactual models are constructed through manual annotation and back translation. Manual annotation has a limited scope, and back translation is essentially equivalent to synonym substitution. This results in incomplete training samples for target classifier models and target counterfactual models, poor reasoning ability of target classifier models and target counterfactual models, and consequently, a deterioration in the model editing effect of large language models. Summary of the Invention
[0005] This application provides a training method, apparatus, storage medium, device, and program product for a large language model, applicable to scenarios related to large language models, including model training, model editing, semantic reasoning, and text generation. It enables the updating of the knowledge base of a large language model without retraining, reducing the cost of updating the knowledge base of a large language model.
[0006] On one hand, embodiments of this application provide a training method for a large-scale language model. The method includes: acquiring a first knowledge sample and a second knowledge sample from original samples, wherein the first knowledge sample is a negative sample of the second knowledge sample; performing reasoning on the first knowledge sample according to the original model of the large-scale language model to obtain a first reasoning sample; performing reasoning on the second knowledge sample according to the original model of the large-scale language model to obtain a second reasoning sample; and using the first knowledge sample, the second knowledge sample, the first reasoning sample, and the second reasoning sample as training samples to train the original model of the large-scale language model to obtain a target model of the large-scale language model.
[0007] On the other hand, embodiments of this application provide a training apparatus for a large-scale language model, including a first acquisition module, a first processing module, a second processing module, and a third processing module. The first acquisition module is configured to acquire a first knowledge sample and a second knowledge sample from the original samples, wherein the first knowledge sample is a negative sample of the second knowledge sample; the first processing module is configured to perform reasoning on the first knowledge sample based on the original model of the large-scale language model to obtain a first reasoning sample; the second processing module is configured to perform reasoning on the second knowledge sample based on the original model of the large-scale language model to obtain a second reasoning sample, wherein the training samples of the target classifier model include the first knowledge sample, the second knowledge sample, the first reasoning sample, and the second reasoning sample; the third processing module is configured to use the first knowledge sample, the second knowledge sample, the first reasoning sample, and the second reasoning sample as training samples to train the original model of the large-scale language model to obtain the target model of the large-scale language model.
[0008] In some embodiments, the third processing module is configured to: input the first knowledge sample and the first inference sample into the original classifier model respectively; classify the first inference sample to obtain a first classification value, wherein the first classification value represents whether the first inference sample is a positive or negative sample of the first knowledge sample; input the first knowledge sample and the second inference sample into the original classifier model respectively; classify the second inference sample to obtain a second classification value, wherein the second classification value represents whether the second inference sample is a positive or negative sample of the first knowledge sample; construct a classification loss function based on the first classification value and the second classification value; and train the original classifier model based on the classification loss function to obtain the target classifier model.
[0009] In some embodiments, the third processing module is configured to: construct a first classification loss function based on the first classification value and the third classification value, wherein the third classification value represents that the first knowledge sample is a positive sample of the first inference sample; and construct a second classification loss function based on the second classification value and the fourth classification value, wherein the fourth classification value represents that the first knowledge sample is a negative sample of the second inference sample.
[0010] In some embodiments, the third processing module is configured to: train the original classifier model according to the classification loss function, and the resulting target classifier model achieves the minimization of the difference between the first classification value and the third classification value, and the minimization of the difference between the second classification value and the fourth classification value.
[0011] In some embodiments, the third processing module is configured to: input the first knowledge sample and the first inference sample into the original counterfactual model respectively, predict whether the first knowledge sample and the first inference sample are related, and obtain a first predicted value; input the first knowledge sample and the second inference sample into the original counterfactual model respectively, predict whether the first knowledge sample and the second inference sample are related, and obtain a second predicted value; construct a prediction loss function based on the first predicted value and the second predicted value; and train the original counterfactual model based on the prediction loss function to obtain the target counterfactual model.
[0012] In some embodiments, the third processing module is configured to: construct a first prediction loss function based on the first predicted value and the third predicted value, wherein the third predicted value represents that the first knowledge sample and the first inference sample are correlated; and construct a second prediction loss function based on the second predicted value and the fourth predicted value, wherein the fourth predicted value represents that the first knowledge sample and the second inference sample are not correlated.
[0013] In some embodiments, the third processing module is configured to: train the original counterfactual model according to the prediction loss function, and the resulting target counterfactual model achieves the minimization of the difference between the first predicted value and the third predicted value, and the minimization of the difference between the second predicted value and the fourth predicted value.
[0014] In some embodiments, the training device further includes a second acquisition module, a third acquisition module, a fourth acquisition module, a fourth processing module, and a fifth processing module. The second acquisition module is configured to acquire knowledge to be modified through edit records; the third acquisition module is configured to acquire knowledge to be updated through the knowledge base to be updated in the large language model; the fourth acquisition module is configured to acquire the target model of the trained large language model; the fourth processing module is configured to input the knowledge to be modified and the knowledge to be updated into the target model of the large language model to obtain edited knowledge; and the fifth processing module is configured to update the knowledge to be updated in the large language model according to the edited knowledge.
[0015] In some embodiments, the fifth processing module is configured to: input the knowledge to be modified and the knowledge to be updated into the target classifier model in the target model of the large language model, classify the knowledge to be updated to determine whether the knowledge to be updated is a positive sample or a negative sample of the knowledge to be modified; if the knowledge to be updated is a positive sample of the knowledge to be modified, then input the knowledge to be modified into the target counterfactual model in the target model of the large language model for prediction to obtain the edited knowledge.
[0016] In some embodiments, the fifth processing module is configured to: input the knowledge to be modified and the knowledge to be updated into the target classifier model in the target model of the large language model, classify the knowledge to be updated to determine whether the knowledge to be updated is a positive sample or a negative sample of the knowledge to be modified; if the knowledge to be updated is a negative sample of the knowledge to be modified, then input the knowledge to be modified into the basic prediction model for prediction to obtain the edited knowledge.
[0017] In some embodiments, the training apparatus further includes a fifth acquisition module and a sixth processing module. The fifth acquisition module is configured to acquire input text, and the sixth processing module is configured to perform reasoning on the input text based on the knowledge base of a large language model to generate a first reasoned text corresponding to the input text.
[0018] On the other hand, an embodiment of this application provides a computer-readable storage medium, which includes a stored program, wherein the program is executed by a processor to perform the method described in any of the above embodiments.
[0019] On the other hand, an embodiment of this application provides a computer device, which includes a processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute the method described in any of the above embodiments through the computer program.
[0020] On the other hand, an embodiment of this application provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the method described in any of the above embodiments.
[0021] This application embodiment obtains a first knowledge sample and a second knowledge sample from the original samples, where the first knowledge sample is a negative sample of the second knowledge sample; it then infers from the first knowledge sample based on the original model of the large language model to obtain a first inference sample; it further infers from the second knowledge sample based on the original model of the large language model to obtain a second inference sample; finally, it uses the first knowledge sample, the second knowledge sample, the first inference sample, and the second inference sample as training samples to train the original model of the large language model, thus obtaining the target model of the large language model. The training method provided in this application embodiment expands the training samples of the target classifier model and the target counterfactual model by utilizing the inference function of the large language model, improving the inference capabilities of the target classifier model and the target counterfactual model, enhancing the model editing effect of the large language model, and enabling the large language model to update its knowledge base without retraining, thereby reducing the cost of updating the knowledge base of the large language model. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A schematic diagram illustrating model editing of a large language model provided in an embodiment of this application.
[0024] Figure 2 This is a schematic diagram illustrating the implementation of the driven model editing method provided in this application embodiment.
[0025] Figure 3 This is a schematic diagram of the structure of the target classifier model provided in an embodiment of this application.
[0026] Figure 4 A schematic diagram of a training method for a large language model provided in an embodiment of this application.
[0027] Figure 5 This is a schematic diagram illustrating how a large language model, as provided in an embodiment of this application, expands the training samples by inferring from the original samples.
[0028] Figure 6 This is a schematic diagram of the training target classifier model provided in an embodiment of this application.
[0029] Figure 7 This is a schematic diagram of the training objective counterfactual model provided in the embodiments of this application.
[0030] Figure 8 A schematic diagram of the structure of a training device for a large language model provided in an embodiment of this application.
[0031] Figure 9 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0033] This application provides a training method, apparatus, storage medium, device, and program product for a large language model. Exemplarily, the method of this application embodiment can be executed by a computer device, which can be a terminal or a server, etc. The terminal can be a smartphone, tablet, laptop, desktop computer, smart TV, smart speaker, wearable smart device, personal computer (PC), smart vehicle terminal, etc. The terminal can also include a client, which can be a video client, shopping application client, reading application client, browser client, or instant messaging client, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. However, it is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application embodiment.
[0034] The embodiments of this application can be applied to scenarios related to large language models, including scenarios involving model training, model editing, semantic reasoning, and text generation.
[0035] First, some of the nouns or terms that appear in the description of the embodiments of this application are explained as follows:
[0036] Large Language Models (LLMs) are artificial intelligence models built using deep learning techniques. These models typically have hundreds of millions to trillions of parameters and learn complex patterns of association between vocabulary, grammar, and semantics through pre-training on massive amounts of text data (such as books, web pages, and dialogues). These models can understand and generate human language, excel at diverse tasks such as text generation, translation, question answering, and summarization, and demonstrate powerful contextual reasoning and zero-shot / few-shot learning capabilities (they can handle new tasks without specific task training).
[0037] Memory-Based Model Editing: A knowledge update technique for large language models (LLMs). Its core idea is to introduce an external memory (such as a key-value database or vector index) to store domain-specific knowledge or facts that need to be corrected. During model inference, the memory content is dynamically combined with the model's inherent parameters to generate results, thereby achieving knowledge updates or error corrections without directly modifying the model parameters.
[0038] Classifier models: Their core objective is to categorize input data into predefined classes based on its features. Their operation typically involves three steps: feature extraction (extracting key attributes from the input, such as word frequency in text or pixel patterns in images), model training (learning the mapping between features and classes using labeled data, for example, using logistic regression, decision trees, or neural networks to construct classification boundaries), and prediction (calculating the probability of a new input belonging to each class and selecting the class with the highest probability as the output). Range determination is a special application of classifier models; essentially, it's a binary classification task used to determine whether an input belongs to a specific target range (e.g., whether it involves a certain type of knowledge).
[0039] Counterfactual Model: Its working principle is based on reasoning and generating hypothetical conditions. Its core objective is to answer the question, "How would the outcome change if a certain premise were changed?" Specifically, this model dynamically overwrites or corrects the inherent knowledge of the original model (Base Model) while maintaining the original model structure, by explicitly injecting modified knowledge (such as updated facts or rules) or adjusting the logical relationship between input and output.
[0040] The solution provided in this application can be applied to model editing scenarios for large language models. In model editing scenarios, memory-driven model editing introduces an external memory (such as a key-value database or vector index) to store domain-specific knowledge or facts that need to be corrected. During model inference, the memory content is dynamically combined with the model's inherent parameters to generate results, thereby achieving knowledge updates or error corrections without directly modifying the model parameters.
[0041] When the model needs to update a piece of outdated information (e.g., "name of the representative of Country A"), the original output of the large language model can be overwritten by storing the new fact in the memory bank and preferentially retrieving the latest data in the memory when the user poses a question. A knowledge triplet is represented by <s, p, o_old>, for example, <Country A, representative, person name b> represents that the representative of Country A is person name b. Then the goal of model editing is to modify this knowledge triplet to <s, p, o_new>, where o_new refers to person name a. Because the goal of model editing is to realize the change of the model's knowledge from <s, p, o_old> to <s, p, o_new>, that is, to modify the knowledge "the representative of Country A is person name b" to "the representative of Country A is person name a". Of course, such modification is not isolated either, because after modifying this knowledge, other related knowledge also needs to be further adjusted. For example, the knowledge that the representative of Country A belongs to Party Y has to be modified to that the representative of Country A belongs to Party X. Meanwhile, knowledge irrelevant to this modification needs to remain unchanged.
[0042] Figure 1 It is a schematic diagram of model editing for a large language model provided by the embodiment of the present application. As Figure 1 shown, the input text of the large language model (LLM) is "Who is the representative of Country A", and the original output of the large language model (LLM) is person name b. By introducing <Country A, representative, person name a>, symbolic knowledge updates the knowledge base of the large language model (LLM) via path 1 update, and neural knowledge edits the neural network through knowledge editing methods such as insertion, modification and erasure (knowledge editing types: insertion modification erasure), and integrates the editing result of the neural network into the neural network of the large language model (LLM) via path 2 merge.
[0043] The range classifier model analyzes the semantics or keywords of the input text to identify whether it is related to the knowledge to be modified (e.g., involving specific entities or events), so as to determine whether to trigger the counterfactual model. Its accuracy depends on the discrimination of feature design (e.g., context embedding vectors) and the optimization of classification boundaries (e.g., training via cross-entropy loss function), ensuring that implicit correlations can be captured while avoiding over-generalization or missed judgment. When the input involves content to be modified, the counterfactual model preferentially calls the latest information stored externally (e.g., corrected data generated by retrieving the memory bank or adjusting parameters), instead of relying on outdated knowledge in pre-training, so as to generate an answer conforming to new facts.
[0044] For any input, if the classifier model detects a correlation between it and the knowledge to be modified (e.g., involving a specific entity or concept), the counterfactual model is triggered to generate the result; otherwise, the output of the unmodified base model is used. The base model typically refers to the initial model that has not been fine-tuned; it may be a general architecture or a pre-trained model that has not yet been optimized for a specific task or domain.
[0045] Figure 2 This is a schematic diagram illustrating the implementation of the driven model editing method provided in this application embodiment. For example... Figure 2 As shown, knowledge 1 and knowledge 2 are input into the scope classifier model, and the edit memory stores the knowledge to be modified. If the scope classifier model detects that knowledge 1 and the knowledge to be modified are not related, knowledge 1 is input into the original model to generate the result. If the scope classifier model detects that knowledge 2 and the knowledge to be modified are related, the knowledge to be modified is input into the counterfactual model to generate the result.
[0046] In related technologies, range classifier models currently employ manual annotation and back-translation to construct in-scope samples, while out-of-scope samples are determined through nearest neighbor search or manual annotation. These two types of sample data are then used to train the classifier model. However, manual annotation has a limited scope, and back-translation is essentially equivalent to synonym substitution, resulting in an incomplete understanding of both in-scope and out-of-scope samples in the trained classifier model. Range discriminative classifier models are typically trained on limited, manually constructed samples (such as direct keyword matching or simple back-translated data) to form binary classification models, making it difficult to identify implicit knowledge boundaries that require inference. For example, if the editing target is "the current representative of country A is changed from name b to name a," the classifier model may not be able to determine "the wife of the representative of country A is name c" as relevant input based on "the wife of name a is name c" (because the training data lacks inference samples with such logical extensions), leading to missed judgments.
[0047] Figure 3 This is a schematic diagram of the structure of the target classifier model provided in an embodiment of this application. Figure 3 As shown, manually labeled samples have parts that are difficult to cover at the boundaries of the range and outside the range, and the trained classifier model does not have comprehensive coverage of samples within and outside the range.
[0048] Counterfactual models only directly cover explicit knowledge but do not enhance the ability to logically reason about modified knowledge. For example, suppose the objective of a counterfactual model is "The most recent Olympics were held in Tokyo -> The most recent Olympics were held in Paris." The current training objective is to input "The most recent Olympics were held in Paris," and the counterfactual model should correctly answer "Paris." However, when faced with related but reasoning-based questions, such as "Was the most recent Olympics held in Europe?", the counterfactual model may not provide an accurate answer. If the counterfactual model could use modified knowledge to reason: the most recent Olympics were held in Paris + Paris is located in Europe, then the model could provide the correct answer. Currently, the training of counterfactual models does not consider this training objective.
[0049] like Figure 2 As shown, within the framework of the memory-driven model editing method, knowledge related to the modified knowledge is answered using the counterfactual model, while knowledge unrelated to the modified knowledge is answered using the original model. However, existing counterfactual models suffer from a lack of reasoning ability. This problem mainly stems from the fact that the counterfactual model is not explicitly required to apply the given knowledge; instead, it is only required to memorize the knowledge. Therefore, to enhance the reasoning ability of this module, the key lies in explicitly requiring the counterfactual model to apply the given knowledge during training.
[0050] Given that current range classifier models and counterfactual models suffer from insufficient reasoning capabilities, the training method provided in this application expands the training samples of the target classifier model and the target counterfactual model by utilizing the reasoning function of a large language model. This improves the reasoning capabilities of the target classifier model and the target counterfactual model, enhances the model editing effect of the large language model, and enables the large language model to update its knowledge base without retraining, thereby reducing the cost of updating the knowledge base of the large language model.
[0051] The following describes the training method for a large language model provided by an exemplary embodiment of this application, in conjunction with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown for the purpose of facilitating the understanding of the principles of this application, and the implementation of this application is not limited in any way in this respect.
[0052] Figure 4 A schematic diagram of a training method for a large language model provided in an embodiment of this application.
[0053] like Figure 4 As shown, the specific implementation process of this method includes:
[0054] Step 011: Obtain the first knowledge sample and the second knowledge sample from the original sample. The first knowledge sample is a negative sample of the second knowledge sample.
[0055] Specifically, the first knowledge sample usually refers to data containing correct or updatable knowledge (positive sample), such as the knowledge tuple <person name b, representative, representative of country A>. The second knowledge sample is knowledge that is completely unrelated to the first knowledge sample, such as <apple, place of origin, country A>.
[0056] Step 012: Reason about the first knowledge sample based on the original model of the large language model to obtain the first reasoning sample.
[0057] Specifically, Figure 5 A schematic diagram illustrating how a large language model, as provided in this application embodiment, expands the training samples by inferring from the original samples. Figure 5 As shown, large-scale language models can use zero-shot learning to take the first knowledge sample as a given case and directly guide the model to generate samples that meet the requirements through natural language prompts. The model does not require any task-specific training data; it can generate high-quality text output simply by parsing the semantic information in the prompts.
[0058] For example, if the first input knowledge sample is "Country B is located in Europe", the knowledge base based on a large language model can tell us that "the capital of Country B is Beijing". Through reasoning, the first inference sample can be "Beijing is located in Europe". The first input knowledge sample could be "Country C and Country B are connected by land", and the semantic reasoning capability based on a large language model can be used to determine the first inference sample as "Country B and Country C share a land border".
[0059] Furthermore, large-scale language models can use zero-shot learning to reason about the first inference sample, and directly guide the model to generate samples that meet the requirements through natural language prompts, thereby further increasing the number of labeled positive samples.
[0060] Step 013: Reason about the second knowledge sample based on the original model of the large language model to obtain the second reasoning sample.
[0061] For example, the second knowledge sample input could be "Country B and Country D are connected by land". Based on the semantic reasoning ability of a large language model, the second inference sample can be "Country B and Country D share a common land border". Since the first and second knowledge samples are negative samples, the second inference sample and the first knowledge sample are also negative samples. For example, "Country C and Country B are connected by land" cannot be used to infer "Country B and Country D share a common land border".
[0062] The reasoning ability of a large language model can be used to expand the number of first and second reasoning samples. The first reasoning sample can be used as a positive sample of the first knowledge sample, and the second reasoning sample can be used as a negative sample of the first knowledge sample.
[0063] Furthermore, large-scale language models can use zero-shot learning to reason about the second inference sample, and directly guide the model to generate samples that meet the requirements through natural language prompts, thereby further expanding the number of labeled negative samples.
[0064] Step 014: Use the first knowledge sample, the second knowledge sample, the first reasoning sample, and the second reasoning sample as training samples to train the original model of the large language model and obtain the target model of the large language model.
[0065] In some embodiments, the original model of the large language model includes an original classifier model, and the target model of the large language model includes a target classifier model. Step 014: Using the first knowledge sample, the second knowledge sample, the first inference sample, and the second inference sample as training samples, train the original model of the large language model to obtain the target model of the large language model, including:
[0066] Step 0151: Input the first knowledge sample and the first inference sample into the original classifier model respectively, classify the first inference sample to obtain the first classification value, which represents whether the first inference sample is a positive or negative sample of the first knowledge sample.
[0067] Specifically, the first classification value output can represent the result of the original classifier model classifying the first inference sample. Assuming that the first classification value output by the classifier model can be 0 or 1 (the distribution of positive or negative samples is uncertain), a first classification value of 1 indicates that the first inference sample is a positive sample of the first knowledge sample, and a first classification value of 0 indicates that the first inference sample is a negative sample of the first knowledge sample.
[0068] Step 0152: Input the first knowledge sample and the second reasoning sample into the original classifier model respectively, classify the second reasoning sample to obtain the second classification value. The second classification value represents whether the second reasoning sample is a positive sample or a negative sample of the first knowledge sample.
[0069] The output second classification value can characterize the result of the original classifier model classifying the second inference sample. Assuming that the output second classification value of the classifier model can be 0 or 1 (the distribution of positive or negative samples is uncertain), a second classification value of 1 indicates that the second inference sample is a positive sample of the first knowledge sample, and a second classification value of 0 indicates that the second inference sample is a negative sample of the first knowledge sample.
[0070] Step 0153: Construct a classification loss function based on the first classification value and the second classification value.
[0071] Specifically, Figure 6 A schematic diagram of the training target classifier model provided in the embodiments of this application, as shown below. Figure 6 As shown, in some embodiments, step 0153: the classification loss function includes a first classification loss function and a second classification loss function. Constructing the classification loss function based on the first classification value and the second classification value includes:
[0072] The first classification loss function is constructed based on the first classification value and the third classification value. The third classification value represents the first knowledge sample as a positive sample of the first inference sample.
[0073] The third classification value can be 1 (in the case of being determined as a positive sample). Since the first inference sample is a positive sample labeled by the first knowledge sample, we hope that the output first classification value represents a positive sample. That is to say, the goal of training based on the first classification loss function is to make the first classification value close to the third classification value.
[0074] A second classification loss function is constructed based on the second classification value and the fourth classification value. The fourth classification value represents the first knowledge sample as a negative sample of the second inference sample.
[0075] The fourth classification value can be 0 (in the case of a negative sample). Since the second inference sample is a negative sample labeled by the first knowledge sample, we hope that the output second classification value represents a negative sample. In other words, the goal of training based on the second classification loss function is to make the second classification value close to the fourth classification value.
[0076] Step 0154: Train the original classifier model according to the classification loss function to obtain the target classifier model.
[0077] Specifically, in some embodiments, step 0154: training the original classifier model according to the classification loss function to obtain the target classifier model includes:
[0078] The original classifier model is trained based on the classification loss function, and the resulting target classifier model minimizes the difference between the first and third classification values, as well as the difference between the second and fourth classification values.
[0079] Specifically, the loss can be calculated separately for the difference between the two sets of classification values, and then combined by weights to form a joint loss function. For example: λ1*L1+λ2*L2. Here, L1 is the difference loss between the first and third classification values, L2 is the difference loss between the second and fourth classification values, and λ1 and λ2 can be weight coefficients used to balance the importance of the two sets of losses. The goal of training is to make the output of the first classification value represent positive samples and the output of the second classification value represent negative samples.
[0080] The original model of the large language model includes an original counterfactual model, and the target model of the large language model includes a target counterfactual model. The first inference sample and the second inference sample can simultaneously serve as training samples for both the target classifier model and the target counterfactual model. In some embodiments, step 014: using the first knowledge sample, the second knowledge sample, the first inference sample, and the second inference sample as training samples to train the original model of the large language model to obtain the target model of the large language model, further includes:
[0081] Step 0161: Input the first knowledge sample and the first reasoning sample into the original counterfactual model respectively, and predict whether the first knowledge sample and the first reasoning sample are related to obtain the first predicted value.
[0082] Specifically, the first predicted value output can characterize the result of the original counterfactual model's prediction of the first inference sample. Assuming that the first predicted value output by the predictor model can be 0 or 1 (the distribution of whether the first knowledge sample and the first inference sample are related is uncertain), a first predicted value of 1 indicates that the first inference sample is related to the first knowledge sample, and a first predicted value of 0 indicates that the first inference sample is not related to the first knowledge sample.
[0083] Step 0162: Input the first knowledge sample and the second reasoning sample into the original counterfactual model respectively, and predict whether the first knowledge sample and the second reasoning sample are related to obtain the second predicted value.
[0084] Specifically, the output second predicted value can characterize the result of the original counterfactual model's prediction of the second inference sample. Assuming the second predicted value output by the predictor model can be 0 or 1 (the distribution of whether the first knowledge sample and the second inference sample are related is uncertain), a second predicted value of 1 indicates that the second inference sample is related to the first knowledge sample, and a second predicted value of 0 indicates that the second inference sample is not related to the first knowledge sample.
[0085] Step 0163: Construct a prediction loss function based on the first and second predicted values.
[0086] Specifically, Figure 7 A schematic diagram of the training objective counterfactual model provided in the embodiments of this application, as shown below. Figure 7 As shown in FIG, in some embodiments, step 0163: constructing a prediction loss function according to a first prediction value and a second prediction value comprises: constructing a first prediction loss function according to the first prediction value and a third prediction value, wherein the third prediction value represents that a first knowledge sample and a first inference sample are correlated with each other; constructing a second prediction loss function according to the second prediction value and a fourth prediction value, wherein the fourth prediction value represents that the first knowledge sample and a second inference sample are not correlated with each other.
[0087] Constructing the first prediction loss function according to the first prediction value and the third prediction value, wherein the third prediction value represents that the first knowledge sample is a positive sample of the first inference sample. The third prediction value may be 1 (a case where it is determined as a positive sample). Since the first inference sample is a positive sample calibrated by the first knowledge sample, the result represented by the output first prediction value is expected to be a positive sample, that is, the training objective based on the first prediction loss function is to make the first prediction value close to the third prediction value.
[0088] Constructing the second prediction loss function according to the second prediction value and the fourth prediction value, wherein the fourth prediction value represents that the first knowledge sample is a negative sample of the second inference sample. The fourth prediction value may be 0 (a case where it is determined as a negative sample). Since the second inference sample is a negative sample calibrated by the first knowledge sample, the result represented by the output second prediction value is expected to be a negative sample, that is, the training objective based on the second prediction loss function is to make the second prediction value close to the fourth prediction value.
[0089] For example, <Country C is land-connected to Country B> may be the first knowledge sample, <Country B and Country C share a common land border> is the first inference sample, and <Country B is land-connected to Country D> is the second inference sample, and the expected training objectives of the counterfactual model are:
[0090] Objective 1: input "think step by step, assuming Country C is land-connected to Country B, the two countries share a common land border, is this correct?", the expected output of the counterfactual model is "correct", and when the prediction value output by the counterfactual model is 1, it represents that the output is "correct";
[0091] Objective 2: input "think step by step, assuming Country C is land-connected to Country B, then Country B is land-connected to Country D, is this correct?", the expected output of the counterfactual model is "incorrect", and when the prediction value output by the counterfactual model is 0, it represents that the output is "incorrect".
[0092] Step 0164: training the original counterfactual model according to the prediction loss function to obtain a target counterfactual model.
[0093] In some embodiments, step 0164: training the original counterfactual model according to the prediction loss function to obtain the target counterfactual model, comprising: training the original counterfactual model according to the prediction loss function, such that the obtained target counterfactual model minimizes the difference between the first prediction value and the third prediction value, and minimizes the difference between the second prediction value and the fourth prediction value.
[0094] Specifically, losses can be calculated separately for the differences between the two sets of prediction values, and then a joint loss function is formed through weighted combination. For example: λ3*L3+λ4*L4. Wherein, L3 is the difference loss between the first prediction value and the third prediction value, L4 is the difference loss between the second prediction value and the fourth prediction value, λ3 and λ4 can be weight coefficients for balancing the importance of the two sets of losses, and the training objective is to make the result represented by the output first prediction value be a positive sample, and the result represented by the output second prediction value be a negative sample.
[0095] In some embodiments, the training method further comprises:
[0096] Step 021: Obtain knowledge to be modified through edit records.
[0097] Specifically, the edit records may be operation logs or change histories recorded by the system, which include modification histories of model parameters or training data, and these records can help locate outdated or wrong information in the model. Knowledge points that need to be adjusted can be extracted by analyzing these records, so as to update the knowledge base of the large language model.
[0098] Step 022: Obtain knowledge to be updated through the knowledge base to be updated in the large language model.
[0099] Specifically, the knowledge to be updated may be the knowledge in the large language model before updating. For example Figure 2 As shown, Knowledge 1 and Knowledge 2 may be knowledge to be updated derived from the knowledge base to be updated in the large language model, and knowledge to be modified is stored in the edit memory. For example, <Country D is contiguous with Country B by land> can be Knowledge 1, <Country B is contiguous with Country C by land> can be Knowledge 2, and <Country B and Country C share a common land border> can be the knowledge to be modified.
[0100] Step 023: Obtain the target model of the large language model obtained through training.
[0101] Specifically, the target model of the large language model can be obtained based on the training of the above steps 011 to 014, which will not be repeated herein.
[0102] Step 024: Input the knowledge to be modified and the knowledge to be updated into the target model of the large language model to obtain edited knowledge.
[0103] In some embodiments, step 024: inputting the knowledge to be modified and the knowledge to be updated into the target model of a large language model to obtain edited knowledge comprises:
[0104] inputting the knowledge to be modified and the knowledge to be updated into a target classifier model in the target model of the large language model, and classifying the knowledge to be updated to determine whether the knowledge to be updated is a positive sample or a negative sample of the knowledge to be modified.
[0105] Specifically, as Figure 2 shown, knowledge 1 and knowledge 2 are input into the target classifier model trained by a first inference sample and a second inference sample, to determine whether the knowledge to be updated is a positive sample or a negative sample of the knowledge to be modified. For example, the target classifier model determines that <Country C is land-connected with Country B> is a positive sample of <Country B and Country C share a common land border>, and <Country B and Country D are land-connected> is a negative sample of <Country B and Country C share a common land border>, and the knowledge to be modified may be <Country B and Country C do not share a common land border>.
[0106] Specifically, as Figure 2 shown, if a range classifier model detects that knowledge 2 is associated with the knowledge to be modified, knowledge 2 is input into a counterfactual model to generate a result, the generated result is first prediction knowledge, and the knowledge base to be updated in the large language model is updated according to the first prediction knowledge. For example, <Country B and Country C do not share a common land border> is input into a counterfactual model to generate a result, the generated result is <Country B and Country C are not connected by land>, and the original knowledge <Country C is land-connected with Country B> is updated to <Country B and Country C are not connected by land>.
[0107] In some embodiments, step 024: inputting the knowledge to be modified and the knowledge to be updated into the target model of the large language model to obtain edited knowledge further comprises:
[0108] if the knowledge to be updated is a negative sample of the knowledge to be modified, inputting the knowledge to be modified into a basic prediction model for prediction to obtain edited knowledge.
[0109] Specifically, as Figure 2 shown, if the range classifier model detects that knowledge 2 is not associated with the knowledge to be modified, knowledge 2 is input into an original model to generate a result, the generated result is second prediction knowledge, and the knowledge base to be updated in the large language model is updated according to the second prediction knowledge. For example, <Country B and Country D are land-connected> is input into a Base prediction model to generate a result, and the generated result is <Country B and Country D share a common land border>.
[0110] In some embodiments, after step 024 of updating the knowledge to be updated in the large language model according to the edited knowledge, the training method further comprises:
[0111] Step 031: Obtain the input text.
[0112] Specifically, the input text for large language models is the core basis for the model's response generation. It includes both raw information provided by the user (such as questions, instructions, or context) and structured information used to guide the model's behavior (such as instruction templates, examples, or special tags). Essentially, it defines the task boundaries for the model, activates relevant knowledge, and constrains the generation logic through text sequences. For example, in a question-answering scenario, the input text might be "Who is the representative of country A?".
[0113] Step 032: Perform reasoning on the input text based on the knowledge base of the large language model to generate the first reasoned text corresponding to the input text.
[0114] When reasoning on input text using a knowledge base based on a large language model, the model first extracts core entities, relationships, and implicit intentions from the input through semantic parsing. Then, it activates the associated conceptual network in the knowledge base and gradually derives coherent conclusions through logical connections (such as causal reasoning and analogical reasoning) and contextual integration (such as combining scientific principles with examples), ultimately generating the first inferred text. For example, when the input is "Who is the spouse of the representative of country A?", the model extracts the knowledge of "Who is the representative of country A?" from the knowledge base, constructs the logical chain "Who is the representative of country A? → Who is the spouse of the representative of country A?", and outputs explanatory text.
[0115] Among them, the knowledge base of the large language model is updated based on the processing results of the target classifier model and the target counterfactual model on the knowledge to be modified. During the training process of the target classifier model and the target counterfactual model, the large language model infers from the original samples to obtain the training samples of the target classifier model and the target counterfactual model.
[0116] Specifically, the knowledge base update process of large-scale language models is achieved through a collaborative target classifier model and a target counterfactual model. The target classifier model is responsible for locating the knowledge scope that needs modification. For example, if the knowledge of "current representative of country A" needs to be updated, the classifier model will learn to identify segments in the input text that involve attributes such as "representative identity, term of office, and relationship," and delineate the boundaries of the entities and relationships that need modification. The target counterfactual model is responsible for generating the corrected knowledge content. For example, when the classifier model determines that the input involves a replacement of the representative's identity ("name b → name a"), the counterfactual model needs to generate a logical association that conforms to the new fact (e.g., "name b's spouse is name d" instead of name c in the original data).
[0117] The semantic reasoning capabilities of large-scale language models can be used to expand the training samples for target classifier models and target counterfactual models. For example, if the original sample is "the current representative of country A is named a", the semantic reasoning capabilities of large-scale language models can determine "the spouse of the current representative of country A is named c". Both "the current representative of country A is named a" and "the spouse of the current representative of country A is named c" can be used as training samples for target classifier models and target counterfactual models.
[0118] This application embodiment obtains a first knowledge sample and a second knowledge sample from the original samples, where the first knowledge sample is a negative sample of the second knowledge sample; it then infers from the first knowledge sample based on the original model of the large language model to obtain a first inference sample; similarly, it infers from the second knowledge sample based on the original model of the large language model to obtain a second inference sample; finally, it uses the first knowledge sample, the second knowledge sample, the first inference sample, and the second inference sample as training samples to train the original model of the large language model, thus obtaining the target model of the large language model. The training method provided in this application embodiment expands the training samples of the target classifier model and the target counterfactual model by utilizing the inference function of the large language model, improving the inference capabilities of the target classifier model and the target counterfactual model, enhancing the model editing effect of the large language model, and enabling the large language model to update its knowledge base without retraining, thereby reducing the cost of updating the knowledge base of the large language model.
[0119] To facilitate better implementation of the training method of this application embodiment, this application embodiment also provides a training device for a large-scale language model. Please refer to... Figure 8 , Figure 8 This is a schematic diagram of the structure of a large-scale language model training device 400 provided in an embodiment of this application. The large-scale language model training device 400 includes a first acquisition module 411, a first processing module 421, a second processing module 422, and a third processing module 423. The first acquisition module 411 is configured to acquire a first knowledge sample and a second knowledge sample from the original samples, wherein the first knowledge sample is a negative sample of the second knowledge sample; the first processing module 421 is configured to perform reasoning on the first knowledge sample based on the original model of the large-scale language model to obtain a first reasoning sample; the second processing module 422 is configured to perform reasoning on the second knowledge sample based on the original model of the large-scale language model to obtain a second reasoning sample; the training samples of the target classifier model include the first knowledge sample, the second knowledge sample, the first reasoning sample, and the second reasoning sample; the third processing module 423 is configured to use the first knowledge sample, the second knowledge sample, the first reasoning sample, and the second reasoning sample as training samples to train the original model of the large-scale language model to obtain the target model of the large-scale language model.
[0120] In some embodiments, the third processing module 423 is configured to: input the first knowledge sample and the first inference sample into the original classifier model respectively; classify the first inference sample to obtain a first classification value, wherein the first classification value represents whether the first inference sample is a positive or negative sample of the first knowledge sample; input the first knowledge sample and the second inference sample into the original classifier model respectively; classify the second inference sample to obtain a second classification value, wherein the second classification value represents whether the second inference sample is a positive or negative sample of the first knowledge sample; construct a classification loss function based on the first classification value and the second classification value; and train the original classifier model based on the classification loss function to obtain the target classifier model.
[0121] In some embodiments, the third processing module 423 is configured to: construct a first classification loss function based on a first classification value and a third classification value, wherein the third classification value represents a positive sample of the first knowledge sample as a first inference sample; and construct a second classification loss function based on a second classification value and a fourth classification value, wherein the fourth classification value represents a negative sample of the first knowledge sample as a second inference sample.
[0122] In some embodiments, the third processing module 423 is configured to: train the original classifier model according to the classification loss function, and obtain the target classifier model to minimize the difference between the first classification value and the third classification value, and minimize the difference between the second classification value and the fourth classification value.
[0123] In some embodiments, the third processing module 423 is configured to: input the first knowledge sample and the first inference sample into the original counterfactual model respectively, predict whether the first knowledge sample and the first inference sample are related, and obtain a first predicted value; input the first knowledge sample and the second inference sample into the original counterfactual model respectively, predict whether the first knowledge sample and the second inference sample are related, and obtain a second predicted value; construct a prediction loss function based on the first predicted value and the second predicted value; and train the original counterfactual model based on the prediction loss function to obtain a target counterfactual model.
[0124] In some embodiments, the third processing module 423 is configured to: construct a first prediction loss function based on a first predicted value and a third predicted value, wherein the third predicted value represents that the first knowledge sample and the first inference sample are correlated; and construct a second prediction loss function based on a second predicted value and a fourth predicted value, wherein the fourth predicted value represents that the first knowledge sample and the second inference sample are not correlated.
[0125] In some embodiments, the third processing module 423 is configured to: train the original counterfactual model according to the prediction loss function, and the resulting target counterfactual model achieves the minimization of the difference between the first prediction value and the third prediction value, and the minimization of the difference between the second prediction value and the fourth prediction value.
[0126] In some embodiments, the training device 400 further includes a second acquisition module 412, a third acquisition module 413, a fourth acquisition module 414, a fourth processing module 424, and a fifth processing module 425. The second acquisition module 412 is configured to acquire knowledge to be modified through edit records; the third acquisition module 413 is configured to acquire knowledge to be updated through a knowledge base to be updated in a large language model; the fourth acquisition module 414 is configured to acquire the target model of the trained large language model; the fourth processing module 424 is configured to input the knowledge to be modified and the knowledge to be updated into the target model of the large language model to obtain edited knowledge; and the fifth processing module 425 is configured to update the knowledge to be updated in the large language model according to the edited knowledge.
[0127] In some embodiments, the fifth processing module 425 is configured to: input the knowledge to be modified and the knowledge to be updated into the target classifier model in the target model of the large language model, classify the knowledge to be updated to determine whether the knowledge to be updated is a positive sample or a negative sample of the knowledge to be modified; if the knowledge to be updated is a positive sample of the knowledge to be modified, then input the knowledge to be modified into the target counterfactual model in the target model of the large language model for prediction to obtain the edited knowledge.
[0128] In some embodiments, the fifth processing module 425 is configured to: input the knowledge to be modified and the knowledge to be updated into the target classifier model in the target model of the large language model, classify the knowledge to be updated to determine whether the knowledge to be updated is a positive sample or a negative sample of the knowledge to be modified; if the knowledge to be updated is a negative sample of the knowledge to be modified, then input the knowledge to be modified into the basic prediction model for prediction to obtain the edited knowledge.
[0129] In some embodiments, the training device 400 further includes a fifth acquisition module 415 and a sixth processing module 426. The fifth acquisition module 415 is configured to acquire input text, and the sixth processing module 426 is configured to perform reasoning on the input text based on the knowledge base of a large language model to generate a first reasoned text corresponding to the input text.
[0130] It should be noted that the functions of each module in the training device 400 for the large language model in this application embodiment can be referred to the specific implementation of any embodiment in the above method embodiments, and will not be repeated here.
[0131] Each unit in the above-described device can be implemented entirely or partially through software, hardware, or a combination thereof. Each unit can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each unit.
[0132] For example, the training device 400 for a large language model can be integrated into a terminal or server that has storage and a processor and thus computing power, or the training device 400 for the large language model can be the terminal or server.
[0133] In some embodiments, this application also provides a computer device, which includes a processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to perform the methods of any of the above embodiments through the computer program.
[0134] Figure 9 A schematic structural diagram of the computer device provided in the embodiments of this application, such as Figure 9 As shown, the computer device 500 may include: a communication interface 501, a memory 502, a processor 503, and a communication bus 504. The communication interface 501, memory 502, and processor 503 communicate with each other via the communication bus 504. The communication interface 501 is used for data communication between the device 500 and external devices. The memory 502 can be used to store software programs and modules, and the processor 503 runs the software programs and modules stored in the memory 502, such as the software programs for the corresponding operations in the aforementioned method embodiments.
[0135] In some embodiments, the processor 503 may invoke software programs and modules stored in the memory 502 to execute the above methods.
[0136] In some embodiments, the computer device 500 may be integrated into a terminal or server that has storage and a processor and thus computing power, or the computer device 500 may be the terminal or server.
[0137] This application also provides a computer-readable storage medium including a stored program, wherein the program is executed by a processor to perform the methods of any of the above embodiments, which will not be described in detail here for the sake of brevity.
[0138] This application also provides a computer program product, which includes a computer program / instructions that, when executed by a processor, implement the methods of any of the above embodiments, and will not be described in detail here for the sake of brevity.
[0139] This application also provides a computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding processes in the methods described above in the embodiments of this application. For brevity, these details will not be elaborated further here.
[0140] It should be understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0141] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0142] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0143] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0144] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0145] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0146] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0147] In addition, the functional units in the embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0148] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer or a server) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0149] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A training method for a large-scale language model, characterized in that, The method includes: Obtain the first knowledge sample and the second knowledge sample from the original sample, wherein the first knowledge sample is a negative sample of the second knowledge sample; The first knowledge sample is inferred based on the original model of the large language model to obtain the first inference sample; The second knowledge sample is inferred based on the original model of the large language model to obtain the second inference sample; The first knowledge sample, the second knowledge sample, the first reasoning sample, and the second reasoning sample are used as training samples to train the original model of the large language model, thereby obtaining the target model of the large language model.
2. The training method for a large language model as described in claim 1, characterized in that, The original model of the large language model includes an original classifier model, and the target model of the large language model includes a target classifier model. The step of training the original model of the large language model using the first knowledge sample, the second knowledge sample, the first inference sample, and the second inference sample as training samples to obtain the target model of the large language model includes: The first knowledge sample and the first inference sample are respectively input into the original classifier model to classify the first inference sample to obtain a first classification value. The first classification value represents whether the first inference sample is a positive or negative sample of the first knowledge sample. The first knowledge sample and the second inference sample are respectively input into the original classifier model to classify the second inference sample to obtain a second classification value. The second classification value represents whether the second inference sample is a positive or negative sample of the first knowledge sample. Construct a classification loss function based on the first classification value and the second classification value; The original classifier model is trained according to the classification loss function to obtain the target classifier model.
3. The training method as described in claim 2, characterized in that, The classification loss function includes a first classification loss function and a second classification loss function. Constructing the classification loss function based on the first classification value and the second classification value includes: A first classification loss function is constructed based on the first classification value and the third classification value, wherein the third classification value represents that the first knowledge sample is a positive sample of the first inference sample; A second classification loss function is constructed based on the second classification value and the fourth classification value, wherein the fourth classification value represents that the first knowledge sample is a negative sample of the second inference sample.
4. The training method as described in claim 3, characterized in that, The step of training the original classifier model according to the classification loss function to obtain the target classifier model includes: The original classifier model is trained according to the classification loss function, and the resulting target classifier model minimizes the difference between the first classification value and the third classification value, as well as the difference between the second classification value and the fourth classification value.
5. The training method as described in claim 1, characterized in that, The original model of the large language model includes an original counterfactual model, and the target model of the large language model includes a target counterfactual model. The process of training the original model of the large language model using the first knowledge sample, the second knowledge sample, the first inference sample, and the second inference sample as training samples to obtain the target model of the large language model includes: The first knowledge sample and the first reasoning sample are respectively input into the original counterfactual model to predict whether the first knowledge sample and the first reasoning sample are related, so as to obtain the first predicted value. The first knowledge sample and the second reasoning sample are respectively input into the original counterfactual model to predict whether the first knowledge sample and the second reasoning sample are related, so as to obtain a second predicted value. Construct a prediction loss function based on the first and second predicted values; The original counterfactual model is trained based on the prediction loss function to obtain the target counterfactual model.
6. The training method as described in claim 5, characterized in that, The prediction loss function includes a first prediction loss function and a second prediction loss function. Constructing the prediction loss function based on the first and second prediction values includes: A first prediction loss function is constructed based on the first predicted value and the third predicted value, wherein the third predicted value represents the correlation between the first knowledge sample and the first inference sample. A second prediction loss function is constructed based on the second and fourth prediction values, wherein the fourth prediction value indicates that the first knowledge sample and the second inference sample are not related to each other.
7. The training method as described in claim 6, characterized in that, The step of training the original counterfactual model based on the prediction loss function to obtain the target counterfactual model includes: The original counterfactual model is trained according to the prediction loss function, and the resulting target counterfactual model minimizes the difference between the first predicted value and the third predicted value, as well as the difference between the second predicted value and the fourth predicted value.
8. The training method according to any one of claims 1-7, characterized in that, include: Obtain the knowledge to be modified by editing the records; The knowledge to be updated is obtained from the knowledge base to be updated in the large language model. Obtain the target model of the large language model obtained through training; The knowledge to be modified and the knowledge to be updated are input into the target model of the large language model to obtain the edited knowledge; Update the knowledge to be updated in the large language model based on the edited knowledge.
9. The training method as described in claim 8, characterized in that, The process of inputting the knowledge to be modified and the knowledge to be updated into the target model of the large language model to obtain edited knowledge includes: The knowledge to be modified and the knowledge to be updated are input into the target classifier model in the target model of the large language model to classify the knowledge to be updated, so as to determine whether the knowledge to be updated is a positive sample or a negative sample of the knowledge to be modified. If the knowledge to be updated is a positive sample of the knowledge to be modified, then the knowledge to be modified is input into the target counterfactual model in the target model of the large language model for prediction, and the edited knowledge is obtained.
10. The training method as described in claim 8, characterized in that, The process of inputting the knowledge to be modified and the knowledge to be updated into the target model of the large language model to obtain edited knowledge includes: The knowledge to be modified and the knowledge to be updated are input into the target classifier model in the target model of the large language model to classify the knowledge to be updated, so as to determine whether the knowledge to be updated is a positive sample or a negative sample of the knowledge to be modified. If the knowledge to be updated is a negative sample of the knowledge to be modified, then the knowledge to be modified is input into the basic prediction model for prediction to obtain the edited knowledge.
11. The training method as described in claim 8, characterized in that, After updating the knowledge to be updated in the large language model based on the edited knowledge, the method further includes: Get the input text; The input text is inferred based on the knowledge base of a large language model to generate a first inferred text corresponding to the input text.
12. A training device for a large-scale language model, characterized in that, The device includes: The first acquisition module is configured to acquire a first knowledge sample and a second knowledge sample from the original sample, wherein the first knowledge sample is a negative sample of the second knowledge sample; The first processing module is configured to reason about the first knowledge sample based on the original model of the large language model to obtain the first reasoning sample. The second processing module is configured to reason about the second knowledge sample based on the original model of the large language model to obtain a second reasoning sample. The third processing module is configured to use the first knowledge sample, the second knowledge sample, the first reasoning sample, and the second reasoning sample as training samples to train the original model of the large language model and obtain the target model of the large language model.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed by a processor, performs the method according to any one of claims 1 to 11.
14. A computer device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 11 through the computer program.
15. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method described in any one of claims 1 to 11.