Method, device and product for updating parameters of large language model
By determining the fact sets and unverified information in the generation process of large language models and using reward data to update model parameters, the problems of factual gaps and fictional information in long-form generation are solved, and the accuracy and comprehensiveness of the generated content are improved.
Patent Information
- Application Number
- CN202510740702.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
AI Technical Summary
Existing large language models have problems with factual gaps or fabricated facts during the long-form generation process, resulting in insufficient accuracy and comprehensiveness of the generated content.
By determining the set of facts that the target text generation process should be based on, the reward data is used to update the parameters of the large language model, including obtaining text generation requests, determining the target fact set, filtering unverified information, calculating reward data, and updating model parameters through reinforcement learning.
It improves the factual accuracy and comprehensiveness of text generated by large language models, effectively avoids factual gaps or fabricated facts in the long-form generation process, and ensures the accuracy and completeness of the generated content.
Smart Images

Figure CN120671824A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as large language models and natural language understanding, and more particularly to a method and device for updating parameters of a large language model, a text generation method and device based on a large language model, an electronic device, a storage medium, and a computer program product, which can be applied in text generation scenarios. Background Art
[0002] With the rapid development of generative AI, applications based on LLMs (Large Language Models) are becoming increasingly popular, encompassing areas such as intelligent question-answering, research assistance, and content creation. Large language models are required for accurate, long-form generation in many application scenarios, including factual writing (such as biographies and long reports on specific professional fields) and long-form factual question-answering (such as introductions to things and comparative analysis of different things). However, existing large language models often generate content that is factually inaccurate or fabricated during long-form generation. Summary of the Invention
[0003] The present disclosure provides a method and apparatus for updating parameters of a large language model, a method and apparatus for generating text based on the large language model, an electronic device, a storage medium, and a computer program product.
[0004] According to a first aspect, a method for updating parameters of a large language model is provided, comprising: generating a target text using the large language model according to a text generation request; determining, according to the text generation request, facts on which the generation process of the target text should be based, to obtain a target fact set; determining reward data according to the target fact set and an information set including unverified information in the target text; and updating the parameters of the large language model according to the reward data.
[0005] According to a second aspect, a text generation method based on a large language model is provided, comprising: obtaining a text generation request; and generating a target text using a large language model according to the text generation request, wherein the large language model is updated by any implementation method of the first aspect.
[0006] According to a third aspect, a device for updating parameters of a large language model is provided, comprising: a text generation unit configured to generate a target text using the large language model according to a text generation request; a fact determination unit configured to determine, according to the text generation request, facts on which the generation process of the target text should be based, and obtain a target fact set; a reward determination unit configured to determine reward data based on the target fact set and an information set including unverified information in the target text; and a parameter updating unit configured to update the parameters of the large language model according to the reward data.
[0007] According to a fourth aspect, a text generation device based on a large language model is provided, comprising: a request acquisition unit configured to acquire a text generation request; a request processing unit configured to generate a target text using the large language model according to the text generation request, wherein the large language model is updated by any implementation method of the third aspect.
[0008] According to the fifth aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so as to enable the at least one processor to execute the method described in any implementation of the first aspect or the second aspect.
[0009] According to a sixth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method described in any one of the implementations of the first and second aspects.
[0010] According to a seventh aspect, a computer program product is provided, comprising: a computer program, which implements the method described in any implementation manner of the first aspect or the second aspect when executed by a processor.
[0011] According to the technology disclosed herein, a method and apparatus for updating the parameters of a large language model are provided. Based on the facts that the text generation process should be based on, an effective and accurate reward signal is provided for the factual accuracy and factual comprehensiveness of the text generated by the large language model. This can effectively avoid the problems of factual voids or factual fabrications in the text generation process of the large language model, especially in the generation of long texts, and help improve the factual accuracy and factual comprehensiveness of the text generated by the large language model.
[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0014] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;
[0015] Figure 2 is a flowchart of one embodiment of a method for updating parameters of a large language model according to the present disclosure;
[0016] Figure 3 Schematic diagram of data processing flow according to this embodiment;
[0017] Figure 4 is a schematic diagram of a process for determining a fact set according to this embodiment;
[0018] Figure 5 is a schematic diagram of an application scenario of the method for updating parameters of a large language model according to this embodiment;
[0019] Figure 6 is a flowchart of another embodiment of a method for updating parameters of a large language model according to the present disclosure;
[0020] Figure 7 is a flowchart of an embodiment of a text generation method based on a large language model according to the present disclosure;
[0021] Figure 8 is a structural diagram of an embodiment of an apparatus for updating parameters of a large language model according to the present disclosure;
[0022] Figure 9 is a structural diagram of an embodiment of a text generation device based on a large language model according to the present disclosure;
[0023] Figure 10 It is a schematic diagram of the structure of a computer system suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION
[0024] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0025] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0026] Figure 1 An exemplary architecture 100 is shown to which the method and apparatus for updating parameters of a large language model, and the method and apparatus for generating text based on a large language model according to the present disclosure can be applied.
[0027] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. The communication connections between terminal devices 101, 102, and 103 constitute a topological network, and network 104 is used to provide a medium for communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0028] Terminal devices 101, 102, and 103 can be hardware devices or software that support network connection for data interaction and data processing. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices that support network connection, information acquisition, interaction, display, processing, and other functions, including but not limited to smartphones, tablet computers, e-book readers, laptop computers, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules, for example, to provide distributed services, or they can be implemented as a single software or software module. No specific limitations are given here.
[0029] Server 105 can be a server that provides various services, for example, a backend processing server that receives text generation requests sent by terminal devices 101, 102, and 103 and generates target text corresponding to the text generation requests using a large language model. As an example, server 105 can be a cloud server.
[0030] It should be noted that the server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules (e.g., software or software modules for providing distributed services), or as a single software or software module. No specific limitations are given here.
[0031] It should also be noted that the methods for updating parameters of a large language model and the methods for generating text based on a large language model provided in the embodiments of the present disclosure are generally executed by a server, but the possibility of execution by a terminal device, or by a server and a terminal device cooperating with each other, is not excluded. Accordingly, the various components (e.g., various units) included in the apparatus for updating parameters of a large language model and the apparatus for generating text based on a large language model can be entirely located in the server, entirely located in the terminal device, or separately located in the server and the terminal device.
[0032] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the system is merely illustrative. Depending on the implementation requirements, any number of terminal devices, networks, and servers may be provided. When the electronic device on which the method for updating the parameters of the large language model and the method for generating text based on the large language model are running does not need to transmit data with other electronic devices, the system architecture may only include the electronic device (e.g., terminal device or server) on which the method for updating the parameters of the large language model and the method for generating text based on the large language model are running.
[0033] Please refer to Figure 2 , Figure 2 A flow chart of a method for updating parameters of a large language model provided by an embodiment of the present disclosure. Figure 3 , Figure 3 A data processing flow diagram of a method for updating parameters of a large language model provided in an embodiment of the present disclosure. In process 200, the following steps are included:
[0034] Step 201: Generate target text using a large language model according to a text generation request.
[0035] In this embodiment, the execution body of the method for updating the parameters of the large language model (for example, Figure 1 The server in the example may obtain a text generation request from a remote location or a local location via a wired network connection or a wireless network connection, and generate a target text based on the text generation request using a large language model.
[0036] The text generation request can be issued by a target object, such as a person, other smart device, artificial intelligence assistant, etc.
[0037] A large language model is a large-scale language model built based on deep learning technology, primarily used for natural language processing tasks. It learns language patterns and structures by training on large amounts of text data, enabling it to generate natural language text or understand natural language input. A large language model typically includes the following modules:
[0038] Input module: Receives text data input from the target object, such as questions, instructions, or text generation requests such as dialogue content.
[0039] Preprocessing module: Preprocesses the input text, including word segmentation, stop word removal, text cleaning, and other operations, to convert the text into a form that can be processed by the model.
[0040] Encoding module: This module encodes the preprocessed text into a vector form that the model can understand and process. Common encoding methods include word embeddings and encoders in the Transformer architecture. Examples of word embeddings include Word2Vec (Words to Vector) and GloVe (Global Vectors for Word Representation).
[0041] Model module: The core component, typically based on a deep learning architecture (such as the Transformer), is responsible for processing encoded text vectors for language understanding and generation. The model learns complex patterns and semantic relationships in language through a multi-layered neural network structure.
[0042] Decoding module: Decodes the model's output vector into natural language text, generating responses or processing results for user input. Decoding methods can include greedy decoding and beam search.
[0043] Output module: Outputs the decoded text in a user-readable form, such as text displayed on the screen or speech after speech synthesis.
[0044] Large language models can be applied to various text generation scenarios. Taking short-text generation as an example, these include but are not limited to automatic replies, copywriting generation, summary generation, email replies, question-and-answer systems, tweet generation, and other scenarios. Taking long-text generation as an example, these include but are not limited to factual writing such as reports, technical documents, user manuals, biographies, and news reports, as well as long-form factual question-and-answer tasks such as introductions to things and comparative analysis between different things.
[0045] In this embodiment, the text generation request can be input into the large language model as a prompt of the large language model, and the large language model outputs the target text.
[0046] Step 202: According to the text generation request, facts that the target text generation process should be based on are determined to obtain a target fact set.
[0047] In this embodiment, the execution subject may determine the facts that the target text generation process should be based on according to the text generation request, and obtain a target fact set.
[0048] The facts that the target text generation process should be based on refer to the facts that exist in the text corresponding to the generated text generation request and have factual comprehensiveness and factual accuracy.
[0049] As an example, first, natural language understanding is performed on the text generation request to determine the scenario to which it belongs; then, based on the scenario, a method for determining the fact set is determined; finally, the determined determination method is used to determine the fact set that the target text generation process should be based on based on the text generation request.
[0050] The correspondence between the scenario and the determination method can be pre-set. For example, in the scenario of writing a biography, the corresponding determination method is to obtain facts related to the target person's life from various official websites involved in the target person to obtain a fact set.
[0051] As another example, first, the text generation request is subjected to lexical analysis, syntactic analysis, and semantic role labeling to extract structured information such as key entities, relationships, and events in the request. The pre-trained language model is used to perform semantic understanding on the text generation request and generate a semantic embedding vector to capture the deep meaning of the request. Then, an existing knowledge graph is constructed or utilized, which contains a large number of entities and the relationships between them, covering multiple fields and topics. The entities and relationships in the knowledge graph are classified and labeled to facilitate subsequent precise queries. Finally, based on the extracted structured information, path search and subgraph matching are performed in the knowledge graph to find entities and relationships related to the text generation request. Graph algorithms (such as the shortest path algorithm, PageRank algorithm, etc.) are used to sort the information in the knowledge graph by importance and extract the facts most relevant to the request. The extracted facts are semantically fused and deduplicated to ensure the simplicity and comprehensiveness of the target fact set, and a fact set is obtained.
[0052] In some optional implementations of this embodiment, the execution entity may perform step 202 as follows:
[0053] In the first step, a plurality of preset determination methods are used to determine the facts that the target text generation process should be based on according to the text generation request, and an initial fact set is obtained.
[0054] Specifically, for each of the multiple preset determination methods, the preset method is adopted to determine the facts that the target text generation process should be based on according to the text generation request, and an initial fact set is obtained; the initial fact sets corresponding to the multiple preset determination methods are combined to obtain multiple initial fact sets.
[0055] The preset determination method and the number of preset determination methods can be specifically set according to actual conditions, for example, including a determination method based on facts retrieved by a retrieval system and a determination method based on facts generated by other large models.
[0056] In the second step, a target fact set is obtained based on a plurality of initial fact sets corresponding one-to-one to a plurality of preset determination methods.
[0057] As an example, for multiple facts belonging to multiple initial fact sets, pairwise matching is performed to determine fact groups, each of which includes mutually matching facts; facts represented by fact groups whose number is greater than a set threshold in the fact group are added to the target fact set.
[0058] In this implementation, a target fact set is obtained based on a plurality of initial fact sets corresponding one-to-one to a plurality of preset determination methods, thereby improving the comprehensiveness and accuracy of the facts in the target fact set.
[0059] Continue to refer Figure 4 , which shows a schematic diagram of the process of determining a fact set.
[0060] In some optional implementations of this embodiment, the execution subject may perform the first step as follows: using at least two of the following preset determination methods to generate an initial fact set respectively:
[0061] Method 1: Use a large language model to generate multiple response texts corresponding to the text generation request, and filter out preliminary facts from the multiple response texts to obtain an initial fact set.
[0062] By inputting a text generation request into the large language model multiple times, or by inputting a text generation request and a prompt word instructing the large language model to generate multiple reply texts into the large language model, multiple reply texts can be obtained; all candidate facts included in the multiple reply texts are clustered and merged to obtain an initial fact set.
[0063] Method 2: Using a large language model, preliminary facts are screened out from the search results generated by the retrieval system for text generation requests to obtain an initial fact set.
[0064] By inputting a text generation request into the retrieval system, multiple retrieval results can be obtained; by inputting multiple retrieval results and prompt words representing the extraction of preliminary facts into the large language model, the large language model's natural language understanding and logical reasoning capabilities are used to screen out preliminary facts from the multiple retrieval results to obtain an initial fact set.
[0065] In order to improve the reliability of the retrieval results, the final retrieval system can be determined based on the evaluation indicators of multiple dimensions such as reliability, data comprehensiveness and stability of each candidate retrieval system.
[0066] In order to further improve the reliability of the retrieval results, multiple retrieval systems can be determined from multiple candidate retrieval systems, and preliminary facts can be screened out from the retrieval results generated by multiple retrieval systems for text generation requests through a large language model. The screened preliminary facts can then be clustered and merged to obtain an initial fact set.
[0067] Method three: Generate an initial fact set based on the received input operations.
[0068] As an example, an initial fact set may be generated based on receiving preliminary facts written by a target person (eg, a human expert) on a display interface.
[0069] In this implementation, the execution order of the various preset determination methods is not limited. The multiple preset determination methods can be executed sequentially according to the preset order or simultaneously.
[0070] In this implementation, a specific preset determination method is provided, which combines a large language model, a retrieval system and manual labor to improve the comprehensiveness of the preliminary facts in the initial fact set.
[0071] In some optional implementations of this embodiment, the execution entity may perform the second step in the following manner:
[0072] First, preliminary facts in multiple initial fact sets are clustered and merged to obtain a merged set.
[0073] For example, extract features from each preliminary fact, such as entity, relationship, time, and location. Select an appropriate clustering algorithm, such as K-Means or DBSCAN (Density-Based Spatial Clustering of Applications with Noise). Cluster the preliminary facts from multiple sets of initial facts based on the extracted features. Within each cluster, merge similar preliminary facts. A voting mechanism can be employed, with the majority of preliminary facts agreeing on the content as the merged fact body.
[0074] Then, the preliminary facts in the merged set are screened and checked to obtain the target fact set.
[0075] As an example, the execution entity may adopt the following methods:
[0076] Internal consistency check: Check whether there are logical contradictions between the preliminary facts in the merged set.
[0077] External fact matching: Verify preliminary facts against authoritative data sources.
[0078] Multi-source cross-validation: If the preliminary facts involve multiple information sources, perform multi-source cross-validation.
[0079] In this implementation, the comprehensiveness and accuracy of the facts in the target fact set are ensured by clustering, merging, screening and verifying the preliminary facts in the multiple initial fact sets.
[0080] In some optional implementations of this embodiment, the execution entity may screen and verify the preliminary facts to obtain the target fact set in the following manner:
[0081] First, the preliminary facts in the merged set are filtered through the large language model to obtain the filtered set.
[0082] As an example, the preliminary facts in the merged set and the prompt words that prompt the large language model to perform fact screening are input into the large language model, and the preliminary facts screened by the large language model are added to the screened set.
[0083] Then, for the preliminary facts in the filtered set, a query request corresponding to the preliminary facts is generated, and a query result of the query request is generated through the retrieval system.
[0084] For each preliminary fact in the filtered set, a query request corresponding to the preliminary fact is generated; the query request is input into the retrieval system, and the retrieval system outputs a query result of the query request.
[0085] To improve the accuracy of query results, for each preliminary fact, the category to which it belongs can be determined. A retrieval system corresponding to the category is then used to generate query results for the query request corresponding to the preliminary fact. The retrieval system has high retrieval accuracy and comprehensiveness for queries of the corresponding category.
[0086] Then, the preliminary facts in the filtered set are verified based on the query results to obtain a verified set.
[0087] As an example, for each preliminary fact in the filtered set, the consistency between the preliminary fact and the corresponding query result is determined; in response to determining that the two are consistent, the preliminary fact is added to the verified set.
[0088] Finally, the target fact set is obtained based on the screening operation on the preliminary facts in the verified set.
[0089] The execution entity may display the preliminary facts in the verified set to a target person (eg, a human expert) through a display interface, and receive a screening operation on the preliminary facts in the verified set to obtain a target fact set.
[0090] In this implementation, the two-level screening of the large language model and the manual screening operation are combined to further improve the accuracy of the facts in the target set.
[0091] Step 203 : determining reward data based on the target fact set and the information set including the unverified information in the target text.
[0092] In this embodiment, the execution entity may determine the reward data based on the target fact set and the information set including the unverified information in the target text.
[0093] As an example, first, based on the natural language understanding results of the target text, the target text is divided into checkpoints, and at least one unverified piece of information corresponding to at least one checkpoint is obtained, thereby forming an information set. Then, each fact in the target fact set is matched with each piece of unverified information in the information set to obtain the number of matches. Finally, reward data is determined based on the number of matches. The number of matches and the degree of reward represented by the reward data are positively correlated.
[0094] As another example, based on determining the number of matches, the number of errors in the unverified information can also be determined, and the reward data can be determined based on the number of errors and the number of matches. The reward level represented by the reward data is positively correlated with the number of matches and negatively correlated with the number of errors.
[0095] In some optional implementations of this embodiment, the execution entity may perform step 203 as follows:
[0096] The first step is to determine the accuracy and comprehensiveness of the unverified information in the information set based on the target fact set.
[0097] Accuracy represents the proportion of correct unverified information in the information set, and comprehensiveness represents the ratio of the number of unverified information in the information set to the number of facts in the target fact set.
[0098] The second step is to determine the reward data based on accuracy and comprehensiveness.
[0099] As an example, the accuracy and comprehensiveness may be summed and weighted to obtain reward data.
[0100] In this implementation, reward data is clearly defined based on the accuracy and comprehensiveness of unverified information in the information set, providing fine-grained reward data, which helps to further improve the factual accuracy and factual comprehensiveness of text generated by the large language model.
[0101] In some optional implementations of this embodiment, the execution entity may perform the first step to determine the accuracy and comprehensiveness in the following manner:
[0102] First, a first quantity value of accurate facts consistent with unverified information, a second quantity value of contradictory facts inconsistent with the unverified information, and a third quantity value of omitted facts not covered by the unverified information in the target fact set are determined.
[0103] Accurate facts: This fact appears in the information collection and is consistent with an unverified information in the information collection.
[0104] Contradictory fact: This fact contradicts an unverified piece of information in the information set.
[0105] Missing fact: This fact does not appear in the information set.
[0106] Then, the accuracy and comprehensiveness are determined based on the first quantity value, the second quantity value, and the third quantity value.
[0107] As an example, a first score representing the accuracy and a second score representing the comprehensiveness may be determined according to the first quantity value, the second quantity value, and the third quantity value, respectively.
[0108] In this implementation, the specific meanings of accurate facts, contradictory facts, and omitted facts are clarified, as well as the specific method of determining accuracy and comprehensiveness based on the quantitative values corresponding to each type of fact. This can quickly and accurately determine the accuracy and comprehensiveness of unverified information in an information set.
[0109] In some optional implementations of this embodiment, the above-mentioned execution entity can calculate the accuracy in the following manner: first, combine the first quantity value and the second quantity value to obtain a fourth quantity value; then, determine the accuracy based on the first quantity value and the fourth quantity value.
[0110] Specifically, the first quantity value and the second quantity value are added to obtain a fourth value; and the ratio of the first quantity value to the fourth quantity value is used as the accuracy.
[0111] In this implementation, a specific method for determining accuracy is provided, which helps to more accurately determine the accuracy of unverified information in an information set.
[0112] In some optional implementations of this embodiment, the above-mentioned execution entity can calculate the comprehensiveness in the following manner: first, combine the first quantity value and the second quantity value to obtain a fourth quantity value; then, combine the first quantity value, the second quantity value and the third quantity value to obtain a fifth quantity value; finally, determine the comprehensiveness based on the fourth quantity value and the fifth quantity value.
[0113] Specifically, the first quantity value and the second quantity value are added to obtain a fourth quantity value; the first quantity value, the second quantity value and the third quantity value are added to obtain a fifth quantity value; and the ratio of the fourth quantity value to the fifth quantity value is used as comprehensiveness.
[0114] In this implementation, a specific method for determining comprehensiveness is provided, which helps to more accurately determine the comprehensiveness of unverified information in an information set.
[0115] For example, when the text generation request is "Introduce the seventh national census," assume that the target text output by the large language model is: "The seventh national census was conducted in 2020. The standard time point for this census was 00:00 on October 1, 2020. The results show that the total population of China is 1,443,497,378."
[0116] The target text generation process should be based on the following facts:
[0117] The seventh national census is a national population census conducted in 2020.
[0118] The standard time point for the seventh national census is 0:00 on November 1, 2020.
[0119] The census results show that the total population of the country is approximately 1.44 billion.
[0120] The census results show that the population of the Hong Kong Special Administrative Region is approximately 7.47 million.
[0121] The census results show that the population of the Macao Special Administrative Region is approximately 680,000.
[0122] Determine based on the facts in the target fact set and the unverified information in the information set:
[0123] Accurate facts: The target text mentions that "the seventh national census is a national census conducted in 2020", which is consistent with the fact "conducted in 2020" in the target fact set.
[0124] Contradictory facts: The target text mentions that "the standard time point for the census is 0:00 on October 1, 2020", which contradicts a fact in the target fact set, "November 1, 2020".
[0125] Accurate facts: The target text states that "the total population of China announced in the seventh national census is 1,443,497,378", which is consistent with the statement of "approximately 1.44 billion people" in the target fact set.
[0126] Missing facts: The target text does not mention the population of the Hong Kong Special Administrative Region.
[0127] Missing fact: The target text does not mention the population of the Macao Special Administrative Region.
[0128] At this point, we will calculate the fact accuracy score = 2 / 3 and the fact completeness score = 3 / 5. The score obtained by weighted average of the two scores will be used as the reward signal for reinforcement learning to guide the parameter update of the large language model.
[0129] Step 204: Update the parameters of the large language model based on the reward data.
[0130] In this embodiment, the execution entity may update the parameters of the large language model according to the reward data.
[0131] After determining the reward signal, we can use RL (Reinforcement Learning) methods to update the parameters of the large language model, using the large language model as the policy model in reinforcement learning. We optimize the policy function through gradient ascent and calculate the gradient of the reward data with respect to the parameters to update the parameters of the large language model.
[0132] By iteratively executing the above steps 201-204 until the preset end condition is reached, a trained large language model is obtained, wherein the preset end condition is, for example, the number of training times exceeds a preset number threshold, the training time exceeds a preset time threshold, and the loss of the large language model converges.
[0133] In some optional implementations of this embodiment, the execution entity may perform step 201 in the following manner: generating multiple target texts using a large language model according to a text generation request.
[0134] As an example, a text generation request and instruction data representing the generation of multiple target texts are input into a large language model, and the large language model generates the multiple target texts.
[0135] For each generated target text, the execution entity may determine the reward data corresponding to the target text according to the target fact set corresponding to the target text and the information set including the unverified information in the target text.
[0136] In this implementation, the execution entity may perform step 204 in the following manner: updating the parameters of the large language model according to the plurality of reward data corresponding one-to-one to the plurality of target texts.
[0137] The multiple reward data corresponding one-to-one to the multiple target texts are generally different. The factual accuracy and factual comprehensiveness of the target text are positively correlated with the reward data corresponding to the target text. Updating the parameters of the large language model based on the multiple reward data corresponding one-to-one to the multiple target texts can make the large language model more inclined to generate target texts with higher factual accuracy and factual comprehensiveness, which helps to further improve the parameter updating efficiency of the large language model and the factual accuracy and factual comprehensiveness of the text generated by the large language model after the update.
[0138] Continue to see Figure 5 , Figure 5 500 is a schematic diagram of an application scenario of the method for updating the parameters of a large language model according to this embodiment. A large language model is deployed in a server 501, and is required to be able to generate long texts with factual accuracy and comprehensiveness. To achieve this goal, first, a target text is generated using the large language model based on a text generation request, where the text generation request is, for example, a request to generate a task biography or a technical report. Then, based on the text generation request, the facts that the target text generation process should be based on are determined to obtain a target fact set. Then, based on the target fact set and an information set including unverified information in the target text, reward data is determined. Finally, based on the reward data, the parameters of the large language model are updated.
[0139] In this embodiment, a method for updating the parameters of a large language model is provided. Based on the facts that the text generation process should be based on, an effective and accurate reward signal is provided for the factual accuracy and factual comprehensiveness of the text generated by the large language model. This can effectively avoid the problems of factual gaps or factual fabrications in the text generation process of the large language model, especially in the generation of long texts, and help improve the factual accuracy and factual comprehensiveness of the text generated by the large language model.
[0140] Continue to refer Figure 6 , shows a schematic process 600 of another embodiment of a method for updating parameters of a large language model according to the present disclosure. In process 600, the following steps are included:
[0141] Step 601: Generate multiple target texts using a large language model according to a text generation request.
[0142] In step 602, a large language model is used to generate multiple reply texts corresponding to the text generation request, and preliminary facts are screened out from the multiple reply texts to obtain an initial fact set. The large language model is used to screen preliminary facts from the search results generated by the retrieval system for the text generation request to obtain an initial fact set. The initial fact set is generated based on the received input operation.
[0143] Step 603: cluster and merge the preliminary facts in the multiple initial fact sets to obtain a merged set.
[0144] Step 604 : Filter the preliminary facts in the merged set using the large language model to obtain a filtered set.
[0145] Step 605: For the preliminary facts in the filtered set, a query request corresponding to the preliminary facts is generated, and a query result of the query request is generated through a retrieval system.
[0146] Step 606: Verify the preliminary facts in the filtered set based on the query result to obtain a verified set.
[0147] Step 607: Obtain a target fact set based on a screening operation on the preliminary facts in the verified set.
[0148] Step 608: Determine the accuracy and completeness of the unverified information in the information set based on the target fact set.
[0149] Step 609: Determine reward data based on accuracy and comprehensiveness.
[0150] Step 610 : Update the parameters of the large language model according to the multiple reward data corresponding to the multiple target texts.
[0151] Compared with the above-mentioned process 200, the process 600 of the method for updating the parameters of the large language model in this embodiment specifically describes the process of generating the target fact set, the process of determining the reward data, and the parameter updating process based on multiple reward data corresponding to multiple target texts, thereby further improving the factual accuracy and comprehensiveness of the text generated by the large language model.
[0152] Continue to refer Figure 7 , shows a flow chart of a text generation method based on a large language model provided by an embodiment of the present disclosure. In process 700, the following steps are included:
[0153] Step 701: Obtain a text generation request.
[0154] In this embodiment, the execution subject of the text generation method based on the large language model (for example, Figure 1 The server in the server) can generate a request for text from a remote location or locally via a wired network connection or a wireless network connection.
[0155] Text generation requests can be requests in text generation scenarios. For example, short-form generation includes, but is not limited to, automated replies, copywriting, summary generation, email replies, question-and-answer systems, and tweet generation. For long-form generation, this includes, but is not limited to, factual writing such as reports, technical documentation, user manuals, biographies, and news reports, as well as long-form factual question-and-answer services such as introductions and comparative analysis of different things.
[0156] Step 702: Generate target text using a large language model according to the text generation request.
[0157] In this embodiment, a text generation request may be input into a large language model, and the large language model generates a target text based on its powerful natural language understanding and logical reasoning capabilities.
[0158] The large language model performs parameter updates based on the above embodiments 200 and 600 to ensure that the generated target text has factual accuracy and factual comprehensiveness. It should be noted that during the application of the large language model, the parameters of the large language model can also be updated according to the above embodiments 200 and 600.
[0159] In this embodiment, the large language model ensures that the generated target text has factual accuracy and factual comprehensiveness.
[0160] Continue to refer Figure 8 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for updating the parameters of a large language model. Figure 2 Corresponding to the method embodiment shown, the system can be specifically applied to various electronic devices.
[0161] like Figure 8 As shown, the device 800 for updating the parameters of the large language model includes: a text generation unit 801, configured to generate a target text using the large language model according to a text generation request; a fact determination unit 802, configured to determine the facts that the generation process of the target text should be based on according to the text generation request, and obtain a target fact set; a reward determination unit 803, configured to determine reward data based on the target fact set and an information set including unverified information in the target text; and a parameter updating unit 804, configured to update the parameters of the large language model according to the reward data.
[0162] In some optional implementations of this embodiment, the fact determination unit 802 is further configured to: adopt multiple preset determination methods, respectively determine the facts that the target text generation process should be based on according to the text generation request, and obtain an initial fact set; obtain a target fact set based on multiple initial fact sets that correspond one-to-one to the multiple preset determination methods.
[0163] In some optional implementations of this embodiment, the fact determination unit 802 is further configured to: generate an initial fact set using at least two of the following preset determination methods: generate multiple reply texts corresponding to the text generation request through a large language model, and filter out preliminary facts from the multiple reply texts to obtain an initial fact set; filter out preliminary facts from the retrieval results generated by the retrieval system for the text generation request through a large language model to obtain an initial fact set; generate an initial fact set based on the received input operation.
[0164] In some optional implementations of this embodiment, the fact determination unit 802 is further configured to: cluster and merge preliminary facts in multiple initial fact sets to obtain a merged set; and screen and verify the preliminary facts in the merged set to obtain a target fact set.
[0165] In some optional implementations of this embodiment, the fact determination unit 802 is further configured to: filter the preliminary facts in the merged set through a large language model to obtain a filtered set; generate a query request corresponding to the preliminary facts in the filtered set, and generate a query result of the query request through a retrieval system; verify the preliminary facts in the filtered set based on the query result to obtain a verified set; and obtain a target fact set based on the filtering operation on the preliminary facts in the verified set.
[0166] In some optional implementations of this embodiment, the reward determination unit 803 is further configured to: determine the accuracy and comprehensiveness of the unverified information in the information set based on the target fact set; and determine reward data based on the accuracy and comprehensiveness.
[0167] In some optional implementations of this embodiment, the reward determination unit 803 is further configured to: determine a first quantitative value of accurate facts in the target fact set that are consistent with the unverified information, a second quantitative value of contradictory facts that are inconsistent with the unverified information, and a third quantitative value of omitted facts that are not covered by the unverified information; and determine accuracy and comprehensiveness based on the first quantitative value, the second quantitative value, and the third quantitative value.
[0168] In some optional implementations of this embodiment, the reward determination unit 803 is further configured to: combine the first quantity value and the second quantity value to obtain a fourth quantity value; and determine the accuracy based on the first quantity value and the fourth quantity value.
[0169] In some optional implementations of this embodiment, the reward determination unit 803 is further configured to: combine the first quantity value and the second quantity value to obtain a fourth quantity value; combine the first quantity value, the second quantity value and the third quantity value to obtain a fifth quantity value; and determine comprehensiveness based on the fourth quantity value and the fifth quantity value.
[0170] In some optional implementations of this embodiment, the text generation unit 801 is further configured to: generate multiple target texts using a large language model according to a text generation request; and the parameter updating unit 804 is further configured to: update the parameters of the large language model according to multiple reward data corresponding one-to-one to the multiple target texts.
[0171] In this embodiment, a device for updating the parameters of a large language model is provided. Based on the facts that the text generation process should be based on, an effective and accurate reward signal is provided for the factual accuracy and factual comprehensiveness of the text generated by the large language model. This can effectively avoid the problems of factual gaps or factual fabrications in the text generation process of the large language model, especially in the generation of long texts, and help improve the factual accuracy and factual comprehensiveness of the text generated by the large language model.
[0172] Continue to refer Figure 9 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a text generation device based on a large language model. Figure 7 Corresponding to the method embodiment shown, the system can be specifically applied to various electronic devices.
[0173] like Figure 9 As shown, the text generation apparatus 900 based on the large language model includes: a request acquisition unit 901, configured to acquire a text generation request; a request processing unit 902, configured to generate a target text using the large language model according to the text generation request.
[0174] The large language model performs parameter updates based on the above-described embodiment 800 to ensure that the generated target text has factual accuracy and factual comprehensiveness.
[0175] In this embodiment, the large language model ensures that the generated target text has factual accuracy and factual comprehensiveness.
[0176] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the method for updating the parameters of the large language model and the text generation method based on the large language model described in any of the above embodiments when executing.
[0177] According to an embodiment of the present disclosure, the present disclosure also provides a readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to implement the method for updating the parameters of a large language model and the text generation method based on a large language model described in any of the above embodiments when executed.
[0178] The embodiments of the present disclosure provide a computer program product. When executed by a processor, the computer program can implement the method for updating the parameters of a large language model and the text generation method based on a large language model described in any of the above embodiments.
[0179] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0180] like Figure 10 As shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0181] Various components in device 1000 are connected to I / O interface 1005, including an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0182] The computing unit 1001 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the method for updating the parameters of the large language model. For example, in some embodiments, the method for updating the parameters of the large language model can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the method for updating the parameters of the large language model described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured in any other appropriate manner (for example, by means of firmware) to execute the method for updating the parameters of the large language model.
[0183] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0184] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable device for updating the parameters of a large language model, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0185] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0186] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0187] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0188] A computer system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and virtual private servers (VPS). It may also be a server in a distributed system or a server integrated with blockchain.
[0189] According to the technical solution of the embodiments of the present disclosure, a method and device for updating the parameters of a large language model are provided. Based on the facts that the text generation process should be based on, an effective and accurate reward signal is provided for the factual accuracy and factual comprehensiveness of the text generated by the large language model. This can effectively avoid the problems of factual voids or factual fabrications in the text generation process of the large language model, especially in the generation process of long texts, and help to improve the factual accuracy and factual comprehensiveness of the text generated by the large language model.
[0190] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided by this disclosure can be achieved. This is not a limitation herein.
[0191] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for updating parameters of a large language model, comprising: Based on the text generation request, a large language model is used to generate the target text; According to the text generation request, determining facts that the target text generation process should be based on to obtain a target fact set; determining reward data based on the target fact set and an information set including unverified information in the target text; The parameters of the large language model are updated according to the reward data.
2. The method according to claim 1, wherein The step of determining, based on the text generation request, facts that the target text generation process should be based on, and obtaining a target fact set, includes: Using a plurality of preset determination methods, respectively according to the text generation request, determining facts that the target text generation process should be based on, and obtaining an initial fact set; The target fact set is obtained based on the multiple initial fact sets corresponding one-to-one to the multiple preset determination methods.
3. The method according to claim 2, wherein: The method adopts multiple preset determination methods to determine the facts that the target text generation process should be based on according to the text generation request, and obtains an initial fact set, including: The initial fact set is generated using at least two of the following preset determination methods: generating, by means of the large language model, a plurality of reply texts corresponding to the text generation request, and filtering preliminary facts from the plurality of reply texts to obtain the initial fact set; Using the large language model, preliminary facts are screened from search results generated by a search system for the text generation request to obtain the initial fact set; The initial fact set is generated according to the received input operation.
4. The method according to claim 2, wherein: The step of obtaining the target fact set based on the multiple initial fact sets corresponding to the multiple preset determination methods includes: Clustering and merging preliminary facts in the plurality of initial fact sets to obtain a merged set; The preliminary facts in the merged set are screened and verified to obtain the target fact set.
5. The method according to claim 4, wherein The screening and verification of the preliminary facts in the merged set to obtain the target fact set includes: Filtering the preliminary facts in the merged set using the large language model to obtain a filtered set; For the preliminary facts in the filtered set, generating a query request corresponding to the preliminary facts, and generating a query result of the query request through a retrieval system; Verifying preliminary facts in the filtered set based on the query result to obtain a verified set; The target fact set is obtained according to a screening operation on the preliminary facts in the verified set.
6. The method according to any one of claims 1 to 5, wherein The determining reward data according to the target fact set and the information set including the unverified information in the target text includes: determining the accuracy and comprehensiveness of the unverified information in the information set based on the target fact set; The reward data is determined based on the accuracy and completeness.
7. The method according to claim 6, wherein: Determining the accuracy and comprehensiveness of the unverified information in the information set based on the target fact set includes: determining a first number of accurate facts in the target fact set that are consistent with the unverified information, a second number of contradictory facts that are inconsistent with the unverified information, and a third number of omitted facts that are not covered by the unverified information; The accuracy and the comprehensiveness are determined based on the first quantity value, the second quantity value, and the third quantity value.
8. The method according to claim 7, wherein: The determining the accuracy according to the first quantity value, the second quantity value, and the third quantity value includes: combining the first quantity value and the second quantity value to obtain a fourth quantity value; The accuracy is determined based on the first quantity value and the fourth quantity value.
9. The method according to claim 7, wherein: Determining the comprehensiveness according to the first quantity value, the second quantity value, and the third quantity value includes: combining the first quantity value and the second quantity value to obtain a fourth quantity value; combining the first quantity value, the second quantity value, and the third quantity value to obtain a fifth quantity value; The comprehensiveness is determined based on the fourth quantity value and the fifth quantity value.
10. The method according to any one of claims 1 to 9, wherein Generating target text using a large language model according to a text generation request includes: Generate a plurality of target texts using the large language model according to the text generation request; and The updating of the parameters of the large language model according to the reward data includes: The parameters of the large language model are updated according to the plurality of reward data corresponding one-to-one to the plurality of target texts.
11. A text generation method based on a large language model, comprising: Get text generation request; Generate target text according to the text generation request using a large language model, wherein the large language model is updated according to any one of claims 1-10.
12. A device for updating parameters of a large language model, comprising: a text generation unit, configured to generate a target text using a large language model according to a text generation request; a fact determination unit configured to determine, according to the text generation request, facts that the target text generation process should be based on, and obtain a target fact set; a reward determining unit configured to determine reward data based on the target fact set and an information set including unverified information in the target text; A parameter updating unit is configured to update parameters of the large language model according to the reward data.
13. A text generation device based on a large language model, comprising: a request obtaining unit configured to obtain a text generation request; The request processing unit is configured to generate a target text according to the text generation request through a large language model, wherein the large language model is updated according to claim 12.
14. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.
15. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 11.
16. A computer program product comprising: A computer program which, when executed by a processor, implements the method according to any one of claims 1 to 11.