Language model training method and apparatus, and computer program product
By collecting feedback results and correcting text during the dialogue between the language model and the user, and iteratively training the preset language model, the problems of poor user experience and inaccurate feedback in the existing technology are solved, and a more natural and accurate model reply is achieved.
Patent Information
- Application Number
- PCT/CN2024/137880
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-30
- Filing Date
- 2024-12-09
- Publication Date
- 2025-07-03
AI Technical Summary
When collecting user feedback, the prior art leads to poor user experience through likes/clicks or selections from multiple replies, and cannot accurately know the user's dissatisfaction with the details of the replies, resulting in the trained language model being unable to output more accurate and text that meets user needs.
During the dialogue between the preset language model and the user, collect user feedback results and model reply correction text, build model training samples, and iteratively train the preset language model to generate more natural and reasonable reply text.
It reduces the complexity of user feedback operations, collects more detailed user feedback information, and improves the accuracy of model responses and meets user needs.
Smart Images

Figure CN2024137880_03072025_PF_FP_ABST
Abstract
Description
A language model training method, device and computer program product
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This disclosure claims priority to Chinese patent application number 2023118595627, filed with the Chinese Patent Office on December 30, 2023, entitled “A Language Model Training Method, Device and Computer Program Product,” the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0003] The present disclosure relates to the field of computer technology, and in particular to a language model training method, apparatus, and computer program product. Background Art
[0004] A large language model is an artificial intelligence model designed to understand and generate human language. It typically collects user feedback on their satisfaction with responses and iteratively trains the model, hoping to output more accurate responses that meet user needs. To this end, some embodiments of this specification provide a language model training method designed to more efficiently and accurately obtain user feedback and use it to train the model. Summary of the Invention
[0005] The embodiments of the present disclosure at least provide a language model training method and related equipment. The related equipment may include a language model training device, an electronic device, a computer-readable storage medium and a computer program product. The model training text can be obtained by reducing the complexity of user feedback operations, and the preset language model can be automatically iteratively trained, so that the preset language model can output text that is more accurate and meets user needs.
[0006] To achieve at least one of the above objectives, the technical solutions adopted in this application are as follows:
[0007] A first aspect: The present disclosure provides a language model training method, comprising:
[0008] Obtain target text corresponding to one or more rounds of dialogue between a preset language model and a user, the target text including user input text and model reply text; analyze the target text using the preset language model to obtain user feedback results for the one or more rounds of dialogue; obtain model reply correction text for the user input text in the target text based on the user feedback results; construct a model training sample based on the target text, the user feedback results, and the model reply correction text; and iteratively train the preset language model based on the model training sample.
[0009] In an optional implementation, obtaining the target text corresponding to one or more rounds of dialogue between the preset language model and the user includes:
[0010] Based on the user input text that has been obtained, analyze whether the user is dissatisfied with the model reply text that has been output or obtain the confidence level corresponding to the model reply text that has been output; if dissatisfied or the confidence level is less than a threshold, generate a model inquiry text; the target text includes the user input text that has been obtained, the model reply text that has been output, the model inquiry text, and the user input text for the model inquiry text.
[0011] In an optional implementation, the model inquiry text includes a statement for inquiring about satisfaction and / or a statement for inquiring about improvement suggestions.
[0012] In an optional implementation, the confidence level corresponding to the model reply text is used to represent the satisfaction of the preset language model with the model reply text.
[0013] In an optional embodiment, calling a preset language model to analyze the target text and obtaining user feedback results for the one or more rounds of dialogue includes:
[0014] Obtain one or more preset feedback inquiry texts; the one or more preset feedback inquiry texts have different question granularity; input the one or more preset feedback inquiry texts into the preset language model to obtain the user feedback result output by the preset language model.
[0015] In an optional embodiment, obtaining a model reply correction text for the user input text in the target text based on the user feedback result includes:
[0016] Obtain a preset revised query text; input the revised query text into the preset language model to obtain a model reply revised text output by the preset language model.
[0017] In an optional embodiment, before obtaining the target text corresponding to one or more rounds of dialogue between the preset language model and the user, the method further includes:
[0018] Acquire a dialogue training text and a guidance training text; the dialogue training text includes question sentences related to a preset topic, and the guidance training text includes a sentence for instructing the model to inquire about satisfaction and / or a sentence for instructing the model to inquire about improvement suggestions; input the dialogue training text into the preset language model to obtain a training reply text output by the preset language model; input the guidance training text into the preset language model so that the preset language model can output a sentence for inquiring about satisfaction and / or a sentence for inquiring about improvement suggestions when conversing with a user.
[0019] According to the language model training method provided in some embodiments of the present specification, the guided training text includes a statement for instructing the model to inquire about satisfaction when it is determined based on the dialogue training text that the other party is dissatisfied or the confidence of the training reply text is less than a set threshold, and / or a statement for instructing the model to inquire about improvement suggestions when it is determined based on the dialogue training text that the other party is dissatisfied or the confidence of the training reply text is less than a set threshold.
[0020] In an optional implementation, the iteratively training the preset language model based on the model training sample includes:
[0021] Inputting the target text and the user feedback results in the model training sample as input data into the preset language model, and obtaining an optimized model reply text output by the preset language model based on the target text and the user feedback results;
[0022] Using the model reply correction text in the model training sample as the supervisory signal, the model parameters of the preset language model are adjusted so that the difference between the optimized model reply text output by the preset language model and the model reply correction text as the supervisory signal is reduced, thereby obtaining the trained preset language model.
[0023] In an optional implementation, the iteratively training the preset language model based on the model training sample includes:
[0024] Replacing one or more model response texts corresponding to the target text in the model training sample with the model response correction text to obtain a positive sample dialogue text;
[0025] The target texts with unsatisfactory user feedback results in the model training samples are used as conversation texts of negative samples;
[0026] A reward model is obtained based on the positive sample or negative sample training, and the target text obtained by the preset language model during the conversation with the user is input into the reward model to obtain a reward score, and the model parameters of the preset language model are adjusted to maximize the reward score to obtain a trained preset language model.
[0027] In an optional implementation, after inputting the guidance training text into the preset language model, the method further includes:
[0028] A target training reply text is determined from the two or more candidate training reply texts output by the preset language model and is input into the preset language model.
[0029] Second aspect: The present disclosure also provides a language model training device, comprising:
[0030] A target text acquisition unit configured to acquire target text corresponding to one or more rounds of dialogue between a preset language model and a user, wherein the target text includes user input text and model reply text;
[0031] an analyzing unit configured to analyze the target text using a preset language model to obtain user feedback results for the one or more rounds of dialogue;
[0032] a correction text acquisition unit configured to acquire a model reply correction text for the user input text in the target text based on the user feedback result;
[0033] A construction unit configured to construct a model training sample based on the target text, the user feedback result, and the model reply correction text;
[0034] The training unit is configured to iteratively train the preset language model based on the model training samples.
[0035] In an optional embodiment, the target text acquisition unit can be specifically configured to analyze whether the user is dissatisfied with the output model reply text based on the user input text that has been acquired or to obtain the confidence level corresponding to the output model reply text; if dissatisfied or the confidence level is less than a threshold, a model inquiry text is generated; the target text includes the user input text that has been acquired, the model reply text that has been output, the model inquiry text, and the user input text of the user for the model inquiry text.
[0036] In an optional implementation, the model inquiry text includes a statement for inquiring about satisfaction and / or a statement for inquiring about improvement suggestions.
[0037] In an optional embodiment, the analysis unit can be specifically configured to obtain one or more preset feedback inquiry texts; the one or more preset feedback inquiry texts have different question granularity; input one or more preset feedback inquiry texts into the preset language model, and obtain the user feedback result output by the preset language model.
[0038] In an optional implementation, the correction text acquisition unit may be specifically configured to acquire a preset correction inquiry text; input the correction inquiry text into the preset language model, and obtain a model reply correction text output by the preset language model.
[0039] In an optional embodiment, it may also include a training text acquisition unit, a first output unit, and a second output unit, wherein the training text acquisition unit is configured to acquire dialogue training text and guidance training text; the dialogue training text includes question sentences related to a preset topic, and the guidance training text includes sentences for instructing the model to inquire about satisfaction and / or sentences for instructing the model to inquire about improvement suggestions; the first output unit is configured to input the dialogue training text into the preset language model to obtain a training reply text output by the preset language model; the second output unit is configured to input the guidance training text into the preset language model, so that the preset language model can output sentences for inquiring about satisfaction and / or sentences for inquiring about improvement suggestions when conversing with a user.
[0040] In an optional embodiment, the guided training text includes a statement for instructing the model to inquire about satisfaction when it is determined based on the dialogue training text that the other party is dissatisfied or the confidence of the training reply text is less than a set threshold, and / or a statement for instructing the model to inquire about improvement suggestions when it is determined based on the dialogue training text that the other party is dissatisfied or the confidence of the training reply text is less than a set threshold.
[0041] In an optional implementation, a determination unit may be further included, wherein the determination unit is configured to determine a target training reply text from two or more candidate training reply texts output by the preset language model and input the target training reply text into the preset language model.
[0042] Aspect 3: The present disclosure also provides an electronic device, comprising a processor and a memory, wherein the memory stores a plurality of instructions, and the processor loads the instructions to execute the steps in the language model training method provided in the embodiments of this specification.
[0043] Fourth aspect: The present disclosure also provides a computer-readable storage medium on which a computer program is stored, wherein when the computer program is executed by a processor, the steps in the language model training method provided in the embodiments of this specification are implemented.
[0044] Fifth aspect: The present disclosure also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps in the language model training method provided in the embodiments of this specification.
[0045] The present disclosure provides a language model training method and related devices that can obtain target text corresponding to one or more rounds of conversations between a preset language model and a user, the target text including user input text and model response text; analyze the target text using the preset language model to obtain user feedback results for the one or more rounds of conversations; obtain model response correction text for the user input text in the target text based on the user feedback results; construct model training samples based on the target text, user feedback results, and model response correction text; and iteratively train the preset language model based on the model training samples. Some embodiments of this specification can iteratively train the preset language model by collecting user feedback results and model response correction text during the conversation between the preset language model and the user, in the expectation that the preset language model can generate more natural and reasonable model response text during the conversation with the user, thereby collecting content related to user feedback while reducing the complexity of the user feedback operation, and ultimately enabling the preset language model to output text that is more accurate and meets user needs.
[0046] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without inventive effort.
[0048] FIG1 is a scenario diagram of a language model training method provided in some embodiments of this specification.
[0049] FIG2 is a flowchart of a language model training method provided in some embodiments of this specification.
[0050] FIG3 is a flowchart of fine-tuning a preset language model provided by some embodiments of this specification.
[0051] FIG4 is a schematic block diagram of a language model training apparatus provided in some embodiments of this specification.
[0052] FIG5 is a schematic diagram of the structure of an electronic device provided in some embodiments of this specification. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. The components of the embodiments of the present disclosure generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the disclosure for which protection is sought, but merely represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present disclosure.
[0054] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0055] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0056] Natural language processing (NLP) is a key area of research in computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. Natural language processing (NLP) integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language we use in everyday life—and is closely linked to the study of linguistics. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question answering, and knowledge graphs.
[0057] A large language model is an artificial intelligence model designed to understand and generate human language. It typically iterates the model by collecting user feedback on their satisfaction with responses, hoping to output more accurate responses that meet user needs. Therefore, efficiently and accurately obtaining user feedback on response satisfaction significantly impacts the effectiveness of model training.
[0058] In some embodiments, after the conversation between the model and the user ends, the user can be asked to collect user feedback by clicking likes / dislikes, or selecting one from multiple replies, etc., where a like indicates that the user is satisfied with the reply, and a dislike indicates that the user is dissatisfied with the reply. After collecting the feedback, the model is iteratively trained, and replies with likes are regarded as good, and replies with dislikes are regarded as bad, so as to suppress bad replies.
[0059] However, collecting user feedback through likes and dislikes results in a poor user experience and is noisy. Selecting a response from multiple replies also offers a poor and unnatural user experience. Furthermore, related technologies are unable to clearly identify which details in a response a user is dissatisfied with, nor can they determine what the user considers the correct answer. Consequently, language models trained in this way cannot output text that is more accurate and meets user needs.
[0060] To this end, some embodiments of this specification provide a language model training method and related equipment, which collects user feedback results and model reply correction text during the conversation between the preset language model and the user, and iteratively trains the preset language model, in the expectation that the preset language model can generate a more natural and more reasonable model reply text during the conversation with the user. Among them, the related equipment may include a language model training device, an electronic device, a computer-readable storage medium and a computer program product. The language model training device can be specifically integrated into an electronic device, which can be a terminal or a server. The language model training method provided in some embodiments of this specification can be executed on a terminal, on a server, or jointly by a terminal and a server. The above examples should not be understood as limiting the embodiments of this specification.
[0061] The following will clearly and completely describe some embodiments of this specification in conjunction with the accompanying drawings. It should be understood that the described embodiments are only part of the embodiments of the technical solution, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of this specification.
[0062] As shown in Figure 1, a language model training method is performed jointly by a terminal and a server. The language model training system provided in the embodiments of this specification includes a terminal 10 and a server 11, etc. The terminal 10 and the server 11 are connected via a network, such as a wired or wireless network, wherein the language model training device can be integrated into the server.
[0063] The server 11 can be configured to: receive the target text sent by the terminal 10, analyze the target text using a preset language model, and obtain user feedback results for one or more rounds of dialogue; based on the user feedback results, obtain model reply correction text for the user input text in the target text; construct a model training sample based on the target text, user feedback results, and model reply correction text; and iteratively train the preset language model based on the model training samples. The server 11 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. In the language model training method or device disclosed in this specification, multiple servers can be combined into a blockchain, and the server is a node on the blockchain.
[0064] The terminal 10 can be configured to collect target text corresponding to one or more rounds of conversation between a preset language model and a user, and send the collected target text to a server for analysis. The terminal 10 can include a mobile phone, an intelligent voice interaction device, a smart home appliance, an in-vehicle terminal, an aircraft, a tablet computer, a laptop computer, or a personal computer (PC). The terminal 10 can also be configured with a client, which can be an application client or a browser client, among others.
[0065] The following are detailed descriptions. It should be noted that the order of description of the following embodiments does not limit the preferred order of the embodiments. This embodiment will be described from the perspective of a language model training device, which can be integrated into an electronic device, such as a server or a terminal. As shown in Figure 2, the specific process of the language model training method may include the following steps.
[0066] S201, obtaining target text corresponding to one or more rounds of dialogue between a preset language model and a user.
[0067] Among them, the preset language model is a natural language processing tool driven by artificial intelligence technology. The preset language model can generate model reply text based on the patterns and statistical laws seen in the pre-training stage, and can also interact with users according to the context of the conversation.
[0068] The target text includes user input text and model reply text. The user input text refers to the text entered by the user in the preset language model, and the model reply text refers to the reply text generated by the preset language model based on the user input text.
[0069] In some embodiments, a target text consisting of a conversation between a preset language model and a user can be obtained. The target text may include a single round of conversation or multiple rounds of conversation. In some embodiments, a single round of conversation may include a sentence output by one participant (e.g., a question posed) and a sentence output by another participant (e.g., a response to the question). That is, a single round of conversation may be a pair of question-and-answer texts or question-and-answer sentences. Correspondingly, a multiple-round conversation may include multiple pairs of question-and-answer texts or multiple pairs of question-and-answer sentences. For example, the target text may include one round of dialogue, which is specifically: user input text "What is machine learning?", model reply text "Machine learning specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structures to continuously improve their own performance." Or the target text may include two or more rounds of dialogue, which may be: user input text "What is machine learning?", model reply text "Machine learning specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structures to continuously improve their own performance.", user input text "Can you introduce it in more detail?", model reply text "Machine learning is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory and other disciplines. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structures to continuously improve their own performance."
[0070] Optionally, in some embodiments, the step of "obtaining target text corresponding to one or more rounds of conversation between a preset language model and a user" may include, through a preset language model: analyzing, based on the acquired user input text, whether the user is dissatisfied with the output model reply text or obtaining the confidence level corresponding to the output model reply text; if dissatisfied or the confidence level is less than a threshold, generating a model inquiry text.
[0071] Among them, the model inquiry text is an inquiry statement put forward to the user when the preset language model analysis finds that the user is not satisfied with the output model response text, or the preset language model itself is not satisfied with the output model response text, in the expectation of obtaining feedback from the user. Through such a model inquiry text, the preset language model can obtain the user's feedback information in a timely manner during the dialogue with the user, thereby improving the timeliness of obtaining user feedback information. On the other hand, the user can describe the feedback information in the form of natural language dialogue, without the need to perform additional operations outside the dialogue process to provide feedback information, which improves the user experience while also helping to obtain more detailed and specific user feedback information.
[0072] In some embodiments, the target text may include the user input text that has been obtained, the model reply text that has been output, the model query text, and the user input text in response to the model query text, for example, the user input text "What is machine learning?", the model reply text "Machine learning is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance.", the model query text "Is this reply too complicated?", and the user input text in response to the model query text "Yes, this reply is too complicated."
[0073] In some embodiments, a preset language model can be used to analyze the user input text to determine whether the user is dissatisfied with the output model response text. If the preset language model analyzes that the user is dissatisfied with the output model response text, a model query text is generated. For example, if the user input text "What's this response?" is obtained, and the preset language model analyzes the user input text and determines that the user is dissatisfied with the output model response text, a model query text "What's the best response to the current question?" is generated in the hope of obtaining user feedback.
[0074] The confidence level corresponding to the model reply text is used to characterize the degree of satisfaction of the preset language model itself with the model reply text that has been output. In some embodiments, the confidence level corresponding to the model reply text that has been output can also be obtained through the preset language model. For example, when the preset language model outputs a sentence of model reply text, it will also automatically calculate the confidence level corresponding to the model reply text. As an embodiment, the confidence level can be a score, such as a value in the range of 0 to 1, to characterize whether the preset language model itself has confidence in the model reply text. In some embodiments, the greater the confidence level, the more confident the preset model is in the model reply text it generates. Correspondingly, if the preset language model determines that the confidence level corresponding to a certain sentence of model reply text is less than a preset threshold, it means that the preset language model itself is not satisfied or confident enough with the model reply text that has been output, then the preset language model can further generate a model query text. For example, the model response text that has been output is "xxx is really beautiful.", and the confidence corresponding to the model response text is less than the preset threshold. This means that the preset language model itself is not satisfied with or confident in the output model response text. In this case, a model inquiry text "Is there anything unreasonable in my answer just now?" can be generated in the hope of obtaining user feedback.
[0075] The above embodiment is based on the inherent ability of the preset language model to automatically determine whether the user is satisfied with the model response text that has been output, or whether the user has confidence in the model response text of the preset language model itself, so that the preset language model can output the model inquiry text in a timely manner during the conversation with the user to obtain user feedback information.
[0076] Optionally, in some embodiments, the model inquiry text includes a statement for inquiring about satisfaction and / or a statement for inquiring about improvement suggestions.
[0077] In some embodiments, the model inquiry text may also include statements for inquiring satisfaction, such as "Did my answer just now offend you?"; in addition, the model inquiry text may also include statements for asking for improvement suggestions, such as "How can I answer better?", etc.
[0078] By introducing the method of model inquiry text, when the user is not satisfied with the output model response text, or the preset language model is not confident in its own model response text, we can further integrate model inquiry text related to asking for user feedback into the process of dialogue with the user, actively discuss with the user, and thus collect direct feedback from the user in a more natural and reasonable way.
[0079] In some embodiments, the preset language model can be fine-tuned or annotated so that it can output the model query text in a timely manner during the conversation with the user. In some embodiments, the preset language model that is fine-tuned or annotated may already have the ability to automatically determine whether the user is satisfied with the model reply text that has been output, or whether the user has confidence in the model reply text of the preset language model itself. This ability can be obtained during the pre-training stage of the preset language model. For more information on fine-tuning or annotating the preset language model, please refer to the relevant description of Figure 3, which will not be repeated here.
[0080] S202: Analyze the target text using a preset language model to obtain user feedback results for one or more rounds of dialogue.
[0081] In some embodiments, the target text can be analyzed using a preset language model, thereby collecting user feedback results for one or more rounds of conversations during the conversation with the user. The user feedback results can be used to represent the user's satisfaction with the one or more rounds of conversations.
[0082] Optionally, in some embodiments, the step of "calling a preset language model to analyze the target text and obtaining user feedback results for one or more rounds of conversations" may include: obtaining one or more preset feedback inquiry texts; one or more preset feedback inquiry texts having different question granularity; inputting one or more preset feedback inquiry texts into the preset language model to obtain user feedback results output by the preset language model.
[0083] In some embodiments, one or more feedback inquiry texts may be pre-set, wherein the pre-set feedback inquiry texts have different question granularities. For example, a relatively coarse-grained feedback inquiry text such as "Did my answer just now offend you?", "Is there anything unreasonable in my answer just now?", etc. may be pre-set; a relatively fine-grained feedback inquiry text may also be pre-set to further ask in more detail when the user is not satisfied with the reply text, such as "Why do you think the third item is not good?", "How to answer it better?", "What is the best answer to the current question?", etc.
[0084] After obtaining the preset feedback query text, the feedback query text can be input into the preset language model so that the preset language model can learn the query method of the feedback query text and guide the preset language model to obtain user feedback results through natural language interaction.
[0085] In some embodiments, there are multiple ways to pre-set one or more feedback query texts and use them to train a preset language model. This can be done manually by annotators or automatically by a computer program. For example, multiple feedback query texts can be pre-set and recorded in a question library, and a call statement can be written into a script file in advance. When the script file is run, one or more feedback query texts can be automatically selected from the question library and input into the preset language model. The embodiments of this specification do not limit the methods for setting feedback query texts and calling the preset language model to obtain user feedback results.
[0086] Some embodiments of this specification directly invoke a preset language model to analyze the target text corresponding to the conversation between the model and the user to generate user feedback results, further enhancing the preset language model's ability to understand user feedback text. Furthermore, when the target text contains the user's response to the model's query, the language model's ability to summarize and summarize user feedback results can be further effectively enhanced, resulting in more accurate user feedback results.
[0087] S203: Based on the user feedback result, obtain a model reply correction text for the user input text in the target text.
[0088] In some embodiments, after the preset language model obtains the user feedback results, it can know the user's satisfaction with the model reply text in the target text, that is, it can know which model reply text or texts the user is satisfied with and which model reply text or texts the user is dissatisfied with. The preset language model can analyze and integrate the model reply texts that the user is satisfied with based on the user feedback results, so as to generate a model reply text that is more suitable for the user input text in the target text. The model reply text after such analysis and integration is the model reply correction text. By inputting the model reply correction text into the preset language model, the preset language model can learn a more appropriate reply method to enhance the interactive function of the preset language model.
[0089] In some embodiments, there may be multiple ways to obtain model reply correction text based on user feedback results. For example, the annotation personnel may summarize the model reply correction text based on the model reply text in the target text according to the user feedback results; the corresponding language model may also be called to analyze and integrate the model reply text based on the user feedback results to obtain the model reply correction text, and so on.
[0090] In some embodiments, the step of "obtaining a model reply correction text for the user input text in the target text based on the user feedback results" may include: obtaining a preset correction query text; inputting the correction query text into a preset language model to obtain a model reply correction text output by the preset language model.
[0091] For example, a correction query text can be set in advance and input into the preset language model to obtain the model response correction text output by the preset language model, thereby using the set correction query text to guide the preset language model to summarize a more optimized model response correction text.
[0092] In some embodiments, there are many ways to pre-set the corrected query text and use the corrected query text to train the preset language model. It can be achieved by manual setting by the annotation personnel, or it can be automatically achieved by a computer program. As an example, similar to the feedback query text, multiple corrected query texts can be pre-set and recorded in a question library (such as another question library), and the call statement can be written into a script file (which can be another script file) in advance. When the script file is run, one or more corrected query texts can be automatically selected from the question library and input into the preset language model. The embodiments of this specification do not limit the method of setting the corrected query text and obtaining the model reply corrected text.
[0093] S204: Construct a model training sample based on the target text, user feedback results, and model reply correction text.
[0094] In some embodiments, a model training sample can be constructed based on the target text, user feedback results, and model reply correction text to train a preset language model.
[0095] S205: Iteratively train a preset language model based on the model training samples.
[0096] In some embodiments, model parameters in a preset language model can be adjusted using model training samples through supervised learning to achieve the required performance. As an example, after obtaining the target text corresponding to one or more rounds of dialogue and obtaining model training samples based on the target text, the target text and user feedback results can be input into the preset language model as input data, so that the preset language model outputs an optimized model response text based on the target text and user feedback results. The model response correction text is used as a supervision signal, and the model parameters of the preset language model are adjusted so that the difference between the optimized model response text output by the preset language model and the model response correction text used as the supervision signal is reduced, thereby obtaining a trained preset language model.
[0097] In other embodiments, reinforcement learning can be used to provide reward or penalty feedback to the model response text of the preset language model during the dialogue process as the model continuously interacts with the user. This allows the feedback to be quantified, and the model parameters of the preset language model are continuously adjusted based on the feedback, ultimately maximizing the benefits of the overall dialogue effect. As an example, the model response correction text can be the result of correcting one or more model response texts in the target text. Therefore, one or more model response texts corresponding to the target text can be replaced with the model response correction text to obtain a positive sample dialogue text. The target text for which the user feedback result is unsatisfactory is used as a negative sample dialogue text, and a reward model is obtained based on the positive or negative sample training. The target text obtained by the preset language model during the user dialogue process is input into the reward model to obtain a reward score. The model parameters of the preset language model are adjusted to maximize the reward score, thereby obtaining a trained preset language model.
[0098] In some embodiments of this specification, the target text corresponding to the preset language model and the user dialogue can be continuously obtained, and thus the iterative training process will be continuously performed, so that the preset language model can be optimized during the dialogue, and a more intelligent preset language model can be obtained, making the dialogue with the user more natural, and at the same time outputting more accurate responses to the user's questions, thereby improving user satisfaction.
[0099] In some embodiments, the user behavior can also be analyzed during the dialogue between the preset language model and the user to determine whether the user is satisfied with the current reply. For example, the interval between adjacent user input texts, other mouse behavior trajectories or touch and slide trajectories on the terminal or web page, the user asking the same question multiple times, the user questioning the model reply text, and other user behavior data can be collected. By analyzing these user behavior data, feedback results on whether the user is satisfied with the current reply can be obtained. For example, it can be defined that a long interval between adjacent user input texts, a user asking the same question multiple times, or a user questioning the model reply text, etc., indicate that the user is not satisfied with the current model reply text. Furthermore, the judgment result of whether the user is satisfied based on user behavior can be combined with the aforementioned user feedback result. For example, the judgment result and the user feedback result can be weighted based on a preset weight, and the calculation result can be used to replace the user feedback result input to the preset language model in the aforementioned supervised learning process. For example, the calculation result can be used as a supervisory signal or label for training a reward model in reinforcement learning.
[0100] As can be seen from the above, this embodiment can obtain target text corresponding to one or more rounds of conversation between a preset language model and a user, the target text including user input text and model response text; analyze the target text using the preset language model to obtain user feedback results for the one or more rounds of conversation; obtain model response correction text for the user input text in the target text based on the user feedback results; construct model training samples based on the target text, user feedback results, and model response correction text; and iteratively train the preset language model based on the model training samples. Some embodiments of this specification can iteratively train the preset language model by collecting user feedback results and model response correction text during the conversation between the preset language model and the user, in the expectation that the preset language model can generate more natural and reasonable model response text during the conversation with the user, thereby collecting more content related to user feedback while reducing the complexity of user operations, and ultimately enabling the preset language model to output text that is more accurate and meets user needs.
[0101] Optionally, in some embodiments, before the step of "obtaining target text corresponding to one or more rounds of conversation between the preset language model and the user," a process of fine-tuning or annotating the preset language model may be included. As shown in FIG3 , the specific flow of this process may include the following steps.
[0102] S301, obtaining dialogue training text and guidance training text.
[0103] S302: Input a dialogue training text into a preset language model to obtain a training reply text output by the preset language model.
[0104] S303: Inputting a guidance training text into a preset language model so that the preset language model can output a sentence for inquiring about satisfaction and / or a sentence for inquiring about improvement suggestions when talking with the user.
[0105] The conversation training text includes question sentences related to preset topics, aiming to simulate user conversations and interact with users based on the context of the conversation. The guided training text includes sentences instructing the model to inquire about satisfaction and / or suggestions for improvement.
[0106] For example, before obtaining the target text corresponding to one or more rounds of conversation between a preset language model and a user, a step of fine-tuning or annotating the preset language model should also be included. This can involve first obtaining conversation training text and guidance training text, then inputting the conversation training text into the preset language model to obtain training response text output by the preset language model, enabling the preset language model to mimic human conversation. Inputting the guidance training text into the preset language model enables the preset language model to output statements inquiring about user satisfaction and / or statements inquiring about improvement suggestions during conversations with the user, thereby successfully obtaining user satisfaction and enabling the model to learn more appropriate response language.
[0107] Optionally, in some embodiments, the guided training text includes a statement for instructing the model to inquire about satisfaction when it is determined based on the dialogue training text that the other party is dissatisfied or the confidence of the training reply text is less than a set threshold, and / or a statement for instructing the model to inquire about improvement suggestions when it is determined based on the dialogue training text that the other party is dissatisfied or the confidence of the training reply text is less than a set threshold.
[0108] Specifically, the guided training text can be used to instruct the model to inquire about satisfaction when it is determined based on the dialogue training text that the other party is dissatisfied or the confidence of the training reply text is less than a set threshold, and / or to instruct the model to inquire about improvement suggestions when it is determined based on the dialogue training text that the other party is dissatisfied or the confidence of the training reply text is less than a set threshold. This enables the model to more clearly understand the user's satisfaction level and provide a more reasonable response method by inquiring about satisfaction or improvement suggestions when it knows that the user is not satisfied with the model reply text, or the model itself is not satisfied with the model reply text.
[0109] Optionally, in some embodiments, after the step of “inputting the guide training text into the preset language model”, the step may further include: determining a target training reply text from two or more candidate training reply texts output by the preset language model and inputting the target training reply text into the preset language model.
[0110] For example, during the model training process, the preset language model can output multiple candidate training reply texts based on the guide training text. It can determine a more reasonable target training reply text from multiple candidate training reply texts, and input the target training reply text into the preset language model for learning. This will enable the preset language model to learn more accurate reply methods, thereby improving the function of the preset language model.
[0111] In some embodiments, there are many ways to determine a more reasonable target training reply text from multiple candidate training reply texts. For example, the target training reply text can be determined from multiple candidate training reply texts through a trained network model; for another example, the target training reply text can be determined from multiple candidate training reply texts through annotation personnel, and so on. The embodiments of this specification do not provide a restrictive description of this.
[0112] Some embodiments of this specification also provide a language model training device, as shown in Figure 4, the language model training device may include a target text acquisition unit 401, an analysis unit 402, a revised text acquisition unit 403, a construction unit 404, and a training unit 405.
[0113] The target text acquisition unit 401 is configured to acquire target text corresponding to one or more rounds of dialogue between a preset language model and a user, wherein the target text includes user input text and model reply text.
[0114] The analyzing unit 402 is configured to analyze the target text using a preset language model to obtain user feedback results for the one or more rounds of dialogue.
[0115] The corrected text acquisition unit 403 is configured to acquire a model reply corrected text for the user input text in the target text based on the user feedback result.
[0116] The construction unit 404 is configured to construct a model training sample based on the target text, the user feedback result, and the model reply correction text.
[0117] The training unit 405 is configured to iteratively train the preset language model based on the model training samples.
[0118] Optionally, in some embodiments of the present specification, the target text acquisition unit 401 can be specifically configured to analyze whether the user is dissatisfied with the output model reply text based on the user input text that has been acquired or to obtain the confidence level corresponding to the output model reply text; if dissatisfied or the confidence level is less than a threshold, a model inquiry text is generated; the target text includes the user input text that has been acquired, the model reply text that has been output, the model inquiry text, and the user input text of the user for the model inquiry text.
[0119] Optionally, in some embodiments of the present specification, the model inquiry text includes a statement for inquiring about satisfaction and / or a statement for inquiring about improvement suggestions.
[0120] Optionally, in some embodiments of the present specification, the analysis unit 402 can be specifically configured to obtain one or more preset feedback inquiry texts; the one or more preset feedback inquiry texts have different question granularity; input one or more preset feedback inquiry texts into the preset language model, and obtain the user feedback result output by the preset language model.
[0121] Optionally, in some embodiments of the present specification, the correction text acquisition unit 403 can be specifically configured to obtain a preset correction query text; input the correction query text into the preset language model, and obtain the model reply correction text output by the preset language model.
[0122] Optionally, in some embodiments of the present specification, the language model training device may further include a training text acquisition unit, a first output unit, and a second output unit.
[0123] A training text acquisition unit is configured to acquire dialogue training text and guidance training text; the dialogue training text includes question sentences related to a preset topic, and the guidance training text includes sentences for indicating the satisfaction of the model inquiry and / or sentences for indicating the model inquiry improvement suggestions.
[0124] The first output unit is configured to input the dialogue training text into the preset language model to obtain a training reply text output by the preset language model.
[0125] The second output unit is configured to input the guidance training text into the preset language model, so that the preset language model can output a sentence for inquiring about satisfaction and / or a sentence for inquiring about improvement suggestions when talking with the user.
[0126] Optionally, in some embodiments of the present specification, the guided training text includes a statement for instructing the model to inquire about satisfaction when it is determined based on the dialogue training text that the other party is dissatisfied or the confidence of the training reply text is less than a set threshold, and / or a statement for instructing the model to inquire about improvement suggestions when it is determined based on the dialogue training text that the other party is dissatisfied or the confidence of the training reply text is less than a set threshold.
[0127] Optionally, in some embodiments of this specification, the language model training device may further include a determination unit.
[0128] The determining unit is configured to determine a target training reply text from the two or more candidate training reply texts output by the preset language model and input the target training reply text into the preset language model.
[0129] As can be seen from the above, in some embodiments of the present specification, the target text acquisition unit 401 can acquire target text corresponding to one or more rounds of dialogue between a preset language model and a user, the target text including user input text and model reply text; the analysis unit 402 analyzes the target text using the preset language model to obtain user feedback results for one or more rounds of dialogue; the correction text acquisition unit 403 obtains model reply correction text for the user input text in the target text based on the user feedback results; the construction unit 404 constructs model training samples based on the target text, user feedback results, and model reply correction text; and the training unit 405 iteratively trains the preset language model based on the model training samples. In some embodiments of the present specification, the preset language model can be iteratively trained by collecting user feedback results and model reply correction text during the dialogue between the preset language model and the user, in the expectation that the preset language model can generate more natural and reasonable model reply text during the dialogue with the user, thereby collecting content related to user feedback while reducing the complexity of the user feedback operation, and ultimately enabling the preset language model to output text that is more accurate and meets user needs.
[0130] For more information about each unit, please refer to the relevant description of Figures 2 and 3, which will not be repeated here. It should be understood that the system and its modules shown in Figures 2 and 3 can be implemented in various ways. For example, in some embodiments, the system and its modules can be implemented by hardware, software, or a combination of software and hardware. Among them, the hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or specially designed hardware. Those skilled in the art will understand that the above-mentioned methods and systems can be implemented using computer-executable instructions and / or control code contained in a processor, such as a carrier medium such as a disk, CD or DVD-ROM, or a memory of a programmable device. The system and its modules of this specification can be implemented not only by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, but can also be implemented by software executed by various types of processors, or by a combination of the above-mentioned hardware circuits and software (for example, firmware).
[0131] It should be noted that the above description of the system and its modules is for convenience only and does not limit this specification to the scope of the illustrated embodiments. It is understood that those skilled in the art, after understanding the principles of the system, may, without departing from these principles, arbitrarily combine the modules to form subsystems connected to other modules. Alternatively, they may split certain modules to obtain more modules or multiple units within a module. Such variations are within the scope of this specification.
[0132] Some embodiments of this specification also provide an electronic device, as shown in FIG5 , which shows a schematic structural diagram of an electronic device involved in the embodiments of this specification. The electronic device may be a terminal or a server, etc.
[0133] The electronic device may include components such as a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a power supply 503, and an input unit 504. Those skilled in the art will appreciate that the electronic device structure shown in FIG5 does not limit the electronic device, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0134] The processor 501 is the control center of the electronic device, connecting the various components of the electronic device using various interfaces and circuits. It executes the various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 502 and accessing data stored in the memory 502. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, where the application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 501.
[0135] The memory 502 can be configured to store software programs and modules, and the processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 502 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.
[0136] The electronic device also includes a power supply 503 for supplying power to various components. Preferably, the power supply 503 can be logically connected to the processor 501 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 503 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0137] The electronic device may further include an input unit 504 , which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0138] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail here. Specifically in this embodiment, the processor 501 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 502 according to the following instructions, and the processor 501 will run the application programs stored in the memory 502 to implement various functions, which may include: obtaining target text corresponding to one or more rounds of dialogue between a preset language model and a user, the target text including user input text and model reply text; analyzing the target text through the preset language model to obtain user feedback results for one or more rounds of dialogue; obtaining model reply correction text for the user input text in the target text based on the user feedback results; constructing a model training sample based on the target text, the user feedback results, and the model reply correction text; iteratively training the preset language model based on the model training sample.
[0139] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0140] As can be seen from the above, this embodiment can obtain target text corresponding to one or more rounds of conversation between a preset language model and a user, the target text including user input text and model reply text; analyze the target text using the preset language model to obtain user feedback results for one or more rounds of conversation; obtain model reply correction text for the user input text in the target text based on the user feedback results; construct a model training sample based on the target text, user feedback results, and model reply correction text; and iteratively train the preset language model based on the model training samples. This embodiment can iteratively train the preset language model by collecting user feedback results and model reply correction text during the conversation between the preset language model and the user, in the expectation that the preset language model can generate more natural and reasonable model reply text during the conversation with the user, thereby collecting content related to user feedback while reducing the complexity of the user feedback operation, and ultimately enabling the preset language model to output text that is more accurate and meets user needs.
[0141] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0142] To this end, some embodiments of this specification further provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the language model training methods provided in some embodiments of this specification. For example, the instructions can execute the following steps.
[0143] Obtain target text corresponding to one or more rounds of dialogue between a preset language model and a user, where the target text includes user input text and model reply text; analyze the target text using the preset language model to obtain user feedback results for one or more rounds of dialogue; obtain model reply correction text for the user input text in the target text based on the user feedback results; construct a model training sample based on the target text, user feedback results, and model reply correction text; and iteratively train the preset language model based on the model training sample.
[0144] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0145] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0146] Since the instructions stored in the computer-readable storage medium can execute the steps in any language model training method provided in some embodiments of this specification, the beneficial effects that can be achieved by any language model training method provided in some embodiments of this specification can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0147] Some embodiments of this specification also provide a computer program product, which includes computer instructions that may be stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the aforementioned language model training.
[0148] The above is a detailed introduction to a language model training method, a language model training device, and related electronic equipment, computer-readable storage media, and a computer program product provided in some embodiments of this specification. This specification uses specific examples to illustrate the principles and implementation methods of this technical solution. The description of the above embodiments is only used to help understand the method of this technical solution and its core idea; at the same time, for technical personnel in this field, based on the idea of this technical solution, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this technical solution. Industrial Applicability
[0149] By adopting the above scheme, the preset language model is iteratively trained by collecting user feedback results and model response correction text during the conversation between the preset language model and the user, in the expectation that the preset language model can generate more natural and reasonable model response text during the conversation with the user, thereby collecting content related to user feedback while reducing the complexity of user feedback operations, and ultimately enabling the preset language model to output text that is more accurate and meets user needs.
Claims
1. A method for training a language model, characterized in that, Including: Obtaining target text corresponding to one or more rounds of conversations between a preset language model and a user, where the target text includes user input text and model response text; Analyzing the target text through the preset language model to obtain a user feedback result of the user for the one or more rounds of conversations; Based on the user feedback result, obtaining a model response correction text for the user input text in the target text; Constructing a model training sample based on the target text, the user feedback result, and the model response correction text; Iteratively training the preset language model based on the model training sample.
2. The language model training method according to claim 1, wherein The obtaining of the target text corresponding to one or more rounds of conversations between the preset language model and the user includes, through the preset language model: Analyzing whether the user is dissatisfied with the model response text that has been output based on the user input text that has been obtained or obtaining the confidence level corresponding to the model response text that has been output; If dissatisfied or the confidence level is less than a threshold, generating a model inquiry text; The target text includes the user input text that has been obtained, the model response text that has been output, the model inquiry text, and the user input text of the user for the model inquiry text.
3. The language model training method according to claim 2, wherein The model inquiry text includes a statement for asking about satisfaction and / or a statement for asking for improvement suggestions.
4. The language model training method according to claim 2, characterized in that The confidence level corresponding to the model response text is used to represent the satisfaction of the preset language model with the model response text.
5. The language model training method according to claim 1, wherein The invoking of the preset language model to analyze the target text to obtain the user feedback result of the user for the one or more rounds of conversations includes: Obtaining one or more preset feedback inquiry texts; the one or more preset feedback inquiry texts have different question granularities; Inputting the one or more preset feedback inquiry texts into the preset language model to obtain the user feedback result output by the preset language model.
6. The language model training method according to claim 1, wherein The obtaining of the model response correction text for the user input text in the target text based on the user feedback result includes: Obtaining a preset correction inquiry text; Inputting the correction inquiry text into the preset language model to obtain the model response correction text output by the preset language model.
7. The language model training method according to claim 1, wherein Before obtaining the target text corresponding to one or more rounds of conversations between the preset language model and the user, it further includes: Obtaining dialogue training text and guiding training text; the dialogue training text includes question statements related to a preset topic, and the guiding training text includes a statement for instructing the model to ask about satisfaction and / or a statement for instructing the model to ask for improvement suggestions; Inputting the dialogue training text into the preset language model to obtain the training response text output by the preset language model; Inputting the guiding training text into the preset language model so that the preset language model can output a statement for asking about satisfaction and / or a statement for asking for improvement suggestions when conversing with the user.
8. The language model training method according to claim 7, wherein The guiding training text includes statements for instructing the model to ask about satisfaction when it is determined based on the conversation training text that the other party is dissatisfied or the confidence level of the training response text is less than a set threshold, and / or statements for instructing the model to ask for improvement suggestions when it is determined based on the conversation training text that the other party is dissatisfied or the confidence level of the training response text is less than a set threshold.
9. The language model training method according to claim 7, wherein After inputting the guiding training text into the preset language model, it further includes: Determine a target training response text from two or more candidate training response texts output by the preset language model and input it into the preset language model.
10. The language model training method according to any one of claims 1-9, characterized in that Iteratively training the preset language model based on the model training samples includes: Taking the target text and the user feedback result in the model training samples as input data and inputting them into the preset language model to obtain an optimized model response text output by the preset language model based on the target text and the user feedback result; Using the model response correction text in the model training samples as a supervision signal, adjusting the model parameters of the preset language model to reduce the difference between the optimized model response text output by the preset language model and the model response correction text used as the supervision signal, and obtaining the trained preset language model.
11. The language model training method according to any one of claims 1-9, characterized in that Iteratively training the preset language model based on the model training samples includes: Replacing one or more model response texts corresponding to the target text in the model training samples with the model response correction text to obtain the dialogue text of the positive sample; Taking the target text with an unsatisfactory user feedback result in the model training samples as the dialogue text of the negative sample; Training a reward model based on the positive sample or the negative sample, inputting the target text obtained by the preset language model during the conversation with the user into the reward model to obtain a reward score, and adjusting the model parameters of the preset language model to maximize the reward score to obtain the trained preset language model.
12. A language model training device, characterized in that, It includes: A target text acquisition unit configured to acquire the target text corresponding to one or more rounds of conversations between the preset language model and the user, where the target text includes the user input text and the model response text; An analysis unit configured to analyze the target text through the preset language model to obtain the user feedback result of the user for the one or more rounds of conversations; A correction text acquisition unit configured to acquire the model response correction text for the user input text in the target text based on the user feedback result; A construction unit configured to construct model training samples based on the target text, the user feedback result, and the model response correction text; A training unit configured to iteratively train the preset language model based on the model training samples.
13. An electronic device, characterized in that, It includes a memory and a processor; the memory stores an application program, and the processor is configured to run the application program in the memory to perform the operations in the language model method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the language model method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, the steps in the language model method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Intelligent dialog management apparatus and method
CN106558307A
Multi-round dialogue online training method and system based on user interaction
CN110059170A
Generative large language model training method and model-based man-machine voice interaction method
CN116127046A
Language model training method and device and computer program product
CN118036587A
Method and system for routing a question based on analysis of the question content and predicted user satisfaction with answer content before the answer content is generated
US10083213B1
Cited By
Real-time voice interaction method and system based on large model
CN120853551A
Model training method, video generation method, electronic equipment and storage medium
CN120953453A
Multi-round jail break attack evaluation method and device for Text-to-SQL system
CN121071899A