Multi-round dialogue model training method and device based on thinking chain
By constructing a multi-turn dialogue training text sequence and building an attention mask matrix, the problems of low efficiency and poor coherence in multi-turn dialogue training are solved, achieving efficient multi-turn dialogue model training and stable online inference performance.
Patent Information
- Application Number
- CN202511004370.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-25
AI Technical Summary
In interactive dialogue scenarios based on large language models, multi-turn dialogue training with thought chains cannot directly adopt traditional dialogue training methods, resulting in low training efficiency and poor model dialogue coherence. Furthermore, the format of historical messages without thought chains is inconsistent during online inference, affecting the model's application performance.
By acquiring multi-turn dialogue sample data, a training text sequence containing multiple consecutive single-turn dialogues is constructed, and an attention mask matrix is built to mask the attention dependencies between single-turn dialogues. The multi-turn dialogue model is trained using the attention mask matrix to ensure the consistency of the model's input format.
It improves the training efficiency of multi-turn dialogue models, ensures that the dependencies during model training are consistent with online inference, enhances the coherence and response quality of multi-turn dialogues, and stably performs the online inference effect.
Smart Images

Figure CN121009904A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and in particular to a training method and device for a multi-turn dialogue model based on a thought chain, a computer device, a computer readable storage medium, and a computer program product. BACKGROUND
[0002] In an interactive dialogue scenario based on a large language model, the introduction of a thought chain can enable the model to have certain thinking ability, so as to improve the quality of reply information generated by the model.
[0003] However, multi-turn dialogue training with a thought chain cannot directly follow the traditional dialogue training method. In related technologies, multi-turn dialogue is usually disassembled into multiple single-turn samples for training, resulting in low training efficiency and poor coherence of the model dialogue. Moreover, in multi-turn dialogue training, all subsequent rounds of dialogue can know the thought chain of the previous rounds of dialogue, which is inconsistent with the format of historical messages without thought chains when the model is online, affecting the application effect of the model. SUMMARY
[0004] Therefore, it is necessary to provide a training method, device, computer device, computer readable storage medium, and computer program product for a multi-turn dialogue model based on a thought chain, which can improve the training efficiency of a multi-turn dialogue model with a thought chain.
[0005] In a first aspect, the present application provides a training method for a multi-turn dialogue model based on a thought chain, which comprises:
[0006] Obtaining multi-turn dialogue sample data to construct a training text sequence containing continuous multiple single-turn dialogues; each single-turn dialogue in the training text sequence includes user interaction text and model reply text of the current round, and the model reply text has corresponding thought chain data;
[0007] According to the user interaction text and model reply text of each single-turn dialogue in the training text sequence and the corresponding thought chain data, an attention mask matrix is constructed; the model reply text of any single-turn dialogue in the attention mask matrix does not have an attention dependency relationship with the thought chain data of the single-turn dialogue before the any single-turn dialogue;
[0008] Based on the attention mask matrix, the multi-turn dialogue model is trained to obtain a trained multi-turn dialogue model; the model input format of the trained multi-turn dialogue model during inference matches the model input format of the multi-turn dialogue model during training.
[0009] In one embodiment, the method further comprises:
[0010] The training text sequence is subjected to model input preprocessing to obtain a text input token sequence and a training label sequence; the training label sequence is used to supervise the model reply text of each single-turn dialogue and the corresponding thinking chain data thereof.
[0011] In one of the embodiments, the model training of the multi-turn dialogue model based on the attention mask matrix comprises:
[0012] The text input token sequence, the training label sequence, and the attention mask matrix are input into the multi-turn dialogue model.
[0013] The multi-turn dialogue model is subjected to model training by calculating attention weights to obtain the trained multi-turn dialogue model.
[0014] In one of the embodiments, the attention mask matrix is constructed according to the user interaction text and the model reply text of each single-turn dialogue in the training text sequence and the corresponding thinking chain data thereof, and comprises:
[0015] The training text sequence is subjected to tokenization processing to generate a basic attention mask.
[0016] According to target dependency relationship shielding information, a supplementary attention mask is constructed according to the training text sequence; the target dependency relationship shielding information is used to indicate that the thinking chain data of a target single-turn dialogue in the training text sequence is shielded, and the dependency relationship between the user interaction text, the model reply text, and the thinking chain data thereof of the single-turn dialogue after the target single-turn dialogue.
[0017] The basic attention mask and the supplementary attention mask are superimposed to obtain the attention mask matrix.
[0018] In one of the embodiments, the supplementary attention mask is constructed according to the training text sequence, and comprises:
[0019] In the training text sequence, the sequence positions of the model reply text and the corresponding thinking chain data of each single-turn dialogue are identified to obtain the end position of the model reply text of each single-turn dialogue and the start position and the end position of the thinking chain data of each single-turn dialogue.
[0020] According to the end position of the model reply text of each single-turn dialogue and the start position and the end position of the thinking chain data of each single-turn dialogue, the supplementary attention mask is constructed.
[0021] In one of the embodiments, the method further comprises:
[0022] In the attention mask matrix, the target positions are marked in an attention dependence shielding manner by using a preset marking manner; the target positions include the thinking chain data of the any single-turn dialogue, and the positions between the user interaction text and the model reply text of the single-turn dialogue after the any single-turn dialogue and the thinking chain data thereof.
[0023] In a second aspect, the present application also provides a training device of a multi-turn dialogue model based on thinking chain, the device comprising:
[0024] A multi-turn dialogue sequence construction module is configured to acquire multi-turn dialogue sample data and construct a training text sequence containing continuous multiple single-turn dialogues; each single-turn dialogue in the training text sequence includes user interaction text and model reply text of the current turn, and the model reply text has corresponding thinking chain data;
[0025] An attention mask matrix construction module is configured to construct an attention mask matrix according to the user interaction text and the model reply text of each single-turn dialogue in the training text sequence and the corresponding thinking chain data thereof; the model reply text of any single-turn dialogue in the attention mask matrix has no attention dependence relationship with the thinking chain data of the single-turn dialogue before the any single-turn dialogue;
[0026] A model training module is configured to perform model training on the multi-turn dialogue model based on the attention mask matrix to obtain a trained multi-turn dialogue model; the model input format of the trained multi-turn dialogue model during inference matches the model input format of the multi-turn dialogue model during training.
[0027] In a third aspect, the present application also provides a computer device comprising a memory and a processor, the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0028] In a fourth aspect, the present application also provides a computer readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the steps of the above method.
[0029] In a fifth aspect, the present application also provides a computer program product comprising a computer program, and the computer program is executed by a processor to implement the steps of the above method.
[0030] The training method, device, computer device, computer readable storage medium, and computer program product of the multi-turn dialogue model based on the thought chain, by obtaining multi-turn dialogue sample data, construct a training text sequence containing continuous multiple single-turn dialogues, each single-turn dialogue in the training text sequence includes user interaction text and model reply text of the current turn, and the model reply text has corresponding thought chain data, then according to the user interaction text and the model reply text of each single-turn dialogue in the training text sequence and the corresponding thought chain data, a attention mask matrix is constructed, the model reply text of any single-turn dialogue in the attention mask matrix does not have an attention dependency relationship with the thought chain data of the single-turn dialogue before any single-turn dialogue, and then the multi-turn dialogue model is trained based on the attention mask matrix, to obtain the trained multi-turn dialogue model, the model input format of the trained multi-turn dialogue model during reasoning matches the model input format of the multi-turn dialogue model during training, the training optimization of the multi-turn dialogue model with the thought chain is realized, the training efficiency can be improved by training the multi-turn dialogue as a single sample and adding special masks, the dependency relationship during model training is consistent with online reasoning, the training efficiency and model effect are considered, and the model can maintain the reply quality advantage brought by the thought chain while enhancing the coherence of the multi-turn dialogue, and ensure stable online reasoning effect. BRIEF DESCRIPTION OF DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other related drawings without creative labor based on these drawings.
[0032] Figure 1 A flowchart of a training method of a multi-turn dialogue model based on a thought chain in an embodiment;
[0033] Figure 2 A schematic diagram of model training connection relationship in an embodiment;
[0034] Figure 3 A schematic diagram of dependency relationship between multi-turn dialogues in an embodiment;
[0035] Figure 4 A flowchart of a training method of a multi-turn dialogue model based on a thought chain in another embodiment;
[0036] Figure 5 A structural block diagram of a training device of a multi-turn dialogue model based on a thought chain in an embodiment;
[0037] Figure 6 Fig. 1 is a schematic diagram of an internal structure of a computer device according to an embodiment. DETAILED DESCRIPTION
[0038] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not intended to limit the present application.
[0039] In an exemplary embodiment, as shown in Figure 1 A method for training a multi-turn dialogue model based on a thought chain is provided. In this embodiment, the method is applied to a terminal. It should be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and can be implemented through the interaction of the terminal and the server. In this embodiment, the method includes the following steps 101 to 103. Wherein:
[0040] Step 101, obtaining multi-turn dialogue sample data, and constructing a training text sequence containing continuous multiple single-turn dialogues.
[0041] As an example, each single-turn dialogue in the training text sequence can include a user interaction text of the current turn and a model reply text, and the model reply text has corresponding thought chain data.
[0042] In actual application, multi-turn dialogue original data with a thought chain can be obtained as multi-turn dialogue sample data. By analyzing the multi-turn dialogue sample data, a training text sequence containing continuous multiple single-turn dialogues can be constructed.
[0043] Exemplarily, for an interactive dialogue scene based on a large language model, taking a role-playing type in social chatting as an example (also applicable to an interactive dialogue with system prompts or without system settings), the following multi-turn dialogue original data with a thought chain can be obtained:
[0044] Dialogue(u, g, n) = system + history(u, g, n-1) + interact(u, g~, n);
[0045] Wherein, u (user) is a user side message, g is an intelligent assistant side message based on a large language model, and g~ is an intelligent assistant side message with a thought chain (cot); the above data can represent n-turn dialogue: composed of system setting + previous n-1 turn dialogue history + n-turn dialogue user and gpt~ (with cot) interaction.
[0046] history(u, g, n-1) = interact(u, g, 1) + interact(u, g, 2) +... + interact(u, g, n-1);
[0047] The above data can represent the history of the previous n-1 rounds of dialogue: composed of the interactions of the n-1 rounds of dialogue user and gpt (without cot).
[0048] interact(u, g~, n) = message(u, n) + message(g~, n);
[0049] The above data can represent the interaction of the n-th round of dialogue user and gpt~ (with cot): composed of the n-th round of dialogue user message + n-th round of dialogue intelligent assistant side (with cot) message.
[0050] Optionally, by parsing the data structure of each round of dialogue, the start and end range of each part of the message can be determined to obtain structured multi-round dialogue data, such as marking the positions of user messages, thought chains, and intelligent assistant messages in the training text sequence to provide a basis for subsequent mask processing; then multiple single-round dialogues can be integrated into one training sample (i.e., a training text sequence), and the sample structure can be:
[0051] system + history(u, g, n-1) + interact(u, g~, n)
[0052] = system + message(u, 1) + ~message(g~, 1) + message(u, 2) + ~message(g~, 2) +... + message(u, n-1) + ~message(g~, n-1) + message(u, n) + ~message(g~, n);
[0053] By constructing an attention mask matrix to shield the attention dependency relationship, n~message(g~, k) can be trained based on the training sample: the intelligent assistant side message of the k-th round of dialogue can know the thought chain of the current round of dialogue, but cannot know the thought chain of the previous k-1 rounds of dialogue, so as to conform to the model online reasoning format. Thus, it is not necessary to disassemble into single-round dialogue samples, the integrity of the multi-round dialogue can be maintained, and the training efficiency can be improved.
[0054] In an example, the model online reasoning format is:
[0055] system + history(u, g, n-1) + interact(u, g~, n)
[0056] = system + interact(u, g, 1) + interact(u, g, 2) +... + interact(u, g, n-1) + interact(u, g, n)
[0057] = system + message(u, 1) + message(g, 1) + message(u, 2) +... + message(u, n-1) + message(g, n-1) + message(u, n) + message(g, n);
[0058] Wherein, message(g, n) is the content learned and generated by the model according to the existing information (system setting, historical dialogue and the nth round of user side message); Since the historical dialogue does not contain the thought chain content, the length of the input prompt word can be greatly reduced, and the reasoning speed of the model can be improved. The traditional single round training method needs to disassemble a single multi-round dialogue into n training samples to adapt to the online reasoning format, so that the model needs to learn n rounds of dialogue content respectively, which is low in training efficiency; Or the traditional multi-round training method, although n message(g, k) are trained in one sample, but the nth round of dialogue intelligent assistant side message can know the thought chain of the previous n-1 rounds of dialogue intelligent assistant side message, which does not conform to the online reasoning format, and affects the training effect.
[0059] Step 102, constructing an attention mask matrix according to the user interaction text and the model reply text of each single round dialogue in the training text sequence and the corresponding thought chain data thereof.
[0060] Wherein, the model reply text of any single round dialogue in the attention mask matrix does not have an attention dependency relationship with the thought chain data of the single round dialogue before it.
[0061] In specific implementation, the training text sequence can be tokenized to generate a basic attention mask, and the information can be shielded according to the target dependency relationship, and a supplementary attention mask can be additionally determined, and then the basic attention mask and the supplementary attention mask can be superimposed to construct the attention mask matrix. Based on the attention mask matrix, the token mark in the subsequent dialogue in the multi-round dialogue can have no dependency relationship with the cot thought chain in the previous dialogue, and does not participate in the weight calculation of the token mark in the subsequent dialogue, so that the attention dependency relationship can be automatically shielded.
[0062] Step 103, model training the multi-round dialogue model based on the attention mask matrix to obtain the trained multi-round dialogue model.
[0063] As an example, the multi-turn dialogue model can be a large language model (LLM) model, which is a natural language processing model based on deep learning technology, which can generate natural and fluent text through learning and understanding of massive text data.
[0064] The model input format of the trained multi-turn dialogue model during inference (such as the online inference format) matches the model input format of the multi-turn dialogue model during training. Thus, by keeping consistent with the online inference format, i.e., ensuring that the model input distribution during inference is close to the input distribution during training, the model can stably exert application effects.
[0065] In the training method of the multi-turn dialogue model based on the thought chain described above, the multi-turn dialogue sample data is obtained to construct a training text sequence containing a plurality of consecutive single-turn dialogues. Then, an attention mask matrix is constructed according to the user interaction text and the model reply text of each single-turn dialogue in the training text sequence and the corresponding thought chain data, and the multi-turn dialogue model is trained based on the attention mask matrix to obtain a trained multi-turn dialogue model, which realizes the training and optimization of the multi-turn dialogue model with the thought chain. By training the multi-turn dialogue as a single sample and adding special masks, the training efficiency can be improved, the dependency relationship during model training can be ensured to be consistent with online inference, the training efficiency and model effect can be considered, and the model can maintain the reply quality advantage brought by the thought chain while enhancing the coherence of the multi-turn dialogue, ensuring stable online inference effects.
[0066] In an exemplary embodiment, the following steps can also be included:
[0067] The training text sequence is subjected to model input preprocessing to obtain a text input token sequence and a training label sequence; the training label sequence is used to supervise the model reply text of each single-turn dialogue and the corresponding thought chain data.
[0068] In an example, input ids (i.e., a text input token sequence) and labels (i.e., a training label sequence) can be generated through preprocessing of model input data; for example, the training text sequence can be converted into a token sequence recognizable by the model to obtain input ids, and labels can be generated as a training target, which can be used to supervise the intelligent assistant side output message (i.e., the model reply text and the corresponding thought chain data). The labels of the user side message and the labels of the system setting message can also be configured not to participate in loss calculation.
[0069] In this embodiment, by performing model input preprocessing on the training text sequence, the text input token sequence and the training label sequence are obtained, which can maintain the structure correspondence between the text input and the label sequence, and provide a coherent basis for subsequent attention mask processing.
[0070] In an exemplary embodiment, the model training of the multi-turn dialogue model based on the attention mask matrix can include the following steps:
[0071] The text input token sequence, the training label sequence, and the attention mask matrix are input into the multi-turn dialogue model; and the multi-turn dialogue model is trained by calculating attention weights to obtain the trained multi-turn dialogue model.
[0072] Exemplarily, the input ids text input token sequence, the labels training label sequence, the attention mask basic attention mask, and the patch attention mask supplementary attention mask calculated additionally can be superimposed on the attention scores score for calculation during training, so that the model can learn content generation based on a single sample according to the strategy of automatically shielding the thought chain in the previous n-1 turns of dialogue in multi-turn dialogue training, and then the trained multi-turn dialogue model can be consistent with the online inference format.
[0073] In this embodiment, by inputting the text input token sequence, the training label sequence, and the attention mask matrix into the multi-turn dialogue model, and then training the multi-turn dialogue model by calculating attention weights, the trained multi-turn dialogue model is obtained, which can realize accurate learning of multi-turn dialogue dependency relationship based on attention weight calculation, so that the model can simultaneously master multi-turn reply logic in a single sample, and the training efficiency is improved.
[0074] In an exemplary embodiment, the attention mask matrix can be constructed according to the user interaction text and model reply text of each single-turn dialogue in the training text sequence and the corresponding thought chain data, which can include the following steps:
[0075] The basic attention mask is generated according to the tokenization processing of the training text sequence; the supplementary attention mask is constructed according to the training text sequence according to the target dependency relationship shielding information; the target dependency relationship shielding information is used to indicate that the thought chain data of the target single-turn dialogue in the training text sequence is shielded, and the dependency relationship between the user interaction text, the model reply text, and the thought chain data of the single-turn dialogue after the target single-turn dialogue; and the attention mask matrix is obtained by superimposing the basic attention mask and the supplementary attention mask.
[0076] In practical applications, by preprocessing the model input data, an attention mask (i.e., a basic attention mask) can be generated. This mask can be used to distinguish between real tokens and padding tokens, and to shield the influence of padding tokens on attention calculation. Additionally, a patch attention mask (i.e., a supplementary attention mask) can be calculated. The attention mask and the patch attention mask can then be superimposed to obtain an attention mask matrix.
[0077] For example, such as Figure 2 As shown, based on the constructed attention mask matrix, the connection relationships between multi-turn dialogue data can be controlled during LLM model training; where, Figure 2 U_MSG1 is the user-side message (i.e., user interaction text) in the first round of dialogue, G_COT1 is the cot1 (i.e., thought chain data) of the intelligent assistant in the first round of dialogue, and G_ANS1 is the answer1 (i.e., model response text) of the intelligent assistant in the first round of dialogue; the connection between G_COT1 in the first round of dialogue and G_COT2 and G_ANS2 in the second round of dialogue has been masked.
[0078] For example, such as Figure 3 As shown, during model training, the model can learn to generate G_COT1, G_ANS1, G_COT2, and G_ANS2. Among them, G_ANS1 is dependent on G_COT1, and G_ANS2 is dependent on G_COT2. That is, there is a dependency between the model's response text in the current dialogue and the thought chain data in the current dialogue. By blocking the connection, G_COT2 and G_ANS2 cannot know G_COT1.
[0079] In this embodiment, a basic attention mask is generated by tokenizing the training text sequence, information is masked according to the target dependency relationship, and a supplementary attention mask is constructed based on the training text sequence. Then, by superimposing the basic attention mask and the supplementary attention mask, an attention mask matrix is obtained, which can achieve precise masking of historical thought chains. The superimposed attention mask matrix can not only retain the masking function, but also meet the special dependency requirements of multi-turn dialogue training, effectively improving the model training effect.
[0080] In an exemplary embodiment, constructing a supplementary attention mask based on the training text sequence may include the following steps:
[0081] In the training text sequence, a sequence position of the model reply text of each single-turn dialogue and corresponding thought chain data is identified, an end position of the model reply text of each single-turn dialogue is obtained, and a start position and an end position of the thought chain data of each single-turn dialogue are obtained; and the supplementary attention mask is constructed according to the end position of the model reply text of each single-turn dialogue and the start position and the end position of the thought chain data of each single-turn dialogue.
[0082] In a specific implementation, the supplementary attention mask can be constructed by identifying the start and end positions of the thought chain and the end position of the reply message in each round of dialogue, and the supplementary attention mask can be used to mask the attention dependence of the intelligent assistant side reply in the nth round of dialogue on the thought chain in the previous n-1 rounds of dialogue.
[0083] In this embodiment, by identifying the sequence position of the model reply text of each single-turn dialogue and corresponding thought chain data in the training text sequence, obtaining the end position of the model reply text of each single-turn dialogue, and obtaining the start position and the end position of the thought chain data of each single-turn dialogue, and then constructing the supplementary attention mask according to the end position of the model reply text of each single-turn dialogue and the start position and the end position of the thought chain data of each single-turn dialogue, the sequence positions of each round of reply and thought chain can be accurately positioned, which provides an accurate basis for constructing the supplementary attention mask, and ensures that the mask can accurately mask the interference of the historical thought chain on the subsequent rounds.
[0084] In one exemplary embodiment, the following steps can also be included:
[0085] In the attention mask matrix, a target position is marked for attention dependence masking in a preset marking manner; the target position includes the thought chain data of the any single-turn dialogue, and a position between the user interaction text, the model reply text and the thought chain data thereof of the single-turn dialogue after the any single-turn dialogue.
[0086] In an example, a special attention matrix, i.e., an attention mask matrix, can be constructed. The attention matrix is a lower triangular matrix, which is based on Figure 2 and Figure 3 As shown in the above, since G_COT1 is not known for tokens after G_ANS1, a preset marking manner (such as a 0 mark with color) can be used in a special position in the matrix. An example of the attention mask matrix is as follows:
[0087]
[0088] Wherein, G_COT1 and U_MSG2, G_COT1 and G_COT2, G_COT1 and G_ANS2 can be based on 0 marks with color.
[0089] In yet another example, the attention calculation based on the attention mask matrix can be performed as follows:
[0090]
[0091] Wherein, X is the input matrix, representing the embedding vector of n tokens in the sequence; is the learnable weight matrix of the i-th attention head, used to generate Query, Key, Value; Q i , K i , V i is the Query, Key, Value matrix of the i-th attention head, obtained by linear transformation; A i is the attention score matrix, which can calculate the dot product of Query and Key and scale; is the scaling factor, which can prevent the dot product value from being too large to cause unstable gradient.
[0092]
[0093] Wherein, H i is the output of the i-th attention head, obtained by Softmax and weighted summation of Value, Softmax is normalized along each row (for each Query) to obtain the attention weight; O is the final output of multi-head attention, which is linearly projected W O after splicing the outputs of all attention heads.
[0094]
[0095] Wherein, is the mask operation on the attention score matrix A; is the additional attention mask, which can realize that the thought chain of the previous n-1 rounds of dialogue does not contribute to the decoding of the current round of dialogue when decoding the n-th round of dialogue.
[0096] In this embodiment, by using the preset marking method to mark the target position for attention dependence shielding in the attention mask matrix, the precise position shielding mark can be used to avoid the interference of the historical thought chain on the subsequent round of learning, ensuring that the attention distribution during training is consistent with the reasoning scenario, which helps to improve the continuity and stability of multi-round dialogue training.
[0097] In one exemplary embodiment, as Figure 4As shown in FIG. 6, another flowchart of a training method of a multi-turn dialogue model based on a thought chain is provided. In this embodiment, the method comprises the following steps:
[0098] In step 401, multi-turn dialogue sample data is obtained, and a training text sequence containing continuous multiple single-turn dialogues is constructed. In step 402, model input preprocessing is performed on the training text sequence to obtain a text input token sequence and a training label sequence. In step 403, tokenization processing is performed according to the training text sequence to generate a basic attention mask, and information is shielded according to a target dependency relationship. A supplementary attention mask is constructed according to the training text sequence. In step 404, the basic attention mask and the supplementary attention mask are superimposed to obtain an attention mask matrix. In step 405, the text input token sequence, the training label sequence, and the attention mask matrix are input into the multi-turn dialogue model. In step 406, the multi-turn dialogue model is trained by calculating attention weights to obtain a trained multi-turn dialogue model.
[0099] It should be noted that the specific limitations of the above steps can refer to the specific limitations of the training method of the multi-turn dialogue model based on the thought chain described above, and will not be described here.
[0100] It should be understood that although each step in the flowchart involved in each of the above embodiments is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each of the above embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.
[0101] Based on the same inventive concept, the present embodiment also provides a training device for the multi-turn dialogue model based on the thought chain involved in the above-mentioned training method of the multi-turn dialogue model based on the thought chain. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, and therefore the specific limitations in one or more training device embodiments for the multi-turn dialogue model based on the thought chain provided below can refer to the limitations of the training method of the multi-turn dialogue model based on the thought chain described above, and will not be described here.
[0102] In one exemplary embodiment, as Figure 5 shown in FIG. 6, a training device for a multi-turn dialogue model based on a thought chain is provided, comprising:
[0103] The multi-turn dialogue sequence construction module 501 is configured to acquire multi-turn dialogue sample data and construct a training text sequence containing continuous multiple single-turn dialogues; each single-turn dialogue in the training text sequence includes user interaction text of the current turn and model reply text, and the model reply text has corresponding thought chain data;
[0104] The attention mask matrix construction module 502 is configured to construct an attention mask matrix according to the user interaction text and the model reply text of each single-turn dialogue in the training text sequence and the corresponding thought chain data; the model reply text of any single-turn dialogue in the attention mask matrix does not have an attention dependency relationship with the thought chain data of the single-turn dialogue before the any single-turn dialogue;
[0105] The model training module 503 is configured to perform model training on the multi-turn dialogue model based on the attention mask matrix to obtain a trained multi-turn dialogue model; the model input format of the trained multi-turn dialogue model during inference matches the model input format of the multi-turn dialogue model during training.
[0106] In one of the embodiments, the device further includes:
[0107] The data preprocessing module is configured to perform model input preprocessing on the training text sequence to obtain a text input token sequence and a training label sequence; the training label sequence is used to supervise the model reply text of each single-turn dialogue and the corresponding thought chain data.
[0108] In one of the embodiments, the model training module 503 is specifically configured to input the text input token sequence, the training label sequence, and the attention mask matrix into the multi-turn dialogue model; and perform model training on the multi-turn dialogue model by calculating attention weights to obtain the trained multi-turn dialogue model.
[0109] In one of the embodiments, the attention mask matrix construction module 502 is specifically configured to perform tokenization processing on the training text sequence to generate a basic attention mask; construct a supplementary attention mask according to the training text sequence and target dependency relationship shielding information; the target dependency relationship shielding information is used to indicate shielding of dependency relationships between thought chain data of a target single-turn dialogue in the training text sequence and user interaction text, model reply text, and thought chain data of single-turn dialogues after the target single-turn dialogue; and obtain the attention mask matrix by superimposing the basic attention mask and the supplementary attention mask.
[0110] In one of the embodiments, the attention mask matrix construction module 502 is further configured to identify, in the training text sequence, a sequence position of the model reply text of each single-turn dialogue and corresponding thought chain data thereof, obtain an ending position of the model reply text of each single-turn dialogue, and a starting position and a terminal position of the thought chain data of each single-turn dialogue; and construct the supplementary attention mask according to the ending position of the model reply text of each single-turn dialogue, and the starting position and the terminal position of the thought chain data of each single-turn dialogue.
[0111] In one of the embodiments, the device further comprises:
[0112] The shielding marking module is configured to mark the target positions in the attention mask matrix in a preset marking manner in an attention-dependent manner; the target positions include the thought chain data of the any single-turn dialogue, and positions between the user interaction text, the model reply text, and the thought chain data of the single-turn dialogue after the any single-turn dialogue.
[0113] The above-mentioned various modules in the training device of the multi-turn dialogue model based on thought chains can be realized by software, hardware, and combinations thereof, in whole or in part. The above-mentioned various modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above-mentioned various modules.
[0114] In one exemplary embodiment, a computer device, which can be a terminal, is provided, and an internal structure diagram of the computer device can be as shown in Figure 6 The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to perform wired or wireless communication with external terminals, and the wireless communication can be achieved through WIFI, mobile cellular network, near field communication (NFC), or other technologies. The computer program is executed by the processor to implement a training method of a multi-turn dialogue model based on thought chains.
[0115] Those skilled in the art can understand that Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0116] In an exemplary embodiment, a computer device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program:
[0117] Obtaining multi-turn dialogue sample data to construct a training text sequence containing continuous multiple single-turn dialogues; each single-turn dialogue in the training text sequence includes user interaction text and model reply text of the current turn, and the model reply text has corresponding thought chain data;
[0118] According to the user interaction text and the model reply text of each single-turn dialogue in the training text sequence and the corresponding thought chain data, an attention mask matrix is constructed; there is no attention dependency relationship between the model reply text of any single-turn dialogue in the attention mask matrix and the thought chain data of the single-turn dialogue before the any single-turn dialogue;
[0119] Based on the attention mask matrix, the multi-turn dialogue model is trained to obtain a trained multi-turn dialogue model; the model input format of the trained multi-turn dialogue model during inference matches the model input format of the multi-turn dialogue model during training.
[0120] In an embodiment, the processor further implements the steps of the training method of the thought chain-based multi-turn dialogue model in the other embodiments described above when executing the computer program.
[0121] In an embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0122] Obtaining multi-turn dialogue sample data to construct a training text sequence containing continuous multiple single-turn dialogues; each single-turn dialogue in the training text sequence includes user interaction text and model reply text of the current turn, and the model reply text has corresponding thought chain data;
[0123] According to the user interaction text and the model reply text of each single-turn dialogue in the training text sequence and the corresponding thought chain data, an attention mask matrix is constructed; there is no attention dependency relationship between the model reply text of any single-turn dialogue in the attention mask matrix and the thought chain data of the single-turn dialogue before the any single-turn dialogue;
[0124] perform model training on the multi-turn dialogue model based on the attention mask matrix to obtain a trained multi-turn dialogue model; the model input format of the trained multi-turn dialogue model during inference matches the model input format of the multi-turn dialogue model during training.
[0125] In one embodiment, the computer program, when executed by a processor, also implements the steps of the training method of the multi-turn dialogue model based on the thought chain in the other embodiments described above.
[0126] In one embodiment, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the following steps:
[0127] obtain multi-turn dialogue sample data to construct a training text sequence containing a plurality of consecutive single-turn dialogues; each single-turn dialogue in the training text sequence includes user interaction text and model reply text of the current turn, and the model reply text has corresponding thought chain data;
[0128] construct an attention mask matrix according to the user interaction text and model reply text of each single-turn dialogue in the training text sequence and the corresponding thought chain data; in the attention mask matrix, the model reply text of any single-turn dialogue does not have an attention dependency relationship with the thought chain data of the single-turn dialogue before the any single-turn dialogue;
[0129] perform model training on the multi-turn dialogue model based on the attention mask matrix to obtain a trained multi-turn dialogue model; the model input format of the trained multi-turn dialogue model during inference matches the model input format of the multi-turn dialogue model during training.
[0130] In one embodiment, the computer program, when executed by a processor, also implements the steps of the training method of the multi-turn dialogue model based on the thought chain in the other embodiments described above.
[0131] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0132] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.
[0133] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.
[0134] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.
Claims
1. A method for training a multi-turn dialogue model based on a thought chain, characterized in that, The method comprises: obtaining multi-turn dialogue sample data, and constructing a training text sequence containing continuous multiple single-turn dialogues; each single-turn dialogue in the training text sequence comprises user interaction text and model reply text of the current turn, and the model reply text has corresponding thought chain data; constructing an attention mask matrix according to the user interaction text and the model reply text of each single-turn dialogue in the training text sequence and the corresponding thought chain data; in the attention mask matrix, the model reply text of any single-turn dialogue has no attention dependency relationship with the thought chain data of the single-turn dialogue before the any single-turn dialogue; based on the attention mask matrix, model training is performed on the multi-turn dialogue model to obtain a trained multi-turn dialogue model; the model input format of the trained multi-turn dialogue model during inference matches the model input format of the multi-turn dialogue model during training.
2. The method of claim 1, wherein, The method further comprises: performing model input preprocessing on the training text sequence to obtain a text input token sequence and a training label sequence; the training label sequence is used to supervise the model reply text of each single-turn dialogue and the corresponding thought chain data.
3. The method of claim 2, wherein, The model training of the multi-turn dialogue model based on the attention mask matrix to obtain the trained multi-turn dialogue model comprises: inputting the text input token sequence, the training label sequence, and the attention mask matrix into the multi-turn dialogue model; performing model training on the multi-turn dialogue model by calculating attention weights to obtain the trained multi-turn dialogue model.
4. The method of claim 1, wherein, The construction of the attention mask matrix according to the user interaction text and the model reply text of each single-turn dialogue in the training text sequence and the corresponding thought chain data comprises: performing tokenization processing on the training text sequence to generate a basic attention mask; according to the training text sequence, constructing a supplementary attention mask by shielding information according to a target dependency relationship; the target dependency relationship shielding information is used to indicate that the dependency relationship between the thought chain data of a target single-turn dialogue in the training text sequence and the user interaction text, the model reply text and the thought chain data of the single-turn dialogue after the target single-turn dialogue is shielded; by superimposing the basic attention mask and the supplementary attention mask, the attention mask matrix is obtained.
5. The method of claim 4, wherein, The construction of the supplementary attention mask according to the training text sequence comprises: in the training text sequence, identifying the sequence positions of the model reply text and the corresponding thought chain data of each single-turn dialogue, obtaining the end position of the model reply text of each single-turn dialogue, and the start position and the end position of the thought chain data of each single-turn dialogue; according to the end position of the model reply text of each single-turn dialogue and the start position and the end position of the thought chain data of each single-turn dialogue, the supplementary attention mask is constructed.
6. The method of claim 1, wherein, The method further comprises: In the attention mask matrix, a preset marking manner is used to mark the target positions for attention dependence shielding; the target positions include the thinking chain data of the any single-turn dialogue, and the positions between the user interaction text and the model reply text of the single-turn dialogue after the any single-turn dialogue and the thinking chain data thereof. 7.A device for training a multi-turn dialogue model based on a chain-of-thought, characterized in that, The device comprises: a multi-turn dialogue sequence construction module, configured to acquire multi-turn dialogue sample data and construct a training text sequence containing continuous multiple single-turn dialogues; each single-turn dialogue in the training text sequence comprises user interaction text and model reply text of the current turn, and the model reply text has corresponding thinking chain data; an attention mask matrix construction module, configured to construct an attention mask matrix according to the user interaction text and the model reply text of each single-turn dialogue in the training text sequence and the corresponding thinking chain data thereof; the model reply text of any single-turn dialogue in the attention mask matrix has no attention dependence relationship with the thinking chain data of the single-turn dialogue before the any single-turn dialogue; a model training module, configured to perform model training on the multi-turn dialogue model based on the attention mask matrix to obtain a trained multi-turn dialogue model; the model input format of the trained multi-turn dialogue model during inference matches the model input format of the multi-turn dialogue model during training. 8.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-7. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.