Method for generating multi-round dialogue corpora and training and testing large language model

By constructing a vector database and using a large language model to generate multiple rounds of dialogue corpus, the problem of limited number of multi-round dialogue corpus in the existing technology is solved, and high-quality multi-round dialogue corpus generation is achieved, meeting the training and evaluation needs of large models.

CN120216640APending Publication Date: 2025-06-27ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510285984.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The number of multi-round dialogue corpus obtained in the prior art is limited and the scenarios are limited, making it difficult to meet the needs of large-scale model training and evaluation.

Method used

By constructing a vector database, using the above vector generated by the question-and-answer conversation sequence of real users online, vector recall is performed, problem-generating constraint information is generated, and multiple rounds of dialogue corpus are generated using a large language model.

Benefits of technology

The generation of a large number of multi-round dialogue corpus is achieved, which improves the consistency between the second question information generated by the large language model and the dialogue sequence where the first question information is located, improves the quality of the multi-round dialogue corpus generated, and meets the training and evaluation needs of the large model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216640A_ABST
    Figure CN120216640A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device and equipment for generating a multi-round dialogue corpus and training and testing a large language model. The method for generating the multi-round dialogue corpus comprises the following steps: acquiring first question information; vectorizing the first question information to obtain a first question vector; querying a target preceding text vector of which the similarity with the first question vector meets a preset similarity condition from a vector database in which a plurality of groups of data pairs are stored; any group of data pair comprises a preceding text vector generated according to a preceding text of a question of the user in a historical interaction process with the language model, and a following text of the question in the historical interaction process; obtaining question generation constraint information based on the question following text corresponding to the target preceding text vector; calling the large language model to generate second question information after the first question information by taking the question generation constraint information as a constraint condition; and generating a multi-round dialogue corpus based on the first question information and the second question information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method for generating multi-turn dialogue corpus, a method for training a large language model, and a method for testing a large language model. This application also relates to a device for generating multi-turn dialogue corpus, a device for training a large language model, a device for testing a large language model, a device for generating multi-turn dialogue corpus, a device for training a large language model, and a device for testing a large language model. Background Art

[0002] The multi-turn dialogue ability is an important ability of large models. In the process of training or evaluating large language models, multi-turn dialogue corpus is usually required. However, in reality, the number of multi-turn dialogue corpus obtained is limited and the scenarios are limited, making it difficult to meet the needs of large model training and evaluation. Based on this, how to generate multi-turn dialogue corpus has become an urgent technical problem to be solved. Summary of the Invention

[0003] In view of this, embodiments of this application provide a method, device, and equipment for generating multi-turn dialogue corpus to solve the problem that the number of multi-turn dialogue corpus obtained in reality is limited and the scenarios are limited, making it difficult to meet the needs of large model training and evaluation.

[0004] According to the first aspect of the embodiments of this application, a method for generating multi-turn dialogue corpus is provided, including:

[0005] Obtain the first question information;

[0006] Perform vectorization processing on the first question information to obtain a first question vector;

[0007] Query from the vector database a target previous context vector whose similarity with the first question vector meets a preset similarity condition; multiple data pairs are stored in the vector database; for any one of the data pairs, it includes a previous context vector generated according to the previous context of the user's questions in the historical interaction process with the language model, and the subsequent question in the historical interaction process;

[0008] Based on the subsequent question corresponding to the target previous context vector, obtain question generation constraint information;

[0009] Call the large language model to generate a second question information after the first question information with the question generation constraint information as a constraint condition;

[0010] Based on the first question information and the second question information, generate multi-turn dialogue corpus.

[0011] According to the second aspect of the embodiments of this application, a method for training a large language model is provided, including:

[0012] Use the multi-turn dialogue corpus generated by the above method for generating a multi-turn dialogue corpus to train a large language model to be trained.

[0013] According to the third aspect of the embodiments of the present application, there is provided a method for testing a large language model, including:

[0014] Use the multi-turn dialogue corpus generated by the above method for generating a multi-turn dialogue corpus to test the trained large language model.

[0015] According to the fourth aspect of the embodiments of the present application, there is provided a device for generating a multi-turn dialogue corpus, including:

[0016] A question information acquisition module, configured to acquire first question information;

[0017] A vectorization processing module, configured to perform vectorization processing on the first question information to obtain a first question vector;

[0018] A vector recall module, configured to query a target previous context vector in a vector database whose similarity to the first question vector meets a preset similarity condition; multiple data pairs are stored in the vector database; for any one of the data pairs, it includes a previous context vector generated according to the previous context of the user's questions in the historical interaction process with the language model, and the subsequent question in the historical interaction process;

[0019] A constraint information generation module, configured to obtain question generation constraint information based on the subsequent question corresponding to the target previous context vector;

[0020] A question information generation module, configured to call a large language model to generate a second question information after the first question information with the question generation constraint information as a constraint condition;

[0021] A dialogue corpus generation module, configured to generate a multi-turn dialogue corpus based on the first question information and the second question information.

[0022] According to the fifth aspect of the embodiments of the present application, there is provided a device for training a large language model, including:

[0023] A training corpus acquisition module, configured to acquire the multi-turn dialogue corpus generated by the above method for generating a multi-turn dialogue corpus;

[0024] A model training module, configured to use the acquired multi-turn dialogue corpus to train a large language model to be trained.

[0025] According to the sixth aspect of the embodiments of the present application, there is provided a device for testing a large language model, including:

[0026] A test corpus acquisition module for acquiring the multi-turn dialogue corpus generated by the method for generating a multi-turn dialogue corpus described above;

[0027] A model testing module for testing the trained large language model by using the acquired multi-turn dialogue corpus.

[0028] According to the seventh aspect of the embodiments of the present application, there is provided a device for generating a multi-turn dialogue corpus, including:

[0029] At least one processor; and,

[0030] A memory communicatively connected to the at least one processor; wherein,

[0031] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to:

[0032] Obtain first question information;

[0033] Perform vectorization processing on the first question information to obtain a first question vector;

[0034] Query a target previous context vector in a vector database whose similarity to the first question vector meets a preset similarity condition; multiple data pairs are stored in the vector database; for any one of the data pairs, it includes a previous context vector generated according to the previous context of the user's questions in the historical interaction process with the language model, and the subsequent question in the historical interaction process;

[0035] Based on the subsequent question corresponding to the target previous context vector, obtain question generation constraint information;

[0036] Call the large language model to generate a second question information after the first question information by using the question generation constraint information as a constraint condition;

[0037] Based on the first question information and the second question information, generate a multi-turn dialogue corpus.

[0038] According to the eighth aspect of the embodiments of the present application, there is provided a device for training a large language model, including:

[0039] At least one processor; and,

[0040] A memory communicatively connected to the at least one processor; wherein,

[0041] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to:

[0042] The multi-turn dialogue corpus generated by using the method for generating a multi-turn dialogue corpus described above is used to train a large language model to be trained.

[0043] According to a ninth aspect of the embodiments of the present application, there is provided a device for testing a large language model, including:

[0044] At least one processor; and,

[0045] A memory communicatively connected to the at least one processor; wherein,

[0046] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can:

[0047] Use the multi-turn dialogue corpus generated by using the method for generating a multi-turn dialogue corpus described above to test the trained large language model.

[0048] At least one embodiment of this specification can achieve at least the following beneficial effects: By using the data pair including the above text vector and the question following text generated from the Q&A dialogue sequence of real online users to construct a vector database, when it is necessary to construct or expand the dialogue corpus, based on the obtained first question information, the question following text corresponding to the question above text with a high similarity to the first question information can be recalled from the vector database by means of vector recall, and the recalled question following text is used as a constraint condition to make the large language model generate the second question information after the first question information, and a large amount of multi-turn dialogue corpus can be obtained.

[0049] In addition, using real online dialogue data as the basis and constraint for the output of the large language model can also improve the coherence of the second question information generated by the large language model with the dialogue sequence where the first question information is located, improve the quality of the generated multi-turn dialogue corpus, and meet the training and evaluation of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 is a schematic diagram of an application scenario of a method for generating a multi-turn dialogue corpus provided by an embodiment of this specification;

[0052] Figure 2 is a schematic flowchart of a method for generating a multi-turn dialogue corpus provided by an embodiment of this specification;

[0053] Figure 3 It is a swimlane diagram of a method for generating multi-turn dialogue corpus provided by an embodiment of this specification;

[0054] Figure 4 It is a schematic structural diagram of a device for generating multi-turn dialogue corpus provided by an embodiment of this specification;

[0055] Figure 5 It is a schematic structural diagram of a device for training a large language model provided by an embodiment of this specification;

[0056] Figure 6 It is a schematic structural diagram of a device for training a large language model provided by an embodiment of this specification;

[0057] Figure 7 It is a schematic structural diagram of a device for generating multi-turn dialogue corpus provided by an embodiment of this specification;

[0058] Figure 8 It is a schematic structural diagram of a device for training a large language model provided by an embodiment of this specification;

[0059] Figure 9 It is a schematic structural diagram of a device for generating multi-turn dialogue corpus provided by an embodiment of this specification. Detailed implementation manners

[0060] Many specific details are set forth in the following description in order to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this application. Therefore, this application is not limited by the specific implementations disclosed below.

[0061] The terms used in one or more embodiments of this application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this application. The singular forms "a", "the", and "said" used in one or more embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this application refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0062] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0063] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in the relevant region, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0064] The following will, with reference to the accompanying drawings, detail the technical solutions provided by each embodiment of this specification.

[0065] Large language models can interact with users as chatbots, and the interaction mode between large language models and users can be in the form of multi-round conversations. To facilitate the training or evaluation of the interaction ability between large language models and users, relatively realistic multi-round conversations are required. Therefore, a method for generating multi-round conversation corpora is needed. In the prior art, general large models can be used to generate multi-round conversation corpora. However, due to the insufficient context memory of general large models, there are some problems with the multi-round conversation corpora generated by general large models. For example, what was said above is repeated below, the content is relatively straightforward, and it is difficult to maintain semantic coherence in continuous multi-round conversations, resulting in poor context relevance and fluency in the overall multi-round conversations. In addition, user large models can also be trained for each user, and the trained user large models are used to output user conversation corpora. However, the number of users is extremely large, resulting in a very high cost of training user large models, and thus a relatively high cost of using user large models to output user conversation corpora.

[0066] To solve the defects in the prior art, the following embodiments are given in this solution.

[0067] Figure 1 It is a schematic diagram of the application scenario of a method for generating multi-round conversation corpora provided by an embodiment of this specification.

[0068] As Figure 1As shown, the solution may include a user terminal 1, a server 1, and a large language model 3. The multi-turn dialogue corpus requester can input the dialogue sequence to be amplified through the user terminal 1. The user terminal 1 can send the dialogue sequence to be amplified input by the user to the server 2. The server 2 can respond to the dialogue corpus generation requirement of the dialogue corpus requester, obtain the first question information from the dialogue sequence, and recall the question context corresponding to the question with high similarity to the first question information from the vector database in a vector recall manner. The server 2 can also call the large language model 3 and input the recalled question context as a constraint condition to the large language model 3, so that the large language model 3 generates the second question information after the first question information. In the scenario as Figure 1 shown, the large language model 3 can be carried on the server 2 or on other servers communicatively connected to the server 2.

[0069] Although Figure 1 shown, after receiving the dialogue sequence to be amplified input by the user, the user terminal 1 can send the dialogue sequence to be amplified to the server 2. In actual application, if the computing resources of the user terminal 1 meet the conditions for running the method of generating multi-turn dialogue corpus, the solution of generating the second question information based on the obtained first question information can also be executed on the user terminal 1. In this case, the called deep learning model 3 can be deployed on the user terminal 1 or on other terminals or servers communicatively connected to the user terminal 1.

[0070] In the application scenario as Figure 1 shown, the server can connect one or more terminal devices through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Figure 1 The server in Figure 1 may include but is not limited to any device, equipment, platform, device cluster, etc. with computing and processing capabilities.

[0071] In this application, a method for generating a multi-turn dialogue corpus, a method for training a large language model, and a method for testing a large language model are provided. This application is also related to a device for generating a multi-turn dialogue corpus, a device for training a large language model, a device for testing a large language model, a device for generating a multi-turn dialogue corpus, a device for training a large language model, and a device for testing a large language model. They will be described in detail one by one in the following embodiments.

[0072] Figure 2 It is a flowchart of a method for generating a multi-turn dialogue corpus provided by an embodiment of this specification.

[0073] From a program perspective, the execution entity of the process can be a program running on an application server or an application terminal. It can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities.

[0074] As Figure 2 shown, the process may include the following steps:

[0075] Step 202: Obtain the first question information.

[0076] In the embodiments of this specification, the first question information may be the question information in the dialogue sequence to be amplified. Specifically, the first question information may be the question information proposed by the user role. The first question information may be a question about any field and any type. For example, the first question information may be question information about the weather or question information about insurance.

[0077] Step 204: Perform vectorization processing on the first question information to obtain a first question vector.

[0078] The first question vector may be a semantic feature vector of the first question information and can be used to represent the first question information. Using the first question vector to represent the first question information can enable the model to more accurately understand the meaning of the first question information, thus facilitating subsequent applications.

[0079] In the embodiments of this specification, a text vectorization model may be used to perform vectorization processing on the first question information. Among them, the text vectorization model may include a term frequency-inverse document frequency model, a bag-of-words model, a word embedding model, etc. In practical applications, other models or methods may also be used to perform vectorization processing on the first question information.

[0080] Step 206: Query a target previous context vector in the vector database whose similarity to the first question vector meets a preset similarity condition; multiple data pairs are stored in the vector database; for any one of the data pairs, it includes a previous context vector generated according to the previous context of the user's questions in the historical interaction process with the language model, and the subsequent context of the questions in the historical interaction process.

[0081] In practical applications, the vector database may be pre-constructed. Multiple data pairs may be stored in the vector database, where each data pair may include a previous context vector of the previous context of the question and the subsequent context of the question. Specifically, the form of each data pair may be Pair(V(U1),U2), where U1 represents the previous context of the question, V(U1) represents the previous context vector corresponding to the previous context U1 of the question, and U2 represents the subsequent context of the question adjacent to the previous context U1 of the question.

[0082] Among them, the upper context of the question and the lower context of the question can be the question information output by the user role. The upper context vector can be the semantic feature vector of the upper context of the question. Specifically, it can be the upper context vector obtained by vectorizing the upper context of the question using a text vectorization model.

[0083] To ensure that the first question information and the upper context of the question are transformed into the same vector space for similarity matching between the first question vector and the upper context vector, a model of the same type as the model used for vectorizing the first question information can be used to vectorize the upper context of the question; further, the model or method directly used for vectorizing the first question information can be used to vectorize the upper context of the question.

[0084] In the embodiments of this specification, the upper context of the question and the lower context of the question can be the question information output by real online users obtained based on the conversation data of real online users. For example, the upper context of the question and the lower context of the question can be the question information output by real online users during the interaction with the language model. Or, the upper context of the question and the lower context of the question can also be the question information raised during the Q&A conversation between real users and human customer service. Among them, the interaction between real online users and the language model can be understood as the user directly interacting with the language model, or the user interacting with a dialogue product based on the language model. For example, the user invokes the language model in the dialogue window to interact with the dialogue agent. The dialogue agent can include a financial management assistant, an insurance planning assistant, etc., and is not limited to these examples.

[0085] In practical applications, the question information output by the user can be obtained during the interaction between the user and the language model or between the user and the human customer service, so as to obtain the upper context of the question and the lower context of the question; or after the interaction between the user and the language model or between the user and the human customer service is completed, the interaction content between the user and the language model or between the user and the human customer service can be obtained, and the question information output by the user in the interaction content can be extracted to obtain the upper context of the question and the lower context of the question.

[0086] In addition, the upper context of the question and the lower context of the question can also be the question information raised by the user role obtained based on the conversation data written by experts. The upper context of the question and the lower context of the question obtained based on the conversation data written by experts or the conversation data of real users have better coherence.

[0087] The following question can be the question information output by the user after outputting the previous question; optionally, the following question can be the question information adjacent to the previous question. The previous question can include at least one question information. If the previous question includes multiple question information, the following question being adjacent to the previous question can mean that after the user outputs the last question information in the previous question, the next question information output is the question information in the following question. That is, there is no other question information between the last question information in the previous question and the question information in the following question.

[0088] In the embodiments of this specification, the target previous context vector can be a previous context vector with a relatively high similarity to the first question vector. In practical applications, a similarity algorithm can be used to calculate the similarity between each previous context vector in the vector database and the first question vector, so as to screen out the target previous context vector based on the preset similarity conditions. Among them, the similarity algorithm can be one or more of the cosine similarity algorithm, Euclidean distance algorithm, edit distance algorithm, and Manhattan distance algorithm.

[0089] In the embodiments of this specification, the preset similarity condition can refer to the condition set based on the preset similarity threshold, or can refer to the condition set based on the ranking of similarities. The target previous context vectors determined based on the preset similarity conditions can be multiple. Specifically, the previous context vectors with a similarity greater than the preset similarity threshold to the first question vector can be used as the target previous context vectors; alternatively, the similarities between each previous context vector and the first question vector can be sorted from large to small, and the top preset number of previous context vectors can be selected as the target previous context vectors.

[0090] In the embodiments of this specification, during retrieval and matching, by only considering the user's question information and not considering the model's response information, overfitting can be avoided.

[0091] Step 208: Obtain question generation constraint information based on the following question corresponding to the target previous context vector.

[0092] In practical applications, the question generation constraint information can be a restriction on the model output, which is used to ensure that the model output conforms to specific rules, standards, or expectations, and can improve the reliability of the model and ensure the accuracy of the model output.

[0093] In the embodiments of this specification, the question generation constraint information can be text data composed of the following questions corresponding to multiple target previous context vectors.

[0094] Step 210: Call the large language model to generate the second question information after the first question information with the question generation constraint information as the constraint condition.

[0095] In the embodiments of this specification, since the preceding question text and the succeeding question text stored in the vector database are coherent, by using the succeeding question text adjacent to the preceding question text with a high similarity to the first question information to generate question generation constraint information, and using this question generation constraint information as the constraint condition when the large language model generates the second question information, thus, the second question information generated by the large language model has better coherence with the first question information.

[0096] Step 212: Generate multi-turn dialogue corpus based on the first question information and the second question information.

[0097] In the embodiments of this specification, in the multi-turn dialogue corpus generated based on the first question information and the second question information, the second question information can be adjacent to the first question information, that is, there may be no other question information between the first question information and the second question information.

[0098] In practical applications, the method flow as shown in Figure 2 can be further adopted to obtain the third question information based on the second question information. Similarly, in the multi-turn dialogue corpus generated based on the first question information, the second question information and the third question information, the third question information can be adjacent to the second question information, that is, there may be no other question information between the second question information and the third question information.

[0099] It should be understood that in the method described in one or more embodiments of this specification, the order of some steps can be adjusted according to actual needs, or some steps can be omitted.

[0100] Figure 2 In the method of , by using the data pairs including the preceding vector and the succeeding question text generated from the Q&A dialogue sequence of real online users to construct the vector database, when it is necessary to construct or expand the dialogue corpus, based on the obtained first question information, the succeeding question text corresponding to the preceding question text with a high similarity to the first question information can be recalled from the vector database by vector recall, and the recalled succeeding question text is used as the constraint condition to make the large language model generate the second question information after the first question information, and a large number of multi-turn dialogue corpus can be obtained.

[0101] In addition, using real online dialogue data as the basis and constraint for the output of the large language model can also improve the coherence between the second question information generated by the large language model and the dialogue sequence where the first question information is located, improve the quality of the generated multi-turn dialogue corpus, and meet the needs of the training and evaluation of the large language model.

[0102] Based on Figure 2 the method, the embodiments of this specification also provide some improved implementation manners of this method, which are described below.

[0103] In one or more embodiments of this specification, optionally, the obtaining of the first question information may specifically include:

[0104] Obtain a conversation sequence; the conversation sequence contains P rounds of conversation information sorted by occurrence time; in any round of the conversation information, there is a user question information and a model reply information generated in response to the user question information; P is a positive integer greater than or equal to 2.

[0105] Determine the user question information in the last Q rounds of conversation information in the conversation sequence as the first question information; Q is a positive integer less than or equal to P and less than or equal to the round threshold.

[0106] In the embodiments of this specification, the conversation sequence may be a conversation sequence that needs to be amplified. The conversation sequence may include at least two rounds of conversation information. Among them, one round of conversation may refer to "one question and one answer", that is, a user question and a reply in response to the user question. One round of conversation information may refer to a user's question information output and the information replied by the model or artificial customer service in response to the user's question information.

[0107] In the embodiments of this specification, the first question information may be used as the question context of the second question information.

[0108] In practical applications, the more question information contained in the first question information, the higher the degree of repetition between the target context vector queried from the vector database when generating question information in the (n - 1)th round and the target context vector queried from the vector database when generating question information in the nth round. As a result, the similarity between the question generation constraint information obtained when generating question information in the (n - 1)th round and the question generation constraint information obtained when generating question information in the nth round will be higher, and then the similarity between the question information generated in the (n - 1)th round and the question information generated in the nth round will be higher, ultimately resulting in a higher degree of repetition of the question information in the generated multi-round conversation corpus. Next, an example is used to illustrate. When considering multiple question information [U1, U2,..., Un] simultaneously in a conversation corpus generation, when n is relatively large, the similarity between the question vectors used for calculating the similarity with each context vector in the vector database in the (n - 1)th round and the question vectors used for calculating the similarity with each context vector in the vector database in the nth round is very high. As a result, the data pairs Pair(V(U1), U2),..., Pair(V(Un), Un+1) containing the target context vector and the question context obtained in the (n - 1)th round and the nth round of search are very likely to be the same, that is, the question generation constraint information obtained when generating question information in the (n - 1)th round and the question generation constraint information obtained when generating question information in the nth round are very likely to be the same, thereby resulting in a very high similarity between the question information generated in the (n - 1)th round and the question information generated in the nth round.

[0109] To solve the above problems, in the embodiments of this specification, the first question information may only include the question information of the user in the last no more than the turn threshold times in the dialogue sequence, so as to ensure the coherence and diversity of the question context recalled from the vector database corresponding to the question next context, and further ensure the coherence and diversity of the second prompt information generated based on the question next context. In the embodiments of this specification, the turn threshold may be, for example, 1, 2, or 3, etc.

[0110] It can be understood that, in order to facilitate screening the target context vector from the vector database, correspondingly, the question context corresponding to the context vector in the vector database may only include the last Q question information adjacent to the question next context before the question next context.

[0111] As a specific implementation manner, in the embodiments of this specification, optionally, determining the user question information in the last Q rounds of dialogue information of the dialogue sequence as the first question information may specifically include:

[0112] Determining the user question information in the last round of dialogue information of the dialogue sequence as the first question information.

[0113] In the embodiments of this specification, the first question information may only include the user's last question information, which can further ensure the diversity of the question next context corresponding to the question context recalled from the vector database, and further ensure the diversity of the second prompt information generated based on the question next context.

[0114] In order to facilitate screening the target context vector from the vector database, correspondingly, the question context corresponding to the context vector in the vector database may only include the last question information adjacent to the question next context before the question next context. That is, the data pairs in the vector database are generated based on binary question groups similar to [the (n - 1)th user question information, the nth user question information].

[0115] For ease of understanding, in the embodiments of this specification, a specific description is also made for the solution of determining the problem generation constraint information.

[0116] In one or more embodiments of this specification, optionally, the target context vector may specifically include M target context vectors, where M is a positive integer; obtaining the problem generation constraint information based on the question next context corresponding to the target context vector may specifically include:

[0117] Obtaining M question next contexts corresponding to the M target context vectors from the vector database; the M question next contexts correspond to the M target context vectors one by one.

[0118] For any one of the M question follow-ups in the following text, splice the one question follow-up to the end of the conversation sequence to which the first question information belongs, to obtain a to-be-determined conversation sequence corresponding to the one question follow-up.

[0119] Calculate the perplexity of the to-be-determined conversation sequences respectively corresponding to each of the M question follow-ups in the following text; the perplexity is used to characterize the coherence degree of the to-be-determined conversation sequence; the lower the perplexity, the more coherent the to-be-determined conversation sequence is.

[0120] Determine N question follow-ups from the M question follow-ups, whose perplexity meets a preset perplexity condition; N is a positive integer less than or equal to M.

[0121] Based on the N question follow-ups, obtain question generation constraint information.

[0122] In the embodiments of this specification, the above text vector and the question follow-up may be stored associatively, and based on the target above text vector, the question follow-up corresponding to the target above text vector can be obtained.

[0123] The to-be-determined conversation sequence may be a conversation sequence after splicing one of the M question follow-ups. The M question follow-ups can be respectively spliced to the end of the conversation sequence where the first question information is located, so as to obtain M to-be-determined conversation sequences. For example, the conversation sequence to which the first question information belongs is [A1, B1, A2, B2,..., An, Bn], where A1, A2...An may be the question information in the above conversation sequence to which the first question information belongs, and B1, B2...Bn may be the model reply information in the above sequence. Assume that the obtained M question follow-ups are U1, U2...Um; the to-be-determined conversation sequences after splicing the question follow-ups are respectively [A1, B1, A2, B2,..., An, Bn, U1]; [A1, B1, A2, B2,..., An, Bn, U2]; ……; [A1, B1, A2, B2,..., An, Bn, Um].

[0124] The perplexity can be calculated for each to-be-determined conversation sequence respectively, so as to determine, according to the perplexity, the conversation sequences in which the spliced question follow-ups are still smooth, coherent, and natural.

[0125] Perplexity (PPL for short in English) can be a text metric method, also called confusion degree, chaos degree, etc., which can be used to reflect the uncertainty of the text, characterize the smoothness degree of the text, or say the coherence degree, naturalness degree, etc. The greater the perplexity of the to-be-determined conversation sequence, the less smooth, less coherent, and less natural the to-be-determined conversation sequence can be; on the contrary, the smaller the perplexity of the to-be-determined conversation sequence, the smoother, more coherent, and more natural the to-be-determined conversation sequence can be.

[0126] In the embodiments of this specification, the subsequent questions that meet the preset perplexity condition may refer to all subsequent questions with perplexity values lower than the preset perplexity threshold; or it may also refer to a preset number of subsequent questions with the lowest perplexity values among the M subsequent questions.

[0127] In the embodiments of this specification, by using subsequent questions screened based on perplexity that have better coherence with the dialogue sequence to which the first question information belongs, and using the question generation constraint information generated based on the above-screened subsequent prompts as the constraint conditions for the large language model, the coherence between the second question information generated by the large language model and the dialogue sequence where the first question information is located can be further improved.

[0128] In the embodiments of this specification, relevant content for calculating the perplexity of the pending dialogue sequence is further provided.

[0129] In one or more embodiments of this specification, optionally, calculating the perplexity of each of the M subsequent questions corresponding to the respective pending dialogue sequences may specifically include:

[0130] For any one of the pending dialogue sequences, perform word segmentation on the pending dialogue sequence to obtain a character sequence.

[0131] Determine the occurrence probability of each character included in the character sequence.

[0132] Based on the occurrence probabilities of the characters in the character sequence, calculate the perplexity of the pending dialogue sequence.

[0133] Word segmentation is the process of splitting a continuous text string into individual words or phrases.

[0134] The hidden Markov model, neural network models such as convolutional neural network CNN, recurrent neural network RNN, long short-term memory network LSTM, etc. can be used to extract features and classify the text, thereby realizing word segmentation. Tools such as jieba, ICTCLAS (NLPIR), HanLP, etc. can also be used to perform word segmentation on the pending dialogue sequence to obtain a character sequence.

[0135] Taking the pending dialogue sequence "What's the weather like in Hangzhou today? Hangzhou is sunny. What about tomorrow?" as an example for illustration, the character sequence obtained after word segmentation can be: 'ˋˋˋtoday / Hangzhou / weather / what / how / ? / Hangzhou / is / sunny / . / that / tomorrow / what / ?ˋˋˋ'.

[0136] In the embodiments of this specification, the occurrence probability of each character may refer to the possibility of each character appearing.

[0137] In the embodiments of this specification, specific content for determining the occurrence probabilities of each character in a character sequence is also provided.

[0138] In one or more embodiments of this specification, optionally, determining the occurrence probabilities of each character included in the character sequence may specifically include:

[0139] Based on the character sequence, generating probability determination prompt information for input to a large language model; the probability determination prompt information includes the conditional probability expressions of each character in the character sequence and probability determination task information; the probability determination task information is used to indicate calculating the occurrence probabilities of each character in the character sequence according to the conditional probability expressions.

[0140] Inputting the probability determination prompt information into the large language model to obtain the occurrence probabilities of each character included in the character sequence output by the large language model.

[0141] In practical applications, one or more of an N-Gram model, a large language model, a neural network model, and a regression model can be used to calculate the occurrence probabilities of each character, and specific limitations are not made here.

[0142] When generating the probability determination prompt information, a probability determination prompt template can be obtained first, and then the conditional probability expressions are inserted into the probability determination template. The probability determination task information can be content already in the probability determination template. Among them, the conditional probability expressions can be generated based on the character sequence, for example, the conditional probability expressions can be automatically generated by writing scripts.

[0143] The conditional probability expressions of each character can be used to represent the probabilities of each character occurring under preset conditions. The preset conditions can refer to a specific context, or the entire corpus, etc.

[0144] In the embodiments of this specification, the conditional probability expressions of each character can be expressions of the probabilities of each character occurring based on the occurrence of its previous character; or, they can also be expressions of the probabilities of each character occurring based on the simultaneous occurrence of its previous character and its subsequent character.

[0145] Continuing with the above example for illustration, the conditional probability expression of the two characters "Hangzhou" can be P(Hangzhou|Today), indicating the probability of the two characters "Hangzhou" occurring based on the occurrence of the two characters "Today".

[0146] The conditional probability expressions of each character can be as follows:

[0147] P(Today);

[0148] P(Hangzhou|Today);

[0149] P(Weather | Hangzhou today);

[0150] P(How | Weather in Hangzhou today);

[0151] P(? | How about the weather in Hangzhou today);

[0152] P(Hangzhou | How about the weather in Hangzhou today?);

[0153] P(Is | How about the weather in Hangzhou today? Hangzhou);

[0154] P(Sunny | How about the weather in Hangzhou today? Hangzhou is).

[0155] The probability determination task information can be an instruction input to the large language model, used to instruct the large language model to calculate the occurrence probability of each character according to the conditional probability expression. For example, the probability determination task information can be: "Output the above probabilities respectively and give specific probability values."

[0156] As an example, the occurrence probabilities of each character output by the large language model can be as follows:

[0157] P(Today) = 0.05;

[0158] P(Hangzhou | Today) = 0.03;

[0159] P(Weather | Hangzhou today) = 0.04;

[0160] P(How | Weather in Hangzhou today) = 0.02;

[0161] P(? | How about the weather in Hangzhou today) = 0.01;

[0162] P(Hangzhou | How about the weather in Hangzhou today?) = 0.03;

[0163] P(Is | How about the weather in Hangzhou today? Hangzhou) = 0.06;

[0164] P(Sunny | How about the weather in Hangzhou today? Hangzhou is) = 0.07.

[0165] In the embodiments of this specification, the large language model for generating the occurrence probability of each character can be the same type of large language model as the large language model for generating the second prompt information in the previous text. Further, the large language model for generating the occurrence probability of each character can be the same as the large language model for generating the second prompt information in the previous text.

[0166] In the embodiments of this specification, using the large language model to generate the occurrence probability of each character is convenient and fast.

[0167] In the embodiments of this specification, specific content for calculating the perplexity of a to-be-determined dialogue sequence is also provided.

[0168] Calculating the perplexity of the to-be-determined dialogue sequence based on the occurrence probabilities of the respective characters in the character sequence may specifically include:

[0169] Calculating the perplexity of the to-be-determined dialogue sequence according to the following formula:

[0170]

[0171] where PPL represents perplexity; N represents the number of characters in the character sequence; P(ω i ω <i ) represents the occurrence probability of the i-th character in the character sequence.

[0172] In the embodiments of this specification, the exponential function is used to convert the average log probability into perplexity, making the coherence of the to-be-determined dialogue sequence more intuitive.

[0173] In the embodiments of this specification, the method of using the exponential function to convert the average log probability into perplexity is only a specific way of calculating the perplexity of the to-be-determined dialogue sequence. In practical applications, other ways can also be used to calculate the perplexity of the to-be-determined dialogue sequence. For example, the perplexity of the to-be-determined dialogue sequence can be calculated using the following formula:

[0174]

[0175] P(ω1ω2…ω N ) = P(ω1)P(ω2|ω1)...P(ω N ω1ω2…ω N-1 );

[0176] where P(ω1ω2…ω N ) is the occurrence probability of the to-be-determined dialogue sequence, and P(ω1), P(ω2|ω1), …, P(ω N |ω1ω2…ω N-1 ) are all the occurrence probabilities of the characters in the to-be-determined dialogue sequence, and N is the number of characters in the to-be-determined dialogue sequence.

[0177] After determining the occurrence probabilities of at least two to-be-determined dialogue sequences based on the occurrence probabilities of the respective characters in the to-be-determined dialogue sequence, the perplexity PPL of the to-be-determined dialogue sequence can be further determined.

[0178] In the embodiments of this specification, specific content regarding problem generation constraint information is also provided.

[0179] In one or more embodiments of this specification, obtaining the question generation constraint information based on the question context corresponding to the target above context may specifically include:

[0180] Generating question generation constraint information including the question context and constraint task information; the constraint task information is used to indicate generating question information similar to the question context.

[0181] In the embodiments of this specification, the question context and the constraint task information may be concatenated to generate the question generation constraint information. Specifically, the constraint task information may be set before the question context or after the question context. No specific limitation is made here.

[0182] In addition, the question context included in the question generation constraint information may be the M hint contexts determined in the above context, or the N hint contexts further determined from the M hint contexts.

[0183] Continuing with the above example for illustration, if the obtained hint context is "1. Is it suitable to dry clothes today? 2. Do I need to bring an umbrella when going out today? 3. Will it be sunny for the next week?" The generated question generation constraint information may be as follows:

[0184] "The generated new content is similar to these 3 pieces of conversation content:

[0185] 1. Is it suitable to dry clothes today?

[0186] 2. Do I need to bring an umbrella when going out today?

[0187] 3. Will it be sunny for the next week?"

[0188] Among them, "The generated new content is similar to these 3 pieces of conversation content" may be used as the constraint task information.

[0189] In the embodiments of this specification, using the question context and the constraint task information to generate the question generation constraint information, and using the question generation constraint information as a constraint condition to constrain the large language model to generate the second question information can make the second question information generated by the large language model have better coherence with the conversation sequence where the first question information is located.

[0190] In the embodiments of this specification, the specific content of calculating the perplexity of the to-be-determined conversation sequence is also provided.

[0191] In one or more embodiments of this specification, optionally, after calling the large language model to generate the second question information with the question generation constraint information as a constraint condition, it may specifically include:

[0192] Generate constraint information based on the above problems, and generate question generation prompt information for input into the large language model; the question generation prompt information includes the conversation sequence to which the first question information belongs, the question generation constraint information, and question generation task information; the question generation task information is used to indicate generating user question information after the conversation sequence with the question generation constraint information as the constraint condition.

[0193] Input the question generation prompt information into the large language model to obtain the second question information after the first question information output by the large language model.

[0194] In the embodiments of this specification, the question generation prompt information can be understood as the prompt words in the prompt word template. In practical applications, when generating the question generation prompt information, the question generation prompt template can be obtained first, and then the question generation constraint information can be inserted into the question generation prompt template. The question generation task information can be the content already existing in the question generation prompt template. The conversation sequence to which the first question information included in the question generation prompt information belongs can include all-round conversation information containing user question information and model reply information.

[0195] Continuing with the above example, the obtained question generation prompt information can be:

[0196] Example:

[0197] # Conversation sequence:

[0198] User question (U1): What's the weather like in Hangzhou today?

[0199] Model reply (R1): It's sunny in Hangzhou today.

[0200] # Question generation constraint information

[0201] The newly generated content is similar to these 3 pieces of conversation content:

[0202] 1. Is it suitable to dry clothes today?

[0203] 2. Do I need to bring an umbrella when going out today?

[0204] 3. Will it be sunny for the next week?

[0205] # Question generation task information

[0206] Please use the 'question generation constraint information' as the constraint condition to generate user question information that may appear after the 'conversation sequence'.

[0207] In the embodiments of this specification, the large language model for generating the second question information is of the same type as the large language model for determining the occurrence probability of each character in the foregoing text. Specifically, they can both be neural network models based on transformers. For example, they can both be GPT series models, Tongyi Qianwen models, etc., without limitation. Further, the large language model for generating the second question information and the large language model for determining the occurrence probability of each character in the foregoing text can be the same one, or they can also be different ones.

[0208] In the embodiments of this specification, using the question generation prompt information including the dialogue sequence to which the first question information belongs, the question generation constraint information, and the question generation task information as the prompt words of the large language model can make the second question information generated by the large language model have better coherence with the dialogue sequence where the first question information is located.

[0209] In the embodiments of this specification, for the convenience of application, the foregoing text vector can be obtained in advance before querying the target foregoing text vector, and further, the foregoing text vector can be stored in the vector database.

[0210] In one or more embodiments of this specification, optionally, before querying from the vector database for the target foregoing text vector whose similarity with the first question vector meets the preset similarity condition, it can further include:

[0211] Obtain the historical question sequence; the user question sequence includes at least two historical question information sorted by the question time.

[0212] From the historical question sequence, determine a historical question subsequence composed of adjacent preset numbers of historical question information; the preset number of questions is less than or equal to the round threshold; the historical question subsequence includes the question foregoing text and the question following text; the question following text includes the last question information in the historical question subsequence; the question foregoing text includes other question information in the historical question subsequence except the last question information.

[0213] Perform vectorization processing on the question foregoing text to obtain the foregoing text vector.

[0214] Associate and store the foregoing text vector with the question following text in the vector database.

[0215] In the embodiments of this specification, the historical question information can be the question information output by the user role.

[0216] The historical question subsequence may include a preset number of historical question messages. For example, if the historical question sequence is [U1, U2, U3, U4, ..., Un], the historical question subsequence may be a sequence including 3 question messages, such as [U1, U2, U3], [U2, U3, U4]... [Un-2, Un-1, Un]; of course, the historical question subsequence may also be a sequence including other numbers of question messages less than or equal to the round threshold.

[0217] The last question message in the historical question subsequence can be used as the subsequent question, and the other question messages except the last one can be used as the preceding questions. For example, in [U1, U2, U3] of the historical question subsequence, [U3] can be the subsequent question; [U1, U2] can be the preceding questions.

[0218] After determining the preceding questions and the subsequent question, the preceding questions can be vectorized to obtain a preceding question vector. In the embodiments of this specification, in order to ensure that the preceding questions and the first question message are converted into the same vector space to facilitate the similarity matching between the first question vector and the preceding question vector, a model of the same type as the model for vectorizing the first question message can be used to vectorize the preceding questions. Further, the first question message and the preceding questions can be vectorized using the same model or method.

[0219] In the embodiments of this specification, the preceding question vector and the subsequent question can be associated and stored to form a data pair.

[0220] In the embodiments of this specification, the round threshold and the number of user question messages used to obtain the first question message in the preceding questions may correspond. For example, if the number of user question messages in the obtained first question message is n, the historical question messages included in the historical question subsequence may be n + 1.

[0221] Limiting the number of historical question messages in the historical question subsequence can avoid the problem of high repetition of the subsequent questions obtained based on the first question message in different rounds.

[0222] In the embodiments of this specification, a specific solution for obtaining the historical question sequence is also provided.

[0223] In one or more embodiments of this specification, optionally, the obtaining of the historical question sequence may specifically include:

[0224] Obtain the historical conversation sequence generated by the user during the historical interaction with the language model; the historical conversation sequence includes at least two rounds of historical conversation information sorted by the question time; in any round of the historical conversation information, there is a historical question information and a model historical reply information generated in response to the historical question information.

[0225] Extract from the historical conversation sequence a user question sequence composed of the historical question information in each round of historical conversation information.

[0226] In the embodiments of this specification, the historical conversation sequence may be a conversation sequence generated during the interaction between the user and the language model. As an implementation manner, the historical conversation sequence may also be a conversation sequence generated during the interaction between the user and the artificial customer service.

[0227] The historical question sequence may be a sequence of question information sorted in chronological order of questions extracted from the historical conversation sequence. For example, the historical conversation sequence between the user and the model or the user and the artificial customer service is:

[0228] User (U1): What's the weather like in Hangzhou today?

[0229] Model (R1): It's sunny in Hangzhou today.

[0230] User (U2): I'm going to play in the West Lake this weekend. What's the weather like this weekend?

[0231] Model (R2): The weather forecast for this weekend shows that it will be sunny on Saturday and there may be scattered light rain on Sunday. You can choose to go to the West Lake on Saturday, and the weather will be better.

[0232] User (U3): What time is it more suitable to go out on Saturday morning?

[0233] Model (R3): The weather between 8:00 and 10:00 on Saturday morning is very suitable. The temperature is moderate and the sun is not too strong, which is suitable for playing.

[0234] ……

[0235] User (Un): ……?

[0236] It is possible to extract each historical question information output by the user in the above historical conversation sequence, and sort each question information in the order in the historical conversation sequence to obtain the historical question sequence. For example, the historical question sequence obtained based on the above historical conversation sequence may be [U1, U2, U3,..., Un].

[0237] In one or more embodiments of this specification, optionally, determining, from the historical question sequence, a historical question subsequence composed of historical question information of a preset number of adjacent questions may specifically include:

[0238] Determine a historical question information group composed of two adjacent historical question information from the historical question sequence.

[0239] Assume the historical question sequence is [U1, U2, U3, U4,..., Un]. The historical question subsequence can be a sequence including two adjacent question information, such as [U1, U2], [U2, U3]... [Un-1, Un]. Multiple historical question information groups can be constructed from the historical question sequence. Among them, in each historical question information group, according to the question time sequence, the question information proposed first can be used as the question context above, and the question information proposed later can be used as the question context below. For example, in the question information group [U1, U2], U1 can be used as the question context above of U2, and U2 can be used as the question context below of U1; in the question information group [U2, U3], U2 can be used as the question context above of U3, and U3 can be used as the question context below of U2.

[0240] In the embodiments of this specification, the question context above can include only one question information. In this scenario, the number of question information included in the first question information can be consistent with the number of question information included in the question context above. By performing similarity matching between the context vector of the question context above that includes only one question information and the first question vector of the first question information that includes only one question information, the diversity of the question context below corresponding to the question context above recalled from the vector database can be further improved, and thus the diversity of the second prompt information generated based on the question context below can be further improved.

[0241] As an implementation manner, in one or more embodiments of this specification, before determining a historical question subsequence composed of historical question information with a preset number of adjacent questions from the historical question sequence, it may further include:

[0242] Adopt a preset data screening rule to screen out deletable dialogue information from the historical question sequence; the data screening rule is used to screen out data whose quality does not meet the preset quality conditions.

[0243] Delete the deletable dialogue information from the historical question sequence to obtain a processed historical question sequence.

[0244] The determining a historical question subsequence composed of historical question information with a preset number of adjacent questions from the historical question sequence may specifically include:

[0245] Determine a historical question subsequence composed of historical question information with a preset number of adjacent questions from the processed historical question sequence.

[0246] In the embodiments of this specification, at least one of methods such as keyword matching method and regular matching method can be used to screen out data that does not meet the preset quality conditions. Among them, the data that does not meet the preset quality conditions can be, for example, question information that is repeated with the previous question information, or, for another example, question information that is semantically incoherent with the question information in the previous or subsequent text, or, for yet another example, question information that only contains modal particles. It is not limited to these examples, and in actual applications, the screening conditions for deletable dialogue information can be set according to business requirements.

[0247] In the embodiments of this specification, a historical question subsequence can be constructed based on the historical question sequence after deleting the data that does not meet the preset quality conditions.

[0248] In the embodiments of this specification, by deleting the data that does not meet the preset quality conditions in the historical question sequence and constructing a historical question subsequence based on the historical question sequence after deleting the data that does not meet the preset quality conditions, the situation where the semantics of the previous and subsequent parts of the question in the question subsequence are incoherent and the question information is repeated can be avoided, which can effectively improve the diversity of the second prompt information generated based on the historical question subsequence, and can also improve the coherence between the first question information and the dialogue sequence to which the first question information belongs.

[0249] As an implementation manner, in one or more embodiments of this specification, optionally, generating multi-turn dialogue corpus based on the first question information and the second question information may specifically include:

[0250] Input the second question information into a large language model to obtain a second reply information output by the large language model for the second question information.

[0251] Concatenate the second question information and the second reply information after the dialogue sequence to which the first question information belongs to obtain multi-turn dialogue corpus.

[0252] In the embodiments of this specification, the second question information and the second reply information generated using the large language model can be used as a new round of dialogue information and added after the dialogue sequence, so as to obtain an updated dialogue sequence as multi-turn dialogue corpus.

[0253] Furthermore, after obtaining the updated dialogue sequence, steps similar to those for generating the second question information in the previous text can be executed again to continue generating a third question information based on the second question information. A third reply information for generating the third question information based on the third question information can also be generated using the large language model, so as to update the dialogue sequence again using the third question information and the third reply information to obtain multi-turn dialogue corpus with more rounds.

[0254] In the embodiments of this specification, a method for training a large language model is also provided.

[0255] Optionally, a method for training a large language model may include:

[0256] Use the multi-turn dialogue corpus generated by the method for generating multi-turn dialogue corpus described above to train the large language model to be trained.

[0257] In the embodiments of this specification, the multi-turn dialogue corpus generated by using the method for generating multi-turn dialogue corpus above may be used as a training set to train the large language model. The specific training process is not limited herein.

[0258] In the embodiments of this specification, a method for testing a large language model is also provided.

[0259] Optionally, a method for testing a large language model may include:

[0260] Use the multi-turn dialogue corpus generated by the method for generating multi-turn dialogue corpus described above to test the trained large language model.

[0261] In the embodiments of this specification, the multi-turn dialogue corpus generated by using the method for generating multi-turn dialogue corpus above may be used as a test set to test the large language model. The specific test process is not limited herein.

[0262] In the embodiments of this specification, the training set that can be used to train the large language model and the test set that can be used to test the large language model may be the same or different.

[0263] The various technical features in the above embodiments can be combined arbitrarily as long as there is no conflict or contradiction between the features. However, due to space limitations, they are not described one by one. Therefore, any combination of the various technical features in the above embodiments also falls within the scope disclosed in this specification.

[0264] To more clearly illustrate a method for generating a multi-turn dialogue corpus provided in the embodiments of this specification, Figure 3 A swimlane diagram of a method for generating a multi-turn dialogue corpus provided in the embodiments of this specification. As Figure 3 shown, the process of the method for generating a multi-turn dialogue corpus can be executed by execution entities such as a server and a large language model. The process of generating the multi-turn dialogue corpus may include a data preparation stage and a multi-turn dialogue corpus generation stage.

[0265] In the data preparation stage, the execution entity may include a server, and specifically may include the following steps:

[0266] Step 302: Obtain the historical dialogue sequence.

[0267] The historical dialogue sequence can be a dialogue sequence generated by the user during the historical interaction with the language model. As an implementation, the historical dialogue sequence can also be a dialogue sequence generated during the interaction between the user and the human customer service.

[0268] The historical dialogue sequence can include at least two rounds of historical dialogue information sorted by the question time; in any round of the historical dialogue information, there can be a historical question information and a model historical reply information generated in response to the historical question information.

[0269] Step 304: Extract from the historical dialogue sequence a user question sequence composed of the historical question information in each round of historical dialogue information.

[0270] The historical question sequence can be a sequence of question information sorted in chronological order of questions extracted from the historical dialogue sequence. For example, the historical dialogue sequence of the user with the model and then with the human customer service is as follows:

[0271] User (U1): Hello, I'm very interested in finance. Can you give me some advice?

[0272] Model (R1): Of course! For beginners, you can start by reading some basic finance books to understand some basic concepts.

[0273] User (U2): Can I make a financial plan first?

[0274] Model (R2): Yes, you need to consider your financial situation, risk tolerance, investment goals, and time frame comprehensively.

[0275] User (U3): That sounds complicated. Is there any simple finance method suitable for beginners?

[0276] Model (R3): For beginners, you can start with simple time deposits or buying money market funds.

[0277] ……

[0278] User (Un): ……?

[0279] It is possible to extract each historical question information output by the user in the above historical dialogue sequence, sort each question information in the order in the historical dialogue sequence, and obtain the historical question sequence. For example, the historical question sequence obtained based on the above historical dialogue sequence can be [U1, U2, U3,..., Un].

[0280] Step 306: Adopt a preset data screening rule to screen out deletable dialogue information from the historical question sequence; the data screening rule is used to screen out data whose quality does not meet the preset quality conditions.

[0281] In the embodiments of this specification, at least one of methods such as keyword matching method and regular matching method can be used to screen out data that does not meet the preset quality conditions. Among them, the data that does not meet the preset quality conditions can be question information that is repeated with the previous question information, or question information that is semantically incoherent with the question information in the previous or subsequent text, etc.

[0282] Step 308: Delete the deletable dialogue information from the historical question sequence to obtain a processed historical question sequence.

[0283] Step 310: Determine a historical question information group composed of two adjacent historical question information from the processed historical question sequence.

[0284] Assume that the processed historical question sequence is [U1, U2, U3, U4,..., Un]. Optionally, the historical question information group can be a sequence including two adjacent question information, such as [U1, U2], [U2, U3]... [Un-1, Un]. Among them, in each historical question information group, according to the question time sequence, the question information proposed first can be used as the question above text, and the question information proposed later can be used as the question below text. For example, in the question information group [U1, U2], U1 can be used as the question above text of U2, and U2 can be used as the question below text of U1; in the question information group [U2, U3], U2 can be used as the question above text of U3, and U3 can be used as the question below text of U2.

[0285] Step 306 and Step 308 can be executed or not. If Step 306 and Step 308 are not executed, during the execution of Step 310, based on the user question sequence extracted in Step 304, a historical question information group composed of two adjacent historical question information can be determined.

[0286] Step 312: Perform vectorization processing on the question above text in the historical question information group to obtain an above text vector.

[0287] The text vectorization model can be used to perform vectorization processing on the question above text. Among them, the text vectorization model can include one-hot encoding, bag-of-words model, word embedding model, etc. In practical applications, other models or methods can also be used to perform vectorization processing on the first question information.

[0288] Step 314: Associatively store the above text vector and the question below text into the vector database.

[0289] In the multi-round dialogue corpus generation stage, the execution entity may include the server and the device where the large language model is located. Specifically, the following steps may be included: The server in the multi-round dialogue corpus generation stage and the server in the data preparation stage may be the same server or different servers.

[0290] Step 316: Obtain the first question information.

[0291] The first question information may be the question information in the dialogue sequence that needs to generate new dialogue corpus. Specifically, the first question information may be the question information proposed by the user role.

[0292] Step 318: Perform vectorization processing on the first question information to obtain a first question vector.

[0293] To ensure that the first question information and the previous question text are transformed into the same vector space for subsequent processing, a model of the same type as the model used for vectorizing the previous question text can be used to perform vectorization processing on the first question information; further, the model or method directly used for vectorizing the previous question text can be used to perform vectorization processing on the first question information.

[0294] Step 320: Query the target previous text vector in the vector database whose similarity with the first question vector meets the preset similarity condition.

[0295] The target previous text vector may be the previous text vector with a relatively high similarity to the first question vector. In practical applications, a similarity algorithm can be used to calculate the similarity between each previous text vector in the vector database and the first question vector, and then the target previous text vector can be selected based on the preset similarity condition. Among them, the similarity algorithm can be one or more of the cosine similarity algorithm, Euclidean distance algorithm, edit distance algorithm, and Manhattan distance algorithm.

[0296] Step 322: Obtain question generation prompt information based on the question next text corresponding to the target previous text vector.

[0297] In the embodiments of this specification, the question generation prompt information can be understood as the prompt words in the prompt word template. The question generation prompt information may include the dialogue sequence to which the first question information belongs, question generation constraint information, and question generation task information; the question generation task information can be used to instruct the large language model to generate the user question information after the dialogue sequence with the question generation constraint information as the constraint condition. The question generation constraint information may include the question next text and constraint task information; among them, the constraint task information can be used to indicate generating question information similar to the question next text.

[0298] In practical applications, when generating question generation prompt information, the question generation prompt template can be obtained first, and then the question generation constraint information can be inserted into the question generation prompt template. The question generation task information can be the content already in the question generation prompt template. The conversation sequence to which the first question information included in the question generation prompt information belongs can include all rounds of conversation information containing the user's question information and the model's response information.

[0299] Step 324: After the first large language model generates the second question information based on the question generation prompt information, the first question information is generated.

[0300] Step 326: The second large language model outputs a second response information for the second question information based on the second question information.

[0301] Among them, the second large language model and the first large language model can be the same or different.

[0302] Step 328: The second question information and the second response information are spliced after the conversation sequence to which the first question information belongs to obtain multi-turn conversation corpus.

[0303] The second question information and the second response information generated by the large language model can be used as a new round of conversation information and added after the conversation sequence to obtain an updated conversation sequence as the multi-turn conversation corpus.

[0304] Furthermore, after obtaining the updated conversation sequence, steps similar to those for generating the second question information in the previous text can be executed again to continue generating the third question information based on the second question information. The third response information for generating the third question information can also be generated using the large language model, so as to update the conversation sequence again using the third question information and the third response information to obtain a multi-turn conversation corpus with more rounds.

[0305] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method.

[0306] Figure 4 It is a schematic structural diagram of a device for generating multi-turn conversation corpus provided by the embodiments of this specification.

[0307] As Figure 4 shown, the device may include:

[0308] The question information acquisition module 402 is used to acquire the first question information;

[0309] The vectorization processing module 404 is used to perform vectorization processing on the first question information to obtain a first question vector;

[0310] The vector recall module 406 is configured to query, from a vector database, target previous context vectors whose similarity to the first query vector meets a preset similarity condition; multiple data pairs are stored in the vector database; for any one of the data pairs, it includes a previous context vector generated according to the previous context of the user's query during the historical interaction with the language model, and the subsequent context of the query during the historical interaction.

[0311] The constraint information generation module 408 is configured to obtain question generation constraint information based on the subsequent context corresponding to the target previous context vector.

[0312] The query information generation module 410 is configured to call a large language model to generate a second query information after the first query information with the question generation constraint information as a constraint condition.

[0313] The dialogue corpus generation module 412 is configured to generate multi-round dialogue corpus based on the first query information and the second query information.

[0314] Based on Figure 4 For the device of [description], the embodiments of this specification also provide some specific implementation schemes of this method, which will be described below.

[0315] Optionally, the query information acquisition module 402 may specifically include:

[0316] The first acquisition unit is configured to acquire a dialogue sequence; the dialogue sequence contains P rounds of dialogue information sorted by occurrence time; for any one of the rounds of dialogue information, it contains a user query information and a model reply information generated in response to the user query information; P is a positive integer greater than or equal to 2

[0317] The first determination unit is configured to determine the user query information in the last Q rounds of dialogue information in the dialogue sequence as the first query information; Q is a positive integer less than or equal to P and less than or equal to the round threshold.

[0318] Optionally, the first determination unit may specifically be configured to:

[0319] Determine the user query information in the last round of dialogue information in the dialogue sequence as the first query information.

[0320] Optionally, the target previous context vector specifically includes M target previous context vectors, and M is a positive integer;

[0321] Optionally, the constraint information generation module 408 may specifically include:

[0322] The second acquisition unit is configured to acquire M subsequent contexts corresponding to the M target previous context vectors from the vector database; the M subsequent contexts correspond to the M target previous context vectors one by one.

[0323] A splicing unit, which is used for any one of the question contexts among the M question contexts, to splice the any one of the question contexts after the conversation sequence to which the first question information belongs, so as to obtain a to-be-determined conversation sequence corresponding to the any one of the question contexts.

[0324] A calculation unit, which is used to calculate the perplexity of the to-be-determined conversation sequences respectively corresponding to the respective question contexts among the M question contexts; the perplexity is used to characterize the coherence degree of the to-be-determined conversation sequence; the lower the perplexity, the more coherent the to-be-determined conversation sequence is.

[0325] A second determination unit, which is used to determine N question contexts whose perplexity meets a preset perplexity condition from among the M question contexts; N is a positive integer less than or equal to M.

[0326] Based on the N question contexts, question generation constraint information is obtained.

[0327] Optionally, the calculation unit may specifically include:

[0328] A marking subunit, which is used for any one of the to-be-determined conversation sequences to perform word segmentation processing on the to-be-determined conversation sequence to obtain a character sequence.

[0329] A determination subunit, which is used to determine the occurrence probability of each character included in the character sequence.

[0330] A calculation subunit, which is used to calculate the perplexity of the to-be-determined conversation sequence based on the occurrence probability of each character in the character sequence.

[0331] Optionally, the determination subunit may specifically be used for:

[0332] Based on the character sequence, generate probability determination prompt information for input to a large language model; the probability determination prompt information includes conditional probability expressions of each character in the character sequence and probability determination task information; the probability determination task information is used to indicate calculating the occurrence probability of each character in the character sequence according to the conditional probability expression.

[0333] Input the probability determination prompt information into the large language model to obtain the occurrence probability of each character included in the character sequence output by the large language model.

[0334] Optionally, the calculation subunit may specifically be used for:

[0335] Calculate the perplexity of the to-be-determined conversation sequence according to the following formula.

[0336]

[0337] Wherein, PPL represents perplexity; N represents the number of characters in the character sequence; P(ω i ω <i ) represents the occurrence probability of the i-th character in the character sequence.

[0338] Optionally, the constraint information generation module 408 may specifically be configured to:

[0339] Generate question generation constraint information including the follow-up of the question and constraint task information; the constraint task information is used to indicate generating question information similar to the follow-up of the question.

[0340] Optionally, the question information generation module 410 may specifically be configured to:

[0341] Generate question generation prompt information for input to the large language model based on the question generation constraint information; the question generation prompt information includes the dialogue sequence to which the first question information belongs, the question generation constraint information, and question generation task information; the question generation task information is used to indicate generating user question information after the dialogue sequence with the question generation constraint information as a constraint condition.

[0342] Input the question generation prompt information into the large language model to obtain the second question information after the first question information output by the large language model.

[0343] Optionally, Figure 4 The device may further include:

[0344] A historical question sequence acquisition module, configured to acquire a historical question sequence; the user question sequence includes at least two historical question information sorted by question time.

[0345] A determination module, configured to determine, from the historical question sequence, a historical question subsequence composed of adjacent preset numbers of historical question information; the preset number of questions is less than or equal to the round threshold; the historical question subsequence includes the above-mentioned part of the question and the follow-up of the question; the follow-up of the question includes the last question information in the historical question subsequence; the above-mentioned part of the question includes other question information in the historical question subsequence except the last question information.

[0346] A first vectorization processing module, which vectorizes the above-mentioned part of the question to obtain the above-mentioned vector.

[0347] A storage module, configured to associatively store the above-mentioned vector and the follow-up of the question in the vector database.

[0348] Optionally, the historical question sequence acquisition module may specifically be configured to:

[0349] Obtain the historical conversation sequence generated by the user during the historical interaction with the language model; the historical conversation sequence includes at least two rounds of historical conversation information sorted by the question time; in any round of the historical conversation information, there is a historical question information and a model historical reply information generated in response to the historical question information.

[0350] Extract the user question sequence composed of the historical question information in each round of historical conversation information from the historical conversation sequence.

[0351] Optionally, the determining module may specifically be used for:

[0352] Determine a historical question information group composed of two adjacent historical question information from the historical question sequence.

[0353] Optionally, Figure 4 The described device may further include:

[0354] A screening module, configured to screen out deletable conversation information from the historical question sequence by using a preset data screening rule; the data screening rule is used to screen out data that does not meet the preset quality conditions.

[0355] A deletion module, configured to delete the deletable conversation information from the historical question sequence to obtain a processed historical question sequence;

[0356] The determining module may specifically be used for:

[0357] Determine a historical question subsequence composed of a preset number of adjacent historical question information from the processed historical question sequence.

[0358] Optionally, the dialogue corpus generation module 412 may specifically be used for:

[0359] Input the second question information into the large language model to obtain a second reply information output by the large language model for the second question information.

[0360] Concatenate the second question information and the second reply information after the conversation sequence to which the first question information belongs to obtain a multi-round dialogue corpus.

[0361] It can be understood that the above-mentioned modules refer to computer programs or program segments for performing one or more specific functions. In addition, the distinction of the above-mentioned modules does not mean that the actual program codes must also be separated.

[0362] The above is a schematic solution of an apparatus for generating multi-turn dialogue corpus according to this embodiment. It should be noted that the technical solution of the apparatus for generating multi-turn dialogue corpus belongs to the same concept as the technical solution of the method for generating multi-turn dialogue corpus described above. For the details not described in the technical solution of the apparatus for generating multi-turn dialogue corpus, reference can be made to the description of the technical solution of the method for generating multi-turn dialogue corpus above.

[0363] Based on the same idea, the embodiments of this specification also provide an apparatus corresponding to a method for training a large language model.

[0364] Figure 5 FIG. is a schematic structural diagram of an apparatus for training a large language model provided by an embodiment of this specification.

[0365] As Figure 5 shown, the apparatus may include:

[0366] A training corpus acquisition module 502, configured to acquire the multi-turn dialogue corpus generated in the method for generating multi-turn dialogue corpus as described above.

[0367] A model training module 504, configured to train the large language model to be trained by using the acquired multi-turn dialogue corpus.

[0368] Based on the same idea, the embodiments of this specification also provide an apparatus corresponding to a method for testing a large language model.

[0369] Figure 6 FIG. is a schematic structural diagram of an apparatus for training a large language model provided by an embodiment of this specification.

[0370] As Figure 6 shown, the apparatus may include:

[0371] A test corpus acquisition module 602, configured to acquire the multi-turn dialogue corpus generated in the method for generating multi-turn dialogue corpus as described above.

[0372] A model testing module 604, configured to test the large language model to be tested by using the acquired multi-turn dialogue corpus.

[0373] Based on the same idea, the embodiments of this specification also provide a device corresponding to the method for generating multi-turn dialogue corpus described above.

[0374] Figure 7 FIG. is a schematic structural diagram of a device for generating multi-turn dialogue corpus provided by an embodiment of this specification. As Figure 7 shown, the device 700 may include:

[0375] At least one processor 710; and,

[0376] A memory 730 communicatively connected to the at least one processor; wherein,

[0377] The memory 730 stores instructions 720 executable by the at least one processor 710, and when the instructions are executed by the at least one processor 710, the at least one processor 710 is enabled to:

[0378] Obtain first query information.

[0379] Perform vectorization processing on the first query information to obtain a first query vector.

[0380] Query a vector database for target context vectors whose similarity to the first query vector meets a preset similarity condition; the vector database stores multiple pairs of data; for any pair of data, it includes a context vector generated based on the query context during the user's historical interaction with the language model, and the query follow-up text during the historical interaction.

[0381] Obtain question generation constraint information based on the query follow-up text corresponding to the target context vector.

[0382] Invoke a large language model to generate second query information after the first query information using the question generation constraint information as a constraint condition.

[0383] Generate multi-turn dialogue corpus based on the first query information and the second query information.

[0384] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method for training a large language model.

[0385] Figure 8 It is a schematic structural diagram of a device for training a large language model provided by the embodiments of this specification. As Figure 8 shown, the device 800 may include:

[0386] At least one processor 810; and,

[0387] A memory 830 communicatively connected to the at least one processor; wherein,

[0388] The memory 830 stores instructions 820 executable by the at least one processor 810, and when the instructions are executed by the at least one processor 810, the at least one processor 810 is enabled to:

[0389] Train the large language model to be trained using the multi-turn dialogue corpus generated by the method of generating multi-turn dialogue corpus described above.

[0390] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method for testing large language models.

[0391] Figure 9 FIG. is a schematic structural diagram of a device for generating multi-turn dialogue corpus provided by an embodiment of this specification. As Figure 9 shown, the device 900 may include:

[0392] At least one processor 910; and,

[0393] A memory 930 communicatively connected to the at least one processor; wherein,

[0394] The memory 930 stores instructions 920 executable by the at least one processor 910, and when the instructions are executed by the at least one processor 910, the at least one processor 910 is enabled to:

[0395] Use the multi-turn dialogue corpus generated by the method for generating multi-turn dialogue corpus described above to test the trained large language model.

[0396] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for devices and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The devices and equipment provided by the embodiments of this specification correspond to the methods, so the devices and equipment also have beneficial technical effects similar to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding devices and equipment will not be elaborated here.

[0397] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0398] In the 1990s, it was quite obvious to distinguish whether an improvement to a technology was an improvement in hardware (e.g., improvement to the circuit structure of diodes, transistors, switches, etc.) or an improvement in software (improvement to the method flow). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to the hardware circuit structure. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. The designer can program by themselves to "integrate" a digital character system on a piece of PLD without asking the chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not only one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0399] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0400] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0401] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0402] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0403] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block of the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart flow or flows and / or block or blocks. Figure 1 in a flow or flows and / or block or blocks Figure 1 or blocks.

[0404] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flowchart flow or flows and / or block or blocks. Figure 1 in a flow or flows and / or block or blocks Figure 1 or blocks.

[0405] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart flow or flows and / or block or blocks. Figure 1 in a flow or flows and / or block or blocks Figure 1 or blocks.

[0406] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0407] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read only memory (ROM) or flash memory. The memory is an example of a computer-readable medium.

[0408] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0409] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0410] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0411] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A method for generating a multi-round dialogue corpus, comprising: Get the first question information; Performing vectorization processing on the first question information to obtain a first question vector; Querying a target context vector whose similarity to the first question vector satisfies a preset similarity condition from a vector database; the vector database stores a plurality of data pairs; any one of the data pairs includes a context vector generated based on a question context of a user in a historical interaction process with a language model, and a question context in the historical interaction process; Obtaining question generation constraint information based on the question context corresponding to the target context vector; Calling a large language model to generate second question information after the first question information by using the question generation constraint information as a constraint condition; Based on the first question information and the second question information, multiple rounds of dialogue data are generated.

2. The method according to claim 1, wherein obtaining the first question information comprises: Obtaining a dialogue sequence; the dialogue sequence includes P rounds of dialogue information sorted by occurrence time; Any round of the dialogue information includes a user question information and a model reply information generated in response to the user question information; P is a positive integer greater than or equal to 2; The user question information in the last Q rounds of dialogue information of the dialogue sequence is determined as the first question information; Q is a positive integer less than or equal to P and less than or equal to the round threshold.

3. The method according to claim 2, wherein determining the user question information in the last Q rounds of dialogue information of the dialogue sequence as the first question information specifically comprises: The user question information in the last round of dialogue information of the dialogue sequence is determined as the first question information.

4. The method according to claim 1, wherein the target context vectors specifically include M target context vectors, where M is a positive integer; and obtaining the question generation constraint information based on the question context corresponding to the target context vectors specifically includes: Obtaining M question contexts corresponding to the M target context vectors from a vector database; The M question contexts correspond one-to-one to the M target context vectors; For any question context among the M question contexts, the any question context is spliced ​​to the dialogue sequence to which the first question information belongs, so as to obtain a pending dialogue sequence corresponding to the any question context; Calculating the perplexity of the pending dialogue sequence corresponding to each of the M question contexts; The perplexity is used to characterize the coherence of the pending dialogue sequence; The lower the perplexity, the more coherent the pending dialogue sequence is; From the M question contexts, determine N question contexts whose perplexity satisfies a preset perplexity condition; N is a positive integer less than or equal to M; Based on the N question contexts, question generation constraint information is obtained.

5. The method according to claim 4, wherein the calculating the perplexity of the pending dialogue sequence corresponding to each of the M question contexts comprises: For any of the pending dialogue sequences, performing word segmentation processing on the pending dialogue sequence to obtain a character sequence; Determining the occurrence probability of each character contained in the character sequence; The perplexity of the pending dialogue sequence is calculated based on the occurrence probability of each character in the character sequence.

6. The method according to claim 5, wherein determining the occurrence probability of each character contained in the character sequence comprises: Based on the character sequence, generating probability determination prompt information for input into a large language model; The probability determination prompt information includes a conditional probability expression of each character in the character sequence and probability determination task information; The probability determination task information is used to instruct to calculate the occurrence probability of each character in the character sequence according to the conditional probability expression; The probability determination prompt information is input into a large language model to obtain the occurrence probability of each character contained in the character sequence output by the large language model.

7. The method according to claim 5, wherein the step of calculating the perplexity of the pending dialogue sequence based on the occurrence probability of each character in the character sequence comprises: The perplexity of the pending dialogue sequence is calculated according to the following formula: Wherein, PPL represents perplexity; N represents the number of characters in the character sequence; P(ω i |ω <i ) represents the occurrence probability of the i-th character in the character sequence.

8. The method according to claim 1, wherein obtaining the question generation constraint information based on the question context corresponding to the target context vector specifically comprises: Generate question generation constraint information including the question context and constraint task information; The constraint task information is used to instruct to generate question information similar to the question context.

9. The method according to claim 1, wherein calling the large language model to generate the second question information after the first question information using the question generation constraint information as a constraint condition specifically comprises: Based on the question generation constraint information, generating question generation prompt information for input into the large language model; The question generation prompt information includes the dialogue sequence to which the first question information belongs, the question generation constraint information and the question generation task information; The question generation task information is used to indicate the user question information after the dialogue sequence is generated using the question generation constraint information as a constraint condition; The question generation prompt information is input into a large language model to obtain second question information subsequent to the first question information output by the large language model.

10. The method according to claim 1, before searching the vector database for a target context vector whose similarity with the first question vector satisfies a preset similarity condition, further comprising: Get historical question sequence; The user question sequence includes at least two historical question information sorted by question time; From the historical question sequence, determine a historical question subsequence consisting of adjacent historical question information of a preset number of questions; the preset number of questions is less than or equal to the round threshold; the historical question subsequence includes the question context and the question context; the question context includes the last question information in the historical question subsequence; the question context includes other question information in the historical question subsequence except the last question information; Performing vectorization processing on the question context to obtain the context vector; The preceding context vector is associated with the question context and stored in the vector database.

11. The method according to claim 10, wherein obtaining the historical question sequence comprises: Acquire a historical dialogue sequence generated during the historical interaction between the user and the language model; the historical dialogue sequence includes at least two rounds of historical dialogue information sorted by question time; any round of the historical dialogue information includes one historical question information and model historical reply information generated in response to the one historical question information; A user question sequence consisting of historical question information in each round of historical dialogue information is extracted from the historical dialogue sequence.

12. The method according to claim 10, wherein determining, from the historical question sequence, a historical question subsequence consisting of adjacent historical question information of a preset number of questions comprises: A historical question information group consisting of two adjacent historical question information is determined from the historical question sequence.

13. The method according to claim 10, before determining from the historical question sequence a historical question subsequence consisting of adjacent historical question information of a preset number of questions, further comprising: Using a preset data screening rule, screening out the deletable dialogue information from the historical question sequence; The data screening rule is used to screen out data whose quality does not meet the preset quality conditions; deleting the deletable dialogue information from the historical question sequence to obtain a processed historical question sequence; Determining a historical question subsequence consisting of adjacent historical question information of a preset number of questions from the historical question sequence specifically includes: From the processed historical question sequence, a historical question subsequence consisting of adjacent historical question information of a preset number of questions is determined.

14. The method according to claim 1, wherein generating multiple rounds of dialogue data based on the first question information and the second question information comprises: Inputting the second question information into a large language model to obtain second reply information output by the large language model in response to the second question information; After the second question information and the second reply information are spliced ​​into the dialogue sequence to which the first question information belongs, a multi-round dialogue data is obtained.

15. A method for training a large language model, comprising: The large language model to be trained is trained using the multi-round dialogue corpus generated by the method according to any one of claims 1 to 14.

16. A method for testing a large language model, comprising: The trained large language model is tested using the multi-round dialogue corpus generated by the method according to any one of claims 1 to 14.

17. A device for generating a multi-round dialogue corpus, comprising: A question information acquisition module, used to acquire first question information; A vectorization processing module, used for performing vectorization processing on the first question information to obtain a first question vector; A vector recall module, configured to query a target context vector whose similarity to the first question vector satisfies a preset similarity condition from a vector database; the vector database stores a plurality of data pairs; any set of the data pairs includes a context vector generated based on the question context of the user in the historical interaction process with the language model, and the question context in the historical interaction process; A constraint information generation module, used to obtain question generation constraint information based on the question context corresponding to the target context vector; A question information generating module, used for calling a large language model to generate second question information after the first question information by using the question generation constraint information as a constraint condition; The dialogue corpus generation module is used to generate multiple rounds of dialogue corpus based on the first question information and the second question information.

18. A device for training a large language model, comprising: A training corpus acquisition module, used to acquire the multi-round dialogue corpus generated by the method according to any one of claims 1 to 14; The model training module is used to train the large language model to be trained by using the acquired multi-round dialogue corpus.

19. A device for testing a large language model, comprising: A test corpus acquisition module, used to acquire the multi-round dialogue corpus generated by the method according to any one of claims 1 to 14; The model testing module is used to test the trained large language model using the acquired multi-round dialogue corpus.

20. A device for generating a multi-turn dialogue corpus, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Get the first question information; Performing vectorization processing on the first question information to obtain a first question vector; Querying a target context vector whose similarity to the first question vector satisfies a preset similarity condition from a vector database; the vector database stores a plurality of data pairs; any one of the data pairs includes a context vector generated based on a question context of a user in a historical interaction process with a language model, and a question context in the historical interaction process; Obtaining question generation constraint information based on the question context corresponding to the target context vector; Calling a large language model to generate second question information after the first question information by using the question generation constraint information as a constraint condition; Based on the first question information and the second question information, multiple rounds of dialogue data are generated.

21. A device for training a large language model, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: The large language model to be trained is trained using the multi-round dialogue corpus generated by the method according to any one of claims 1 to 14.

22. A device for testing a large language model, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: The trained large language model is tested using the multi-round dialogue corpus generated by the method according to any one of claims 1 to 14.

Citation Information

Cited By

  • Training data generation method, electronic equipment, storage medium and program product

    CN120632467A

  • A training data generation method, electronic equipment, storage medium and program product

    CN120632467B

  • Training data generation method, electronic equipment, storage medium and product

    CN120994799A

  • AI large model multi-round dialogue deduplication method and device and electronic equipment

    CN121166871A