Intelligent dialogue method, related device, equipment and storage medium
Through knowledge distillation technology, the knowledge of the second dialogue model is passed on to the first dialogue model, which solves the problems of high resource requirements and long inference time of intelligent dialogue model, and achieves more efficient intelligent dialogue.
Patent Information
- Application Number
- CN202411904688.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-09
AI Technical Summary
The existing intelligent dialogue model has a long inference time due to the large amount of parameters and high demand for computing resources. How to reduce resource requirements and shorten inference time has become an urgent problem.
Through knowledge distillation technology, the second dialogue model with more parameters and the first dialogue model with less parameters are trained based on sample dialogue data. The first dialogue model is used as student model and the second dialogue model is used as teacher model. The knowledge distillation is used to reduce the burden on training data.
Through knowledge distillation technology, it is possible to reduce resource requirements and shorten reasoning time in the intelligent dialogue process, achieving more efficient intelligent dialogue.
Smart Images

Figure CN119962624A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing, and in particular to an intelligent dialogue method and related devices, equipment and storage media. Background Art
[0002] At present, the emergence of intelligent dialogue models such as large language models has brought revolutionary changes to natural language understanding, thanks to their greatly improved semantic understanding ability, instruction-following ability and multi-round response coherence compared to general machine learning models.
[0003] However, since large language models often have a large number of parameters and high computing resource requirements, the resource requirements of intelligent dialogue are extremely high and the reasoning time is often long. In view of this, how to reduce the resource requirements of intelligent dialogue and shorten the reasoning time of intelligent dialogue as much as possible has become an urgent problem to be solved. Summary of the invention
[0004] The main technical problem solved by the present application is to provide an intelligent dialogue method and related devices, equipment and storage media, which can reduce the resource requirements of intelligent dialogue as much as possible and shorten the reasoning time of intelligent dialogue.
[0005] In order to solve the above technical problems, the first aspect of the present application provides an intelligent dialogue method, including: obtaining a first sentence to be replied; inputting the first sentence into a first dialogue model to obtain an output sentence of the first dialogue model as a second sentence to reply to the first sentence; wherein the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with a second dialogue model having more parameters than the first dialogue model, and the second dialogue model selects candidate dialogue data from a candidate dialogue set as sample dialogue data, the candidate dialogue data including a first sample sentence and a second sample sentence to reply to the first sample sentence, and the second sample sentence is the output sentence after the first sample sentence is input into the second dialogue model.
[0006] In order to solve the above technical problems, the second aspect of the present application provides an intelligent dialogue device, including: a sentence acquisition module and a sentence reply module, the sentence acquisition module is used to acquire a first sentence to be replied; the sentence reply module is used to input the first sentence into a first dialogue model to obtain the output sentence of the first dialogue model as a second sentence to reply to the first sentence; wherein the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with a second dialogue model having more parameters than the first dialogue model, the second dialogue model selects candidate dialogue data from the candidate dialogue set as sample dialogue data, the candidate dialogue data includes a first sample sentence and a second sample sentence to reply to the first sample sentence, and the second sample sentence is the output sentence after the first sample sentence is input into the second dialogue model.
[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, which at least includes a memory and a processor coupled to each other, the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the intelligent dialogue method in the above first aspect.
[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium, which stores program instructions that can be executed by a processor, and the program instructions are used to implement the intelligent dialogue method of the first aspect.
[0009] In the above scheme, a first sentence to be replied is obtained, and the first sentence is input into the first dialogue model to obtain the output sentence of the first dialogue model as the second sentence to reply the first sentence, and the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with a second dialogue model having more parameters than the first dialogue model, and the second dialogue model selects candidate dialogue data from the candidate dialogue set as sample dialogue data, and the candidate dialogue data includes the first sample sentence and the second sample sentence to reply the first sample sentence, and the second sample sentence is the output sentence after the first sample sentence is input into the second dialogue model. Therefore, on the one hand, the sample dialogue data applicable to knowledge distillation is generated by the second dialogue model and further selected, which can reduce the burden of training data in the training process and help reduce the demand for computing resources. On the other hand, since the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with the second dialogue model, and the number of parameters of the first dialogue model is less than that of the second dialogue model, the first dialogue model with fewer parameters can be used to reply the first sentence in the intelligent dialogue process. Therefore, the resource demand of intelligent dialogue can be reduced as much as possible and the reasoning time of intelligent dialogue can be shortened. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is a flow chart of an embodiment of the intelligent dialogue method of the present application; Figure 2 This is a process diagram of an embodiment of the intelligent dialogue method of the present application; Figure 3 It is a schematic diagram of the framework of an embodiment of the intelligent dialogue device of the present application; Figure 4 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application; Figure 5 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0011] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.
[0012] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.
[0013] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the fragment " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.
[0014] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the intelligent dialogue method of the present application. Specifically, it may include the following steps: Step S11: Obtain the first sentence to be replied.
[0015] In one implementation scenario, as a possible example, the first sentence can be input by voice, such as the first sentence can be input by voice through a microphone or other sound pickup device. It should be noted that in the case of input by voice, in order to facilitate the subsequent dialogue model processing, the first sentence in voice form can be firstly subjected to voice recognition to obtain the first sentence in text form, so as to facilitate the subsequent dialogue model to process the first sentence in text form. Of course, the first sentence in voice form can also be processed by the dialogue model later. As another possible example, the first sentence can also be input by text, such as the first sentence can be input by input methods such as Pinyin and Wubi. The above examples are only several possible examples of the input method of the first sentence, and do not limit other possible input methods of the first sentence. The input methods of the first sentence will not be given one by one here.
[0016] In an implementation scenario, the specific content of the first sentence is related to the dialogue task, and the specific content of the first sentence is not limited here. For example, in the case where the dialogue task is a mathematical question and answer, the specific content of the first sentence can be related to mathematics, such as including but not limited to: "Please briefly describe the Pythagorean theorem", "Given the lengths of the three sides of a triangle, how to find the angle between any two sides of the triangle" and so on; or, in the case where the dialogue task is a reasoning question and answer, the specific content of the first sentence can be related to reasoning, such as including but not limited to: "Xiao Zhang, Xiao Wang, and Xiao Li are watching a movie. Xiao Zhang sits next to Xiao Wang, but not next to Xiao Li. Who is sitting next to Xiao Li" and so on; or, in the case where the dialogue task is code generation, the specific content of the first sentence can be related to code, such as including but not limited to: "Please generate a piece of code, requiring the input parameter to be the radius of a circle, and the output to be the area of a circle" and so on; or, in the case where the dialogue task is text generation, the specific content of the first sentence can be related to copywriting, such as including but not limited to: "Please generate a speech on the development of artificial intelligence technology" and so on. The above examples are only several possible examples of the first sentence, and do not limit the specific content of the first sentence when the dialogue task is the above example task or other possible situations. The specific content of the first sentence will not be given one by one here.
[0017] Step S12: inputting a first sentence into the first dialogue model to obtain an output sentence of the first dialogue model as a second sentence in response to the first sentence.
[0018] Specifically, after obtaining the first sentence, the first sentence can be input into the first dialogue model, and the first dialogue model can process the first sentence, so that the output sentence of the first dialogue model can be obtained as the second sentence to reply to the first sentence. It should be noted that the first dialogue model may include but is not limited to: convolutional neural network, related neural network based on Transformer, etc., and the network structure of the first dialogue model is not limited here. For example, in the case where the dialogue task is a math question-answering task, if the first sentence is "Please briefly describe the Pythagorean theorem", the second sentence may include but is not limited to: "In a right triangle, the sum of the squares of the two sides of the right angle is equal to the square of the hypotenuse. Specifically, if the lengths of the two right-angled sides of the right triangle are a and b respectively, and the length of the hypotenuse is c, then the relationship between them can be expressed by the following mathematical formula: a²+ b² = c²"; or, in the case where the dialogue task is a reasoning question-answering task, if the first sentence is "Xiao Zhang, Xiao Wang, and Xiao Li are watching a movie. Xiao Zhang sits next to Xiao Wang, but not next to Xiao Li. Who is sitting next to Xiao Li?", the second sentence may include but is not limited to: "Xiao Wang is sitting next to Xiao Li"; or, in the case where the dialogue task is code generation, if the first sentence is "Please generate a piece of code that requires the input parameter to be the radius of a circle and the output to be the area of a circle", the second sentence may include but is not limited to the following: The following is a simple Python code example that calculates and prints the area of a circle of a given radius. The code uses the area formula for a circle: Area = π * Radius².
[0019] import math def calculate_circle_area(radius): # Calculate the area of the circle area = math.pi * (radius ** 2) return area # Enter the radius of the circle radius = float(input("Please enter the radius of the circle: ")) # Calculate and output the area of the circle area = calculate_circle_area(radius) print(f"The area of a circle with radius {radius} is: {area:.2f}") In this code: 1. We first imported the math module so that we can use the pi constant to represent pi.
[0020] 2. Define a function calculate_circle_area, which accepts a parameter radius (the radius of the circle), And returns the calculated area of the circle.
[0021] 3. Use the input function to get the radius of the circle from the user and convert it to a floating point number.
[0022] 4. Call the calculate_circle_area function to calculate the area of the circle, and use the print function to output the result. Output to the console. When outputting, we use a format string to limit the number of digits after the decimal point to two.
[0023] You can copy this code into your Python environment and run it, and enter a circle radius to see the result. fruit.
[0024] It should be noted that the above examples are merely possible examples of the specific content of the second sentence in several dialogue tasks, and other possible situations will not be given one by one here.
[0025] In the disclosed embodiment, the first dialogue model can be obtained by knowledge distillation training based on sample dialogue data with a second dialogue model having more parameters than the first dialogue model. That is, in the knowledge distillation process, the second dialogue model can be used as a teacher model, and the first dialogue model can be used as a student model. In addition, as a possible example, the second dialogue model can include but is not limited to a large language model, such as the second dialogue model can include but is not limited to open source large models such as LLAMA, Bloom, etc., or the second dialogue model can also be obtained by fine-tuning parameters based on an open source large model based on a specific corpus; or the second dialogue model can also be a custom large model, and the specific source and specific structure of the second dialogue model when the second dialogue model is a large language model are not limited here.
[0026] In the embodiment of the present disclosure, the second dialogue model can select candidate dialogue data from the candidate dialogue set as sample dialogue data, and the candidate dialogue data can specifically include a first sample sentence and a second sample sentence that replies to the first sample sentence, and the second sample sentence is an output sentence after the first sample sentence is input into the second dialogue model. It should be noted that the specific meaning and possible content of the first sample sentence and the second sample sentence can refer to the aforementioned description of the first sentence and the second sentence, which will not be repeated here.
[0027] In one implementation scenario, as a possible example, a first sample sentence can be obtained in advance, such as from the Internet, a related data set, etc., which is not limited here. On this basis, the first sample can be input into the second dialogue model to obtain the output sentence of the second dialogue model as the second sample sentence that replies to the first sample sentence, and the first sample sentence and the second sample sentence that replies to the first sample sentence can be used as candidate dialogue data. Or, as another possible example, for each dialogue task, the first sample sentence can be obtained in advance, and then the first sample sentence can be input into the second dialogue model to obtain the output sentence of the second dialogue model as the second sample sentence that replies to the first sample sentence, and the first sample sentence and the second sample sentence that replies to the first sample sentence can be used as candidate dialogue data under the dialogue task. Exemplarily, still taking various dialogue tasks including math question and answer, reasoning question and answer, code generation, and text generation as an example, a number of candidate dialogue data under the dialogue task "mathematical problem" can be generated based on this, forming a candidate dialogue set under the dialogue task "mathematical problem", and a number of candidate dialogue data under the dialogue task "reasoning question and answer" can be generated based on this, forming a candidate dialogue set under the dialogue task "reasoning question and answer", and a number of candidate dialogue data under the dialogue task "code generation" can be generated based on this, forming a candidate dialogue set under the dialogue task "code generation", and a number of candidate dialogue data under the dialogue task "text generation" can be generated based on this, forming a candidate dialogue set under the dialogue task "text generation".
[0028] In an implementation scenario, after obtaining the candidate dialogue set, for each candidate dialogue data in the selected dialogue set, the first sample sentence can be input into the first dialogue model (i.e., the student model of knowledge distillation) to obtain the output sentence of the first dialogue model as the third sample sentence to reply to the first sample sentence, and the third sample sentence is evaluated based on the second dialogue model (i.e., the teacher model of knowledge distillation) to obtain the sentence score of the third sample sentence, and then based on the sentence score of the third sample sentence, it can be determined whether to select the candidate dialogue data where the first sample sentence to which the third sample sentence replies is located as the sample dialogue data. In the above manner, the third sample sentence output by the first dialogue model (i.e., the student model of knowledge distillation) in reply to the same first sample sentence is evaluated by the second dialogue model (i.e., the teacher model of knowledge distillation), which can reflect the real dialogue level of the first dialogue model, and the sample dialogue data is selected to train the first dialogue model, which can help to improve the dialogue response level of the first dialogue model in a targeted manner. Of course, the above screening process is only a possible example of screening candidate dialogue data in actual application, and other possible ways of screening candidate dialogue data are not limited here. For example, it is also possible to determine whether to filter the sample conversation data based on the decoding probability of each character in the second sample sentence when the second conversation model outputs the second sample sentence (for example, the average value of the decoding probability is lower than the probability threshold, etc.). Other possible ways of filtering the candidate conversation data will not be given one by one here.
[0029] In a specific implementation scenario, as a possible example, the third sample sentence can be evaluated from the perspectives of expression fluency, ambiguity, and pertinence to obtain a sentence score of the third sample sentence; or, as another possible example, as mentioned above, the second dialogue model can be a large language model, then an evaluation instruction can be constructed based on the third sample sentence, and the evaluation instruction is used to instruct the second dialogue model to evaluate the perplexity (PPL) of the third sample sentence. It should be noted that, in general, the lower the perplexity of the model output, the better the performance of the representation model. On the contrary, the second dialogue model (i.e., the teacher model of knowledge distillation) has a strong fitting ability, that is, its perplexity on the Golden data distribution is low, and its perplexity on the non-Golden data distribution is high. Therefore, by analyzing the perplexity level in different situations, the degree of fit of the model to the Golden data can be effectively judged, thereby providing a basis for evaluating the ability of the first dialogue model (i.e., the student model of knowledge distillation). On this basis, the evaluation instruction can be input to the second dialogue model to obtain the perplexity output by the second dialogue model as the sentence score of the third sample sentence. It should be noted that the higher the sentence score, the higher the representation perplexity, and vice versa, the lower the sentence score, the lower the representation perplexity. In addition, for the calculation method of perplexity, please refer to the technical details of perplexity, which will not be repeated here. In the above method, when the second dialogue model is a large language model, an evaluation instruction is constructed in combination with the third sample sentence to instruct the second dialogue model to evaluate the third sample sentence, thereby obtaining the perplexity of the third sample sentence as the sentence score of the third sample sentence, which can provide a reliable basis for the sentence evaluation of the third sample sentence with the help of the large language model. Of course, the above examples are only a few possible examples of evaluating the third sample sentence, and other possible evaluation methods will not be given one by one here.
[0030] In a specific implementation scenario, after obtaining the sentence score, each third sample sentence can be sorted according to the sentence score, and the candidate dialogue data where the first sample sentence is replied by the third sample sentence before a preset ratio is selected as the sample dialogue data. Exemplarily, still taking perplexity as the sentence score as an example, each third sample sentence can be sorted in descending order according to the sentence score, and the candidate dialogue data where the first sample sentence is replied by the third sample sentence before a preset ratio (e.g., the first 10%, 20%, etc.) is selected as the sample dialogue data. For example, if a third sample sentence "Who is sitting next to Xiao Li? Let's ask Xiao Li together" has a higher sentence score (i.e., a higher perplexity) and is located before the preset proportion after sorting, the candidate dialogue data containing the first sample sentence "Xiao Zhang, Xiao Wang, and Xiao Li are watching a movie. Xiao Zhang is sitting next to Xiao Wang, but not next to Xiao Li. Who is sitting next to Xiao Li" to which the third sample sentence responds can be selected as sample dialogue data (i.e., including the first sample sentence "Xiao Zhang, Xiao Wang, and Xiao Li are watching a movie. Xiao Zhang is sitting next to Xiao Wang, but not next to Xiao Li. Who is sitting next to Xiao Li" and the second sample sentence responding to the first sample sentence "Xiao Wang is sitting next to Xiao Li") can be selected as sample dialogue data. Of course, the above example is only a possible example of selecting candidate dialogue data as sample dialogue data based on sentence scores, and other possible situations will not be given examples one by one here. The above method sorts each third sample sentence according to the sentence score, and then selects the candidate dialogue data of the first sample sentence to which the third sample sentence before a preset ratio replies as the sample dialogue data. This can selectively screen out sample dialogue data for the weak capabilities of the first dialogue model (i.e., the student model of knowledge distillation), which helps to improve the dialogue ability of the first dialogue model after training.
[0031] In a specific implementation scenario, as described above, the second dialogue model pre-generates candidate dialogue sets under multiple dialogue tasks (such as math question-answering, reasoning question-answering, code generation, text generation, etc.). For each candidate dialogue set under each dialogue task, the steps for each candidate dialogue data in the candidate dialogue set can be respectively performed to screen and obtain sample dialogue data under the corresponding dialogue task. In other words, the sample dialogue data under each dialogue task can be obtained by referring to the aforementioned screening process. For example, the aforementioned screening process can be referred to to obtain several sample dialogue data under the dialogue task "math question-answering", several sample dialogue data under the dialogue task "reasoning question-answering", several sample dialogue data under the dialogue task "code generation", and several sample dialogue data under the dialogue task "text generation". Of course, the above example is only a possible example of screening sample dialogue data under various dialogue tasks, and other possible situations are not given one by one here. In the above method, the second dialogue model pre-generates candidate dialogue sets for multiple dialogue tasks. For each candidate dialogue set under each dialogue task, the steps for each candidate dialogue data in the candidate dialogue set are executed to screen and obtain sample dialogue data under the corresponding dialogue task. The screening and optimization process can be performed independently for each dialogue task, which helps to focus on improving the most challenging sample performance in various dialogue tasks respectively.
[0032] In an implementation scenario, after obtaining the sample dialogue data, knowledge distillation can be performed based on the sample dialogue data in combination with the first dialogue model (i.e., the student model of knowledge distillation) and the second dialogue model (i.e., the teacher model of knowledge distillation) to train the first dialogue model, so that the first dialogue model learns the knowledge content of the second dialogue model after training and has the dialogue ability level of the second dialogue model as much as possible. Specifically, the first sample sentence can be input to the first dialogue model (i.e., the student model of knowledge distillation) to at least obtain the first probability distribution (i.e., the logits distribution) when the first dialogue model outputs the sentence, and obtain the second probability distribution (i.e., the logits distribution) when the second dialogue model outputs the sentence after the first sample sentence is input to the second dialogue model (i.e., the teacher model of knowledge distillation). On this basis, the network parameters of the first dialogue model (i.e., the student model of knowledge distillation) can be adjusted based on at least the distribution difference between the first probability distribution and the second probability distribution. In the above manner, by constraining the output distribution to be as consistent as possible, the first dialogue model (i.e., the knowledge distillation) can learn the second dialogue model as much as possible.
[0033] In a specific implementation scenario, as a possible example, the difference between the first probability distribution and the second probability distribution can be directly measured by using, for example, KL divergence (Kullback-Leibler Divergence), which can be expressed as: …… (1) In the above formula (1), represents the first probability distribution of the first dialogue model (i.e., the student model of knowledge distillation), Represents the second probability distribution of the second dialogue model (i.e., the teacher model for knowledge distillation). represents the input data (i.e., the first sample sentence). As another possible example, different from the aforementioned measurement method, the first probability distribution and the second probability distribution may be weighted to obtain a weighted probability distribution, and then the network parameters of the first dialogue model may be adjusted based on at least the distribution difference between the weighted probability distribution and the second probability distribution. Exemplarily, the aforementioned KL divergence may be used to measure the difference between the weighted probability distribution and the second probability distribution, which may be expressed as: …… (2) In the above formula (2), represents the first probability distribution of the first dialogue model (i.e., the student model of knowledge distillation), represents the second probability distribution of the second dialogue model (i.e., the teacher model of knowledge distillation), The above method can improve the gradient stability by first weighting the first probability distribution and the second probability distribution to obtain a weighted probability distribution, and then measuring the distribution difference between the weighted probability distribution and the second probability distribution.
[0034] In a specific implementation scenario, in addition to measuring the training loss based on the distribution difference between the first probability distribution and the second probability distribution, the output statement of the first dialogue model after the first sample statement is input into the first dialogue model can be obtained during the training process as a fourth sample statement to reply to the first sample statement, so that the network parameters of the first dialogue model can be adjusted based on the distribution difference between the first probability distribution and the second probability distribution and the difference between the second sample statement and the fourth sample statement. Exemplarily, as described above, the distribution difference between the first probability distribution and the second probability distribution can be measured by, for example, KL divergence, and the difference between the second sample statement and the fourth sample statement can be measured based on a loss function such as cross entropy, so that the training loss L of knowledge distillation can be obtained based on these two losses: …… (3) In the above formula (3), represents the sub-loss obtained by measuring the difference between the second sample sentence and the fourth sample sentence through the cross entropy loss function, represents the sub-loss obtained by measuring the distribution difference between the first probability distribution and the second probability distribution, such as KL divergence, Represents the weighting coefficient of the above two losses.
[0035] In a specific implementation scenario, as described above, in the process of screening sample dialogue data, a number of sample dialogue data can be screened for each dialogue task for training the first dialogue model in the knowledge distillation stage. In this case, in each round of iteration of the first dialogue model, part of the sample dialogue data can be extracted from each dialogue task, so as to train the first dialogue data in this round of iteration based on the sample dialogue data.
[0036] In an implementation scenario, please refer to Figure 2 , Figure 2 This is a process diagram of an embodiment of the intelligent dialogue method of the present application. Figure 2 As shown, the first sample sentence is input to the second dialogue model (i.e., the teacher model of knowledge distillation), and the second sample sentence that responds to the first sample sentence can be obtained. At this time, the first sample sentence and the second sample sentence that responds to the first sample sentence can be used as selected dialogue data. The first dialogue model (i.e., the student model of knowledge distillation) also inputs the same first sample sentence, and the third sample sentence that responds to the first sample sentence can be obtained. Then, the second dialogue model can evaluate the third sample sentence to obtain the sentence score (e.g., perplexity) of the third sample sentence, and then select the candidate dialogue data where the first sample sentence that is responded to by the third sample sentence is located as the sample dialogue data. Based on this sample dialogue data, the first dialogue model (i.e., the student model of knowledge distillation) can be trained. The specific training process can refer to the above-mentioned related description, which will not be repeated here. It can be seen that the above process draws on the "teaching feedback" mechanism in the human education process and optimizes model distillation by imitating the interaction between teachers and students.
[0037] In the above scheme, a first sentence to be replied is obtained, and the first sentence is input into the first dialogue model to obtain the output sentence of the first dialogue model as the second sentence to reply the first sentence, and the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with a second dialogue model having more parameters than the first dialogue model, and the second dialogue model selects candidate dialogue data from the candidate dialogue set as sample dialogue data, and the candidate dialogue data includes the first sample sentence and the second sample sentence to reply the first sample sentence, and the second sample sentence is the output sentence after the first sample sentence is input into the second dialogue model. Therefore, on the one hand, the sample dialogue data applicable to knowledge distillation is generated by the second dialogue model and further selected, which can reduce the burden of training data in the training process and help reduce the demand for computing resources. On the other hand, since the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with the second dialogue model, and the number of parameters of the first dialogue model is less than that of the second dialogue model, the first dialogue model with fewer parameters can be used to reply the first sentence in the intelligent dialogue process. Therefore, the resource demand of intelligent dialogue can be reduced as much as possible and the reasoning time of intelligent dialogue can be shortened.
[0038] See also Figure 3 , Figure 3 It is a schematic diagram of the framework of an embodiment of the intelligent dialogue device of the present application. The intelligent dialogue device 30 includes: a sentence acquisition module 31 and a sentence reply module 32, the sentence acquisition module 31 is used to acquire a first sentence to be replied; the sentence reply module 32 is used to input the first sentence into the first dialogue model to obtain the output sentence of the first dialogue model as the second sentence to reply the first sentence; wherein the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with a second dialogue model having more parameters than the first dialogue model, the second dialogue model selects candidate dialogue data from the candidate dialogue set as sample dialogue data, the candidate dialogue data includes the first sample sentence and the second sample sentence to reply the first sample sentence, and the second sample sentence is the output sentence after the first sample sentence is input into the second dialogue model.
[0039] In the above scheme, the intelligent dialogue device 30 obtains the first sentence to be replied, inputs the first sentence into the first dialogue model, and obtains the output sentence of the first dialogue model as the second sentence to reply the first sentence, and the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with the second dialogue model having more parameters than the first dialogue model, and the second dialogue model selects candidate dialogue data from the candidate dialogue set as sample dialogue data, and the candidate dialogue data includes the first sample sentence and the second sample sentence to reply the first sample sentence, and the second sample sentence is the output sentence after the first sample sentence is input into the second dialogue model. Therefore, on the one hand, the sample dialogue data applicable to knowledge distillation is generated by the second dialogue model and further screened, which can reduce the burden of training data in the training process and help reduce the demand for computing resources. On the other hand, since the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with the second dialogue model, and the number of parameters of the first dialogue model is less than that of the second dialogue model, the first dialogue model with fewer parameters can be used to reply the first sentence in the intelligent dialogue process. Therefore, the resource demand of the intelligent dialogue can be reduced as much as possible and the reasoning time of the intelligent dialogue can be shortened.
[0040] In some disclosed embodiments, the intelligent dialogue device 30 includes a sentence evaluation module, which is used to input a first sample sentence into a first dialogue model for each candidate dialogue data in a candidate dialogue set, so as to obtain an output sentence of the first dialogue model as a third sample sentence that replies to the first sample sentence, and to evaluate the third sample sentence based on the second dialogue model to obtain a sentence score for the third sample sentence; the intelligent dialogue device 30 includes a data selection module, which is used to determine whether to select the candidate dialogue data containing the first sample sentence that is replied to by the third sample sentence as the sample dialogue data based on the sentence score of the third sample sentence.
[0041] In some disclosed embodiments, the data selection module includes a sorting submodule for sorting each third sample sentence according to the sentence score; the data selection module includes a selection submodule for selecting candidate conversation data containing the first sample sentence that is replied to the third sample sentence before a preset proportion as the sample conversation data.
[0042] In some disclosed embodiments, the second dialogue model is a large language model, and the sentence evaluation module includes a construction submodule for constructing an evaluation instruction based on a third sample sentence; wherein the evaluation instruction is used to instruct the second dialogue model to evaluate the perplexity of the third sample sentence; the sentence evaluation module includes a scoring submodule for inputting the evaluation instruction into the second dialogue model to obtain the perplexity output by the second dialogue model as the sentence score of the third sample sentence.
[0043] In some disclosed embodiments, the second dialogue model pre-generates candidate dialogue sets for multiple dialogue tasks, and the sentence evaluation module and the data selection module are specifically used to execute steps for each candidate dialogue data in the candidate dialogue set for each dialogue task, so as to screen and obtain sample dialogue data for the corresponding dialogue task.
[0044] In some disclosed embodiments, the multiple dialogue tasks include at least two of mathematical question answering, reasoning question answering, code generation, and text generation.
[0045] In some disclosed embodiments, the intelligent dialogue device 30 includes a distribution acquisition module, which is used to input a first sample sentence into a first dialogue model to at least obtain a first probability distribution when the first dialogue model outputs a sentence, and to obtain a second probability distribution when the second dialogue model outputs a sentence after the first sample sentence is input into the second dialogue model; the intelligent dialogue device 30 includes a parameter adjustment module, which is used to adjust the network parameters of the first dialogue model based on at least the distribution difference between the first probability distribution and the second probability distribution.
[0046] In some disclosed embodiments, the parameter adjustment module includes a weighting submodule for weighting based on the first probability distribution and the second probability distribution to obtain a weighted probability distribution; the parameter adjustment module includes an adjustment submodule for adjusting the network parameters of the first dialogue model based at least on the distribution difference between the weighted probability distribution and the second probability distribution.
[0047] In some disclosed embodiments, the intelligent dialogue device 30 includes a reply acquisition module, which is used to obtain an output statement of the first dialogue model after the first sample statement is input into the first dialogue model as a fourth sample statement in reply to the first sample statement; the parameter adjustment module is specifically used to adjust the network parameters of the first dialogue model based on the distribution difference between the first probability distribution and the second probability distribution and the difference between the second sample statement and the fourth sample statement.
[0048] See also Figure 4 , Figure 4 : is a schematic diagram of the framework of an embodiment of an electronic device of the present application. The electronic device 40 includes at least a memory 41 and a processor 42 coupled to each other, the memory 41 stores at least program instructions, and the processor 42 is used to execute the program instructions to implement the steps in any of the above-mentioned intelligent dialogue method embodiments. For details, please refer to the aforementioned disclosed embodiments, which will not be repeated here. As a possible example, the electronic device 40 may include but is not limited to an office notebook, an e-book reader, a smart phone, a tablet computer, etc., and the specific type of the electronic device 40 is not limited here.
[0049] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned intelligent dialogue method embodiments. The processor 42 can also be called a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 42 can be implemented by an integrated circuit chip.
[0050] In the above scheme, the electronic device 40 obtains the first sentence to be replied, inputs the first sentence into the first dialogue model, and obtains the output sentence of the first dialogue model as the second sentence to reply the first sentence, and the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with the second dialogue model having more parameters than the first dialogue model, and the second dialogue model selects candidate dialogue data from the candidate dialogue set as sample dialogue data, and the candidate dialogue data includes the first sample sentence and the second sample sentence to reply the first sample sentence, and the second sample sentence is the output sentence after the first sample sentence is input into the second dialogue model. Therefore, on the one hand, the sample dialogue data applicable to knowledge distillation is generated by the second dialogue model and further screened, which can reduce the burden of training data in the training process and help reduce the demand for computing resources. On the other hand, since the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with the second dialogue model, and the number of parameters of the first dialogue model is less than that of the second dialogue model, the first dialogue model with fewer parameters can be used to reply the first sentence in the intelligent dialogue process. Therefore, the resource demand of the intelligent dialogue can be reduced as much as possible and the reasoning time of the intelligent dialogue can be shortened.
[0051] See also Figure 5 , Figure 5 1 is a schematic diagram of a framework of an embodiment of a computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor, and the program instructions 51 are used to implement the steps in any of the above-mentioned intelligent dialogue method embodiments.
[0052] In the above scheme, the computer-readable storage medium 50 obtains a first sentence to be replied, inputs the first sentence into the first dialogue model, and obtains the output sentence of the first dialogue model as the second sentence to reply the first sentence, and the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with a second dialogue model having more parameters than the first dialogue model, and the second dialogue model selects candidate dialogue data from the candidate dialogue set as sample dialogue data, and the candidate dialogue data includes the first sample sentence and the second sample sentence to reply the first sample sentence, and the second sample sentence is the output sentence after the first sample sentence is input into the second dialogue model. Therefore, on the one hand, the sample dialogue data applicable to knowledge distillation is generated by the second dialogue model and further screened, which can reduce the burden of training data in the training process and help reduce the demand for computing resources. On the other hand, since the first dialogue model is obtained by knowledge distillation training based on sample dialogue data with the second dialogue model, and the number of parameters of the first dialogue model is less than that of the second dialogue model, the first dialogue model with fewer parameters can be used to reply the first sentence in the intelligent dialogue process. Therefore, the resource demand of the intelligent dialogue can be reduced as much as possible and the reasoning time of the intelligent dialogue can be shortened.
[0053] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0054] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.
[0055] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0056] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0057] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0058] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0059] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
Claims
1. An intelligent dialogue method, characterized in that: include: Get the first sentence to be replied; inputting the first sentence into a first dialogue model to obtain an output sentence of the first dialogue model as a second sentence in reply to the first sentence; The first dialogue model is obtained by knowledge distillation training based on sample dialogue data with a second dialogue model having more parameters than the first dialogue model, the second dialogue model selects candidate dialogue data from a candidate dialogue set as the sample dialogue data, the candidate dialogue data includes a first sample sentence and a second sample sentence in reply to the first sample sentence, the second sample sentence is an output sentence after the first sample sentence is input into the second dialogue model.
2. The method according to claim 1, characterized in that The step of screening the sample conversation data includes: For each of the candidate dialogue data in the candidate dialogue set, input the first sample sentence into the first dialogue model to obtain an output sentence of the first dialogue model as a third sample sentence in reply to the first sample sentence, and evaluate the third sample sentence based on the second dialogue model to obtain a sentence score of the third sample sentence; Based on the sentence score of the third sample sentence, it is determined whether to select the candidate dialogue data where the first sample sentence to which the third sample sentence replies belongs as the sample dialogue data.
3. The method according to claim 2, characterized in that The step of determining whether to select the candidate dialogue data containing the first sample sentence replied by the third sample sentence as the sample dialogue data based on the sentence score of the third sample sentence comprises: sorting the third sample sentences according to the sentence scores; The candidate dialogue data in which the first sample sentence that is replied to by the third sample sentence before a preset ratio is located is selected as the sample dialogue data.
4. The method according to claim 2, characterized in that: The second dialogue model is a large language model, and the evaluating the third sample sentence based on the second dialogue model to obtain a sentence score of the third sample sentence includes: Based on the third sample sentence, construct an evaluation instruction; wherein the evaluation instruction is used to instruct the second dialogue model to evaluate the perplexity of the third sample sentence; The evaluation instruction is input into the second dialogue model to obtain the perplexity output by the second dialogue model as a sentence score of the third sample sentence.
5. The method according to claim 2, characterized in that: The second dialogue model pre-generates the candidate dialogue sets under multiple dialogue tasks, and the method further includes: For the candidate dialogue set under each of the dialogue tasks, the step of filtering each of the candidate dialogue data in the candidate dialogue set to obtain the sample dialogue data under the corresponding dialogue task.
6. The method according to claim 5, characterized in that The multiple dialogue tasks include at least two of mathematical question answering, reasoning question answering, code generation, and text generation.
7. The method according to claim 1, characterized in that The training step of the first dialogue model includes: Inputting the first sample sentence into the first dialogue model to at least obtain a first probability distribution when the first dialogue model outputs a sentence, and obtaining a second probability distribution when the second dialogue model outputs a sentence after the first sample sentence is input into the second dialogue model; Based at least on a distribution difference between the first probability distribution and the second probability distribution, a network parameter of the first dialogue model is adjusted.
8. The method according to claim 7, characterized in that The adjusting the network parameters of the first dialogue model based at least on the distribution difference between the first probability distribution and the second probability distribution includes: weighting the first probability distribution and the second probability distribution to obtain a weighted probability distribution; Based at least on a distribution difference between the weighted probability distribution and the second probability distribution, a network parameter of the first dialogue model is adjusted.
9. The method according to claim 7, characterized in that: Before adjusting the network parameters of the first dialogue model based at least on the distribution difference between the first probability distribution and the second probability distribution, the method further includes: obtaining, after the first sample sentence is input into the first dialogue model, an output sentence of the first dialogue model as a fourth sample sentence in reply to the first sample sentence; The adjusting the network parameters of the first dialogue model based at least on the distribution difference between the first probability distribution and the second probability distribution includes: Based on the distribution difference between the first probability distribution and the second probability distribution and the difference between the second sample sentence and the fourth sample sentence, the network parameters of the first dialogue model are adjusted.
10. An intelligent dialogue device, characterized in that: include: A sentence acquisition module, used for acquiring a first sentence to be replied; a sentence reply module, configured to input the first sentence into the first dialogue model to obtain an output sentence of the first dialogue model as a second sentence replying to the first sentence; The first dialogue model is obtained by knowledge distillation training based on sample dialogue data with a second dialogue model having more parameters than the first dialogue model, the second dialogue model selects candidate dialogue data from a candidate dialogue set as the sample dialogue data, the candidate dialogue data includes a first sample sentence and a second sample sentence in reply to the first sample sentence, the second sample sentence is an output sentence after the first sample sentence is input into the second dialogue model.
11. An electronic device, characterized in that: The intelligent dialogue method comprises at least a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the intelligent dialogue method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the intelligent dialogue method described in any one of claims 1 to 9.
Citation Information
Cited By
Security reply generation method, related device, equipment and storage medium
CN120408414A
Speech quality evaluation model training method, speech quality evaluation result generating method and related products
CN120748377A