Model distillation method, reply information generation method and device

By dividing the sample reply information of the large language model into inference process and answer information, and combining the prediction information of the small language model for training, the problem of insufficient inference in the processing of complex problems is solved, and the effect of in-depth reasoning and efficient answer generation is achieved.

CN120218182APending Publication Date: 2025-06-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510330312.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing model distillation methods lack deep reasoning capabilities when dealing with complex problems, resulting in efficient use of simple answers but insufficient reasoning when dealing with complex problems.

Method used

By obtaining the sample reply information output by the large language model, it is divided into inference process information and answer information, and using the initial small language model to obtain predictive inference process information and predictive answer information, and training the model with sample information to achieve deep inference and answer generation of the model.

Benefits of technology

The processing efficiency and accuracy of small language models when processing different tasks is improved, so that they can provide detailed inference processes and quick response answers while ensuring the depth of inference, solving the balance problem between efficient inference and deep inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218182A_ABST
    Figure CN120218182A_ABST
Patent Text Reader

Abstract

The invention discloses a model distillation method and a reply information generation method and device, and relates to the technical field of computers, in particular to the artificial intelligence fields of natural language processing, deep learning, large models and the like. According to the specific implementation scheme, sample question information and sample reply information of the sample question information are obtained; wherein the sample reply information comprises sample reasoning process information and sample answer information; according to the sample problem information, adopting an initial small language model to obtain a prediction reasoning process corresponding to the sample problem information; wherein the model scale of the initial small language model is smaller than that of the large language model; according to the sample question information, adopting an initial small language model to obtain prediction answer information of the sample question information; and training the initial small language model according to the sample reasoning process information, the sample answer information, the predicted reasoning process information and the predicted answer information to obtain a trained small language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, especially to artificial intelligence fields such as natural language processing, deep learning, and large models. Specifically, it relates to a model distillation method, a response information generation method, and a device. Background Art

[0002] In the field of artificial intelligence, knowledge of large language models can be transferred to small language models through model distillation, aiming to improve the inference efficiency and performance of the models while maintaining performance similar to that of large language models. Summary of the Invention

[0003] This application provides a model distillation method, a response information generation method, and a device. The specific solutions are as follows:

[0004] According to one aspect of this application, a model distillation method is provided, including:

[0005] Obtain sample question information and sample response information corresponding to the sample question information; wherein, the sample response information includes sample reasoning process information and sample answer information, and the sample response information is output by a large language model after processing the sample question information;

[0006] According to the sample question information, use an initial small language model to obtain predicted reasoning process information corresponding to the sample question; wherein, the model scale of the initial small language model is smaller than that of the large language model;

[0007] According to the sample question information, use the initial small language model to obtain predicted answer information for the sample question information;

[0008] Train the initial small language model according to the sample reasoning process information, the sample answer information, the predicted reasoning process information, and the predicted answer information to obtain a trained small language model.

[0009] According to another aspect of this application, a response information generation method is provided, including:

[0010] Obtain information of the question to be processed;

[0011] According to the information of the question to be processed, use the inference algorithm of the small language model to obtain response information for the information of the question to be processed; wherein, the small language model is obtained by using the above distillation method.

[0012] According to another aspect of this application, a model distillation device is provided, including:

[0013] A first acquisition module, configured to acquire sample question information and sample response information of the sample question information; wherein, the sample response information includes sample reasoning process information and sample answer information, and the sample response information is output by a large language model after processing the sample question information;

[0014] A second acquisition module, configured to acquire predicted reasoning process information corresponding to the sample question information by using an initial small language model according to the sample question information; wherein, the model scale of the initial small language model is smaller than that of the large language model;

[0015] A third acquisition module, configured to acquire predicted answer information of the sample question information by using the initial small language model according to the sample question information;

[0016] A training module, configured to train the initial small language model according to the sample reasoning process information, the sample answer information, the predicted reasoning process information and the predicted answer information, so as to obtain a trained small language model.

[0017] According to another aspect of the present application, there is provided a response information generation device, including:

[0018] A first acquisition module, configured to acquire information of a question to be processed;

[0019] A second acquisition module, configured to acquire response information of the question to be processed by using an inference algorithm of a small language model according to the information of the question to be processed; wherein, the small language model is obtained by using the above-mentioned distillation method.

[0020] According to another aspect of the present application, there is provided an electronic device, including:

[0021] At least one processor; and

[0022] A memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the above embodiments.

[0024] According to another aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described in the above embodiments.

[0025] According to another aspect of the present application, there is provided a computer program product, including a computer program, and the computer program realizes the steps of the method described in the above embodiments when executed by a processor.

[0026] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings are used to better understand the present solution and do not constitute a limitation to the present application. Among them:

[0028] Figure 1 is a schematic flowchart of the model distillation method provided by an embodiment of the present application;

[0029] Figure 2 is a schematic flowchart of the model distillation method provided by another embodiment of the present application;

[0030] Figure 3 is a schematic flowchart of the model distillation method provided by another embodiment of the present application;

[0031] Figure 4 is a schematic flowchart of the reply information generation method provided by an embodiment of the present application;

[0032] Figure 5 is a schematic flowchart of the reply information generation method provided by another embodiment of the present application;

[0033] Figure 6 is a schematic structural diagram of the model distillation device provided by an embodiment of the present application;

[0034] Figure 7 is a schematic structural diagram of the reply information generation device provided by an embodiment of the present application;

[0035] Figure 8 is a block diagram of the electronic device for implementing the model distillation method of the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The following makes an explanation of the exemplary embodiments of the present application in conjunction with the drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted below.

[0037] It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solution of the present application all comply with the relevant regulations of national laws and regulations and do not violate public order and good customs.

[0038] The model distillation method, response information generation method, device, electronic device, and storage medium according to the embodiments of the present application will be described below with reference to the accompanying drawings.

[0039] In some embodiments, the model can be simplified by learning the answer part output by the large language model. However, this distillation method can make the small language model more efficient in providing simple answers, but lacks sufficient depth reasoning ability when dealing with complex problems.

[0040] Based on this, the embodiments of the present application propose a model distillation method. Figure 1 It is a schematic flowchart of the model distillation method provided by an embodiment of the present application.

[0041] The model distillation method according to the embodiments of the present application can be executed by the model distillation device according to the embodiments of the present application, and this device can be configured in an electronic device.

[0042] Among them, the electronic device can be any device with computing power, such as a personal computer, a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc., which are hardware devices with various operating systems, touch screens, and / or display screens.

[0043] As Figure 1 shown, the model distillation method includes:

[0044] Step 101, obtain sample question information and sample response information of the sample question information.

[0045] Among them, the sample response information can include sample reasoning process information and sample answer information. Exemplarily, the sample response information can be the response information output after processing the sample question information by using a large language model.

[0046] That is to say, in the present application, the output of the large language model can be divided into two parts: the reasoning part and the answer part. Among them, the reasoning part can be the intermediate steps, reasoning links, and internal reasoning logic in the reasoning process of the model. The output of the reasoning part includes how the model reasons through multiple intermediate reasoning steps, which is usually more complex and has a large amount of calculation. The answer part is the concise and clear answer or conclusion of the final output. The answer part is usually compressed and refined after a long time of reasoning, directly answering the user's query, and has a simpler structure.

[0047] Step 102, according to the sample question information, use the initial small language model to obtain the predicted reasoning process information corresponding to the sample question information.

[0048] Among them, the model scale of the small language model is smaller than that of the large language model. Exemplarily, the model scale can be measured by the number of parameters of the model, and the number of parameters of the small language model is less than that of the large language model. Optionally, the number of parameters of the small language model and the large language model can both be greater than a preset number.

[0049] In this application, the training tasks of the small language model can include an inference task and an answer task, and the small language model can be trained for these two tasks in parallel. Exemplarily, the inference task can be called a thinking task.

[0050] Among them, the inference task can refer to that during the training process, the small language model has to learn the inference link of the large language model, such as how to gradually generate intermediate inference steps from the input information and then derive the final conclusion.

[0051] The answer task can require the small language model to learn how to generate the final concise answer from the output answers of the large language model. The focus of the answer task training is on learning how to quickly generate accurate conclusions without the intermediate steps of in-depth reasoning.

[0052] For the inference task, in this application, an initial small language model can be used to obtain the solution steps of the sample question information, that is, to obtain the predicted inference process information. Among them, the predicted inference process information can refer to the inference process output by the small language model for the sample question information.

[0053] Exemplarily, hint information for indicating the output of the inference process information can be obtained based on the sample question information, and the small language model can be used to process the hint information to obtain the predicted inference process information.

[0054] For example, by inputting "Please output the solution steps or problem-solving ideas of question Q1" into the small language model, the predicted inference process corresponding to Q1 can be obtained.

[0055] Step 103, according to the sample question information, use the initial small language model to obtain the predicted answer information of the sample question information.

[0056] Among them, the predicted answer information can refer to the answer output by the small language model for the sample question information.

[0057] For the reply task, in this application, the sample question information can be directly input into the small language model to enable the small language model to answer the sample question information and output the predicted answer information, or alternatively, hint information for indicating the small language model to output the answer information can be obtained based on the sample question information, and the small language model can be used to process the hint information to obtain the predicted answer information.

[0058] Step 104: Train the initial small language model based on the sample inference process information, sample answer information, predicted inference process information, and predicted answer information to obtain a trained small language model.

[0059] In this application, the model loss can be determined based on the predicted inference process information, predicted answer information, sample inference process information, and sample answer information. According to the model loss, the parameters of the initial small language model are adjusted, and the small language model with adjusted parameters is continuously trained until the training end condition is met, obtaining a trained small language model, thereby realizing the distillation of the large language model. Thus, the small language model can learn the inference algorithm of the large language model and the ability to generate answers.

[0060] Among them, the training end condition can be that the number of training times reaches a preset number, or the model loss is less than a preset threshold, or other conditions, which are not limited herein.

[0061] The model distillation method of this application can be widely applied to multiple fields. For example, the trained small language model can be used to construct a more intelligent and accurate question answering system, or applied to an intelligent search engine, or an intelligent customer service system, etc.

[0062] For example, in a question answering scenario, the question input by the user can be obtained, and the trained small language model is used to process the question input by the user to obtain the reply information output by the small language model.

[0063] Again, the text-form question input by the user can be obtained, and the trained small language model is used to process the text used to describe the question to obtain the reply information output by the small language model.

[0064] Furthermore, the question input by the user can be obtained, and the description of the question includes an image. The trained small language model can be used to recognize the image and process the recognized content to obtain the reply information output by the small language model.

[0065] In the embodiments of the present application, the sample response information output by the large language model for the sample question information is divided into sample inference process information and sample answer information. For the sample question information, the initial small language model is used to obtain the predicted inference process information and the predicted answer information respectively. Based on the predicted inference process information, the predicted answer information, the sample inference process information and the sample answer information, the initial small language model is trained. Thus, by separately obtaining the predicted inference process information and the predicted answer information through the initial small language model, and combining the sample inference process information and the sample answer information to train the initial small language model, the small language model can not only learn the inference process information of the sample question information, but also learn the sample answer information of the sample question information. Therefore, through multi-task learning, the small language model can achieve deep inference and answer generation, improving the accuracy of the response information output by the model.

[0066] Figure 2 It is a schematic flowchart of the model distillation method provided by another embodiment of the present application.

[0067] As Figure 2 shown, the model distillation method includes:

[0068] Step 201, obtain the sample question information and the sample response information of the sample question information.

[0069] Step 202, according to the sample question information, use the initial small language model to obtain the predicted inference process information corresponding to the sample question information.

[0070] Step 203, according to the sample question information, use the initial small language model to obtain the predicted answer information of the sample question information.

[0071] In the present application, steps 201 - 203 can adopt any implementation manner in the embodiments of the present application, so details are not described herein again.

[0072] Step 204, determine the first loss according to the difference between the predicted inference process information and the sample inference process information.

[0073] In the present application, for the inference task, the first loss corresponding to the inference task can be determined according to the predicted inference process information and the sample inference process information by using a loss function. Among them, the loss function can be, for example, mean square error or cross-entropy loss, etc., or other loss functions, which can be determined according to actual needs.

[0074] Step 205, determine the second loss according to the difference between the predicted answer information and the sample answer information.

[0075] In this application, for the answer task, the second loss corresponding to the answer task can be determined using a loss function based on the sample inference process information and the predicted inference process information. The loss functions used to calculate the second loss and the first loss can be the same or different, and this is not limited.

[0076] Step 206: Train the initial small language model based on the first loss and the second loss to obtain a trained small language model.

[0077] As a possible implementation, the first loss and the second loss can be added together to obtain the sum of the two losses. Based on the sum of the losses, the parameters of the initial small language model are adjusted, and the small language model with adjusted parameters is continuously trained until the training end condition is met, obtaining a trained small language model.

[0078] Considering that the weights of the inference task and the answer task in model training may be different, as another possible implementation, the first weight corresponding to the first loss and the second weight corresponding to the second loss can be determined. Based on the first weight and the second weight, the first loss and the second loss are weighted to obtain the total loss. Based on the total loss, the initial small language model is trained to obtain a trained small language model. Thus, by adjusting the weights of the first loss and the second loss, the learning focus of the model can be adjusted, thereby meeting different training requirements and improving the accuracy of the model.

[0079] Among them, the first weight and the second weight may be the same or different, and this is not limited.

[0080] Since different problem complexities may be different, the inference process of some problems may be relatively simple, and the inference process of some problems may be relatively complex. Therefore, exemplarily, the first attribute information of the sample problem information can be determined, and based on the first attribute information, the first complexity of the sample problem information is determined. Then, based on the first complexity, the first weight and the second weight are determined.

[0081] Exemplarily, a mapping relationship between complexity and the two losses can be established in advance, and the mapping relationship is queried based on the first complexity to determine the first weight and the second weight.

[0082] Exemplarily, the first attribute information may include but is not limited to the type, length, number of sub - problems, etc. of the sample problem information. For example, if the type of the sample problem information is a query type, the first complexity is relatively low, and the first weight can be less than the second weight. Another example is that if the sample problem is relatively long, the first complexity is relatively high, and the first weight can be greater than the second weight, etc.

[0083] Therefore, based on the attribute information of the sample question information, the complexity of the sample question information is determined. Based on the complexity of the sample question information, the weights of the loss of the inference task and the weights of the loss of the answer task are determined, improving the accuracy of the weights and thus improving the accuracy of the model.

[0084] Due to different application scenarios, the requirements for the inference ability of the model may vary. For example, in scenarios such as mathematical calculation, the requirements for the inference ability of the model are relatively high, while in simple query scenarios, the requirements for the inference ability of the model are not high. Therefore, exemplarily, the application scenario corresponding to the initial small language model can be determined, and based on the application scenario, the first weight and the second weight are determined.

[0085] For example, a mapping relationship between different application scenarios and the two losses can be established in advance. Then, based on this mapping relationship, the first weight and the second weight can be determined.

[0086] Therefore, by determining the weights of the two losses according to the application scenario of the small language model, the accuracy of the weights can be improved, and the requirements for the small language model in this application scenario can be met.

[0087] In the implementation of this application, through training the thinking task based on the prediction inference process information and the sample inference process information, and training the answer task based on the prediction answer information and the sample answer information, through the multi-task training of the thinking task and the answer task, the small language model can improve the processing efficiency of different tasks while ensuring the inference depth, and achieve the dual capabilities of providing a detailed inference process and quickly responding in the same model.

[0088] Moreover, through the multi-task distillation scheme of the inference task and the answer task, the adaptability of the model in multi-task and multi-scenario can be effectively improved, and the business access speed and efficiency of the model are greatly improved.

[0089] Figure 3 It is a schematic flowchart of the model distillation method provided by another embodiment of this application.

[0090] As Figure 3 shown, the model distillation method includes:

[0091] Step 301, obtain the sample question information and the sample reply information of the sample question information.

[0092] In this application, step 301 can adopt any implementation manner in the various embodiments of this application, so it will not be elaborated here.

[0093] Step 302, fill the inference task prompt template according to the sample question information and the sample answer information to generate inference task prompt information.

[0094] Among them, the inference task prompt template can refer to the prompt template used to indicate the information of the inference process output by the model. Exemplarily, the inference task prompt template can include slot information such as question information and answers. For example, the inference task prompt template is "The answer to the question [] is [], please explain how to derive this answer".

[0095] In this application, according to the sample question information and sample answer information, the corresponding slots in the inference task prompt template can be filled to obtain the inference task prompt information.

[0096] Among them, the inference task prompt information can be used to indicate that the model outputs the inference process for the sample question information.

[0097] For example, the inference task prompt information is "The answer to the question Q1 is D1, please explain how to derive this answer".

[0098] Another example, the inference task prompt information is "The answer to the question Q1 is D1, please describe the specific steps to obtain this answer".

[0099] To improve accuracy, exemplarily, in addition to the sample question information and sample answer information, the inference task prompt information can also include information such as inference examples and model output requirements. Among them, the inference examples can include reference questions, answers to reference questions, and inference processes, etc., so that the initial small language model can output the inference process information corresponding to the sample question information by referring to the inference examples.

[0100] Step 303, use the initial small language model to process the inference task prompt information to obtain the predicted inference process information.

[0101] In this application, the inference task prompt information can be input into the initial small language model, and the initial small language model can be used to identify the inference task prompt information to determine the processing purpose, and retrieve based on the sample question information and sample answer information in the inference task prompt information. Based on the retrieved knowledge content and processing purpose, the predicted inference process information is obtained.

[0102] Step 304, according to the sample question information, use the initial small language model to obtain the predicted answer to the sample question information.

[0103] In this application, step 304 can adopt any implementation manner in the various embodiments of this application, so it will not be elaborated here.

[0104] To improve the processing efficiency of the model, exemplarily, according to the sample question information, the answer task prompt template can be filled to obtain the answer task prompt information, and the answer task prompt information is input into the initial small language model. The initial small language model is used to process the answer task prompt information to obtain the predicted answer information.

[0105] Among them, the answer task prompt template can be a prompt template for instructing the model to output the answer to the question. Exemplarily, the answer task prompt template can include slot information such as questions.

[0106] Exemplarily, according to the sample question information, the corresponding slots in the answer task prompt model can be filled to obtain the answer task prompt information.

[0107] Among them, the answer task prompt information can be used to instruct the model to output the answer information of the sample question information. For example, the answer task prompt information is "Please give the answer to question Q1".

[0108] Optionally, to improve the accuracy, the answer task prompt information can include, in addition to the sample question information, information such as answer examples and model output requirements. Among them, the answer examples can include reference questions and the answers to the reference questions, so that the initial small language model can use the answer examples as references.

[0109] Thus, based on the answer task prompt information, the initial small language model can determine the requirements for outputting answers, improving the accuracy of the output results.

[0110] Step 305, train the initial small language model according to the sample inference process information, sample answer information, predicted inference process information, and predicted answer information to obtain a trained small language model.

[0111] In this application, step 305 can adopt any implementation manner in the embodiments of this application, so it will not be elaborated here.

[0112] Since the inference process is the process of deriving the answer, optionally, the parameters of the initial small language model can also be adjusted according to the difference between the predicted inference process information and the sample inference process information, the difference between the predicted answer information and the sample answer information, and the difference between the answer information derived from the predicted inference process information and the sample answer information, thereby improving the accuracy of the model.

[0113] In the embodiments of the present application, by filling in the inference task prompt template according to the sample question information and the sample answer information, the inference task prompt information is obtained. Based on the inference task prompt information, the initial small language model outputs the predicted inference process information for the sample question information. Thus, through the inference task prompt information, the small language model is guided to focus on the intermediate inference process and output the inference process information, which can facilitate the small language model to better learn the inference process and improve the inference ability of the model. For example, it can improve the ability to gradually generate intermediate inference steps from the input information of the small language model.

[0114] Figure 4 It is a schematic flowchart of a response information generation method provided by an embodiment of the present application.

[0115] As Figure 4 shown, the response information generation method includes:

[0116] Step 401, obtain the problem information to be processed.

[0117] In the present application, the user can input a question in the interaction interface of the artificial intelligence service based on the small language model, so as to obtain the problem information to be processed. Alternatively, the problem information to be processed can also be obtained by other means, which is not limited herein.

[0118] Step 402, according to the problem information to be processed, adopt the inference algorithm of the small language model to obtain the response information for the problem information to be processed.

[0119] Among them, the small language model can be trained by using the distillation method described in any of the above embodiments, so that the small language model can learn the inference algorithm of the large language model. Among them, the inference algorithm can be used to gradually generate intermediate inference steps from the input information of the small language model, and then derive the final conclusion. That is to say, the small language model can learn the inference ability of the large language model.

[0120] In the present application, the problem information to be processed can be input into the small language model, and the inference algorithm of the small language model can be used to answer the problem information to be processed, so as to obtain the response information for the problem information to be processed.

[0121] Among them, the response information can include the answer information for the problem information to be processed, or include the inference process information and the answer information, etc.

[0122] Exemplarily, the small language model can be used to identify the problem information to be processed to determine the type of the problem information to be processed, and according to the type of the problem information to be processed, the problem information to be processed is processed to output the response information matching the type of the problem information to be processed.

[0123] For example, if the type of the problem information to be processed is a query type, the small language model can directly output the answer. If the type of the problem information to be processed is a mathematical calculation type, the small language model can output the inference process information and the answer information.

[0124] In the embodiment of the present application, the small language model trained by the above distillation method can perform in-depth reasoning. Therefore, by using the inference algorithm of the small language model to process the problem information to be processed, the accuracy of the reply information can be improved.

[0125] Figure 5 It is a schematic flowchart of a reply information generation method provided by another embodiment of the present application.

[0126] Such as Figure 5 shown, the reply information generation method includes:

[0127] Step 501, obtain the problem information to be processed.

[0128] In the present application, step 501 can adopt any implementation manner in the embodiments of the present application, so it will not be elaborated here.

[0129] Step 502, determine the processing mode of the problem information to be processed.

[0130] Among them, the processing mode may include an in-depth reasoning mode, a quick response mode, etc.

[0131] Among them, in the in-depth reasoning mode, the small language model performs reasoning and outputs the reasoning process and the answer. The in-depth reasoning mode can be applicable to scenarios that require explanations or reasoning processes, such as the solution of complex problems or the processing of multi-round conversations, etc.

[0132] Among them, in the quick response mode, the small language model can directly generate and output the answer to the question. The quick response mode can be applicable to scenarios such as simple queries or direct requirements.

[0133] As a possible implementation manner, the operation information for the target control in the interaction interface can be obtained, and the processing mode can be determined according to the operation information.

[0134] Among them, the target control can be a control related to the processing mode.

[0135] Exemplarily, the target control can be a control related to the in-depth reasoning mode. If a trigger operation on the target control is detected, the processing mode can be determined to be the in-depth reasoning mode. If no trigger operation on the target control is detected, the processing mode can be determined to be the quick response mode.

[0136] For example, after the user enters a question in the interaction interface and triggers the "Deep Thinking" control in the interaction interface, it can be determined that the processing mode of the question is the deep reasoning mode. If the user enters a question in the interaction interface and triggers the question submission control without triggering the "Deep Thinking" control, it can be determined that the processing mode of the question is the quick response mode.

[0137] Exemplarily, the target control may include a control related to the deep reasoning mode and a control related to the quick response mode. If a triggering operation on the control related to the deep reasoning mode is detected, it can be determined that the processing mode is the deep reasoning mode. If a triggering operation on the control related to the quick response mode is detected, it can be determined that the processing mode is the quick response mode.

[0138] Thus, by different operation information of the target control, switching between different processing modes is realized, so that the small language model can dynamically select the processing mode according to the operation information of the target control, and different question processing requirements can be met.

[0139] As another possible implementation, the second attribute information of the problem information to be processed can be determined, and according to the second attribute information, the second complexity of the problem information to be processed can be determined. According to the second complexity, the processing mode of the problem information to be processed can be determined.

[0140] Among them, the second attribute information may include but is not limited to the type, length, number of sub-questions, etc. of the problem information to be processed. For example, if the type of the problem information to be processed is a query type, the second complexity is relatively low, and the processing mode is determined to be the quick response mode. Another example is that if the problem to be processed is relatively long and the second complexity is relatively high, the processing mode is determined to be the deep reasoning mode.

[0141] Thus, based on the attribute information of the problem information to be processed, the complexity of the problem information to be processed is determined, and based on the complexity, the processing mode is determined, so that the processing mode can be flexibly adjusted according to the complexity of the problem, and different question processing requirements can be met.

[0142] Step 503, according to the processing mode and the problem information to be processed, use the inference algorithm of the small language model to obtain the reply information.

[0143] In this application, different processing modes may have different prompt templates. According to the problem information to be processed and the prompt template corresponding to the processing mode, the reply generation prompt information corresponding to the processing mode can be obtained, and the inference algorithm of the small language model is used to process the reply generation prompt information to obtain the reply information in the processing mode.

[0144] Among them, the prompt template corresponding to the processing mode may be a prompt template for instructing the model to output reply information matching the processing mode.

[0145] Exemplarily, according to the information of the problem to be processed, processing examples corresponding to the processing mode, etc., corresponding slots in the prompt template corresponding to the processing mode can be filled to obtain the reply generation prompt information corresponding to the processing mode.

[0146] Among them, the processing examples may include reference questions, reply information of reference questions, etc.

[0147] Among them, the reply generation prompt information corresponding to the processing mode can be used to instruct the model to output the reply information of the problem information to be processed in the processing mode.

[0148] In some embodiments, if the processing mode is the deep inference mode, according to the problem information to be processed and the deep inference mode, the inference algorithm of the small language model can be adopted to obtain the inference process information for the problem information to be processed, and according to the inference process information, the reply information can be obtained.

[0149] Exemplarily, the answer information can be determined according to the inference process information, and the reply information can be determined according to the inference process information and the answer information. Among them, the reply information may include the inference process information and the reply information.

[0150] Exemplarily, if the processing mode is the deep inference mode, according to the problem information to be processed and the prompt template corresponding to the deep inference mode, the reply generation prompt information corresponding to the deep inference mode can be obtained, and the small language model is used to process the reply generation prompt information to obtain the inference process information, and the answer information is obtained according to the inference process information. Among them, the answer information may include the inference process information and the reply information of the answer.

[0151] In some embodiments, if the processing mode of the problem information to be processed is the quick response mode, such as a query-type problem, the small language model can directly generate the answer information of the problem information to be processed without calling the inference algorithm and output the answer information, which can improve the response speed and save resources. Thus, the small language model can not only implement deep inference under complex tasks but also improve the response speed under simple tasks. Therefore, based on the reply generation prompt information corresponding to the processing mode, the small language model can be used to obtain the reply information matching the processing mode, thereby improving the accuracy of the reply information in different processing modes.

[0152] In the embodiments of the present application, by determining the processing mode of the problem information to be processed, according to the problem information to be processed, combined with the processing mode, and using the small language model to obtain the reply information of the problem information to be processed, the accuracy of the reply information can be improved.

[0153] Optionally, to improve the accuracy of the response information, if the complexity of the problem information to be processed is relatively low, such as query-type problems, etc., the inference algorithm of the small language model can also be used to obtain the inference process information for the problem information to be processed, determine the answer information for the problem information to be processed according to the inference process information, and output the answer information.

[0154] The model distillation method of the embodiments of the present application, through the multi-task distillation scheme of the inference task and the answer task, can not only separate the deep inference ability and the fast response ability in the model inference process, but also flexibly switch the processing mode to adapt to different application scenarios, ensuring that the model can provide accurate answers when deep inference is required and efficient answers when fast response is required, and solving the balance problem between efficient inference and deep inference.

[0155] In addition, based on the multi-task training method of the inference task and the answer task, a mode switching mechanism can be introduced, enabling the model to dynamically select the processing mode to meet different business requirements.

[0156] If the problem requires higher accuracy and more complex inferences, the model can enter the deep inference mode, where the model can learn and perform multi-step inferences, and finally output a detailed answer after deep inference. If the problem is relatively simple or requires a quick response, the model can enter the fast response mode, quickly generate and output the answer, and may not perform complex inferences. The introduction of this mode switching mechanism can flexibly adjust the processing strategy according to the requirements in different task environments, thereby improving the application efficiency and response speed of the model.

[0157] To implement the above embodiments, the embodiments of the present application also propose a model distillation device. Figure 6 It is a schematic structural diagram of the model distillation device provided by an embodiment of the present application.

[0158] As Figure 6 shown, the model distillation device 600 includes:

[0159] A first acquisition module 610, configured to acquire sample problem information and sample response information of the sample problem information; wherein, the sample response information includes sample inference process information and sample answer information, and the sample response information is output by a large language model after processing the sample problem information;

[0160] A second acquisition module 620, configured to use an initial small language model to acquire predicted inference process information corresponding to the sample problem information according to the sample problem information; wherein, the model scale of the initial small language model is smaller than the model scale of the large language model;

[0161] A third acquisition module 630, configured to obtain predicted answer information for the sample question information by using the initial small language model according to the sample question information;

[0162] A training module 640, configured to train the initial small language model according to the sample inference process information, the sample answer information, the predicted inference process information, and the predicted answer information, so as to obtain a trained small language model.

[0163] Optionally, the training module 640 is configured to:

[0164] Determine a first loss according to the difference between the predicted inference process information and the sample inference process information;

[0165] Determine a second loss according to the difference between the predicted answer information and the sample answer information;

[0166] Train the initial small language model according to the first loss and the second loss, so as to obtain the trained small language model.

[0167] Optionally, the training module 640 is configured to:

[0168] Determine a first weight corresponding to the first loss and a second weight corresponding to the second loss;

[0169] Weight the first loss and the second loss according to the first weight and the second weight to obtain a total loss;

[0170] Train the initial small language model according to the total loss, so as to obtain the trained small language model.

[0171] Optionally, the training module 640 is configured to:

[0172] Determine first attribute information of the sample question information;

[0173] Determine a first complexity of the sample question information according to the first attribute information;

[0174] Determine the first weight and the second weight according to the first complexity.

[0175] Optionally, the training module 640 is configured to:

[0176] Determine an application scenario corresponding to the initial small language model;

[0177] Determine the first weight and the second weight according to the application scenario.

[0178] Optionally, the second acquisition module 620 is configured to:

[0179] Fill the inference task prompt template according to the sample question information and the sample answer information to generate inference task prompt information;

[0180] Use the initial small language model to process the inference task prompt information to obtain the predicted inference process.

[0181] Optionally, the third acquisition module 630 is used for:

[0182] Fill the answer task prompt template according to the sample question information to obtain answer task prompt information;

[0183] Use the initial small language model to process the answer task prompt information to obtain the predicted answer information.

[0184] It should be noted that the above explanation of the model distillation method embodiment also applies to the model distillation device of this embodiment, so it will not be repeated here.

[0185] In the embodiment of the present application, by dividing the sample reply information output by the large language model for the sample question information into sample inference process information and sample answer information, for the sample question information, the initial small language model is respectively used to obtain the predicted inference process information and the predicted answer information, and based on the predicted inference process information, the predicted answer information, the sample inference process information and the sample answer information, the initial small language model is trained. Thus, by separately obtaining the predicted inference process information and the predicted answer information through the initial small language model, and combining the sample inference process information and the sample answer information to train the initial small language model, the small language model can not only learn the inference process information of the sample question information, but also learn the sample answer information of the sample question information. Therefore, through multi-task learning, the small language model can achieve in-depth inference and answer generation, improving the accuracy of the reply information output by the model.

[0186] To implement the above embodiment, the embodiment of the present application also proposes a reply information generation device. Figure 7 It is a structural schematic diagram of a reply information generation device provided in an embodiment of the present application.

[0187] As Figure 7 shown, the reply information generation device 700 includes:

[0188] The first acquisition module 710 is used to acquire the problem information to be processed;

[0189] A second acquisition module 720, configured to obtain a reply message for the to-be-processed problem information by using an inference algorithm of a small language model, where the small language model is obtained by using the distillation method described in any of the foregoing embodiments.

[0190] Optionally, the second acquisition module 720 is configured to:

[0191] Determine a processing mode for the to-be-processed problem information;

[0192] In response to the processing mode being a deep inference mode, obtain inference process information for the to-be-processed problem information by using the inference algorithm of the small language model according to the deep inference mode and the to-be-processed problem information;

[0193] Obtain the reply message according to the inference process information.

[0194] Optionally, the second acquisition module 720 is configured to:

[0195] Obtain operation information for a target control in an interaction interface;

[0196] Determine the processing mode according to the operation information.

[0197] Optionally, the second acquisition module 720 is configured to:

[0198] Determine second attribute information for the to-be-processed problem information;

[0199] Determine a second complexity for the to-be-processed problem information according to the second attribute information;

[0200] Determine the processing mode according to the second complexity.

[0201] Optionally, the second acquisition module 720 is configured to:

[0202] Generate reply generation prompt information corresponding to the deep inference mode according to the to-be-processed problem information and a prompt template corresponding to the deep inference mode;

[0203] Process the reply generation prompt information by using the inference algorithm of the small language model to obtain inference process information.

[0204] It should be noted that the explanations of the foregoing embodiments of the reply message generation method are also applicable to the reply message generation device of this embodiment, and thus will not be elaborated herein.

[0205] In the embodiments of the present application, the small language model trained by the above distillation method can perform in-depth reasoning. Therefore, by adopting the reasoning algorithm of the small language model to process the problem information to be processed, the accuracy of the reply information can be improved.

[0206] According to an embodiment of the present application, the present application further provides an electronic device, a readable storage medium, and a computer program product.

[0207] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present application described and / or claimed herein.

[0208] As Figure 8 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 802 or a computer program loaded from a storage unit 808 into a RAM (Random Access Memory) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An I / O (Input / Output) interface 805 is also connected to the bus 804.

[0209] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0210] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the model distillation method. For example, in some embodiments, the model distillation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the model distillation method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the model distillation method by any other suitable means (e.g., by means of firmware).

[0211] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, FPGA (Field Programmable Gate Array), ASIC (Application-Specific Integrated Circuit), ASSP (Application Specific Standard Product), SOC (System On Chip), CPLD (Complex Programmable Logic Device), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0212] The program code for implementing the methods of this application can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0213] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0214] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0215] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend, middleware, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0216] A computer system can include clients and servers. The clients and servers are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (Virtual Private Server). The server can also be a server of a distributed system, or a server combined with blockchain.

[0217] It should be noted that the electronic device for implementing the reply information generation method of the embodiments of the present application is similar to the above-mentioned electronic device, so it will not be elaborated here.

[0218] According to an embodiment of the present application, the present application also provides a computer program product, which, when executed by an instruction processor in the computer program product, executes the model distillation method or the reply information generation method proposed in the above embodiments of the present application.

[0219] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present application can be achieved, and no limitation is imposed herein.

[0220] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.

Claims

1. A model distillation method, comprising: Acquire sample question information and sample answer information of the sample question information; wherein the sample answer information includes sample reasoning process information and sample answer information, and the sample answer information is output after the large language model processes the sample question information; According to the sample question information, an initial small language model is used to obtain prediction reasoning process information corresponding to the sample question information; wherein the model scale of the initial small language model is smaller than the model scale of the large language model; According to the sample question information, using the initial small language model, obtaining predicted answer information of the sample question information; The initial small language model is trained according to the sample reasoning process information, the sample answer information, the predicted reasoning process information and the predicted answer information to obtain a trained small language model.

2. The method of claim 1, wherein: The training of the initial small language model according to the sample reasoning process information, the sample answer information, the predicted reasoning process information and the predicted answer information to obtain a trained small language model includes: Determining a first loss according to a difference between the prediction reasoning process information and the sample reasoning process information; determining a second loss according to a difference between the predicted answer information and the sample answer information; The initial small language model is trained according to the first loss and the second loss to obtain the trained small language model.

3. The method of claim 2, wherein: The step of training the initial small language model according to the first loss and the second loss to obtain the trained small language model includes: Determine a first weight corresponding to the first loss and a second weight corresponding to the second loss; According to the first weight and the second weight, weighting the first loss and the second loss to obtain a total loss; The initial small language model is trained according to the total loss to obtain the trained small language model.

4. The method of claim 3, wherein: The determining a first weight corresponding to the first loss and a second weight corresponding to the second loss includes: Determine first attribute information of the sample question information; Determining a first complexity of the sample question information according to the first attribute information; The first weight and the second weight are determined according to the first complexity.

5. The method of claim 3, wherein: The determining a first weight corresponding to the first loss and a second weight corresponding to the second loss includes: Determine an application scenario corresponding to the initial small language model; According to the application scenario, the first weight and the second weight are determined.

6. The method according to any one of claims 1 to 5, wherein: The step of using an initial small language model according to the sample question information to obtain prediction reasoning process information corresponding to the sample question information includes: Filling a reasoning task prompt template according to the sample question information and the sample answer information to generate reasoning task prompt information; The initial small language model is used to process the reasoning task prompt information to obtain the predictive reasoning process information.

7. The method according to any one of claims 1 to 5, wherein: The step of acquiring predicted answer information of the sample question information by using the initial small language model according to the sample question information includes: Filling the answer task prompt template according to the sample question information to obtain answer task prompt information; The initial small language model is used to process the answer task prompt information to obtain the predicted answer information.

8. A method for generating a reply message, comprising: Get information about pending issues; According to the question information to be processed, an inference algorithm of a small language model is used to obtain answer information of the question information to be processed; wherein the small language model is obtained by using the method described in any one of claims 1-7.

9. The method of claim 8, wherein: The step of using a small language model to obtain answer information of the problem information to be processed according to the problem information to be processed includes: Determine a processing mode for the problem information to be processed; In response to the processing mode being the deep reasoning mode, according to the deep reasoning mode and the problem information to be processed, using the reasoning algorithm of the small language model, obtaining reasoning process information for the problem information to be processed; The answer information is obtained according to the reasoning process information.

10. The method of claim 9, wherein: The determining of the processing mode of the problem information to be processed includes: Get the operation information of the target control in the interactive interface; The processing mode is determined according to the operation information.

11. The method of claim 9, wherein: The determining of the processing mode of the problem information to be processed includes: Determine second attribute information of the problem information to be processed; Determining a second complexity of the problem information to be processed according to the second attribute information; The processing mode is determined according to the second complexity.

12. The method of claim 9, wherein: The acquiring, according to the deep reasoning mode and the problem information to be processed, reasoning process information for the problem information to be processed using the reasoning algorithm of the small language model includes: Generate prompt information for answer generation corresponding to the deep reasoning mode according to the problem information to be processed and the prompt template corresponding to the deep reasoning mode; The answer generation prompt information is processed using the inference algorithm of the small language model to obtain the inference process information.

13. A model distillation apparatus comprising: A first acquisition module is used to acquire sample question information and sample answer information of the sample question information; wherein the sample answer information includes sample reasoning process information and sample answer information, and the sample answer information is output after the large language model processes the sample question information; A second acquisition module is used to acquire prediction and reasoning process information corresponding to the sample question information by using an initial small language model according to the sample question information; wherein the model scale of the initial small language model is smaller than the model scale of the large language model; A third acquisition module is used to acquire predicted answer information of the sample question information by using the initial small language model according to the sample question information; A training module is used to train the initial small language model according to the sample reasoning process information, the sample answer information, the predicted reasoning process information and the predicted answer information to obtain a trained small language model.

14. A reply information generating device, comprising: The first acquisition module is used to obtain information about problems to be processed; The second acquisition module is used to use the inference algorithm of the small language model to obtain the answer information of the question to be processed according to the question information to be processed; wherein the small language model is obtained by using the method described in any one of claims 1-7.

15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.

17. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.

Citation Information

Cited By

  • Training data synthesis method and device based on error extrapolation and inference chain analysis, medium and program product

    CN120611192A

  • Smart park information generation method and device, equipment, storage medium and product

    CN121636695A