Question answering method, device and equipment based on task-oriented question answering model cluster distillation system

By obtaining target device parameters and task guidance factors, the training of student model clusters is solved, and the deployment problem of multi-question and answer tasks on resource-constrained devices of large language models is solved, achieving efficient compression and improved computing efficiency of the model.

CN120354862BActive Publication Date: 2025-08-19TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510828315.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-08-19
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Large language models are difficult to efficiently deploy on resource-constrained devices in multi-question and answer tasks scenarios. The existing distillation method has problems such as application limitations of single task, lack of task guidance mechanisms, and inefficiency in resource-constrained environments in multi-task learning.

Method used

By obtaining the target device parameter values, determining the target parameter information, and performing cluster distillation training on the student model cluster based on the task guidance factor of the multi-question task teacher model, the student model is optimized to adapt to the needs of multi-tasks, and the model compression and computing efficiency improvement are achieved.

Benefits of technology

The computing efficiency and deployment feasibility of multi-question and answer tasks are improved on resource-constrained devices, and the generalization ability and computing performance of student model clusters in multi-task scenarios are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354862B_ABST
    Figure CN120354862B_ABST
Patent Text Reader

Abstract

The present application proposes a question-answering method, apparatus and equipment based on a task-oriented guided question-answering model cluster distillation system, which relates to the field of large language model question-answering technology. By obtaining the device parameter value of the target device of the multi-question-answering task target model cluster to be deployed, and determining the target parameter information of the multi-question-answering task target model cluster, further based on the multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster, according to the target parameter information, the student model cluster to be trained is cluster distilled and trained to obtain the multi-question-answering task target model cluster, and finally the target multi-question-answering task from the target device is input into the multi-question-answering task target model cluster to obtain the answer to the target multi-question-answering task. This method optimizes the student model guided by each task by combining the knowledge of multiple teacher models and the task guidance factors, and realizes efficient model deployment under the premise of maintaining the collaborative reasoning ability of multiple question-answering tasks, thereby improving the computational efficiency and deployment feasibility of the question-answering model cluster in a resource-constrained environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of large language model question answering technology, and in particular to a question answering method, apparatus and equipment based on a task-oriented guided question answering model cluster distillation system. Background Art

[0002] Large language models (LLMs) are widely used in customer service, education, recommendation systems, healthcare, and other fields. They often need to simultaneously process multiple question-answering tasks based on user input. This multi-task reasoning requires efficient parallel computing capabilities, but this comes with high computational and storage requirements, making their deployment on resource-constrained devices such as mobile terminals and embedded systems a significant challenge.

[0003] How to achieve efficient model deployment while maintaining the collaborative reasoning capability of multiple question-answering tasks has become an urgent problem that needs to be solved. Summary of the Invention

[0004] The present application provides a question-answering method, apparatus, and device based on a task-oriented question-answering model cluster distillation system to solve the problem of achieving efficient model deployment while maintaining the collaborative reasoning capability of multiple question-answering tasks.

[0005] In a first aspect, the present application proposes a question answering method based on a task-oriented question answering model cluster distillation system, the method comprising:

[0006] Obtaining device parameter values of a target device of a target model cluster for a multi-question-answering task to be deployed, wherein the device parameter values include at least one of the following: storage space size and computing speed;

[0007] Determining target parameter information of the multi-question-answering task target model cluster according to the device parameter value, wherein the target parameter information includes at least one of the following: model parameter quantity, reasoning response speed, and reasoning calculation amount;

[0008] Based on multiple task guidance factors of each teacher model in the multi-question answering task teacher model cluster, cluster distillation training is performed on the to-be-trained student model cluster according to the target parameter information to obtain the multi-question answering task target model cluster;

[0009] The target multi-question-answering task from the target device is input into the multi-question-answering task target model cluster to obtain an answer to the target multi-question-answering task.

[0010] Optionally, the multi-question-answering task teacher model cluster includes K teacher models; and the method further includes:

[0011] Input the test multi-question answering task into the k-th teacher model, and obtain T answers output by the k-th teacher model, where the t-th answer is the answer to the t-th test question answering task in the test multi-question answering task;

[0012] determining an importance score of the k-th teacher model for the t-th test question-answering task based on a difference between the t-th answer and the t-th correct answer corresponding to the t-th test question-answering task;

[0013] According to the importance score of the k-th teacher model for the t-th test question-answering task and the importance score of the k-th teacher model for T test question-answering tasks, the guidance factor of the k-th teacher model for the t-th test question-answering task is determined.

[0014] Optionally, based on multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster, cluster distillation training is performed on the to-be-trained student model cluster according to the target parameter information, including:

[0015] Input the t-th sample question-answering task in the sample multi-question-answering task into the multi-question-answering task teacher model cluster, and obtain the answer of the k-th teacher model for the t-th sample question-answering task;

[0016] Obtaining, based on the answer of the k-th teacher model to the t-th sample question-answering task and the guidance factor of the k-th teacher model to the t-th test question-answering task, the answer of the multi-question-answering task teacher model cluster to the t-th sample question-answering task, wherein the t-th sample question-answering task and the t-th test question-answering task are of the same task type;

[0017] With the goal of learning the answers of the multi-question-answering task teacher model cluster to the sample multi-question-answering task, cluster distillation training is performed on the student model cluster to be trained according to the target parameter information.

[0018] Optionally, based on multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster, cluster distillation training is performed on the to-be-trained student model cluster according to the target parameter information, including:

[0019] Determining a distillation loss value for the t-th sample question-answering task based on a difference between an answer of the multi-question-answering task teacher model cluster for the t-th sample question-answering task and an answer of the multi-question-answering task student model cluster for the t-th sample question-answering task;

[0020] The total distillation loss is obtained based on the distillation loss of the t-th sample question-answering task and the guidance factors of the K teacher models for the t-th test question-answering task.

[0021] According to the total distillation loss value and the target parameter information, cluster distillation training is performed on the student model cluster to be trained.

[0022] Optionally, the multi-question-answering task student model cluster includes K student models; and the method further includes:

[0023] Share the guidance factor of the k-th teacher model for the t-th test question-answering task with the k-th student model;

[0024] Inputting the t-th sample question-answering task in the sample multi-question-answering task into the multi-question-answering task student model cluster, and obtaining the answer of the k-th student model to the t-th sample question-answering task;

[0025] According to the answer of the kth student model to the tth sample question-answering task and the guidance factor of the kth student model to the tth test question-answering task, the answer of the multi-question-answering task student model cluster to the tth sample question-answering task is obtained.

[0026] Optionally, inputting the target multi-question-answering task from the target device into the multi-question-answering task target model cluster to obtain an answer to the target multi-question-answering task includes:

[0027] Sharing the guidance factor of the k-th student model for the t-th test question-answering task with the k-th target model in the multi-question-answering task target model cluster;

[0028] Input the multi-question-answering task from the target device into the multi-question-answering task target model cluster, and obtain the answer output by the k-th target model for the t-th target question-answering task in the target multi-question-answering task;

[0029] According to the answer output by the kth target model for the tth target question-answering task and the guidance factor of the kth target model for the tth test question-answering task, the answer of the multi-question-answering task target model cluster for the tth target question-answering task is obtained.

[0030] Optionally, performing cluster distillation training on the student model cluster to be trained according to the total distillation loss value and the target parameter information, including:

[0031] According to the total distillation loss value, updating the model parameters of the student model cluster to be trained multiple times until the training is completed;

[0032] Check whether the parameter information of the trained student model cluster meets the target parameter information;

[0033] When the parameter information of the trained student model cluster does not meet the target parameter information, optimizing the trained student model cluster to obtain an optimized student model cluster, wherein the optimization includes pruning and model parameter quantization;

[0034] The optimized student model cluster is used as the student model cluster to be trained and returns to step: for multiple task guidance factors of each teacher model in the multi-question and answering task teacher model cluster, cluster distillation training is performed on the student model cluster to be trained according to the target parameter information to obtain the multi-question and answering task target model cluster.

[0035] Optionally, the method further includes:

[0036] Input each sample question in the sample question-answering dataset into the multi-question-answering task teacher model cluster to obtain the answer of the multi-question-answering task teacher model cluster to each sample question;

[0037] Based on the answers of the multi-question answering task teacher model cluster to each sample question and combined with the task types of the T test question answering tasks, a multi-question answering task dataset is constructed;

[0038] A portion of the multi-question-answering task dataset is used as the test multi-question-answering task, and the remaining portion is used as the sample multi-question-answering task.

[0039] In a second aspect, the present application proposes a question-answering device based on a task-oriented question-answering model cluster distillation system, the device comprising:

[0040] An acquisition module is used to obtain device parameter values of a target device of a target model cluster of a multi-question-answering task to be deployed, wherein the device parameter values include at least one of the following: storage space size and computing speed;

[0041] A determination module, configured to determine target parameter information of the target model cluster of the multi-question-answering task based on the device parameter value, wherein the target parameter information includes at least one of the following: model parameter quantity, reasoning response speed, and reasoning calculation amount;

[0042] A training module is used to perform cluster distillation training on the student model cluster to be trained based on multiple task guidance factors of each teacher model in the multi-question answering task teacher model cluster and the target parameter information to obtain the multi-question answering task target model cluster;

[0043] The task question-answering module is used to input the target multi-question-answering task from the target device into the multi-question-answering task target model cluster to obtain the answer to the target multi-question-answering task.

[0044] In a third aspect, the present application proposes an electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement a question-answering method based on a task-oriented question-answering model cluster distillation system as described in any one of the first aspects above.

[0045] The present application includes the following advantages: The present application proposes a question-answering method, apparatus and device based on a task-oriented guided question-answering model cluster distillation system, which obtains the device parameter value of the target device of the multi-question-answering task target model cluster to be deployed, and then determines the target parameter information of the multi-question-answering task target model cluster based on the device parameter value, and further based on the multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster, performs cluster distillation training on the student model cluster to be trained according to the target parameter information to obtain the multi-question-answering task target model cluster, and finally inputs the target multi-question-answering task from the target device into the multi-question-answering task target model cluster to obtain the answer to the target multi-question-answering task. This method improves the computational efficiency and accuracy of the target model cluster by combining the knowledge of multiple teacher models and the task guidance factors to optimize the student model for each task guidance, so that the target model cluster can also realize the reasoning and calculation of multi-question-answering tasks in resource-constrained devices, thereby improving the feasibility of deploying the target model cluster on resource-constrained devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] Figure 1 This is a flowchart of the steps of a question-answering method based on a task-oriented guided question-answering model cluster distillation system proposed in an embodiment of the present application;

[0048] Figure 2 is a schematic diagram of a teacher model cluster provided in an embodiment of the present application;

[0049] Figure 3 is a schematic diagram of a student model cluster provided in an embodiment of the present application;

[0050] Figure 4 This is a schematic diagram of a model cluster distillation process provided in an embodiment of the present application;

[0051] Figure 5Schematic diagram of the functional modules of a question-answering device based on a task-oriented guided question-answering model cluster distillation system provided in an embodiment of the present application;

[0052] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0054] Large language models excel in natural language processing tasks, particularly in text generation, machine translation, and question-answering systems. However, these models often contain billions or even tens of billions of parameters, placing significant demands on computing resources and storage during training and inference. This is especially true in real-world applications, especially on resource-constrained devices.

[0055] The Model Distillation method has been proposed to solve this problem, but it only achieves certain results in scenarios with a single task and a single teacher model.

[0056] However, existing model distillation methods have many limitations in multi-task learning scenarios. These limitations are mainly reflected in the limitation of single-task application, the lack of effective task guidance mechanisms, and low efficiency in resource-constrained environments. These problems prevent them from fully realizing their performance and efficiency in complex multi-task applications:

[0057] (1) Single-task limitation: Existing model distillation methods mostly focus on single-task scenarios and cannot effectively handle the challenges of multi-task learning. In a multi-task environment, the heterogeneity and interdependence between different tasks make it difficult for traditional distillation methods to fully mine and share knowledge between tasks.

[0058] (2) Lack of task guidance: Although existing cluster distillation methods can utilize the complementarity of multiple teacher models, in multi-task scenarios, how to effectively guide multiple teacher models to work together to adapt to the specific needs of each task remains an unresolved problem.

[0059] (3) Inefficiency in resource-constrained environments: Although existing cluster distillation methods can reduce the size of the model, the efficiency and performance improvement of the distillation process are still limited in the multi-task learning scenario. In particular, on resource-constrained devices, it is impossible to fully utilize task guidance to optimize the model, resulting in the deployment of large models still facing high computing and storage overhead.

[0060] Based on this, this application proposes a question-answering method based on a task-oriented guided question-answering model cluster distillation system, which aims to enable multiple teacher models to be optimized according to specific tasks and work together in the distillation process by introducing a task guidance mechanism. This method not only helps the student model to obtain better generalization capabilities in multi-task learning, but also effectively reduces the computational and storage overhead of the model. Through task guidance, the student model cluster can learn the knowledge associations between different tasks, thereby improving its performance on multiple tasks, especially in resource-constrained environments. Show higher efficiency. In this application, the teacher model cluster is a multi-question-answering task teacher model cluster, and the student model cluster is a multi-question-answering task student model cluster.

[0061] In the first aspect, the embodiment of the present application proposes a question answering method based on a task-oriented question answering model cluster distillation system, see Figure 1 , Figure 1 : This is a flowchart of a question-answering method based on a task-oriented question-answering model cluster distillation system proposed in an embodiment of the present application. The method includes the following steps:

[0062] Step 101: Obtain device parameter values of target devices of the target model cluster for the multi-question-answering task to be deployed.

[0063] The target devices mentioned above are used to deploy the target model cluster for the multi-question answering task. These target devices can be hardware devices, such as servers and computers. They can also be resource-constrained devices. Devices with any of the following resource constraints can be identified as target devices: unable to host standard-scale models, unable to meet the computing power requirements for real-time inference, limited battery capacity, or limited communication resources. Examples include mobile and embedded devices with ≤4 GB of memory and ≤1 TFLOP of computing power (such as smartphones, tablets, and IoT terminals); edge computing devices with a computing power of 1-10 TFLOPS (such as industrial edge gateways and in-vehicle ECUs); and low-power, specialized AI-optimized hardware with an energy efficiency ratio greater than 10 GFLOPS (Giga Floating-Point Operations Per Second) / W, such as AI accelerator chips and FPGA terminals (Field-Programmable Gate Arrays).

[0064] This embodiment performs cluster distillation training on the student model cluster to be trained based on the teacher model of the multi-question-answering task teacher model cluster, obtains the distilled student model cluster as the multi-question-answering task target model cluster, and deploys the model cluster on the target device to achieve compression of the model cluster and improve computational efficiency.

[0065] In this embodiment, the device parameter values of the target device on which the multi-question-answering task target model cluster is to be deployed are first obtained. The device parameter values of the target device include at least one of the following: storage space size, computing speed, and may also include parameters such as memory bandwidth. These device parameter values may be factory settings for the target device or historical or real-time operating data of the target device during operation.

[0066] Step 102: Determine target parameter information of the multi-question-answering task target model cluster based on the device parameter value.

[0067] In this embodiment, the target device's device parameter values are used to determine target parameter information for the target model cluster for the multi-question-answering task. This target parameter information includes at least one of the following: model parameter quantity, inference response speed, and inference computational effort. Further, when distilling the target model cluster for the multi-question-answering task, the student model cluster to be trained can be distilled according to the calculated target parameter information, thereby achieving compression of the model cluster while improving its computational efficiency.

[0068] Step 103: Based on the multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster, cluster distillation training is performed on the student model cluster to be trained according to the target parameter information to obtain the multi-question-answering task target model cluster.

[0069] Considering that the existing distillation method does not consider the knowledge sharing between tasks in the multi-task scenario, resulting in insufficient knowledge sharing between tasks. Based on this, this embodiment proposes to realize the knowledge transfer and output exchange between different teacher models in the multi-question answering task teacher model cluster through the task guidance factor. The task guidance factor adjusts the influence of the teacher model cluster on the student model cluster according to the importance of the task (the schematic diagram of the teacher model cluster and the student model cluster can be referred to Figure 2 and Figure 3 The introduction of the task guidance factor allows different tasks to obtain different weights during the distillation training process, thereby helping the teacher model cluster to provide targeted guidance to the student model cluster according to task requirements in multi-task scenarios.

[0070] Based on this, based on the multiple task guidance factors of each teacher model in the multi-question answering task teacher model cluster and according to the target parameter information, the student model cluster to be trained is cluster distilled to obtain the multi-question answering task target model cluster.

[0071] The task-based guidance factor ensures that knowledge between different tasks can be effectively shared during multi-task learning. At the same time, the task-based guidance factor adjusts the influence of the teacher model cluster on the student model cluster based on the importance and priority of the task, allowing the student model cluster to acquire more appropriate knowledge across multiple tasks and improve its performance on each task.

[0072] Step 104: Input the target multi-question-answering task from the target device into the multi-question-answering task target model cluster to obtain an answer to the target multi-question-answering task.

[0073] The multi-question-answering task target model cluster obtained through training in this embodiment is used to answer the multi-question-answering task on the target device.

[0074] Specifically, the multi-question and answer task target model cluster is deployed on the target device, the target multi-question and answer task from the target device is input into the multi-question and answer task target model cluster, and the answer to the target multi-question and answer task is obtained through the interaction in the multi-question and answer task target model cluster.

[0075] The embodiment of the present application proposes a question-answering method based on a task-guided question-answering model cluster distillation system, which obtains the device parameter value of the target device of the multi-question-answering task target model cluster to be deployed, and then determines the target parameter information of the multi-question-answering task target model cluster based on the device parameter value, and further based on the multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster, performs cluster distillation training on the student model cluster to be trained according to the target parameter information to obtain the multi-question-answering task target model cluster, and finally inputs the target multi-question-answering task from the target device into the multi-question-answering task target model cluster to obtain the answer to the target multi-question-answering task. This method optimizes the student model guided by each task by combining the knowledge of multiple teacher models and the task guidance factor, while achieving model compression and improved computational efficiency, thereby improving the computational efficiency and deployment feasibility of the model in a resource-constrained environment.

[0076] Based on the above embodiment, in an optional implementation manner, in the above step 102, determining the target parameter information of the multi-question-answering task target model cluster according to the device parameter value specifically includes the following steps:

[0077] Step 1021: Obtain device parameter values of the target device. The device parameter values include storage space size (expressed as M, in GB, such as CPU video memory), computing speed (expressed as P, TFLOPS (Tera Floating-Point Operations Per Second, a measurement of trillion floating-point operations per second, such as GPU peak computing power) and memory bandwidth (expressed as B, in GB / s).

[0078] For example, the target device's computing speed can be calculated using the following process:

[0079] First, by running the LINPACK benchmark, we measured the target device's actual computing speed at 78% of its theoretical peak. We then calculated its actual computing speed. For example, 1 TFLOPS of computing power can support a model with 1B parameters to complete a single inference within 100 ms. The calculation formula is: 1B parameters × 2 FLOPs / parameter / 10¹² FLOPs / s = 0.002 s.

[0080] Step 1022: Based on the historical operation data or real-time operation data of the target device, establish a constraint relationship between the device parameter value of the target device and the target parameter information of the multi-question-answering task target model cluster, and calculate the target parameter information through the constraint relationship. In the embodiment of the present application, the target parameter information may include model parameter quantity, inference response speed, and inference calculation amount. The historical operation data and real-time operation data are both real device parameter values when the target device is actually running, and can be directly obtained from files such as the operation log of the target device.

[0081] The constructed constraints are as follows:

[0082] Model parameter quantity: Set the storage space of the target device to the maximum model parameter quantity of the target model cluster of the multi-question-answering task. The constraint relationship can be: Maximum model parameter quantity N max = Storage space size M × compression coefficient l / parameter byte number s. The compression coefficient is determined according to the different compression methods. For example, through half-precision floating-point quantization, the storage requirement is reduced by 50% (i.e., using FP16 quantization compression model), and the compression coefficient is set to 0.5; after removing 30% of the convolution kernels (i.e., using structured pruning compression model), the model parameter storage requirement is reduced to 70%, and the compression coefficient is set to 0.7. For example, when the storage space size M = 8GB, FP16 is used, the parameter byte number s = 2, and the compression coefficient l = 0.5, the maximum model parameter number N max =2 B parameter.

[0083] Inference computing capacity: Set the computing speed of the target device to the upper limit of the inference computing capacity of the target model cluster of the multi-question answering task. The constraint relationship can be: Maximum computing capacity of a single inference (TFLOPS) Fmax = Peak computing power of the device P × safety factor α. The safety factor can be determined based on the real-time requirements of the task and the thermal design power consumption of the hardware. Specifically, according to the mission criticality classification: the safety factor of critical tasks is ≤0.5, and the safety factor of non-critical tasks is ≤0.8; according to the dynamic adjustment method: the safety factor is adjusted in real time according to the temperature of the target device, such as reducing it by 0.1 when the temperature is ≥80°C. In addition, the safety factor can also be determined based on the hardware margin test, task delay constraints, power consumption limits, and multi-task preemption of the target device. For example, for a target device with a peak computing power of P=10TFLOPS, set the safety factor α=0.6, and get F max =6 TFLOPS.

[0084] Inference response speed: Based on the actual computing power (FLOPS) of the target device F, the effective computing power of the device B, the data handling volume D, and the I / O fixed overhead T io Establish a constraint relationship for the inference response speed of the computing model. The constraint relationship can be: inference response speed T = F / device peak computing power P + D / B + T io The actual computing capacity (FLOPS) of the target device is F, the effective computing power of the device is B, the data handling capacity is D, and the I / O fixed overhead is T. io It can be obtained from historical operation data or real-time operation data.

[0085] According to the above constraints, the target parameter information of the multi-question-answering task target model cluster can be calculated based on the device parameter values of the target device. Based on this target parameter information, when the subsequent multi-question-answering task target model cluster is deployed on the target device, the above target parameter information must be met.

[0086] Based on the above embodiment, in an optional implementation manner, the present application further provides a question-answering method based on a task-oriented guided question-answering model cluster distillation system, in which the guiding factor for testing the question-answering task is obtained based on the importance score of the task, and the method specifically further includes the following steps:

[0087] Step 105: Input the test multi-question answering task into the k-th teacher model to obtain T answers output by the k-th teacher model.

[0088] In this embodiment, the test multi-question-answering task can be a multi-question-answering task directly obtained from an existing database, or can be a multi-question-answering task that has been tested by other models.

[0089] In this implementation, the above-mentioned multi-question and answer task teacher model cluster includes K teacher models, K≥2, 1≤k≤K, and the test multi-question and answer task includes T test question and answer tasks, T≥2, 1≤t≤T. The test multi-question and answer task is input into the kth model in the multi-question and answer task teacher model cluster, and T answers output by the kth model for the test multi-question and answer task can be obtained, where the tth answer is the answer to the tth test question and answer task in the test multi-question and answer task.

[0090] Step 106: Determine the importance score of the k-th teacher model for the t-th test question-answering task based on the difference between the t-th answer and the t-th correct answer corresponding to the t-th test question-answering task.

[0091] In this embodiment, the importance score of the teacher model for each test question-answering task can be determined based on the difference between the answer given by each teacher model for each test question-answering task and the correct answer to the test question-answering task.

[0092] It should be noted that each teacher model in the multi-question-answering task teacher model cluster has a different importance score for each test question-answering task. The importance score is used to measure the priority of the test question-answering task or the contribution of the test question-answering task to the overall model performance.

[0093] Specifically, the importance score of the k-th teacher model for the t-th test question-answering task can be determined by the difference between the t-th answer and the t-th correct answer corresponding to the t-th test question-answering task.

[0094] Step 107: Determine the guidance factor of the kth teacher model for the tth test question-answering task based on the importance score of the kth teacher model for the tth test question-answering task and the importance scores of the kth teacher model for T test question-answering tasks.

[0095] In this embodiment, the guidance factor of each teacher model for each test question-answering task can be calculated based on the importance score of each teacher model for each test question-answering task and the total importance score of the teacher model for all test question-answering tasks.

[0096] It should be noted that different task guidance factors can enable different test question-answering tasks to obtain different weights in the distillation process, further helping the teacher model to provide targeted guidance and optimization to the student model cluster according to the requirements of different tasks in multi-task scenarios.

[0097] Specifically, the guidance factor of the kth teacher model for the tth test question-answering task can be determined based on the ratio of the importance score of the kth teacher model for the tth test question-answering task and the importance score of the kth teacher model for the Tth test question-answering task. The specific calculation formula is as follows:

[0098] ,

[0099] in, Indicates that the k-th teacher model is for the The guiding factor for the test question answering task, Indicates the total number of tasks for testing multiple question-answering tasks, Indicates that the k-th teacher model is for the Importance scores for the test question answering tasks.

[0100] Based on the above embodiment, in an optional implementation, the present application further provides a question-answering method based on a task-oriented guided question-answering model cluster distillation system. In this method, in the above step 103, based on multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster, cluster distillation training is performed on the to-be-trained student model cluster according to the target parameter information, specifically including the following steps:

[0101] Step 1031: input the t-th sample question-answering task in the sample multi-question-answering task into the multi-question-answering task teacher model cluster to obtain the answer of the k-th teacher model to the t-th sample question-answering task.

[0102] In this embodiment, the sample multi-question-answering task can be a multi-question-answering task directly obtained from an existing database, or a multi-question-answering task that has been tested by other models can be used as the sample multi-question-answering task.

[0103] The above-mentioned multi-sample question-answering task includes T sample question-answering tasks. By inputting the tth sample question-answering task in the multi-sample question-answering task into the multi-question-answering task teacher model cluster, the answer of the kth teacher model for the tth sample question-answering task can be obtained. Each teacher model obtains T answers for the T sample multi-question-answering tasks.

[0104] Step 1032: Based on the answer of the kth teacher model to the tth sample question-answering task and the guidance factor of the kth teacher model to the tth test question-answering task, obtain the answer of the multi-question-answering task teacher model cluster to the tth sample question-answering task.

[0105] In this application, in order to achieve model cluster distillation, by using multiple teacher models The cluster guides the learning of the student model cluster, and each teacher model in the teacher model cluster generates a predicted output (i.e., answer) for the corresponding task based on the task guidance factor.

[0106] Specifically, based on the answer of the kth teacher model to the tth sample question-answering task and the guidance factor of the kth teacher model to the tth test question-answering task, the answers of all teacher models in the multi-question-answering task teacher model cluster to the tth test question-answering task are integrated to obtain the answer of the multi-question-answering task teacher model cluster to the tth sample question-answering task. The calculation process is as follows:

[0107]

[0108] in, Indicates the The teacher model is for The answers to the sample question-answering tasks, Indicates the The teacher model is for The guiding factor of the test question answering task , represents the total number of teacher models in the multi-question answering task teacher model cluster, Indicates that the teacher model cluster of the multi-question answering task is targeted at the Answers to sample question-answering tasks.

[0109] In this way, the contribution of each teacher model in the multi-question answering task teacher model cluster to different multi-question answering tasks is dynamically adjusted, and it can provide different levels of guidance according to the different needs and importance of the multi-question answering tasks.

[0110] It should be noted that the task type in the sample question-answering test is the same as the task type in the test question-answering task, that is, the task type of the t-th sample question-answering task must be the same as the task type of the t-th test question-answering task. For example, if the task type of the t-th sample question-answering task is an image analysis task, then the task type of the t-th test question-answering task is also an image analysis task. If the task type of the t-th sample question-answering task is a text analysis task, then the task type of the t-th test question-answering task is also a text analysis task. Correspondingly, The teacher model is for The guiding factor of the sample question answering task is The guiding factors for the three test question-answering tasks are the same.

[0111] Step 1033: With the goal of learning the answers of the multi-question-answering task teacher model cluster to the sample multi-question-answering task, cluster distillation training is performed on the student model cluster to be trained according to the target parameter information.

[0112] In this embodiment, the answers of the multi-question-answering task teacher model cluster to the sample multi-question-answering task are used as the training target, and cluster distillation training is performed on the student model cluster to be trained according to the calculated target parameter information.

[0113] The multi-question-answering task teacher model cluster in this method provides different degrees of guidance and optimization for the answers to sample multi-question-answering tasks according to the different needs and importance of the multi-question-answering tasks. The student model cluster trained by this method can obtain more accurate answers for different question-answering tasks. At the same time, the student model cluster is distilled with target parameter information to obtain a student model cluster with reduced computational and storage overhead of the model, which can improve the computational efficiency and deployment feasibility of the trained student model cluster in resource-constrained environments.

[0114] Furthermore, through the cluster distillation method, multiple teacher models can work together according to task requirements, rather than relying on a single teacher model. This method fully utilizes the knowledge complementarity of multiple teacher models, enhances the knowledge transfer effect during the distillation process, and thus improves the performance of the student model cluster.

[0115] Based on the above embodiment, in an optional implementation, the present application further proposes a question-answering method based on a task-oriented guided question-answering model cluster distillation system. In this method, in the above step 103, based on multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster, cluster distillation training is performed on the student model cluster to be trained according to the target parameter information, and specifically further includes the following steps:

[0116] Step 1034: Determine the distillation loss value of the t-th sample question-answering task based on the difference between the answer of the multi-question-answering task teacher model cluster for the t-th sample question-answering task and the answer of the multi-question-answering task student model cluster for the t-th sample question-answering task.

[0117] To effectively guide the student model to learn the knowledge of the teacher cluster, this implementation uses a distillation loss function to balance the differences between the responses of the teacher and student model clusters. This distillation loss function combines the Kullback-Leibler (KL) divergence and the mean squared error (MSE) to balance the differences between the responses of the teacher and student model clusters.

[0118] Specifically, the distillation loss value of the t-th sample question answering task is determined based on the difference between the answer of the multi-question answering task teacher model cluster for the t-th sample question answering task and the answer of the multi-question answering task student model cluster for the t-th sample question answering task.

[0119] For example, the answer of the multi-question answering task student model cluster for the t-th sample question answering task is expressed as , the answer of the multi-question answering task teacher model cluster for the t-th sample question answering task is expressed as , distillation loss function It is expressed as follows:

[0120]

[0121] Among them, KL represents Kullback-Leibler divergence, which is used to measure the difference between the answers of the multi-question answering task teacher model cluster and the multi-question answering task student model cluster for the t-th sample question answering task; MSE represents mean square error, which is used to measure the difference between the multi-question answering task student model cluster and the true answer for the t-th sample question answering task; This hyperparameter represents the balance between the responses of the teacher and student models for the t-th sample question answering task. It ranges from [0, 1] and controls the balance between KL divergence and MSE. The distillation loss is calculated using the aforementioned distillation loss function.

[0122] Step 1035: Obtain a total distillation loss value based on the distillation loss value of the t-th sample question-answering task and the guidance factors of the K teacher models for the t-th test question-answering task.

[0123] In the learning process of multiple question answering tasks, the distillation loss of each question answering task depends not only on the task guidance factor , and the losses of all question-answering tasks need to be weighted summed to achieve the final optimization goal.

[0124] Specifically, the total distillation loss value is obtained based on the distillation loss value of the t-th sample question-answering task and the guidance factors of the K teacher models for the t-th test question-answering task.

[0125] Total distillation loss function It is expressed as follows:

[0126]

[0127] Among them, the total distillation loss value is calculated by the total distillation loss function. represents the guidance factor of K teacher models for the t-th test question answering task, = .

[0128] Step 1036: Perform cluster distillation training on the student model cluster to be trained according to the total distillation loss value and the target parameter information.

[0129] The student model cluster to be trained is clustered according to the total distillation loss value, and the student model cluster to be trained is distilled according to the target parameter information, so that the trained student model cluster can not only effectively learn the knowledge in the teacher model cluster, but also achieve compression of the model cluster and improve the computational efficiency.

[0130] This cluster distillation method, based on task guidance, effectively compresses the size of the teacher model cluster, reducing computational and storage overhead. During distillation training, the student model cluster maintains high performance while reducing computational complexity. This approach is particularly suitable for resource-constrained devices and edge computing scenarios, significantly improving the inference efficiency and application feasibility of the student model cluster.

[0131] Based on the above embodiment, in an optional implementation, the present application further proposes a question-answering method based on a task-oriented question-answering model cluster distillation system. In this method, the multi-question-answering task student model cluster includes K student models, and the method further includes the following steps:

[0132] Step 108: Share the guidance factor of the k-th teacher model for the t-th test question-answering task with the k-th student model.

[0133] In this implementation, the contribution of each teacher model in the teacher model cluster to a task is dynamically adjusted, providing varying degrees of guidance based on the varying needs and importance of the task. To enable the student model cluster to fully learn from the knowledge of the teacher model cluster, the guidance factors of each teacher model in the teacher model cluster for different test question-answering tasks are shared with each student model in the corresponding student model cluster.

[0134] Specifically, the guidance factor of the k-th teacher model for the t-th test question-answering task is shared with the k-th student model.

[0135] Step 109: Input the t-th sample question-answering task in the sample multi-question-answering task into the multi-question-answering task student model cluster to obtain the answer of the k-th student model to the t-th sample question-answering task.

[0136] After each student model in the multi-question-answering task student model cluster obtains the guidance factors for different test question-answering tasks, it obtains the answer to the sample question-answering task based on the input sample question-answering task and the corresponding guidance factors.

[0137] Specifically, the tth sample question-answering task in the sample multi-question-answering task is input into the multi-question-answering task student model cluster, and the answer of the kth student model to the tth sample question-answering task is obtained.

[0138] Step 110: Based on the answer of the kth student model to the tth sample question-answering task and the guidance factor of the kth student model to the tth test question-answering task, obtain the answer of the multi-question-answering task student model cluster to the tth sample question-answering task.

[0139] In this embodiment, the multi-question and answer task student model cluster includes multiple student models that are the same as the multi-question and answer task teacher model cluster, wherein each student model obtains the answer of the multi-question and answer task student model cluster to each sample question and answer task and the corresponding guidance factor for each sample question and answer task.

[0140] Specifically, based on the answer of the kth student model to the tth sample question-answering task and the guidance factor of the kth student model to the tth test question-answering task, the answer of the multi-question-answering task student model cluster to the tth sample question-answering task is obtained. The calculation process is as follows:

[0141]

[0142] in, Indicates the The student model is for The answers to the sample question-answering tasks, Indicates the The student model is for The guiding factor of the test question answering task , Represents the total number of student models in the multi-question answering task student model cluster Indicates that the student model cluster for the multi-question answering task is Answers to sample question-answering tasks.

[0143] It should be noted that the task types in the student model cluster are the same as those in the teacher model cluster. The student model is for The guiding factor for the sample question answering task is the same as the The guiding factors for the three test question-answering tasks are the same.

[0144] Based on the above embodiment, in an optional implementation, the present application further proposes a question-answering method based on a task-oriented question-answering model cluster distillation system. In this method, the above step 104 further includes the following steps:

[0145] Step 1041: Share the guidance factor of the k-th student model for the t-th test question-answering task with the k-th target model in the multi-question-answering task target model cluster.

[0146] After the multi-question answering task student model cluster is trained by distillation based on the answers from the multi-question answering task teacher model cluster and the guidance factors for different test question answering tasks, the task guidance factors of each student model in the multi-question answering task student model cluster can be shared with the multi-question answering task target model cluster to enable the multi-question answering task target model cluster to answer the question answering task on the target device. The number of target models in the multi-question answering task target model cluster is the same as the number of student models in the multi-question answering task student model cluster.

[0147] Specifically, the guidance factor of the k-th student model for the t-th test question-answering task is shared with the k-th target model in the multi-question-answering task target model cluster.

[0148] Step 1042: Input the multiple question-answering tasks from the target device into the multiple question-answering task target model cluster to obtain the answer output by the kth target model for the tth target question-answering task in the target multiple question-answering tasks.

[0149] By inputting the multiple question-answering tasks on the target device into the multi-question-answering task target model cluster, the answers of each target model in the multi-question-answering task target model cluster to each question-answering task in the multi-question-answering task can be obtained.

[0150] Specifically, for the tth target question answering task in the target multi-question answering task, the kth target model in the multi-question answering task target model cluster outputs the corresponding answer for the tth target question answering task.

[0151] Step 1043: Based on the answer output by the kth target model for the tth target question-answering task, and the guidance factor of the kth target model for the tth test question-answering task, obtain the answer of the multi-question-answering task target model cluster for the tth target question-answering task.

[0152] Based on the answers of each target model in the multi-question and answering task target model cluster to each target question and answer task in the target multi-question and answering task, and the guidance factors of each target model to each target question and answer task, the answers of the multi-question and answering task target model cluster to each target question and answer task are obtained.

[0153] Specifically, for the tth target question-answering task in the multi-question-answering task, based on the answer output by the kth target model for the tth target question-answering task and the guidance factor of the kth target model for the tth test question-answering task, the answer of the multi-question-answering task target model cluster for the tth target question-answering task can be calculated.

[0154] It should be noted that during the actual multi-answer task, if the target model's guidance factor for a certain type of task is too small, it indicates that the target model is less important for answering that type of task. Therefore, when answering, the target model and the corresponding task guidance factor are not used to answer this type of task.

[0155] For example, suppose a multi-question answering task target model cluster includes three target models, A, B, and C. Target model A has image analysis and text analysis capabilities, target model B has image analysis and text generation capabilities, and target model C has text analysis and text generation capabilities. The input multi-question answering task is: Please analyze the selected area in the image (the image can be any image). The guidance factors for target models A, B, and C for different tasks are all pre-defined guidance factors for tasks of the same type.

[0156] This question-answering task includes multiple tasks: performing image analysis on the selected area in the image, further performing text analysis on the image analysis results, and finally generating a corresponding text description.

[0157] The multi-question answering task target model cluster processes the multi-question answering task as follows: The multi-question answering task is processed using a multi-question answering target model consisting of target models A, B, and C. Target model A supports image analysis and text analysis, target model B supports image analysis and text generation, and target model C supports text analysis and text generation. For the multi-question answering task "Please analyze the selected area in the image," the task is divided into three subtasks: image analysis, text analysis, and text generation. First, the image analysis task is assigned to target models A and B, which have image analysis capabilities. Target models A and B are loaded and processed to process the guiding factors for image analysis tasks, resulting in a description of the selected area in the image, such as "The selected area in the image contains an orange cat sitting on a sofa, surrounded by green plants." Next, the image analysis results are fed into target models A and C, which are loaded and processed to process the guiding factors for text analysis tasks, extracting and analyzing semantic information and achieving a deeper level of text understanding, such as "The scene is indoors, the style is cozy, and the subject is a pet cat." Finally, based on the aforementioned analysis results, target models B and C are loaded simultaneously to process the guiding factors of the text generation task, complete text generation, and output a natural language description: "The selected area of the image shows an orange cat sitting quietly on a sofa, surrounded by green plants. The overall environment appears warm and natural, exuding a leisurely and comfortable atmosphere." This process demonstrates the ability of multi-objective models to collaboratively handle multi-task question and answering, with clear task division and reasonable model scheduling. The guiding factors also enhance the model's generalization and response accuracy.

[0158] Based on the above embodiment, in an optional implementation manner, the above step 1036 further includes the following steps:

[0159] Step 10361: Based on the total distillation loss value, update the model parameters of the student model cluster to be trained multiple times until the training is completed.

[0160] The total distillation loss is calculated to repeatedly update the model parameters of the student model cluster to be trained. This process continuously balances the differences between the responses of the student model cluster to the sample multi-question answering task and the responses of the teacher model cluster to the sample multi-question answering task. Through multiple model parameter updates, this difference is minimized until training is complete. This training process enables the student model cluster to learn the important features and representations of the teacher model cluster through knowledge distillation, reducing computational effort while maintaining high accuracy.

[0161] Step 10362: Check whether the parameter information of the trained student model cluster meets the target parameter information.

[0162] In order to effectively deploy the trained student model cluster in a resource-constrained environment, after the training of the student model cluster is completed, its parameter information needs to be judged to determine whether it meets the target parameter information. If not, the parameters of the trained student model cluster need to be optimized and adjusted (i.e., the student model cluster needs to be compressed) to make it meet the target parameter information.

[0163] Step 10363: When the parameter information of the trained student model cluster does not meet the target parameter information, the trained student model cluster is optimized to obtain an optimized student model cluster, wherein the optimization includes pruning and model parameter quantization.

[0164] If the parameters of the trained student model cluster do not meet the target parameter information, the trained student model cluster needs to be optimized to further reduce the computational complexity and storage overhead of the student model cluster, and reduce the number of parameters to maintain high computing performance. This optimization includes operations such as pruning and model parameter quantization.

[0165] After optimizing the trained student model cluster, an optimized student model cluster is obtained.

[0166] Step 10364: Use the optimized student model cluster as the student model cluster to be trained and return to step: Based on the multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster, cluster distillation training is performed on the student model cluster to be trained according to the target parameter information to obtain the multi-question-answering task target model cluster.

[0167] After the student model cluster is optimized, the optimized student model cluster is used again as the student model cluster to be trained, and is processed again according to step 103 until the student model cluster to be trained can meet the target parameter information.

[0168] This method optimizes the student model cluster during distillation training by performing optimizations such as pruning and quantization. This significantly reduces the number of parameters in the student model cluster, while also reducing the storage requirements and computational complexity of the model cluster while ensuring performance. This makes this method promising for application in resource-limited environments such as mobile devices and embedded systems.

[0169] Based on the above embodiment, in an optional implementation, the present application also proposes a question answering method based on a task-oriented question answering model cluster distillation system. The method can be found in Figure 4 , Figure 4 : is a schematic diagram of a model cluster distillation process proposed in an embodiment of the present application, wherein the method further includes the following steps:

[0170] Step 111: Input each sample question in the sample question-answering dataset into the multi-question-answering task teacher model cluster to obtain the answer of the multi-question-answering task teacher model cluster to each sample question.

[0171] This implementation takes into account the low quality of sample question-and-answer datasets obtained directly from existing databases. Training a multi-question-and-answer task teacher model cluster directly based on this sample question-and-answer dataset results in poor training results. This multi-question-and-answer task teacher model cluster filters the data in the sample question-and-answer dataset to remove incorrect answers or garbled data, thereby improving the quality of the sample question-and-answer dataset.

[0172] Based on this, each sample question in the sample question-answering dataset is first input into the multi-question-answering task teacher model cluster to obtain the answer to each sample question from the multi-question-answering task teacher model cluster. The data in the sample question-answering dataset is filtered based on the answer to each sample question from the multi-question-answering task teacher model cluster.

[0173] Step 112: Based on the answers of the multi-question-answering task teacher model cluster to each sample question and combined with the task types of the T test question-answering tasks, a multi-question-answering task dataset is constructed.

[0174] Based on the answers of the multi-question-answering task teacher model cluster to each sample question, they are compared with the real answers, and the wrong answers or garbled data are removed. From the remaining sample question-answering data, the sample question-answering data with the same task type as the T test question-answering tasks are further screened to construct a multi-question-answering task dataset.

[0175] Among them, the same task type means that if the task types included in the test question-answering task are image analysis, text generation, text analysis, etc., then the task types in the constructed multi-question-answering task dataset should also include the above types.

[0176] Step 113: Use a portion of the multi-question-answering task dataset as the test multi-question-answering task, and the remaining portion as the sample multi-question-answering task.

[0177] Finally, a part of the screened question-answering task dataset is used as a test multi-question-answering task. This part of the test question-answering task is used to determine the guiding factors of different tasks, and the remaining part is used as a sample multi-question-answering task to perform distillation training on the student model cluster to be trained.

[0178] This application proposes a question-answering method based on a task-oriented, guided question-answering model cluster distillation system. By introducing a task-guidance mechanism and cluster distillation technology, it solves the knowledge sharing problem in multi-task learning and optimizes the synergy between teacher models in a teacher model cluster across multiple tasks. This method not only effectively reduces computational and storage overhead, improving the reasoning efficiency of student model clusters in resource-constrained environments, but also enhances the generalization ability of student model clusters, ensuring their excellent performance across multiple tasks. This makes student model clusters more efficient in multi-task scenarios and provides a more feasible solution for model deployment in practical applications.

[0179] In the second aspect, the embodiment of the present application also proposes a question-answering device based on a task-oriented question-answering model cluster distillation system, see Figure 5 , Figure 5 This is a functional module diagram of a question-answering device based on a task-oriented question-answering model cluster distillation system proposed in an embodiment of the present application, the device comprising:

[0180] An acquisition module 501 is configured to acquire device parameter values of a target device of a target model cluster for a multi-question-answering task to be deployed, wherein the device parameter values include at least one of the following: storage space size and computing speed;

[0181] A determination module 502 is configured to determine target parameter information of the multi-question-answering task target model cluster based on the device parameter value, wherein the target parameter information includes at least one of the following: model parameter quantity, reasoning response speed, and reasoning calculation amount;

[0182] A training module 503 is configured to perform cluster distillation training on the to-be-trained student model cluster based on multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster and the target parameter information to obtain the multi-question-answering task target model cluster;

[0183] The task question-answering module 504 is configured to input the target multi-question-answering task from the target device into the multi-question-answering task target model cluster to obtain an answer to the target multi-question-answering task.

[0184] Optionally, the multi-question-answering task teacher model cluster includes K teacher models; and the device further includes:

[0185] A test multi-question-answer task answer acquisition module is used to input the test multi-question-answer task into the k-th teacher model, and obtain T answers output by the k-th teacher model, where the t-th answer is the answer to the t-th test question-answer task in the test multi-question-answer task;

[0186] an importance score determination module, configured to determine an importance score of the kth teacher model for the tth test question-answering task based on a difference between the tth answer and the tth correct answer corresponding to the tth test question-answering task;

[0187] The task guidance factor determination module is used to determine the guidance factor of the kth teacher model for the tth test question and answer task based on the importance score of the kth teacher model for the tth test question and answer task and the importance score of the kth teacher model for T test question and answer tasks.

[0188] Optionally, the training module 503 includes:

[0189] The first acquisition submodule is used to input the t-th sample question-answering task in the sample multi-question-answering task into the multi-question-answering task teacher model cluster, and obtain the answer of the k-th teacher model for the t-th sample question-answering task;

[0190] a second acquisition submodule, configured to obtain an answer of the multi-question-answering task teacher model cluster for the tth sample question-answering task based on the answer of the kth teacher model for the tth sample question-answering task and the guidance factor of the kth teacher model for the tth test question-answering task, wherein the tth sample question-answering task and the tth test question-answering task are of the same task type;

[0191] The first cluster distillation training submodule is used to perform cluster distillation training on the student model cluster to be trained according to the target parameter information, with the goal of learning the answers of the multi-question-answering task teacher model cluster to the sample multi-question-answering task.

[0192] The training module 503 further includes:

[0193] a distillation loss value determination submodule, configured to determine a distillation loss value for the t-th sample question-answering task based on a difference between an answer of the multi-question-answering task teacher model cluster for the t-th sample question-answering task and an answer of the multi-question-answering task student model cluster for the t-th sample question-answering task;

[0194] The total distillation loss value determination submodule is used to obtain the total distillation loss value based on the distillation loss value of the t-th sample question-answering task and the guidance factors of the K teacher models for the t-th test question-answering task;

[0195] The second cluster distillation training submodule is used to perform cluster distillation training on the student model cluster to be trained according to the total distillation loss value and the target parameter information.

[0196] The multi-question-answering task student model cluster includes K student models; the device further includes:

[0197] A sharing module, configured to share the guidance factor of the k-th teacher model for the t-th test question-answering task with the k-th student model;

[0198] A first input module is configured to input the t-th sample question-answering task in the sample multi-question-answering task into the multi-question-answering task student model cluster, and obtain an answer of the k-th student model for the t-th sample question-answering task;

[0199] The student model cluster answer determination module is used to obtain the answer of the multi-question and answer task student model cluster to the tth sample question and answer task based on the answer of the kth student model to the tth sample question and answer task and the guidance factor of the kth student model to the tth test question and answer task.

[0200] The task answering module 504 includes:

[0201] A guidance factor sharing submodule, configured to share the guidance factor of the k-th student model for the t-th test question-answering task with the k-th target model in the multi-question-answering task target model cluster;

[0202] The target model answer determination submodule is used to input the multiple question-answering tasks from the target device into the multiple question-answering task target model cluster, and obtain the answer output by the kth target model for the tth target question-answering task in the target multiple question-answering tasks;

[0203] The answer determination submodule of the target model cluster is used to obtain the answer of the multi-question and answer task target model cluster for the tth target question and answer task based on the answer output by the kth target model for the tth target question and answer task, and the guidance factor of the kth target model for the tth test question and answer task.

[0204] The second cluster distillation training submodule includes:

[0205] A model parameter updating unit, configured to update the model parameters of the student model cluster to be trained multiple times according to the total distillation loss value until the training is completed;

[0206] A parameter information judgment unit is used to detect whether the parameter information of the trained student model cluster meets the target parameter information;

[0207] an optimization processing unit, configured to optimize the trained student model cluster to obtain an optimized student model cluster when parameter information of the trained student model cluster does not meet the target parameter information, wherein the optimization processing includes pruning and model parameter quantization;

[0208] The multi-question and answer task target model cluster determination unit is used to use the optimized student model cluster as the student model cluster to be trained and return to the step: for multiple task guidance factors of each teacher model in the multi-question and answer task teacher model cluster, cluster distillation training is performed on the student model cluster to be trained according to the target parameter information to obtain the multi-question and answer task target model cluster.

[0209] Wherein, the device further comprises:

[0210] The second input module is used to input each sample question in the sample question and answer dataset into the multi-question and answer task teacher model cluster to obtain the answer of the multi-question and answer task teacher model cluster to each sample question;

[0211] A construction module is used to construct a multi-question answering task dataset based on the answers of the multi-question answering task teacher model cluster to each sample question and in combination with the task types of the T test question answering tasks;

[0212] The question-answering task determination module is used to use a part of the multi-question-answering task dataset as the test multi-question-answering task and the remaining part as the sample multi-question-answering task.

[0213] Based on the same application concept, the embodiment of the present application discloses an electronic device in a third aspect. Figure 6 A schematic diagram of an electronic device disclosed in an embodiment of the present application is shown. Figure 6 As shown, the electronic device 100 includes: a memory 110 and a processor 120. The memory of the electronic device is not less than 12G, the main frequency of the processor is not less than 2.4GHz, the memory 110 and the processor 120 are connected through bus communication, and a computer program is stored in the memory 110. The computer program can be run on the processor 120 to implement a question and answer method based on a task-oriented guided question and answer model cluster distillation system disclosed in an embodiment of the present application.

[0214] Based on the same application concept, the fourth aspect of the embodiment of the present application discloses a computer-readable storage medium on which a computer program / instruction is stored. When the computer program / instruction is executed by the processor, a question-answering method based on a task-oriented guided question-answering model cluster distillation system disclosed in the embodiment of the present application is implemented.

[0215] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0216] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, electronic devices, and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0217] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0218] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0219] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0220] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements that are inherent to such process, method, article, or terminal device. In the absence of further restrictions, an element defined by the phrase "comprises a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0221] The above is a detailed introduction to the question-answering method, device and equipment based on a task-oriented guided question-answering model cluster distillation system provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.

Claims

1. A question answering method based on a task-oriented question answering model cluster distillation system, characterized in that: The method comprises: Obtaining device parameter values of a target device of a target model cluster for a multi-question-answering task to be deployed, wherein the device parameter values include at least one of the following: storage space size and computing speed; Determining target parameter information of the multi-question-answering task target model cluster according to the device parameter value, wherein the target parameter information includes at least one of the following: model parameter quantity, reasoning response speed, and reasoning calculation amount; Based on multiple task guidance factors of each teacher model in the multi-question answering task teacher model cluster, cluster distillation training is performed on the to-be-trained student model cluster according to the target parameter information to obtain the multi-question answering task target model cluster; Inputting the target multi-question answering task from the target device into the multi-question answering task target model cluster to obtain an answer to the target multi-question answering task; The multi-question-answering task teacher model cluster includes K teacher models; and the method further includes: Input the test multi-question answering task into the k-th teacher model, and obtain T answers output by the k-th teacher model, where the t-th answer is the answer to the t-th test question answering task in the test multi-question answering task; determining an importance score of the k-th teacher model for the t-th test question-answering task based on a difference between the t-th answer and the t-th correct answer corresponding to the t-th test question-answering task; According to the importance score of the k-th teacher model for the t-th test question-answering task and the importance score of the k-th teacher model for T test question-answering tasks, the guidance factor of the k-th teacher model for the t-th test question-answering task is determined.

2. The question answering method based on the task-oriented guided question answering model cluster distillation system according to claim 1, characterized in that: Based on multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster, cluster distillation training is performed on the to-be-trained student model cluster according to the target parameter information, including: Input the t-th sample question-answering task in the sample multi-question-answering task into the multi-question-answering task teacher model cluster, and obtain the answer of the k-th teacher model for the t-th sample question-answering task; Obtaining, based on the answer of the k-th teacher model to the t-th sample question-answering task and the guidance factor of the k-th teacher model to the t-th test question-answering task, the answer of the multi-question-answering task teacher model cluster to the t-th sample question-answering task, wherein the t-th sample question-answering task and the t-th test question-answering task are of the same task type; With the goal of learning the answers of the multi-question-answering task teacher model cluster to the sample multi-question-answering task, cluster distillation training is performed on the student model cluster to be trained according to the target parameter information.

3. The question answering method based on the task-oriented guided question answering model cluster distillation system according to claim 1, characterized in that: Based on multiple task guidance factors of each teacher model in the multi-question-answering task teacher model cluster, cluster distillation training is performed on the to-be-trained student model cluster according to the target parameter information, including: Determining a distillation loss value for the t-th sample question-answering task based on a difference between an answer of the multi-question-answering task teacher model cluster for the t-th sample question-answering task and an answer of the multi-question-answering task student model cluster for the t-th sample question-answering task; The total distillation loss is obtained based on the distillation loss of the t-th sample question-answering task and the guidance factors of the K teacher models for the t-th test question-answering task. According to the total distillation loss value and the target parameter information, cluster distillation training is performed on the student model cluster to be trained.

4. The question answering method based on the task-oriented guided question answering model cluster distillation system according to claim 3, characterized in that: The multi-question answering task student model cluster includes K student models; the method further includes: Share the guidance factor of the k-th teacher model for the t-th test question-answering task with the k-th student model; Input the t-th sample question-answering task in the sample multi-question-answering task into the multi-question-answering task student model cluster, and obtain the answer of the k-th student model for the t-th sample question-answering task; According to the answer of the kth student model to the tth sample question-answering task and the guidance factor of the kth student model to the tth test question-answering task, the answer of the multi-question-answering task student model cluster to the tth sample question-answering task is obtained.

5. The question answering method based on the task-oriented guided question answering model cluster distillation system according to claim 4, characterized in that: Inputting the target multi-question answering task from the target device into the multi-question answering task target model cluster to obtain an answer to the target multi-question answering task, including: Sharing the guidance factor of the k-th student model for the t-th test question-answering task with the k-th target model in the multi-question-answering task target model cluster; Input the multi-question-answering task from the target device into the multi-question-answering task target model cluster, and obtain the answer output by the k-th target model for the t-th target question-answering task in the target multi-question-answering task; According to the answer output by the kth target model for the tth target question-answering task and the guidance factor of the kth target model for the tth test question-answering task, the answer of the multi-question-answering task target model cluster for the tth target question-answering task is obtained.

6. The question answering method based on the task-oriented guided question answering model cluster distillation system according to claim 3, characterized in that: According to the total distillation loss value and the target parameter information, cluster distillation training is performed on the student model cluster to be trained, including: According to the total distillation loss value, updating the model parameters of the student model cluster to be trained multiple times until the training is completed; Check whether the parameter information of the trained student model cluster meets the target parameter information; When the parameter information of the trained student model cluster does not meet the target parameter information, optimizing the trained student model cluster to obtain an optimized student model cluster, wherein the optimization includes pruning and model parameter quantization; The optimized student model cluster is used as the student model cluster to be trained and returns to step: for multiple task guidance factors of each teacher model in the multi-question and answering task teacher model cluster, cluster distillation training is performed on the student model cluster to be trained according to the target parameter information to obtain the multi-question and answering task target model cluster.

7. The question answering method based on the task-oriented guided question answering model cluster distillation system according to claim 2, characterized in that: The method further comprises: Input each sample question in the sample question-answering dataset into the multi-question-answering task teacher model cluster to obtain the answer of the multi-question-answering task teacher model cluster to each sample question; Based on the answers of the multi-question answering task teacher model cluster to each sample question and combined with the task types of the T test question answering tasks, a multi-question answering task dataset is constructed; A portion of the multi-question-answering task dataset is used as the test multi-question-answering task, and the remaining portion is used as the sample multi-question-answering task.

8. A question answering device based on a task-oriented question answering model cluster distillation system, characterized in that: The device comprises: An acquisition module is used to obtain device parameter values of a target device of a target model cluster of a multi-question-answering task to be deployed, wherein the device parameter values include at least one of the following: storage space size and computing speed; A determination module, configured to determine target parameter information of the target model cluster of the multi-question-answering task based on the device parameter value, wherein the target parameter information includes at least one of the following: model parameter quantity, reasoning response speed, and reasoning calculation amount; A training module is used to perform cluster distillation training on the student model cluster to be trained based on multiple task guidance factors of each teacher model in the multi-question answering task teacher model cluster and the target parameter information to obtain the multi-question answering task target model cluster; A task question-answering module, configured to input a target multi-question-answering task from a target device into the multi-question-answering task target model cluster to obtain an answer to the target multi-question-answering task; The multi-question-answering task teacher model cluster includes K teacher models; the device further includes: A test multi-question-answer task answer acquisition module is used to input the test multi-question-answer task into the k-th teacher model, and obtain T answers output by the k-th teacher model, where the t-th answer is the answer to the t-th test question-answer task in the test multi-question-answer task; an importance score determination module, configured to determine an importance score of the kth teacher model for the tth test question-answering task based on a difference between the tth answer and the tth correct answer corresponding to the tth test question-answering task; The task guidance factor determination module is used to determine the guidance factor of the kth teacher model for the tth test question and answer task based on the importance score of the kth teacher model for the tth test question and answer task and the importance score of the kth teacher model for T test question and answer tasks.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the question answering method based on the task-oriented guided question answering model cluster distillation system according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Large language model reasoning method, device and equipment based on multi-level speculation sampling

    CN119831036A

  • Large-model end-to-end distillation deployment method, device and equipment for low-computing-power equipment and medium

    CN120066803A