Dynamic optimization-based question and answer big language model cluster collaborative question and answer method, system and equipment and medium
By dynamically optimizing the filtering of subsets of Q&A large language models and using the coordinated large language models to integrate candidate answers, the multi-model cluster computing overhead and accuracy problems are solved, and efficient and accurate Q&A services are achieved.
Patent Information
- Application Number
- CN202510828320.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multi-model cluster system has limitations in computing overhead and memory requirements, and the uncertainty of output results caused by most voting mechanisms and excessive consumption of computing resources affects the practical application and accuracy of the model cluster.
The dynamic optimization method is used to filter out a subset of Q&A large language models matching the target task, and integrate candidate answers using the coordinated large language model to improve the computing performance and accuracy of the model cluster.
On the basis of reducing computing overhead, the answer output accuracy and fault tolerance of the model cluster are improved, and the application performance and lightweight deployment capabilities of the model cluster are enhanced.
Smart Images

Figure CN120336496A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to a collaborative question - answering method, system, device, and medium for a question - answering large language model cluster based on dynamic optimization. Background Art
[0002] With the development of Large Language Models (LLMs) technology, current research is exploring how to integrate powerful single models into multi - model clusters to improve the execution efficiency and accuracy of tasks. Related research results show that collaborative work can optimize information processing and decision - making through the complementary characteristics between models. Through advanced scheduling algorithms and load - balancing strategies, models in the cluster can share knowledge and experience in real - time, forming a more powerful intelligent system to adapt to complex and changing practical application scenarios.
[0003] However, most related multi - model clusters adopt a simple integration strategy of majority voting, which cannot fully utilize the complementary advantages between multi - models. In addition, the increase in the scale of the model cluster also leads to an increase in computational overhead and memory requirements, which to a certain extent limits the actual deployment of multi - model clusters. Therefore, it is crucial to design a reasonable collaborative architecture for the model cluster to balance application performance and computational overhead, and to improve the answer output accuracy of the model cluster on the basis of reducing the computational overhead of the model cluster. Summary of the Invention
[0004] Based on the above - mentioned technical problems, the present invention provides a collaborative question - answering method, system, device, and medium for a question - answering large language model cluster based on dynamic optimization, aiming to overcome or at least partially solve the above - mentioned problems.
[0005] In the first aspect of the present invention, a collaborative question - answering method for a question - answering large language model cluster based on dynamic optimization is provided. The method includes: Through multiple rounds of question - answering instructions associated with the target task, a subset of question - answering large language models that match the target task is selected from the question - answering large language model pool. The subset of question - answering large language models includes M question - answering large language models, where M is an integer greater than 1; Through the subset of question - answering large language models, the current question - answering instruction associated with the target task is inferred to obtain M candidate answers; The M candidate answers and the first answer coordination instruction are input into the coordination large language model to obtain the final answer to the current question - answering instruction. The first answer coordination instruction is used to instruct the coordination large language model to understand the M candidate answers to determine the final answer to the current question - answering instruction.
[0006] The second aspect of the present invention provides a collaborative question-and-answer system for a large language model cluster based on dynamic optimization. The system includes: A model dynamic optimization module, configured to screen out a subset of question-and-answer large language models that match the target task from a question-and-answer large language model pool through multiple rounds of question-and-answer instructions associated with the target task. The subset of question-and-answer large language models includes M question-and-answer large language models, where M is an integer greater than 1; A candidate answer acquisition module, configured to infer the current question-and-answer instruction associated with the target task through the subset of question-and-answer large language models to obtain M candidate answers; A model integration module, configured to input the M candidate answers and a first answer coordination instruction into a coordination large language model to obtain the final answer to the current question-and-answer instruction. The first answer coordination instruction is used to instruct the coordination large language model to understand the M candidate answers to determine the final answer to the current question-and-answer instruction.
[0007] The third aspect of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the collaborative question-and-answer method for a large language model cluster based on dynamic optimization as in the first aspect of the present invention.
[0008] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the collaborative question-and-answer method for a large language model cluster based on dynamic optimization as in the first aspect of the present invention.
[0009] In the collaborative question - answering method for a large - language model cluster based on dynamic optimization proposed by the present invention, considering the performance differences of different large - language models for question - answering in the target task, through multiple rounds of dynamic optimization for the target task, based on multiple rounds of question - answering instructions associated with the target task, a subset of large - language models for question - answering with performance advantages and matching the target task is selected from the large - language model pool for question - answering. Thus, through task - adaptive cluster dynamic optimization, the present invention can effectively retain the models with relative performance advantages in the target task, thereby promoting the optimization of the computational performance and task adaptability of the overall model cluster. Then, the subset of large - language models for question - answering after dynamic optimization is used to infer M candidate answers for the current question - answering instruction associated with the target task, and the M candidate answers and the first answer coordination instruction are input into the coordination large - language model to obtain the final answer to the current question - answering instruction. Thus, by introducing a coordination model to integrate the outputs of the dynamically optimized multiple models (i.e., the subset of large - language models for question - answering), the present invention enhances the in - depth understanding of the output results of multiple models, improves the effectiveness and fault - tolerance ability of the integrated results of multiple models, and improves the accuracy of answer output of the model cluster while reducing the computational overhead of the model cluster. Compared with related model clusters, the present invention can give full play to the advantages of each model, optimize the computational overhead of the model cluster, improve the application performance of the model cluster, and achieve lightweight deployment on user terminals on the basis of improving the output performance of the model cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0011] Figure 1 is the basic structure diagram of a model cluster system shown in the related art; Figure 2 is the flowchart of the steps of a collaborative question - answering method for a large - language model cluster based on dynamic optimization shown in an embodiment of the present invention; Figure 3 is a schematic diagram of model output integration shown in an embodiment of the present invention; Figure 4 is the processing flowchart of a collaborative planning system for a model cluster based on dynamic optimization shown in an embodiment of the present invention; Figure 5 is a basic collaborative planning framework diagram of a model cluster shown in an embodiment of the present invention; Figure 6 is a schematic diagram of multiple - round dynamic optimization shown in an embodiment of the present invention; Figure 7 It is a structural block diagram of a collaborative question - answering system for a question - answering large - language model cluster based on dynamic optimization provided by an embodiment of the present invention; Figure 8 It is a schematic diagram of an electronic device shown by an embodiment of the present invention. Detailed implementation manners
[0012] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0013] The current model cluster system, such as Figure 1 shown Figure 1 is a basic structural diagram of a model cluster system shown in the related art. Most current model cluster systems collect the prediction results of multiple independent models and use a majority voting mechanism to determine the final output, reducing the output deviation that may be brought by a single model and achieving certain results in improving the accuracy and robustness of cluster decision - making. However, such systems still have the following limitations: 1. The majority voting mechanism can improve the accuracy and robustness of decision - making, but when a large number of models are integrated, the computational overhead of the system will increase significantly, resulting in high consumption of computing resources. This may lead to a longer response time and an increase in computing costs of the model cluster, restricting the practical application of the system.
[0014] 2. The majority voting mechanism aggregates the prediction results of each model in the cluster and selects the result with the most votes. This process limits the performance ceiling of the integrated result.
[0015] 3. Although the majority voting mechanism can improve the accuracy by integrating the predictions of multiple models, in the case where the outputs of each model diverge greatly, the final voting result may still be unreliable, resulting in a high degree of uncertainty in the output result. This uncertainty may lead to incorrect decisions, thus affecting the overall trust of the model cluster.
[0016] Therefore, in order to at least partially solve one or more of the above problems and other potential problems, an object of the present invention is to propose a new model cluster architecture to achieve reasonable and efficient collaborative planning of multiple model clusters. Based on this, an embodiment of the present invention proposes a collaborative question-answering method for a question-answering large language model cluster based on dynamic optimization. For a target task, on the basis of collaborative planning of the model cluster, in the dynamic optimization stage, through task-adaptive cluster dynamic optimization, models with relative performance advantages in the target task can be effectively retained, thereby promoting the optimization of the computing performance and task adaptability of the overall model cluster, realizing dynamic adjustment and optimization of the model cluster to solve the problem of large consumption of computing resources in related cluster systems; in the model integration stage, an overall planning model is introduced to integrate the outputs of multiple models, enhancing the in-depth understanding of the output results of multiple models, thereby significantly improving the effectiveness and fault tolerance of the model integration results, and thus giving full play to the advantages of each model to improve the answer output accuracy of the model cluster on the basis of reducing the computing overhead of the model cluster.
[0017] Please refer to Figure 2 , Figure 2 which is a flowchart of the steps of a collaborative question-answering method for a question-answering large language model cluster based on dynamic optimization shown in an embodiment of the present invention. As Figure 2 shown, the collaborative question-answering method for a question-answering large language model cluster based on dynamic optimization provided in this embodiment at least includes the following steps: Step S11: Through multiple rounds of question-answering instructions associated with the target task, select a subset of question-answering large language models that match the target task from the question-answering large language model pool, where the subset of question-answering large language models includes M question-answering large language models, and M is an integer greater than 1.
[0018] In this embodiment, a question-answering large language model pool is pre-deployed, and the question-answering large language model pool includes multiple trained question-answering large language models. For the target task, through multiple rounds of question-answering instructions associated with the target task, after multiple rounds of dynamic optimization, M question-answering large language models that match the target task can be selected from the question-answering large language model pool. The M question-answering large language models are all question-answering large language models with relative performance advantages in the target task, so as to form a subset of question-answering large language models based on the M question-answering large language models.
[0019] The target task of this embodiment can be any task to be processed, such as language understanding tasks, information extraction tasks, translation tasks, text generation tasks, reasoning tasks, speech recognition tasks, image processing tasks, multi-modal fusion tasks, etc. This embodiment does not limit the target task.
[0020] Step S12: Through the subset of question-answering large language models, perform reasoning on the current question-answering instruction associated with the target task to obtain M candidate answers.
[0021] In this embodiment, after obtaining the dynamically optimized subset of the question-and-answer large language model, M question-and-answer large language models in the subset of the question-and-answer large language model can be used to respectively reason about the current question-and-answer instruction associated with the target task to obtain M candidate answers. Among them, the current question-and-answer instruction associated with the target task is a question-and-answer instruction under the target task. For example, when the target task is a translation task, the current question-and-answer instruction associated with the target task can be: Please translate the following Chinese sentence into English: "The weather is really nice today."; M question-and-answer large language models can respectively output M candidate answers for this same current question-and-answer instruction; Another example is that when the target task is a reasoning task, the current question-and-answer instruction associated with the target task can be: Solve the equation: 2x + 5 = 17; Please explain the process step by step and verify whether the answer is correct; Similarly, M question-and-answer large language models can respectively output M candidate answers for this same current question-and-answer instruction. The above are only related examples of the question-and-answer instruction, and this embodiment does not impose any restrictions on this.
[0022] Step S13: Input the M candidate answers and the first answer coordination instruction into the coordination large language model to obtain the final answer to the current question-and-answer instruction, where the first answer coordination instruction is used to instruct the coordination large language model to understand the M candidate answers to determine the final answer to the current question-and-answer instruction.
[0023] In this embodiment, a pre-trained coordination large language model is introduced. This coordination large language model can understand and integrate M candidate answers, and combine the first answer coordination instruction to obtain the final answer output by the subset of the question-and-answer large language model for the current question-and-answer instruction. It can be understood that the input of the coordination large language model is M candidate answers and the first answer coordination instruction, and the output is the final answer to the current question-and-answer instruction. The coordination large language model can generate a final result (i.e., the final answer) that synthesizes M candidate answers after deeply understanding the M candidate answers. In this way, by introducing the coordination large language model to integrate the outputs of multiple models, the understanding of the results of different models can be deepened, and the depth of understanding of model integration can be improved. This integration method can more comprehensively grasp the outputs of multiple models, promote the fusion of information, make the final result more accurate and interpretable, and thus improve the effectiveness of cluster decision-making.
[0024] In this embodiment, for a target task, based on multi-round Q&A instructions associated with the target task, a subset of Q&A large language models with performance advantages and matching the target task is selected from the pool of Q&A large language models. Thus, through task-adaptive cluster dynamic optimization, models with relative performance advantages on the target task can be effectively retained, thereby promoting the optimization of the computational performance of the overall model cluster and task adaptability. Then, by integrating the outputs (i.e., M candidate answers) of the large language models for the dynamically optimized multi-models (i.e., the subset of Q&A large language models) through overall planning, the final answer for the current Q&A instruction is obtained. In this way, compared with the traditional majority voting integration method, this embodiment improves the in-depth understanding of the output results of multi-models by introducing an overall planning model for output integration, enhances the effectiveness and fault tolerance of the multi-model integration results, and improves the accuracy of the answer output of the model cluster while reducing the computational overhead of the model cluster. Compared with related model clusters, this embodiment can give full play to the advantages of each model, optimize the computational overhead of the model cluster, improve the application performance of the model cluster, and achieve lightweight deployment on user terminals on the basis of improving the output performance of the model cluster.
[0025] In one embodiment, as Figure 3 shown, Figure 3 is a schematic diagram of model output integration shown in an embodiment of the present invention. In Figure 3 , after completing the dynamic optimization of the model cluster, the M candidate answers output by the subset of Q&A large language models finally optimized can be integrated by introducing an overall planning model (such as the overall planning large language model ) to output the final answer for the current Q&A instruction. Among them, the subset of performance advantage models obtained by model dynamic optimization (i.e., the subset of Q&A large language models) is , and the overall planning large language model receives the outputs of each model in, and combines the task instructions (i.e., the first answer overall planning instruction) to generate the final output of the model cluster (the final answer for the current Q&A instruction). The process of this model integration can be expressed by the following formula: ; where, is the final output of the model cluster (the final answer for the current Q&A instruction), is the large language model as the overall planning model (i.e., the overall planning large language model), is the first answer overall planning instruction of the overall planning large language model, and is the output set of performance advantage models (i.e., M candidate answers). Compared with the traditional majority voting integration method, this embodiment improves the in-depth understanding of the output results of multi-models by introducing an overall planning large language model for output integration, thereby significantly improving the effectiveness and fault tolerance of the model integration results.
[0026] Combined with the above embodiments, in one implementation, the present invention further provides a collaborative question-answering method for a question-answering large language model cluster based on dynamic optimization. In this method, the step of "screening out a subset of question-answering large language models that match the target task from the question-answering large language model pool through multi-round question-answering instructions associated with the target task" in the above step S11 includes at least the following steps S21 to S24: Step S21: Input the initial question-answering instruction and the first-round question-answering instruction in the multi-round question-answering instructions into each question-answering large language model in the question-answering large language model pool to obtain multiple candidate answers for the first-round question-answering instruction.
[0027] In this embodiment, for the first round of multi-round dynamic optimization, the initial question-answering instruction and the first-round question-answering instruction in the multi-round question-answering instructions can be respectively input into each question-answering large language model in the question-answering large language model pool to obtain multiple candidate answers for the first-round question-answering instruction output by multiple question-answering large language models in the question-answering large language model pool. Among them, the initial question-answering instruction refers to the initial target task instruction, such as: the default task instruction corresponding to the target task. In this embodiment, the multi-round question-answering instructions associated with the target task include: question-answering instructions corresponding to each round of multi-round question-answering. For example, the first round of question-answering corresponds to the first-round question-answering instruction, the second round of question-answering corresponds to the second-round question-answering instruction, etc. The question-answering instructions corresponding to different rounds of question-answering can be the same or different.
[0028] Step S22: Determine the question-answering performance parameter values of each question-answering large language model in the question-answering large language model pool for the target task according to the differences between the multiple candidate answers for the first-round question-answering instruction and the correct answer for the first-round question-answering instruction.
[0029] In this embodiment, a correct answer (i.e., the correct answer) is preset for each round of question-answering instructions in the multi-round question-answering instructions. For example, for the first-round question-answering instruction, it corresponds to the correct answer for the first-round question-answering instruction. The question-answering performance parameter values of each question-answering large language model in the question-answering large language model pool for the target task can be determined according to the differences between the multiple candidate answers for the first-round question-answering instruction and the correct answer for the first-round question-answering instruction. Among them, the question-answering performance parameter values can be parameter values representing the performance of the question-answering large language model, such as answer accuracy, answer similarity, etc., and there is no limitation on this.
[0030] For example, for the question-and-answer large language model 1 in the question-and-answer large language model pool, the question-and-answer performance parameter value of the question-and-answer large language model 1 in the question-and-answer large language model pool for the target task can be determined according to the difference between the candidate answer 1 of the first-round question-and-answer instruction output by the question-and-answer large language model 1 and the correct answer of the first-round question-and-answer instruction. For the question-and-answer large language model 5 in the question-and-answer large language model pool, the question-and-answer performance parameter value of the question-and-answer large language model 5 in the question-and-answer large language model pool for the target task can be determined according to the difference between the candidate answer 5 of the first-round question-and-answer instruction output by the question-and-answer large language model 5 and the correct answer of the first-round question-and-answer instruction.
[0031] Step S23: Obtain the target performance parameter threshold configured by the user terminal running the subset of the question-and-answer large language models.
[0032] In this embodiment, the subset of the question-and-answer large language models after dynamic optimization and the coordinated large language model will be deployed on the user terminal, so that the user terminal can output the final answer to the current question-and-answer instruction through the subset of the question-and-answer large language models and the coordinated large language model for the current question-and-answer instruction associated with the target task. In this embodiment, the user terminal can set a target performance parameter threshold for the subset of the question-and-answer large language models to be deployed (i.e., to be run), so as to obtain the target performance parameter threshold configured by the user terminal running the subset of the question-and-answer large language models. Wherein, the target performance parameter threshold characterizes the minimum question-and-answer performance parameter of the question-and-answer large language model that can be deployed on the user terminal. For example, the target performance parameter threshold can be an answer accuracy threshold, an answer similarity threshold, etc., and there is no limitation on this.
[0033] Step S24: Screen out multiple question-and-answer large language models with question-and-answer performance parameter values greater than the target performance parameter threshold from the question-and-answer large language model pool in the order of the question-and-answer performance parameter values from large to small, as the second-round subset of the question-and-answer large language models, for reasoning on the second-round question-and-answer instruction in the multiple rounds of question-and-answer instructions.
[0034] In this embodiment, multiple question-and-answer large language models with question-and-answer performance parameter values greater than the target performance parameter threshold can be screened out from the question-and-answer large language model pool in the order of the question-and-answer performance parameter values of each question-and-answer large language model in the question-and-answer large language model pool from large to small, and the question-and-answer large language models with question-and-answer performance parameter values less than the target performance parameter threshold in the question-and-answer large language model pool are excluded, so as to use the multiple question-and-answer large language models with question-and-answer performance parameter values greater than the target performance parameter threshold as the second-round subset of the question-and-answer large language models. The second-round subset of the question-and-answer large language models is used for reasoning on the second-round question-and-answer instruction in the multiple rounds of question-and-answer instructions.
[0035] Combined with the above embodiments, the present invention also provides a collaborative question-answering method for a question-and-answer large language model cluster based on dynamic optimization. In this method, the step of "selecting a subset of question-and-answer large language models that match the target task from the question-and-answer large language model pool through multi-round question-and-answer instructions associated with the target task" in the above step S11 may further include the following steps S31 to S33: Step S31: Input the initial question-and-answer instruction, the t-th round question-and-answer instruction, and the historical answer library of the (t - 1)-th round into each question-and-answer large language model in the t-th round question-and-answer large language model subset to obtain multiple candidate answers for the t-th round question-and-answer instruction.
[0036] In this embodiment, for the t-th round in multi-round dynamic optimization, where t is an integer greater than or equal to 2, the initial question-and-answer instruction, the t-th round question-and-answer instruction in the multi-round question-and-answer instructions, and the historical answer library of the (t - 1)-th round corresponding to each question-and-answer large language model in the t-th round question-and-answer large language model subset are input into each question-and-answer large language model in the t-th round question-and-answer large language model subset to obtain multiple candidate answers for the t-th round question-and-answer instruction output by multiple question-and-answer large language models in the t-th round question-and-answer large language model subset.
[0037] For example, the t-th round question-and-answer large language model subset contains n question-and-answer large language models , and in each round of question-and-answer, the question-and-answer large language model is given the initial question-and-answer instruction, the current round instruction (i.e., the t-th round question-and-answer instruction), and the answer retrieved from the historical answer library of the (t - 1)-th round as a reference for the current answer, so as to output a candidate answer for the t-th round question-and-answer instruction. By way of example, the formula for the question-and-answer large language model to answer in the round is: ; wherein, is the candidate answer for the t-th round question-and-answer instruction output by the question-and-answer large language model , is the question-and-answer large language model, is the initial question-and-answer instruction, is the t-th round question-and-answer instruction (i.e., the task instruction of the round), is the historical answer of the question-and-answer large language model corresponding to the recent rounds. After completing the answer for the current round, the output of the model will also be stored in the library for subsequent rounds to refer to, so as to maximize the synergistic effect between multiple models and give play to the performance advantages of the multi-model cluster.
[0038] That is to say, after the question-and-answer large language model outputs a candidate answer for each round of question-and-answer instructions, it will store the candidate answer in the historical answer library corresponding to the question-and-answer large language model. For example, after the question-and-answer large language model outputs the candidate answer for the first round of question-and-answer instructions for the first round of question-and-answer instructions, it will store the candidate answer for the first round of question-and-answer instructions in the historical answer library corresponding to the first round of the question-and-answer large language model; after the question-and-answer large language model outputs the candidate answer for the second round of question-and-answer instructions for the second round of question-and-answer instructions, it will store the candidate answer for the second round of question-and-answer instructions in the historical answer library corresponding to the first round of the question-and-answer large language model, so as to update the historical answer library corresponding to the second round of the question-and-answer large language model. And so on, after the question-and-answer large language model outputs the candidate answer for the t-th round of question-and-answer instructions for the t-th round of question-and-answer instructions, it will store the candidate answer for the t-th round of question-and-answer instructions in the historical answer library corresponding to the (t - 1)-th round of the question-and-answer large language model, so as to update the historical answer library corresponding to the t-th round of the question-and-answer large language model.
[0039] Step S32: Determine the question-and-answer performance parameter values of each question-and-answer large language model in the t-th round question-and-answer large language model subset for the target task according to the differences between the multiple candidate answers of the t-th round question-and-answer instructions and the correct answer of the t-th round question-and-answer instructions.
[0040] In this embodiment, a correct answer (i.e., the correct solution) is preset for the t-th round of question-and-answer instructions in the multi-round question-and-answer instructions. For example, for the t-th round of question-and-answer instructions, there is a correct answer corresponding to the t-th round of question-and-answer instructions. The question-and-answer performance parameter values of each question-and-answer large language model in the t-th round question-and-answer large language model subset for the target task can be determined according to the differences between the multiple candidate answers of the t-th round question-and-answer instructions and the correct answer of the t-th round question-and-answer instructions.
[0041] For example, for the question-and-answer large language model y in the t-th round question-and-answer large language model subset, the question-and-answer performance parameter value of the question-and-answer large language model y in the t-th round question-and-answer large language model subset for the target task can be determined according to the difference between the candidate answer y of the t-th round question-and-answer instructions output by the question-and-answer large language model y and the correct answer of the t-th round question-and-answer instructions. It should be noted that for the same question-and-answer large language model, the question-and-answer performance parameter values for the target task in each round can be the same or different. In the current round, the screening is performed according to the question-and-answer performance parameter values for the target task in the current round.
[0042] Step S33: Screen out multiple question-and-answer large language models with the question-and-answer performance parameter values greater than the target performance parameter threshold from the question-and-answer large language model subset of the t-th round in descending order of the question-and-answer performance parameter values, and use them as the question-and-answer large language model subset of the (t + 1)-th round until the number of question-and-answer large language models included in the question-and-answer large language model subset of the (t + 1)-th round is M, so as to obtain a question-and-answer large language model subset matching the target task; the question-and-answer large language model subset of the (t + 1)-th round is used to reason about the question-and-answer instruction of the (t + 1)-th round in the multi-round question-and-answer instructions.
[0043] In this embodiment, multiple question-and-answer large language models with question-and-answer performance parameter values greater than the target performance parameter threshold can be screened out from the question-and-answer large language model subset of the t-th round in descending order of the question-and-answer performance parameter values of each question-and-answer large language model in the question-and-answer large language model subset of the t-th round (i.e., the question-and-answer performance parameter values corresponding to the t-th round), and the question-and-answer large language models with question-and-answer performance parameter values less than the target performance parameter threshold in the question-and-answer large language model subset of the t-th round are excluded, so as to use multiple question-and-answer large language models with the question-and-answer performance parameter values corresponding to the t-th round greater than the target performance parameter threshold as the question-and-answer large language model subset of the (t + 1)-th round. This question-and-answer large language model subset of the (t + 1)-th round is used to reason about the question-and-answer instruction of the (t + 1)-th round in the multi-round question-and-answer instructions, and so on, until the number of question-and-answer large language models included in the obtained question-and-answer large language model subset of the (t + 1)-th round is M. At this time, these M question-and-answer large language models are determined as the question-and-answer large language model subset matching the target task.
[0044] In this embodiment, a question-and-answer large language model subset with performance advantages and matching the target task is determined from the question-and-answer large language model pool through multi-round dynamic optimization. This process helps to improve the output stability of the model cluster, and as the number of models decreases, the response efficiency of the model cluster can be greatly improved. That is, in this embodiment, through the dynamic adjustment and optimization of the model cluster, models that perform excellently in the target task are adaptively retained, and the cluster computing efficiency is improved.
[0045] In this embodiment, through the model evaluation and adjustment in the dynamic optimization stage, models with performance advantages in the target task can be adaptively discriminated and retained, thereby effectively reducing the use of redundant models and reducing unnecessary consumption of computing resources. This targeted model adjustment improves the response efficiency of the model cluster and also reduces the operating cost of the system.
[0046] In addition, in combination with the above embodiments, in another embodiment, after obtaining multiple candidate answers to the Q&A instruction in the t-th round, according to the differences between the candidate answers to the Q&A instruction in the t-th round output by each Q&A large language model in the Q&A large language model subset of the t-th round and the answers in the corresponding historical answer library of the (t - 1)-th round of each Q&A large language model in the Q&A large language model subset of the t-th round, the Q&A performance parameter values of each Q&A large language model in the Q&A large language model subset of the t-th round for the target task can be determined. For example, for the Q&A large language model a in the Q&A large language model subset of the t-th round, according to the differences between the candidate answer a to the Q&A instruction in the t-th round output by the Q&A large language model a and the answers in the historical answer library a of the (t - 1)-th round corresponding to the Q&A large language model a, the Q&A performance parameter value of the Q&A large language model a in the Q&A large language model subset of the t-th round for the target task can be determined, such as the answer similarity is 40%. For the Q&A large language model h in the Q&A large language model subset of the t-th round, according to the differences between the candidate answer h to the Q&A instruction in the t-th round output by the Q&A large language model h and the answers in the historical answer library h of the (t - 1)-th round corresponding to the Q&A large language model h, the Q&A performance parameter value of the Q&A large language model h in the Q&A large language model subset of the t-th round for the target task can be determined, such as the answer similarity is 60% and so on.
[0047] That is to say, in this embodiment, in the t-th round, the candidate answers of each Q&A large language model are evaluated. If the candidate answer of a certain Q&A large language model is significantly different from its answer in the historical round, it indicates that this Q&A large language model lacks stability or advantages in this target task. In this embodiment, this poorly performing model can be excluded in the (t + 1)-th round. Among them, the comparison method of the answers in the current round and the historical round depends on the specific task. If there is a specific answer to the target task, the consistency of the answers can be directly compared; if the answer to the target task is open-ended, the pre-trained language model BERT is used to determine whether there is a significant semantic difference in the answers, such as determining the semantic similarity between the two. When it is determined that the semantic similarity is less than the target performance parameter threshold, it is determined that there is a significant semantic difference, and at this time, it can be excluded.
[0048] In combination with the above embodiments, in one implementation manner, the present invention also provides a collaborative Q&A method for a Q&A large language model cluster based on dynamic optimization. In this embodiment, the step of "screening out a subset of Q&A large language models that match the target task from the Q&A large language model pool through multiple Q&A instructions associated with the target task" in the above step S11 may specifically include the following steps S41 to S45: Step S41: Input the initial Q&A instruction, the t-th Q&A instruction in the multi-round Q&A instructions, and the historical answer library of the (t - 1)-th round into each Q&A large language model in the t-th round Q&A large language model subset to obtain multiple candidate answers for the t-th round Q&A instruction, where t is an integer greater than or equal to 2.
[0049] In this embodiment, for the t-th round in multi-round dynamic optimization, where t is an integer greater than or equal to 2, the initial Q&A instruction, the t-th Q&A instruction in the multi-round Q&A instructions, and the historical answer library of the (t - 1)-th round can be input into each Q&A large language model in the t-th round Q&A large language model subset to obtain multiple candidate answers for the t-th round Q&A instruction output by multiple Q&A large language models in the t-th round Q&A large language model subset. Among them, the initial Q&A instruction refers to the initial target task instruction, such as: the default task instruction corresponding to the target task. In this embodiment, the multi-round Q&A instructions associated with the target task include: Q&A instructions corresponding to each round of Q&A, such as the 1st round Q&A corresponds to the 1st round Q&A instruction, the t-th round Q&A corresponds to the t-th round Q&A instruction, etc. The Q&A instructions corresponding to different rounds of Q&A can be the same or different.
[0050] For the first round in multi-round dynamic optimization, the initial Q&A instruction and the 1st round Q&A instruction in the multi-round Q&A instructions can be respectively input into each Q&A large language model in the Q&A large language model pool to obtain multiple candidate answers for the 1st round Q&A instruction output by multiple Q&A large language models in the Q&A large language model pool. Then, the multiple candidate answers for the 1st round Q&A instruction and the second answer coordination instruction are input into the coordination large language model to obtain the final answer for the 1st round Q&A instruction output by the coordination large language model. Then, according to the difference between the multiple candidate answers for the 1st round Q&A instruction and the final answer for the 1st round Q&A instruction, the Q&A performance parameter values of each Q&A large language model in the Q&A large language model pool for the target task are determined. Finally, in the order from largest to smallest of the Q&A performance parameter values of each Q&A large language model in the Q&A large language model pool for the target task, multiple Q&A large language models with Q&A performance parameter values greater than the target performance parameter threshold are selected from the Q&A large language model pool as the 2nd round Q&A large language model subset for reasoning about the 2nd round Q&A instruction in the multi-round Q&A instructions. That is to say, the 2nd round Q&A large language model subset includes: multiple Q&A large language models selected from the Q&A large language model pool for reasoning about the 2nd round Q&A instruction.
[0051] Step S42: Input the multiple candidate answers for the t-th round Q&A instruction and the second answer coordination instruction into the coordination large language model to obtain the final answer for the t-th round Q&A instruction.
[0052] In this embodiment, multiple candidate answers to the t-th round of Q&A instructions and a second answer coordination instruction can be input into a coordinated large language model. This coordinated large language model can understand and integrate the multiple candidate answers to the t-th round of Q&A instructions, and output the final answer to the t-th round of Q&A instructions in combination with the second answer coordination instruction. The second answer coordination instruction is used to instruct the coordinated large language model to understand the multiple candidate answers to the t-th round of Q&A instructions to determine the final answer to the t-th round of Q&A instructions.
[0053] Step S43: Determine the Q&A performance parameter values of each Q&A large language model in the t-th round Q&A large language model subset for the target task according to the difference between the multiple candidate answers to the t-th round of Q&A instructions and the final answer to the t-th round of Q&A instructions.
[0054] In this embodiment, the Q&A performance parameter values of each Q&A large language model in the t-th round Q&A large language model subset for the target task can be determined according to the difference between the multiple candidate answers to the t-th round of Q&A instructions and the final answer to the t-th round of Q&A instructions. For example, for the Q&A large language model q in the t-th round Q&A large language model subset, the Q&A performance parameter value of the Q&A large language model q in the t-th round Q&A large language model subset for the target task can be determined according to the difference between the candidate answer q to the t-th round of Q&A instructions output by the Q&A large language model q and the final answer to the t-th round of Q&A instructions. It should be noted that, for the same Q&A large language model, its Q&A performance parameter values for the target task in each round can be the same or different. In the current round, screening is performed according to the Q&A performance parameter values for the target task in the current round.
[0055] Step S44: Obtain the target performance parameter threshold configured by the user terminal running the Q&A large language model subset.
[0056] Step S44 in this embodiment is the same as or similar to the above-mentioned step S23, and will not be elaborated here.
[0057] Step S45: Screen out multiple Q&A large language models with Q&A performance parameter values greater than the target performance parameter threshold from the t-th round Q&A large language model subset in descending order of the Q&A performance parameter values as the (t + 1)-th round Q&A large language model subset until the number of Q&A large language models included in the (t + 1)-th round Q&A large language model subset is M, to obtain a Q&A large language model subset matching the target task; the (t + 1)-th round Q&A large language model subset is used to reason about the (t + 1)-th round of Q&A instructions in the multiple rounds of Q&A instructions.
[0058] Step S45 in this embodiment is the same as or similar to the above-mentioned step S33, and will not be elaborated here.
[0059] Combined with the above embodiments, in one implementation, the present invention further provides a collaborative question-answering method for a question-and-answer large language model cluster based on dynamic optimization. In this method, in addition to the above steps, steps S51 to S53 may also be included: Step S51: Store the final answer to the first-round question-and-answer instruction in the historical answer library of the first round; the final answer to the first-round question-and-answer instruction is obtained by inputting multiple candidate answers to the first-round question-and-answer instruction and the second answer coordination instruction into the coordination large language model.
[0060] In this embodiment, the final answer to the first-round question-and-answer instruction can be stored in the historical answer library of the first round.
[0061] Step S52: Determine the capacity L of the historical answer library according to the storage space of the user terminal running the subset of the question-and-answer large language model.
[0062] In this embodiment, the capacity L of the historical answer library can be determined according to the storage space of the user terminal (i.e., the user terminal to which the subset of the question-and-answer large language model is to be deployed) running the subset of the question-and-answer large language model. L represents that the historical answer library can store L answers.
[0063] Step S53: Store the final answers to the question-and-answer instructions from the (t - L)-th round to the t-th round in the historical answer library of the t-th round.
[0064] In this embodiment, when t is not greater than L, the final answers to the question-and-answer instructions from the first round to the t-th round can be stored in the historical answer library of the t-th round; when t is greater than L, the final answers to the question-and-answer instructions from the (t - L)-th round to the t-th round can be stored in the historical answer library of the t-th round.
[0065] For example, when L is 50, when the t-th round does not exceed the 50th round, the final answers to the question-and-answer instructions of each previous round can be stored in the historical answer library of the t-th round. For example, in the 2nd round, the final answers to the 1st and 2nd round question-and-answer instructions are stored in the historical answer library of the 2nd round; in the 49th round, the final answers to the question-and-answer instructions from the 1st round to the 49th round are stored in the historical answer library of the 49th round; in the 50th round, the final answers to the question-and-answer instructions from the 1st round to the 50th round are stored in the historical answer library of the 50th round. When the t-th round exceeds the 50th round, the final answers to the question-and-answer instructions from the (t - L)-th round to the t-th round can be stored in the historical answer library of the t-th round. For example, in the 55th round, the final answers to the question-and-answer instructions from the 5th round to the 55th round can be stored in the historical answer library of the 55th round.
[0066] Combined with the above embodiments, in one implementation, the present invention also provides a collaborative question-answering method for a question-and-answer large language model cluster based on dynamic optimization. In this method, after the above step S12, steps S61 to S63 may further be included, and specifically, "inputting the M candidate answers and the first answer coordination instruction into the coordination large language model to obtain the final answer to the current question-and-answer instruction" in the above step S13 may specifically include step S64: Step S61: Obtain the consistency threshold configured by the user terminal running the subset of the question-and-answer large language model.
[0067] In this embodiment, the user terminal may set a consistency threshold for the subset of the question-and-answer large language model to be deployed (i.e., to be run), so as to obtain the consistency threshold configured by the user terminal running the subset of the question-and-answer large language model. This consistency threshold represents the lowest consistency for which the M candidate answers symbolize the same answer. Among them, when the consistency between any two of the M candidate answers is greater than or equal to this consistency threshold, it indicates that the M candidate answers are very similar, and the M candidate answers can be regarded as the same answer; when the consistency between any two of the M candidate answers is lower than this consistency threshold, it indicates that the M candidate answers are not similar and cannot be regarded as the same answer.
[0068] Step S62: Detect the consistency between the M candidate answers.
[0069] In this embodiment, after obtaining the M candidate answers, the consistency between the M candidate answers can be detected. For example, the consistency between any two of the M candidate answers can be detected to obtain multiple consistencies, and the smallest consistency among the multiple consistencies can be determined as the consistency between the M candidate answers.
[0070] Step S63: Compare the consistency with the consistency threshold.
[0071] In this embodiment, the consistency between the M candidate answers can be compared with the consistency threshold.
[0072] Step S64: In the case where the consistency between the M candidate answers is lower than the consistency threshold, input the M candidate answers and the first answer coordination instruction into the coordination large language model to obtain the final answer to the current question-and-answer instruction.
[0073] In this embodiment, in the case where it is determined that the consistency between the M candidate answers is not lower (i.e., greater than or equal to) the consistency threshold, it is determined that the M candidate answers are very similar and can be regarded as the same answer. At this time, any one of the M candidate answers can be determined as the final answer to the current question-and-answer instruction.
[0074] In the case where the consistency among the M candidate answers is lower than the consistency threshold, it is determined that the M candidate answers are not similar and cannot be regarded as the same answer. At this time, it is necessary to input the M candidate answers and the first answer coordination instruction into the coordinated large language model to obtain the final answer to the current question and answer instruction output by the coordinated large language model.
[0075] Combined with the above embodiments, in one implementation manner, the present invention also provides a coordinated question and answer method for a question and answer large language model cluster based on dynamic optimization. In this method, in addition to the above steps, it may further include step S71 and step S72: Step S71: Obtain the target computing resource consumption and target response speed configured by the user terminal running the subset of the question and answer large language model.
[0076] In this embodiment, the user terminal can set the target computing resource consumption and target response speed for the subset of the question and answer large language model to be deployed (i.e., to be run), so as to obtain the target computing resource consumption and target response speed configured by the user terminal running the subset of the question and answer large language model. Among them, the target computing resource consumption is the maximum computing resource consumption that the user terminal can bear to run this subset of the question and answer large language model, and the target response speed is the slowest response speed allowed by the user terminal when running this subset of the question and answer large language model.
[0077] Step S72: Determine the value of M according to the target computing resource consumption and the target response speed.
[0078] In this embodiment, the value of M can be determined according to the obtained target computing resource consumption and target response speed. In an alternative manner, it can be by testing different numbers of question and answer large language models in the question and answer large language model pool to determine different test round question and answer large language model subsets composed of different numbers of question and answer large language models, and for different test round question and answer large language model subsets, determine the test resource consumption and test response speed respectively corresponding to obtaining the final answers to different test round question and answer instructions. Then, according to the test resource consumption and test response speed, as well as the target computing resource consumption and target response speed, determine the value of M. For example, the number of question and answer large language models corresponding to the test round question and answer large language model subset with test resource consumption less than the target computing resource consumption and test response speed greater than the target response speed can be determined as the value of M.
[0079] For example, the target response speed is G, the target computing resource consumption is J, the test resource consumption corresponding to a test round question and answer large language model subset composed of 6 question and answer large language models is g, and the corresponding test response speed is j, and g < G, j > J, then M can be determined to be 6.
[0080] In one embodiment, as Figure 4 shown, Figure 4 Figure 4 is a processing flow chart of a model cluster collaborative planning system based on dynamic optimization shown in an embodiment of the present invention. In Figure 4 it, the core of the system is a data-driven model cluster. Based on the collaborative planning of the model cluster, the performance of the model cluster is dynamically optimized and adjusted by comparing the model with the historical answer library to evaluate the model performance, and the models with performance advantages in specific tasks are retained to obtain an optimized cluster; in the model integration stage, the system introduces a coordinated large language model to integrate and fuse the multi-model outputs of the optimized cluster, so as to improve the in-depth understanding of the multi-model output results, and significantly enhance the effectiveness and fault tolerance of the model integration results. This architecture design enables the model cluster to have better scalability, is easy to add or replace models according to requirements, and enhances the scalability and fault tolerance of the model cluster. In addition, through the mechanism of retaining excellent models and coordinated integration, the system can maintain a high fault tolerance, effectively ensuring the stability and reliability of the system.
[0081] In summary, this embodiment significantly reduces the computational resource consumption of the model cluster by introducing the dynamic optimization of the model cluster, innovatively introduces a coordinated model to integrate the model outputs, improves the in-depth understanding of the model cluster, enhances the scalability and fault tolerance of the model cluster, and provides a solution for the deployment of the model cluster in complex and changeable real scenarios.
[0082] In one embodiment, as Figure 5 shown, Figure 5 Figure 5 is a basic model cluster collaborative planning framework diagram shown in an embodiment of the present invention. In Figure 5 it, consider a model pool P, which contains n models , and in this framework, multiple models in the cluster are used for multiple rounds (such as the current round, the next round) of answering for a specific task to generate a final result. In each round of answering, the model will be given an initial instruction, the current round instruction, and the answer queried from the historical answer library (that is, the historical answer (updatable) in the figure, including: the common answer in the i-th round, the common answer in the i + 1-th round, etc.; where the common answer can be either the final answer in each round or multiple answers output by multiple models in each round, and there is no limitation on this) as a reference for the current answer.
[0083] In one embodiment, as Figure 6 shown, Figure 6 Figure 6 is a schematic diagram of multi-round dynamic optimization shown in an embodiment of the present invention. In Figure 6 it, in the above Figure 5Based on the shown basic model cluster collaborative planning framework, dynamic optimization of the model cluster is achieved: Considering that there are performance differences among different models for a specific task, it is desired to determine a subset of models with performance advantages (subset of large language models for question answering) from the model pool. , this process helps to improve the output stability of the model cluster, and as the number of models decreases, the response efficiency of the model cluster can be significantly improved. Specifically, the dynamic optimization of the model cluster is implemented as follows: In the current round, the answers of each model are evaluated. If the answer of a certain model significantly differs from its answer in the historical round (i.e., Figure 6 the common answer in the historical answers), it indicates that the model lacks stability or advantages in this task. At this time, this poorly performing model will be excluded in the next round.
[0084] It should be noted that for the method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0085] Based on the same inventive concept, an embodiment of the present invention provides a collaborative question answering system for a large language model cluster for question answering based on dynamic optimization. Refer to Figure 7 , Figure 7 is a structural block diagram of a collaborative question answering system for a large language model cluster for question answering based on dynamic optimization provided by an embodiment of the present invention. As Figure 7 shown, the system includes: A model dynamic optimization module, configured to screen out a subset of large language models for question answering that matches the target task from the large language model pool for question answering through multiple rounds of question answering instructions associated with the target task. The subset of large language models for question answering includes M large language models for question answering, and M is an integer greater than 1; A candidate answer acquisition module, configured to infer the current question answering instruction associated with the target task through the subset of large language models for question answering to obtain M candidate answers; A model integration module, configured to input the M candidate answers and the first answer coordination instruction into a coordinated large language model to obtain the final answer to the current question answering instruction. The first answer coordination instruction is used to instruct the coordinated large language model to understand the M candidate answers to determine the final answer to the current question answering instruction.
[0086] Optionally, the model dynamic optimization module at least includes: A first input module, configured to input an initial Q&A instruction and the first-round Q&A instruction in the multi-round Q&A instruction into each Q&A large language model in the Q&A large language model pool, to obtain multiple candidate answers to the first-round Q&A instruction; A first determination module, configured to determine the Q&A performance parameter values of each Q&A large language model in the Q&A large language model pool for the target task according to the differences between the multiple candidate answers to the first-round Q&A instruction and the correct answer to the first-round Q&A instruction; A first acquisition module, configured to acquire a target performance parameter threshold configured by a user terminal running the subset of Q&A large language models; A first screening module, configured to screen out, from the Q&A large language model pool in descending order of the Q&A performance parameter values, multiple Q&A large language models whose Q&A performance parameter values are greater than the target performance parameter threshold as a second-round Q&A large language model subset for inferring the second-round Q&A instruction in the multi-round Q&A instruction.
[0087] Optionally, the model dynamic optimization module further includes: A second input module, configured to input the initial Q&A instruction, the t-th round Q&A instruction, and the (t - 1)-th round historical answer library into each Q&A large language model in the t-th round Q&A large language model subset to obtain multiple candidate answers to the t-th round Q&A instruction, where t is an integer greater than or equal to 2; A second determination module, configured to determine the Q&A performance parameter values of each Q&A large language model in the t-th round Q&A large language model subset for the target task according to the differences between the multiple candidate answers to the t-th round Q&A instruction and the correct answer to the t-th round Q&A instruction; A second screening module, configured to screen out, from the t-th round Q&A large language model subset in descending order of the Q&A performance parameter values, multiple Q&A large language models whose Q&A performance parameter values are greater than the target performance parameter threshold as a (t + 1)-th round Q&A large language model subset until the (t + 1)-th round Q&A large language model subset contains M Q&A large language models, to obtain a Q&A large language model subset matching the target task; the (t + 1)-th round Q&A large language model subset is used to infer the (t + 1)-th round Q&A instruction in the multi-round Q&A instruction.
[0088] Optionally, the model dynamic optimization module includes: A third input module, configured to input the initial Q&A instruction, the t-th round Q&A instruction in the multi-round Q&A instruction, and the historical answer library of the (t - 1)-th round into each Q&A large language model in the t-th round Q&A large language model subset, to obtain multiple candidate answers for the t-th round Q&A instruction, where t is an integer greater than or equal to 2, and the second round Q&A large language model subset includes: multiple Q&A large language models screened from the Q&A large language model pool for reasoning on the second round Q&A instruction; A third determination module, configured to input the multiple candidate answers for the t-th round Q&A instruction and the second answer coordination instruction into the coordination large language model, to obtain the final answer for the t-th round Q&A instruction, where the second answer coordination instruction is used to instruct the coordination large language model to understand the multiple candidate answers for the t-th round Q&A instruction, so as to determine the final answer for the t-th round Q&A instruction; A fourth determination module, configured to determine the Q&A performance parameter values of each Q&A large language model in the t-th round Q&A large language model subset for the target task according to the difference between the multiple candidate answers for the t-th round Q&A instruction and the final answer for the t-th round Q&A instruction; A second acquisition module, configured to acquire the target performance parameter threshold configured by the user terminal running the Q&A large language model subset; A third screening module, configured to screen out multiple Q&A large language models whose Q&A performance parameter values are greater than the target performance parameter threshold from the t-th round Q&A large language model subset in the order of the Q&A performance parameter values from large to small, as the (t + 1)-th round Q&A large language model subset, until the number of Q&A large language models included in the (t + 1)-th round Q&A large language model subset is M, to obtain a Q&A large language model subset matching the target task; the (t + 1)-th round Q&A large language model subset is used to reason on the (t + 1)-th round Q&A instruction in the multi-round Q&A instruction.
[0089] Optionally, the system further includes: A first storage module, configured to store the final answer for the first round Q&A instruction into the historical answer library of the first round; the final answer for the first round Q&A instruction is obtained by inputting the multiple candidate answers for the first round Q&A instruction and the second answer coordination instruction into the coordination large language model; A fifth determination module, configured to determine the capacity L of the historical answer library according to the storage space of the user terminal running the Q&A large language model subset; A second storage module, configured to store the final answers for the Q&A instructions from the (t - L)-th round to the t-th round into the historical answer library of the t-th round.
[0090] Optionally, the system further includes: A third acquisition module, configured to obtain a consistency threshold configured by a user terminal that runs the subset of the question-and-answer large language model after obtaining M candidate answers; A detection module, configured to detect the consistency among the M candidate answers; A comparison module, configured to compare the consistency with the consistency threshold; A model integration module, including: A model integration sub-module, configured to input the M candidate answers and the first answer coordination instruction into the coordinated large language model when the consistency among the M candidate answers is lower than the consistency threshold, so as to obtain a final answer to the current question-and-answer instruction.
[0091] Optionally, the system further includes: A fourth acquisition module, configured to obtain a target computing resource consumption and a target response speed configured by a user terminal that runs the subset of the question-and-answer large language model; A sixth determination module, configured to determine the value of M according to the target computing resource consumption and the target response speed.
[0092] Wherein, the innovation points of the model dynamic optimization module in this embodiment are: (1) Model cluster collaboration: Multiple models collaborate to make decisions, enhancing the overall ability of the model cluster; (2) Dynamic adjustment and optimization: By comparing with the historical answer library, the system adaptively discriminates and retains models with performance advantages in specific tasks, thereby effectively reducing the use of redundant models and improving the efficiency and performance stability of the model cluster.
[0093] The innovation points of the model integration module in this embodiment are: (1) Coordinated model: Introducing a coordinated model to integrate the outputs of multiple models can deepen the system's understanding of the results of different models and significantly improve the effectiveness of the model cluster decision-making; (2) Enhancement of scalability and fault tolerance: The model integration method abandons the traditional majority voting method, uses a language model to integrate the outputs of multiple models, and combines the dynamic adjustment and optimization of the models, which helps to enhance the scalability and fault tolerance of the model cluster.
[0094] The key points of the entire question-and-answer large language model cluster collaborative question-and-answer system based on dynamic optimization in this embodiment are: (1) Adopting a model dynamic optimization module, for specific tasks, by comparing with the historical answer library, adaptively discriminating and retaining models with performance advantages, thereby improving the efficiency and performance stability of the model cluster; (2) Adopting an efficient model integration module, introducing a coordinated model to integrate the outputs of multiple models, deepening the system's understanding of the results of different models, effectively alleviating the uncertainty of the cluster output results caused by majority voting, and significantly improving the effectiveness of the model cluster decision-making.
[0095] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the collaborative question-answering method of the question-answering large language model cluster based on dynamic optimization as described in any of the above embodiments of the present invention are implemented.
[0096] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, as Figure 8 shown, Figure 8 is a schematic diagram of an electronic device shown in an embodiment of the present invention. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, the steps in the collaborative question-answering method of the question-answering large language model cluster based on dynamic optimization as described in any of the above embodiments of the present invention are implemented.
[0097] For the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment.
[0098] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts among the embodiments, refer to each other.
[0099] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention can take the form of completely hardware embodiments, completely software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0100] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowcharts and / or block diagrams can also be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0101] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 or more processes and / or blocks Figure 1 or more blocks specified in one block or more blocks.
[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process Figure 1 or more processes and / or blocks Figure 1 or more blocks specified in one block or more blocks.
[0103] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0104] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or terminal device comprising the said element.
[0105] The above has introduced in detail a method, system, device and medium for collaborative question and answer of a large language model cluster based on dynamic optimization provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for collaborative question answering in a large language model cluster for question answering based on dynamic optimization, characterized in that, The method includes: Through multi-round Q&A instructions associated with the target task, screening out a subset of Q&A large language models that match the target task from the Q&A large language model pool, where the subset of Q&A large language models includes M Q&A large language models, and M is an integer greater than 1; Through the subset of Q&A large language models, reasoning on the current Q&A instruction associated with the target task to obtain M candidate answers; Inputting the M candidate answers and the first answer coordination instruction into the coordination large language model to obtain the final answer to the current Q&A instruction, where the first answer coordination instruction is used to instruct the coordination large language model to understand the M candidate answers to determine the final answer to the current Q&A instruction.
2. The collaborative question-answering method for the question-and-answer large language model cluster based on dynamic optimization according to claim 1, characterized in that Through multi-round Q&A instructions associated with the target task, screening out a subset of Q&A large language models that match the target task from the Q&A large language model pool, at least including: Inputting the initial Q&A instruction and the first-round Q&A instruction in the multi-round Q&A instructions into each Q&A large language model in the Q&A large language model pool to obtain multiple candidate answers to the first-round Q&A instruction; Determining the Q&A performance parameter values of each Q&A large language model in the Q&A large language model pool for the target task according to the difference between the multiple candidate answers to the first-round Q&A instruction and the correct answer to the first-round Q&A instruction; Obtaining the target performance parameter threshold configured by the user terminal running the subset of Q&A large language models; Screening out multiple Q&A large language models with Q&A performance parameter values greater than the target performance parameter threshold from the Q&A large language model pool in descending order of the Q&A performance parameter values as the second-round subset of Q&A large language models for reasoning on the second-round Q&A instruction in the multi-round Q&A instructions.
3. The collaborative question-answering method for the question-and-answer large language model cluster based on dynamic optimization according to claim 2, wherein Through multi-round Q&A instructions associated with the target task, screening out a subset of Q&A large language models that match the target task from the Q&A large language model pool, further including: Inputting the initial Q&A instruction, the t-th round Q&A instruction, and the (t - 1)-th round historical answer library into each Q&A large language model in the t-th round subset of Q&A large language models to obtain multiple candidate answers to the t-th round Q&A instruction, where t is an integer greater than or equal to 2; Determining the Q&A performance parameter values of each Q&A large language model in the t-th round subset of Q&A large language models for the target task according to the difference between the multiple candidate answers to the t-th round Q&A instruction and the correct answer to the t-th round Q&A instruction; Screening out multiple Q&A large language models with Q&A performance parameter values greater than the target performance parameter threshold from the t-th round subset of Q&A large language models in descending order of the Q&A performance parameter values as the (t + 1)-th round subset of Q&A large language models until the (t + 1)-th round subset of Q&A large language models contains M Q&A large language models, obtaining a subset of Q&A large language models that match the target task; the (t + 1)-th round subset of Q&A large language models is used to reason on the (t + 1)-th round Q&A instruction in the multi-round Q&A instructions.
4. The collaborative question-answering method for the question-and-answer large language model cluster based on dynamic optimization according to claim 1, characterized in that Through multi-round Q&A instructions associated with the target task, a subset of Q&A large language models that match the target task is screened out from the Q&A large language model pool, including: Input the initial Q&A instruction, the t-th round Q&A instruction in the multi-round Q&A instructions, and the historical answer library of the (t - 1)-th round into each Q&A large language model in the t-th round Q&A large language model subset to obtain multiple candidate answers for the t-th round Q&A instruction. t is an integer greater than or equal to 2. The second round Q&A large language model subset includes: multiple Q&A large language models screened out from the Q&A large language model pool for reasoning on the second round Q&A instruction; Input the multiple candidate answers for the t-th round Q&A instruction and the second answer coordination instruction into the coordination large language model to obtain the final answer for the t-th round Q&A instruction. The second answer coordination instruction is used to instruct the coordination large language model to understand the multiple candidate answers for the t-th round Q&A instruction to determine the final answer for the t-th round Q&A instruction; Determine the Q&A performance parameter values of each Q&A large language model in the t-th round Q&A large language model subset for the target task according to the difference between the multiple candidate answers for the t-th round Q&A instruction and the final answer for the t-th round Q&A instruction; Obtain the target performance parameter threshold configured by the user terminal running the Q&A large language model subset; Screen out multiple Q&A large language models in the t-th round Q&A large language model subset whose Q&A performance parameter values are greater than the target performance parameter threshold in descending order of the Q&A performance parameter values as the (t + 1)-th round Q&A large language model subset until the number of Q&A large language models included in the (t + 1)-th round Q&A large language model subset is M, to obtain a subset of Q&A large language models that match the target task; the (t + 1)-th round Q&A large language model subset is used to reason on the (t + 1)-th round Q&A instruction in the multi-round Q&A instructions; 5. The collaborative question-answering method for a question-answering large language model cluster based on dynamic optimization according to claim 4, wherein, The method further includes: Store the final answer of the first round Q&A instruction in the historical answer library of the first round; the final answer of the first round Q&A instruction is obtained by inputting the multiple candidate answers for the first round Q&A instruction and the second answer coordination instruction into the coordination large language model; Determine the capacity L of the historical answer library according to the storage space of the user terminal running the Q&A large language model subset; Store the final answers of the Q&A instructions from the (t - L)-th round to the t-th round in the historical answer library of the t-th round.
6. The collaborative question-answering method for the question-answering large language model cluster based on dynamic optimization according to claim 1, wherein, After obtaining M candidate answers, the method further includes: Obtain the consistency threshold configured by the user terminal running the Q&A large language model subset; Detect the consistency among the M candidate answers; Compare the consistency with the consistency threshold; Input the M candidate answers and the first answer coordination instruction into the coordination large language model to obtain the final answer for the current Q&A instruction, including: In the case where the consistency among the M candidate answers is lower than the consistency threshold, input the M candidate answers and the first answer coordination instruction into the coordination large language model to obtain the final answer for the current Q&A instruction.
7. The collaborative question-answering method for the question-answering large language model cluster based on dynamic optimization according to any one of claims 1 to 6, characterized in that The method further includes: Obtain the target computing resource consumption and target response speed configured by the user terminal for running the subset of the question-and-answer large language model. Determine the value of M according to the target computing resource consumption and the target response speed.
8. A collaborative question-answering system for a large language model cluster of question-answering based on dynamic optimization, characterized in that, The system includes: A model dynamic optimization module, configured to screen out a subset of the question-and-answer large language model that matches the target task from the question-and-answer large language model pool through multiple rounds of question-and-answer instructions associated with the target task. The subset of the question-and-answer large language model includes M question-and-answer large language models, and M is an integer greater than 1. A candidate answer acquisition module, configured to infer the current question-and-answer instruction associated with the target task through the subset of the question-and-answer large language model to obtain M candidate answers. A model integration module, configured to input the M candidate answers and the first answer coordination instruction into the coordinated large language model to obtain the final answer to the current question-and-answer instruction. The first answer coordination instruction is used to instruct the coordinated large language model to understand the M candidate answers to determine the final answer to the current question-and-answer instruction.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the dynamic optimization-based coordinated question-and-answer method for a question-and-answer large language model cluster according to any one of claims 1 to 7.
10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the dynamic optimization-based coordinated question-and-answer method for a question-and-answer large language model cluster according to any one of claims 1 to 7.
Citation Information
Patent Citations
Natural language question answering method and device based on small language model cluster and medium
CN116910217A
Question and answer method and device, related equipment and computer program product
CN118733730A
Question and answer method, device and equipment based on large language model and storage medium
CN119003729A
Language model intelligent question-answering system-oriented multi-objective optimization method
CN119323271A
Question answer recommendation method, storage medium, and electronic device
WO2025025953A1