Interactive question answering system and question answering method based on multi-model parallel reasoning
By introducing multi-model parallel inference and interactive answer integration technology into the large-model collaboration system, the problem of lack of user intervention and traceability functions for answer integration in the existing technology is solved, and a more efficient and accurate question-and-answer process is achieved.
Patent Information
- Application Number
- CN202510593209.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The existing large-model collaboration technology lacks user intervention and answer traceability functions in the answer integration process, resulting in the final answer missing important information and it is difficult for users to trace the source of the answer.
An interactive question-and-answer system based on multi-model parallel inference is designed, combining pre-inference and post-inference integration methods, integrating submodules and answers through weight calculation submodules, setting model weights and integration algorithms according to user preferences, generating the final integrated answers, and providing answer traceability and visualization of differences.
It realizes that users can intervene in the process of answer integration to improve the matching degree of the final answer; at the same time, it provides answer traceability function, reducing the time cost of users when reading and evaluating multiple large model answers, and improving usage efficiency.
Smart Images

Figure CN120144724A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large model applications, and particularly to an interactive question-answering system and a question-answering method based on multi-model parallel inference. Background Art
[0002] The release of GPT-3 has enabled society and the market to recognize the value and potential of large language models (hereinafter referred to as large models). In recent years, hundreds of large model products have been successively launched by companies, laboratories, and research institutions around the world independently or in cooperation. The initially released large models were usually trained based on a large amount of multi-source heterogeneous data, which had a wide range of sources, such as news reports, social media, encyclopedias, etc. The models trained in this way focused on obtaining extensive knowledge and language patterns, and such models became general models. Based on the general models, the training data was concentrated in a specific field. For example, the training data for a medical professional model might come from medical records, medical research papers, clinical images, etc. The models trained in this way would have significantly higher performance in the professional field than the general models, and were called professional models. For ordinary users, large models can provide a lot of help in work, study, and life. However, there are currently a wide variety of large models, and the applicable application scenarios are different. Even for high-performance models, the output fluctuations in some application scenarios are still very large, making it difficult for users to trust. And it is too time-consuming for users to ask questions to numerous large models one by one and then collect and evaluate the answers, which seriously affects the production and learning efficiency of users.
[0003] In order to enable users to quickly obtain answers with relatively high confidence, currently, methods of collaborating multiple large models are adopted, and their implementation methods are divided into three categories: Merge, Ensemble, and Cooperate. Among them, the large model collaboration in the Ensemble method is divided into three types: pre-inference integration, in-inference integration, and post-inference integration:
[0004] (1) Pre-inference integration method
[0005] The pre-inference integration method is to integrate large language models before performing inference operations, and select the most suitable large model for the current task or input from numerous large models. The specific implementation method is: train an external router, and the role of this router is to select a suitable large model for this inference task according to certain criteria or rules before inference.
[0006] (2) In-inference integration method
[0007] The in-inference integration method combines the outputs of multiple large models during the inference process, that is, in the decoding step.
[0008] (3) Post-inference integration method
[0009] The post-inference integration method is to let multiple large models perform inferences separately, generate their respective outputs, and then perform integration operations on these outputs after the inferences are completed. Specifically: sort multiple large models according to performance metrics (such as the number of parameters, BERTScore, BLEURT, BARTScore, etc.), and then perform integration through cascading, selection, or fusion models.
[0010] The existing integration methods have the following deficiencies:
[0011] 1) Although pre-inference integration can screen and make integration decisions for large language models in advance, avoiding poor performance caused by using inappropriate large language models during the inference process, the single large model selected may not be able to give the most comprehensive and accurate answer, and the selection of this large model is greatly affected by the router. When the router performance is poor, it will lead to poor quality of the finally output answer.
[0012] 2) Although the final answer obtained by using the post-inference integration method integrates the answers given by different large models, users cannot intervene in the integration process according to their own needs during the integration process, which may lead to the abandonment of some information important to users. Moreover, users cannot trace the origin of the final answer, resulting in difficulties in follow-up questions. Summary of the Invention
[0013] Aiming at the deficiencies of the prior art, the present invention proposes an interactive question-answering system and question-answering method based on multi-model parallel inference, combining the pre-inference and post-inference integration methods. The specific technical solutions are as follows:
[0014] An interactive question-answering system based on multi-model parallel inference, including an interaction module, a model database, and an answer integration module;
[0015] The interaction module is used to receive the user's input, send it to the answer integration module, and display information;
[0016] The model database is used to store the information sets of n large models;
[0017] The answer integration module includes a weight calculation sub-module, an answer integration sub-module, and a statistical information calculation and visualization sub-module; among them,
[0018] The weight calculation sub-module has a model weighting algorithm adjusted by user preferences, and calculates the weights of each large model for the question proposed by the user based on the question proposed by the user and the information of each large model;
[0019] The answer integration sub-module incorporates an answer integration algorithm regulated by user preferences. Based on the answers provided by each large model to the question raised by the user, the weights of each model transmitted by the weight calculation sub-module, and the user's preference setting parameters, it calculates the final integrated answer.
[0020] The statistical information calculation and visualization sub-module calculates the statistical information of each word in the final integrated answer across all model answers based on the final integrated answer and the answers of each model to the question raised by the user. After visualizing it, it is displayed on the interaction module.
[0021] Furthermore, the specific execution process of the weight calculation sub-module is as follows:
[0022] S1.1: Initialize the user ratings of all large models to ratings without preferences, and adjust the ratings of the corresponding large models to the user input values according to the user input.
[0023] S1.2: Convert each large model label t in the set tags consisting of all large model labels in the model database into a vector representation v i through a machine learning method, and construct a model label vector set tags i ; where tags = {t v , t 1 , …, t 2 , …, t i , …, t x}, tags v = {v 1 , v 2 , …, v i , …, v x}; In addition, for all the labels corresponding to each large model M i in the model database, construct a model label matrix , where a is the number of all labels;
[0024] S1.3: Extract the keyword words in the user question through a neural network-based method in natural language processing, give the classification labels of this question based on the set tags consisting of all large model labels, select the k labels with the highest probabilities according to the sorting of the label probability values in the algorithm result, and construct a question label matrix , and record the probability value of each label ;
[0025] S1.4: Match the classification labels of the extracted user question with the labels owned by all large models in the model database, calculate the similarity matrix S i , and sort the models with successful matches in descending order of the matching degree, and normalize the matching degree to obtain the matching weight value ;
[0026] S1.5: According to the user's selection, use the large model's public score or user-defined score on the network to construct a set of scores for all large models and perform normalization. Denote the model score as score norm = {s 1 , …, s i , …, s n}, calculate the weight w i of the i-th large model for the question raised by the current user, and pass the calculated weight w i to the answer integration sub-module for subsequent calculations.
[0027] Further, in the step S1.4, the specific matching process of matching the classified labels of the extracted user questions with the labels owned by all large models in the model database is as follows:
[0028] Calculate the weighted cosine vector similarity matrix S q between the question label vector V i and the model label matrix of each large model M i in the model database. If there is a value not less than the similarity threshold h i set by the user in each row of S s , it is considered that the model M i is successfully matched, and record the number of values not less than h i in each row of S s as the model matching degree.
[0029] Further, in the S1.5, the calculation formula of the weight w i of each large model is as follows:
[0030]
[0031] where z is the preference coefficient set by the user, which is used to overall control the system's tendency for decision-making.
[0032] Further, the process by which the answer integration sub-module calculates the final integrated answer is as follows:
[0033] S2.1: First, read the answers output by each large model for this question from the model database and perform standardization processing on each one;
[0034] S2.2: Map the answers output by each large model into a high-dimensional semantic space to generate corresponding semantic embedding vectors; then, calculate the semantic similarity between all pairs of answers based on cosine similarity to construct a symmetric similarity matrix, denoted as Sim(i,j), where i and j represent different answers, and Sim(i,j) represents the semantic similarity between answer i and answer j;
[0035] S2.3: Integrate the weights w i of each large model and the semantic similarity to calculate the final integrated weight of each answer:
[0036]
[0037]
[0038] where, represents the integrated weight of the i-th answer, n represents the number of answers, is a hyperparameter for adjusting the influence of semantic similarity, with a value range of [0,1]; is the semantic similarity threshold; is a piecewise function;
[0039] S2.4: Select a large model with context understanding, reasoning, and text generation capabilities as the final answer integration model, generate a logically coherent and grammatically fluent answer by inputting the original answers and final integrated weights of each large model, and receive the keywords related to the question asked by the user, the semantic similarity threshold to adjust the semantic tendency of the final integrated answer; the answer integration sub-module also maps each word in the final integrated answer back to the original answer.
[0040] Furthermore, the interaction module specifically includes:
[0041] Initialization stage: Receive the question asked by the user, the default model data or custom model data selected by the user, and classify and display the model information;
[0042] Interaction stage for integrating answers: Receive the keywords related to the question asked by the user, the semantic similarity threshold, the modified model information, and the preference coefficient, display the integrated answer obtained from the answer integration module, and display the user question parameters and the information of each large model used for the integrated answer;
[0043] After the final integrated answer is determined: Receive the user's word selection operation and display the traceability and statistical information related to the final integrated answer.
[0044] Furthermore, the information set of each large model in the model database is represented by a tuple:
[0045] M i=(name i , source i , parameters i , score_original i , score_client i ,domain_orginal i , domain_client i , tags_original i , tags_client i , weight i ,description i , answer i ,…)
[0046] Among them, name i represents the name of the i-th large model; source i represents the source of the i-th large model (released by which company, institution or university); parameters i represents the number of parameters of the i-th large model; score_original i represents the public score of the i-th large model on the network; score_client i represents the score given by the user to the i-th large model; domain_orginal i represents the application field positioned by the publisher of the i-th large model; domain_client i is the application field that the user believes is suitable for the i-th large model according to their own needs; tags_original i represents the tags marked by the publisher for the i-th large model; tags_client i are the tags marked by the user for the i-th large model according to their own preferences; weight i represents the weight score of the i-th large model, which is obtained by a model weighting algorithm adjusted by user preferences; description i represents the description of the i-th large model by the publisher, and the user can also change it by themselves; answer i is used to store the answer inferred by the i-th large model for the question raised by the user.
[0047] Furthermore, the initial value of the semantic similarity threshold is 0.5. By increasing the value, more relevant answers are screened during the answer integration process. By decreasing the value, more potential relevant answers are introduced during the answer integration process.
[0048] Further, the standardization process specifically includes unifying case, removing punctuation, and removing stop words.
[0049] An interactive question-answering method based on multi-model parallel inference, which is implemented by an interactive question-answering system based on multi-model parallel inference. The method includes the following steps:
[0050] Step 1: The interaction module sets model information, loads default values, selects a model group, receives the user's question, and sends the user's question to all large models in the model cluster for answering;
[0051] Step 2: The weight calculation sub-module in the answer integration module classifies and locates the user's question, and calculates the weight of each large model in the model database for the user's question;
[0052] Step 3: The answer integration sub-module in the answer integration module outputs an integrated answer based on the answers given by each model to the user's question, the weights of each model transmitted by the weight calculation sub-module, and the user's preference setting parameters;
[0053] Step 4: The output integrated answer is displayed to the user in real time through the interaction module, and the user's instruction is received to confirm whether it is the final integrated answer; if so, execute Step 5; if not, receive the user's modified preference setting parameters, and the answer integration sub-module re-outputs the integrated answer;
[0054] Step 5: The statistical information calculation and visualization sub-module in the answer integration module calculates the statistical information of each word in the final integrated answer in all model answers based on the final integrated answer and the answers of each large model, and after visualization processing, it is displayed by the interaction module.
[0055] The beneficial effects of the present invention are as follows:
[0056] (1) For the system and method of the present invention, after the user asks a question each time, each model will receive the user's question and give an inference answer. The system assigns weights to all large models according to the user's question and various parameter indicators of the models in the system. The weight is only for the question asked by the user this time. The present invention can consider and retain valuable answers provided by some models with relatively low overall scores.
[0057] (2) For the system and method of the present invention, the user can personalize and intervene in the answer integration process according to their own needs (such as setting keywords, setting word frequency thresholds, etc.). The integrated answer will change with the user's intervention, making the final answer more in line with the user's needs.
[0058] (3) The system and method of the present invention can identify each word and sentence in the finally integrated answer, facilitating the user to query at any time which model or models any word or sentence in the finally integrated answer comes from, and making it more convenient for the user to trace the source.
[0059] (4) The system and method of the present invention provide a visual display of answer differences for the user, enabling the user to conveniently view information such as the consistency, difference degree, and difference points between the answers of any model and the finally integrated answer, reducing the time cost of reading and evaluating the answers of each large model, and improving the usage efficiency. Brief Description of the Drawings
[0060] Figure 1 It is a schematic diagram of the composition of the interactive question-answering system based on multi-model parallel inference of the present invention and its relationship with the user and the model cluster.
[0061] Figure 2 It is a schematic diagram of the initial state of the interaction module.
[0062] Figure 3 It is a schematic diagram of the interaction process of integrating answers.
[0063] Figure 4 It is a schematic diagram of the answer traceability and visualization process.
[0064] Figure 5 It is a flowchart of the interactive question-answering method of the present invention, where the solid arrow represents the first running process, and the dashed arrow represents the process of the user intervening in the answer according to preferences. Detailed Embodiments
[0065] The present invention will be described in detail below with reference to the drawings and preferred embodiments. The purpose and effects of the present invention will become more apparent. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0066] Technical Term Explanation:
[0067] GloVe: Global Vectors for Word Representation, word embedding of global vectors
[0068] Word2vec: Word to vector, a toolkit for obtaining word vectors (word vector) open-sourced by Google in 2013;
[0069] BERT model, Bidirectional Encoder Representations from Transformers, a pre-trained language model;
[0070] Sentence-BERT: Sentence Bidirectional Encoder Representations from Transformers, a sentence embedding representation model;
[0071] GPT: Generative Pre-trained Transformer, a generative pre-trained transformer.
[0072] On the one hand, the present invention provides an interactive question-answering system based on multi-model parallel inference. As Figure 1 shown, the system includes three parts: an interaction module, a model database, and an answer integration module.
[0073] I. Interaction Module
[0074] The interaction module is used to receive the user's input, send it to the answer integration module, and at the same time obtain the final integrated answer from the answer integration module and display it. The user's input is different at different stages of the question-answering process, which are divided into:
[0075] Initialization stage: The interaction module receives the question proposed by the user, the default model data or custom model data selected by the user, sends it to the answer integration module, and classifies and displays each model and its parameters in the model group.
[0076] Interaction stage for integrated answers: When the user is not satisfied with the integrated answer, the interaction module receives the keywords related to the question proposed by the user, the semantic similarity threshold, the modified model information, the preference coefficient, etc., sends them to the answer integration module, then obtains the regenerated integrated answer from the answer integration module and displays it, and at the same time displays the user question parameters and the information of each model used in the integrated answer.
[0077] After the final integrated answer is determined: The interaction module receives the user's word selection operation and visually displays the traceability information and statistical information of the word selection.
[0078] After the system is started, the interaction module is initialized. As Figure 2As shown in the figure, it is the display of the interactive interface in the initialization stage. Area A is the input interaction area, which serves as the entry for user questions. It adopts the form of a text input box to provide the original input data for the subsequent Q&A process. Area B is used for model database management. It uses a dual-mode switching mechanism and supports two loading modes: default model data and custom model database. Among them, the default model data directly calls the metadata such as model scores and labels preset by the system; the custom model data allows users to modify the basic information of the model (such as score weights, domain labels), and the modification results will be used as input parameters for calculating the model answer weights. Area C is used to display model information, including static information such as model name and version number, as well as configurable information such as domain labels and score weights. Area D is used for model screening, providing screening dimensions such as domain, label, and number of parameters to achieve the function of multi-condition combined screening. By default, all options are selected.
[0079] As Figure 3 shown, in the interactive stage of integrating answers, after the interaction module receives the question raised by the user, it obtains the integrated answer from the answer integration module and displays it. Among them, Area A is used to display the final integrated answer, and the integration process follows the model weights and screening conditions configured by the user. Area B is used for parameter intervention operations. First, it displays the parameters used in the generation process of the final integrated answer, provides a modification function to meet the user's need for parameter modification, and at the same time receives parameters such as keywords related to the user's question, semantic similarity threshold, modified model information, and preference coefficient input by the user. After the parameters are modified, they are sent to the answer integration module for calculation, and the answer content displayed in Area A is updated in real time with the change of parameters.
[0080] As Figure 4 shown, when the final integrated answer is determined, that is, when the user is satisfied, the interaction module provides a function of tracing the origin of the selected word, receives the user's word selection operation, and starts the data tracing and statistics process. Area B is used to display the tracing information and statistical information of the word. At the same time, some appropriate statistical information is visually displayed.
[0081] II. Model Database
[0082] The model database is used to store the information set M = {M 1 , M 2 , M 3 , …, M i , …, M n} of n large models. The information of any large model M i is represented by a tuple:
[0083] M i = (name i , source i , parameters i, score_original i , score_client i ,domain_orginal i , domain_client i , tags_original i , tags_client i , weight i ,description i , answer i ,…)
[0084] , where name i represents the name of the i-th large model; source i represents the source of the i-th large model (released by which company, institution or university); parameters i represents the number of parameters of the i-th large model; score_original i represents the public score of the i-th large model on the network; score_client i represents the score given by the user to the i-th large model; domain_orginal i represents the application field positioned by the publisher for the i-th large model; domain_client i is the application field that the user believes is suitable for the i-th large model according to their own needs; tags_original i represents the tags marked by the publisher for the i-th large model; tags_client i are the tags marked by the user for the i-th large model according to their own preferences; weight i represents the weight score of the i-th large model, obtained by a model weighting algorithm adjusted by user preferences; description i represents the description of the i-th large model by the publisher, which can also be changed by the user himself; answer i is used to store the answer inferred by the i-th large model for the question raised by the user.
[0085] III. Answer Integration Module
[0086] The answer integration module is divided into a weight calculation sub-module, an answer integration sub-module, and a statistical information calculation and visualization sub-module.
[0087] 3.1. Weight Calculation Sub-module
[0088] The input of the weight calculation sub-module is the question raised by the user and the information of each large model, and the output is the weight w of each large model for the question raised by the user. i The weight calculation sub-module internally uses a model weighting algorithm adjusted by user preferences to calculate the weight w of each large model for the question raised by the user. i The specific implementation process is as follows:
[0089] (1) Initialization configuration
[0090] First, initialize the user scores score_client of all large models i to 5.0 (representing no preference), and adjust the score of the corresponding large model according to the user input to the user input value.
[0091] (2) Vector representation of model labels
[0092] Through machine learning methods (such as Word2vec or GloVe), each large model label t in the set tags composed of all large model labels in the model database i is converted into a vector representation v i , and a model label vector set tags is constructed v ; where, tags = {t 1 , t 2 , …, t i , …, t x}, tags v = {v 1 , v 2 , …, v i , …, v x}; In addition, for all the labels corresponding to each large model M in the model database i , a model label matrix is constructed , where a is the number of all labels.
[0093] (3) Question label extraction
[0094] Extract the keyword words in the user's question through a neural network-based method in natural language processing (such as the BERT model), give the classification label of this question based on the set tags composed of all large model labels, and select the k labels with the highest probability according to the label probability values in the algorithm result (k is input by the user, default is 3, ), construct a question label matrix , and record the probability value of each label .
[0095] (4) Label matching
[0096] Match the classification tags of the extracted user questions with all the tags owned by the large models in the model database. Specifically: the question tag vector V q is compared with each model M i in the model database in terms of its model tag matrix to calculate the weighted cosine vector similarity matrix S i . If there is a value not less than the similarity threshold h i set by the user in each row of S s , it is considered that the model M i is successfully matched, and the number of values not less than h i in each row of S s is recorded as the model matching degree. Finally, the successfully matched models are sorted from high to low according to the matching degree, and the matching degree is normalized to obtain the matching weight value .
[0097] Among them, the calculation formula of the similarity matrix is as follows:
[0098]
[0099] Among them, represents the Hadamard product operation, is a matrix with a size of , representing the similarity of each question tag to each model tag.
[0100] (5)Model weight calculation
[0101] According to the user's selection, use the public score of the model on the network or the user-defined score, construct a set of all model scores and perform normalization processing. Denote the model score as score norm ={s 1 ,…,s i ,…,s n}, calculate the weight value w i of the i-th model for the question proposed by the current user, and transfer the calculated weights w i of each model to the answer integration sub-module for subsequent calculations.
[0102] The calculation formula for the weight value w i of the i-th model for the question proposed by the current user is as follows:
[0103]
[0104] Among them, z is the preference coefficient set by the user, which is used to overall control the decision-making tendency of the system.
[0105] 3.2 Answer integration sub-module
[0106] The input of the answer integration sub-module is the answers given by each large model to the question raised by the user, the weights w of each model passed by the weight calculation sub-module i and the user's preference setting parameters. The output is the final integrated answer, which is used to solve the problems that the user cannot intervene in the answer integration process and the final answer cannot be traced. The answer integration sub-module has an answer integration algorithm regulated by user preferences, and the specific execution process is as follows:
[0107] 1) Text preprocessing
[0108] First, read the answers output by each model for this question from the model database and perform standardization processing one by one to eliminate the influence of format and grammar and ensure the accuracy of semantic calculation, including unifying case, removing punctuation, removing stop words, etc.
[0109] 2) Semantic weight calculation
[0110] Use the Sentence-BERT model to map the answers output by each model into a high-dimensional semantic space to generate corresponding semantic embedding vectors. Then, calculate the semantic similarity between all pairs of answers based on cosine similarity to construct a symmetric similarity matrix, denoted as Sim(i,j), where i and j represent different answers, and Sim(i,j) represents the semantic similarity between answer i and answer j.
[0111] 3) Integrated weight calculation
[0112] Integrate the model weights and semantic similarity to calculate the final integrated weight of each answer. Suppose there are n answers, and the integrated weight of the i-th answer is calculated by the following formula:
[0113]
[0114]
[0115] where, w i represents the model weight of the i-th model obtained by the model weighting algorithm regulated by user preferences. The initial value is set to the weight obtained by the weight calculation module, which reflects its reliability. Subsequently, the user can adjust the model weight according to their own preferences to intervene in the content of the final integrated answer.
[0116] is a hyperparameter for adjusting the influence of semantic similarity, and its value range is [0,1]. The user can adjust it according to the task requirements. For tasks that require highly consistent answers (such as calculation problems), increase the value; for tasks that require a comprehensive diversity perspective (such as open-ended questions), decrease the value.
[0117] is the semantic similarity threshold, which is used to filter out answer pairs with low semantic similarity to reduce the interference of irrelevant information. The initial value is recommended to be set to 0.5, and it can be adjusted according to experimental data to meet the specific scenario requirements. Users can control the integration of answers by adjusting the value to improve the value, and the system will more strictly screen highly relevant answers, which is suitable for scenarios with high consistency requirements; reducing the value, the system can introduce more potentially relevant answers, which is suitable for questions that require heuristic discussions.
[0118] 4) Answer generation
[0119] Select a large model with context understanding, reasoning, and text generation capabilities as the final answer integration model (the specific criteria are: large-scale multi-task language understanding is not less than 80%; the context window size is not less than 32kb, and GPT is selected in this embodiment). Generate a logically coherent and grammatically fluent answer by inputting the original answers of each model and the final integration weights, and users can intervene in the semantic tendency of the final integrated answer by setting keywords. In the generated answer, the answer integration sub-module will map each word in the final integrated answer back to the original answer (retaining the mapping relationship between the original answer and the final integrated answer) to implement the answer traceability function, that is, each word in the final integrated answer can be traced back to its main contributing model. Users can understand the impact of each model on the final integrated answer through the traceability information, and can intervene in the content of the final integrated answer according to their own preferences by adjusting various parameters.
[0120] 3.3 Statistical information calculation and visualization sub-module
[0121] The input of the statistical information calculation and visualization sub-module is the final integrated answer and the answers of each model to the question proposed by the user. The output is the statistical information of each word in the final integrated answer in all model answers, such as the percentage of answers using the word, the number of times the word appears in each answer, the number of synonyms / antonyms, etc. For appropriate data forms, corresponding visualization methods such as pie charts, bar charts, word clouds, etc. are used for display.
[0122] On the other hand, the present invention also provides an interactive question-answering method based on multi-model parallel reasoning, as Figure 5 shown. This method includes the following steps:
[0123] Step 1: The interaction module sets the model information, loads the default values, selects the model group, receives the user's question, and sends the user's question to all large models in the model cluster for answering;
[0124] Step 2: The weight calculation sub-module in the answer integration module classifies and locates the user's question, and calculates the weight of each large model in the model database for the user's question;
[0125] Step 3: The answer integration sub-module in the answer integration module outputs the integrated answer according to the answers given by each model for the user's question, the weights of each model transmitted by the weight calculation sub-module, and the user's preference setting parameters;
[0126] Step 4: The output integrated answer is displayed to the user in real time through the interaction module, and the user's instruction is received to confirm whether it is the final integrated answer; if so, proceed to Step 5; if not, receive the modified preference setting parameters of the user, and the answer integration sub-module re-outputs the integrated answer;
[0127] Step 5: The statistical information calculation and visualization sub-module in the answer integration module calculates the statistical information of each word in the final integrated answer in all model answers according to the final integrated answer and the answers of each large model, and after visualization processing, it is displayed by the interaction module.
[0128] Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not used to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, for those skilled in the art, they can still modify the technical solutions described in the foregoing examples, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the invention shall be included within the protection scope of the invention.
Claims
1. An interactive question-answering system based on multi-model parallel reasoning, characterized in that: Includes interactive module, model database and answer integration module; The interactive module is used to receive user input, send it to the answer integration module, and display information; The model database is used to store information sets of n large models; The answer integration module includes a weight calculation submodule, an answer integration submodule and a statistical information calculation and visualization submodule; wherein, The weight calculation submodule has a built-in model weighting algorithm adjusted by user preferences, and calculates the weight of each large model for the question raised by the user based on the question raised by the user and the information of each large model; The answer integration submodule has a built-in answer integration algorithm adjusted by user preferences, which calculates the final integrated answer based on the answers given by each large model to the question raised by the user, the weights of each model delivered by the weight calculation submodule, and the user's preference setting parameters; The statistical information calculation and visualization submodule calculates the statistical information of each word in the final integrated answer in all model answers based on the final integrated answer and the answers of each model to the question raised by the user, and displays it in the interactive module after visualization.
2. The interactive question-answering system based on multi-model parallel reasoning according to claim 1, characterized in that: The specific execution process of the weight calculation submodule is as follows: S1.1: Initialize the user ratings of all large models to unbiased ratings, and adjust the ratings of the corresponding large models to the user input values according to the user input; S1.2: Use machine learning methods to transform each large model tag t in the set tags composed of all large model tags in the model database i Convert to vector representation v i , build model tag vector set tags v ; Among them, tags={t1,t2,…,t i ,…,t x }, tags v ={v1,v2,…, v i ,…,v x }; In addition, for each large model M in the model database i Corresponding to all labels, build the model label matrix , where a is the number of all labels; S1.3: Extract keywords from user questions through neural network-based methods in natural language processing, give classification labels for the questions based on the set tags composed of all large model labels, select the k labels with the highest probability according to the label probability values in the algorithm results, and construct the question label matrix , and record the probability value of each label ; S1.4: Match the extracted classification labels of user questions with the labels of all large models in the model database and calculate the similarity matrix S i , and sort the successfully matched models from high to low according to the matching degree, and normalize the matching degree to obtain the matching weight value ; S1.5: Based on the user's choice, use the public score of the big model on the network or the user-defined score to build a collection of scores of all big models and normalize them. The model score is recorded as score norm ={s1,…,s i ,…,s n }, calculate the weight w of the i-th large model for the question raised by the current user i , and the calculated weight w i Passed to the answer integration submodule for subsequent calculations.
3. The interactive question-answering system based on multi-model parallel reasoning according to claim 2 is characterized in that: In step S1.4, the specific matching process of matching the extracted classification labels of user questions with the labels owned by all large models in the model database is as follows: The question label vector V q With each large model M in the model database i The model label matrix Calculate the weighted cosine vector similarity matrix S i , if S i Each row in the table has a similarity value no less than the user-set threshold h. s The value of , then the model M i Matching is successful, and S is recorded i In each row, there is no less than h s The number of values is taken as the model fit.
4. The interactive question-answering system based on multi-model parallel reasoning according to claim 2, characterized in that: In S1.5, the weights w of each large model i The calculation formula is as follows: ; Among them, z is the preference coefficient set by the user, which is used to control the tendency of the overall control system for decision-making.
5. The interactive question-answering system based on multi-model parallel reasoning according to claim 2, characterized in that: The process by which the answer integration submodule calculates the final integrated answer is as follows: S2.1: First, read the answers output by each large model for this problem from the model database and standardize them one by one; S2.2: Map the answers output by each large model into a high-dimensional semantic space to generate the corresponding semantic embedding vector; then, calculate the semantic similarity between all answer pairs based on cosine similarity and construct a symmetric similarity matrix, denoted as Sim(i,j), where i and j represent different answers and Sim(i,j) represents the semantic similarity between answer i and answer j; S2.3: Comprehensive weights w of each large model i And semantic similarity, calculate the final integrated weight of each answer: ; ; in, represents the integrated weight of the i-th answer, n represents the number of answers, is a hyperparameter for adjusting the influence of semantic similarity, with a value range of [0,1]; is the semantic similarity threshold; is a piecewise function; S2.4: Select a large model with context understanding, reasoning and text generation capabilities as the final answer integration model, generate a logically coherent and grammatically fluent answer by inputting the original answers of each large model and the final integration weight, and receive keywords related to the question entered by the user and the semantic similarity threshold to adjust the semantic tendency of the final integrated answer; the answer integration submodule also maps each word in the final integrated answer back to the original answer.
6. The interactive question-answering system based on multi-model parallel reasoning according to claim 1, characterized in that: The interaction module specifically includes: Initialization phase: receiving questions raised by users, default model data or custom model data selected by users, and displaying model information by category; The interactive stage of answer integration: receiving keywords related to the question, semantic similarity threshold, modified model information, and preference coefficient input by the user, displaying the integrated answer obtained from the answer integration module, and displaying the user's question parameters and the information of each large model used in the integrated answer; After the final integrated answer is determined: receive the user's word-marking operation, and display the traceability and statistical information related to the final integrated answer.
7. The interactive question-answering system based on multi-model parallel reasoning according to claim 1, characterized in that: The information set of each large model in the model database is represented by a tuple: M i =(name i , source i , parameters i , score_original i , score_client i , domain_orginal i , domain_client i , tags_original i , tags_client i , weight i , description i ,answer i ,…) Among them, name i Indicates the name of the i-th large model; source i Indicates the source of the i-th large model (which company, institution or university published it); parameters i Indicates the number of parameters of the i-th large model; score_original i Indicates the public score of the ith large model on the network; score_client i Indicates the user's rating of the i-th large model; domain_orginal i Indicates the application domain targeted by the publisher of the i-th large model; domain_client i It is the application field that the user considers suitable for the i-th large model according to his own needs; tags_original i Indicates the tags that the publisher of the i-th large model has marked for it; tags_client i is the label that the user marks for the i-th large model according to his or her own preference; weight i represents the weight score of the ith large model (derived by the model weighting algorithm adjusted by user preferences); description i It represents the publisher's description of the i-th large model, which can also be modified by the user; answer i Used to store the answers inferred by the i-th large model to questions raised by users.
8. The interactive question-answering system based on multi-model parallel reasoning according to claim 1, characterized in that: Semantic similarity threshold The initial value is 0.5, increasing value, the more relevant answers will be selected during the answer integration process, reducing value, more potential relevant answers are introduced into the answer integration process.
9. The interactive question-answering system based on multi-model parallel reasoning according to claim 5, characterized in that: The standardization process specifically includes unifying upper and lower case, removing punctuation marks, and removing stop words.
10. An interactive question-answering method based on multi-model parallel reasoning, characterized in that: The method is implemented by the interactive question-answering system based on multi-model parallel reasoning according to claim 1, and the method comprises the following steps: Step 1: The interactive module sets model information, loads default values, selects a model cluster, receives user questions, and sends user questions to all large models in the model cluster for answering; Step 2: The weight calculation submodule in the answer integration module classifies and locates the user's questions and calculates the weight of each large model in the model database for the user's questions; Step 3: The answer integration submodule in the answer integration module outputs an integrated answer based on the answers given by each model to the question raised by the user, the weights of each model passed by the weight calculation submodule, and the user's preference setting parameters; Step 4: The output integrated answer is displayed to the user in real time through the interactive module, and the user's instruction is received to confirm whether it is the final integrated answer; If yes, then execute step 5; if no, then receive the preference setting parameters modified by the user, and the answer integration submodule re-outputs the integrated answer; Step 5: The statistical information calculation and visualization submodule in the answer integration module calculates the statistical information of each word in the final integrated answer in all model answers based on the final integrated answer and the answers of each large model, and displays it through the interactive module after visualization.
Citation Information
Patent Citations
Multi-engine intelligent question answering system for multi-type knowledge base
CN115238101A
Answer generation and display method and system
CN118277535A
Data processing method and device, equipment and medium
CN119474961A
Hybrid interaction system based on AI large model
CN119669923A
Fusion question and answer method, device and equipment of mixed expert large language model and medium
CN119692477A
Cited By
Question and answer model cluster collaborative question and answer method, system and device based on dual reinforcement learning and medium
CN120338123A
Model cluster driven hybrid enhanced question-answering system, method and equipment and medium
CN120338124A
Model cluster driven hybrid enhanced question answering system, method, device and medium
CN120338124B
Question and answer method based on large model and data agent
CN121413751A