Intelligent Question Answering Method, Device and System Based on Cognitive Reasoning
By dynamically selecting the optimal large language model and optimizing the call sequence, the problem of the intelligent question-answer system being inefficient and resource consumption when dealing with complex and multi-field cross-questions is solved, and efficient and accurate question-and-answer and multi-field cross-questions are achieved.
Patent Information
- Application Number
- CN202510210521.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-25
AI Technical Summary
When dealing with complex and intersecting problems in multiple fields, existing intelligent question-and-answer systems are not efficient, have high resource consumption, and lack effective processing mechanisms, so they cannot fully utilize the advantages of large language models.
Through the cloud model calling server and question-and-answer container, the optimal large language model is dynamically selected to answer questions, comprehensively considering answer efficiency, resource occupation and communication delay, and optimizing the order and method of the large language model.
It improves the efficiency and accuracy of Q&A, reduces resource consumption and communication delays, supports cross-questioning in multiple fields, and ensures that questions can be answered in a comprehensive and accurate manner.
Smart Images

Figure CN119691139B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence, and in particular to an intelligent question-answering method, device and system based on cognitive reasoning. Background Art
[0002] With the rapid development of information technology, intelligent question-answering systems have become an important bridge connecting people with information and services. Traditional question-answering systems often rely on methods such as keyword matching, preset rules or templates. These methods perform well when dealing with simple, well-structured questions, but their limitations and shortcomings are fully exposed when faced with complex, multi-domain cross-cutting questions that require deep understanding. In recent years, with the continuous advancement of artificial intelligence technology, especially the rise of large language models, new hope has been brought to intelligent question-answering systems. With its powerful natural language processing capabilities and extensive knowledge coverage, large language models can more accurately understand user intentions and provide more accurate and comprehensive answers.
[0003] However, large language models are not omnipotent. Their performance in different fields varies, and problems such as high resource consumption and long response time also limit their widespread application. How to efficiently and reasonably utilize large language model resources to achieve fast and accurate question and answer has become a major challenge facing the current intelligent question and answer system. In addition, with the increasing diversification of user needs, single-field question and answer can no longer meet actual needs, and multi-field cross-question and answer has become a new trend. This requires the intelligent question and answer system to automatically identify the problem domain and dynamically adjust the question and answer strategy to ensure that each question can get the most appropriate answer.
[0004] Although existing technologies have proposed some solutions to the above problems, such as using domain classifiers to identify the problem domain and then calling a large language model in the corresponding domain to answer the problem, these solutions often ignore the differences in resource usage and communication delays of different large language models, resulting in low overall system efficiency. At the same time, existing technologies lack effective processing mechanisms for multi-domain cross-problems, and can often only simply split the problems and handle them separately, ignoring the correlation and integrity between the problems. Summary of the invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide an intelligent question answering method based on cognitive reasoning, comprising the following steps:
[0006] Step 1: The cloud model calls the server to obtain the answer efficiency of the large language model in the corresponding field according to the answer speed and accuracy of different large language models in different fields, and obtains the large language model sequence in the corresponding field according to the answer efficiency of the large language model in the corresponding field. The cloud model calls the server to obtain the question information to be answered sent by the question-answering device and generate a question-answering container;
[0007] Step 2: The Q&A container obtains the field of the question to be answered based on the question information to be answered sent by the Q&A device. If the field of the question to be answered is a single-field question, proceed to Step 3; otherwise, proceed to Step 4;
[0008] Step 3: According to the field of the question to be answered, obtain the sequence of large language models in the corresponding field. According to the answering efficiency of each large language model, the Q&A container sends the question information to be answered to the large language model with the highest answering efficiency to generate the Q&A content, and proceed to Step 6;
[0009] Step 4: According to the fields included in the question to be answered, obtain the sequence of large language models in each corresponding field. According to the resource occupancy of the large language models in the sequence of large language models in the corresponding field, obtain the call priorities of each large language model in the sequence of large language models in the corresponding field, and according to the communication latency between the Q&A device and each large language model in the sequence of large language models in the corresponding field and the large language model call priorities, obtain the large language model Q&A priorities, and obtain the large language model Q&A sequence in the corresponding field according to the Q&A priorities;
[0010] Step 5: According to the large language model with the highest Q&A priority in the large language model Q&A sequence in each corresponding field and the order of the fields of the question to be answered, obtain the large language Q&A model sequence corresponding to the question to be answered. The Q&A container sends the question to be answered to the large language Q&A model sequence corresponding to the question to be answered to generate the Q&A content;
[0011] Step 6: The Q&A container returns the Q&A content to the Q&A device to complete the intelligent Q&A.
[0012] Furthermore, the cloud model call server obtains the answering efficiency of the large language model in the corresponding field according to the answering speed and accuracy rate of different large language models in different fields, including:
[0013] Answering efficiency = γ × Answering speed + δ × Accuracy rate
[0014] where γ represents the weight coefficient of the answering speed, δ represents the weight coefficient of the accuracy rate.
[0015] Furthermore, the method of obtaining the call priorities of each large language model in the sequence of large language models in the corresponding field according to the resource occupancy of the large language models in the sequence of large language models in the corresponding field, and obtaining the large language model Q&A priorities according to the communication latency between the Q&A device and each large language model in the sequence of large language models in the corresponding field and the large language model call priorities, includes:
[0016] The resource occupancy of the large language model includes computing power occupancy, answer load, and answer efficiency. The call priority of the large language model is obtained based on the resource occupancy, using the following formula:
[0017] Call priority = α × Computing power occupancy + β × Answer load + Answer efficiency
[0018] Wherein, α represents the computing power occupancy weight coefficient, β represents the answer load weight coefficient;
[0019] Based on the communication delay between the question - answering device and each large language model in the large language model sequence of the corresponding field and the call priority of the large language model, the question - answering priority of the large language model is obtained:
[0020] Large language model question - answering priority = α × Computing power occupancy + β × Answer load + Answer efficiency + Communication delay between the question - answering device and the large language model.
[0021] Furthermore, the communication delay between the question - answering device and each large language model in the large language model sequence of the corresponding field includes:
[0022] The question - answering device sends a test request with a timestamp to the large language model, records the timestamps of the responses of each large language model, and determines the communication delay by calculating the time difference between the request sending and response receiving.
[0023] Furthermore, the question - answering container sends the question to be answered to the large language question - answering model sequence corresponding to the question to be answered, and generates question - answering content, including:
[0024] The question - answering container decomposes the question to be answered into sub - questions in the corresponding field, obtains the sub - question sequence in the corresponding field according to the order of the fields for answering the questions, sends the sub - questions in the corresponding field to the large language question - answering model sequence corresponding to the question to be answered according to the sub - question sequence in the corresponding field, and sends the question - answering result returned by the large language question - answering model and the next sub - question in the corresponding field to the next large language question - answering model until the question - answering of all sub - questions in the corresponding field is completed to obtain the question - answering content.
[0025] The intelligent question - answering device based on cognitive reasoning applies the intelligent question - answering method based on cognitive reasoning, and includes a communication device, a delay test module, a data processing module, and an information input module; the communication device, the delay test module, and the information input module are respectively connected to the data processing module.
[0026] An intelligent question-answering system based on cognitive reasoning, which applies the intelligent question-answering device based on cognitive reasoning, includes a cloud model call server and a communication device;
[0027] The intelligent question-answering device based on cognitive reasoning is communicatively connected to the cloud model call server through the communication device.
[0028] The beneficial effects of the present invention are as follows: improving the efficiency and accuracy of question answering: by comprehensively considering the answering efficiency and resource occupancy of large language models in different fields, the present invention can dynamically select the optimal large language model to answer questions, thereby significantly improving the efficiency and accuracy of question answering.
[0029] Reducing resource consumption: through refined resource management and scheduling, the present invention avoids unnecessary resource waste. Especially when dealing with cross-domain problems, by optimizing the call order and method of large language models, the overall resource consumption of the system is effectively reduced.
[0030] Reducing communication latency: by measuring and considering communication latency in real time, the present invention can select the large language model with the minimum communication latency to answer questions, thereby reducing the user waiting time and improving the user experience.
[0031] Supporting cross-domain question answering: the present invention can automatically identify and process cross-domain questions. By constructing a sequence of large language question-answering models, it ensures that questions can be comprehensively and accurately answered in the optimal order and manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic flowchart of an intelligent question-answering method based on cognitive reasoning;
[0033] Figure 2 It is a schematic diagram of the principle of an intelligent question-answering device based on cognitive reasoning;
[0034] Figure 3 It is a schematic diagram of the principle of an intelligent question-answering system based on cognitive reasoning. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The technical solution of the present invention will be further described in detail below with reference to the drawings, but the protection scope of the present invention is not limited to the following.
[0036] The features and performance of the present invention will be further described in detail below with reference to the embodiments.
[0037] As Figure 1 shown, the intelligent question-answering method based on cognitive reasoning includes the following steps:
[0038] Step 1: The cloud model call server obtains the answering efficiency of large language models in the corresponding fields based on the answering speed and accuracy of different large language models in different fields, obtains the large language model sequence in the corresponding fields according to the answering efficiency of large language models in the corresponding fields, and the cloud model call server obtains the information of the question to be answered sent by the Q&A device and generates a Q&A container;
[0039] Step 2: The Q&A container obtains the field of the question to be answered according to the information of the question to be answered sent by the Q&A device. If the field of the question to be answered is a single-field question, go to Step 3; otherwise, go to Step 4;
[0040] Step 3: According to the field of the question to be answered, obtain the large language model sequence in the corresponding field. According to the answering efficiency of each large language model, the Q&A container sends the information of the question to be answered to the large prediction model with the highest answering efficiency, generates Q&A content, and enters Step 6;
[0041] Step 4: According to the fields included in the question to be answered, obtain the large language model sequences in each corresponding field. According to the resource occupancy of the large language models in the large language model sequence in the corresponding field, obtain the call priorities of each large language model in the large language model sequence in the corresponding field respectively, and according to the communication delay between the Q&A device and each large language model in the large language model sequence in the corresponding field and the large language model call priority, obtain the large language model Q&A priority, and obtain the large language model Q&A sequence in the corresponding field according to the Q&A priority;
[0042] Step 5: According to the large language model with the highest Q&A priority in each large language model Q&A sequence in the corresponding field and the sequence of the fields of the question to be answered, obtain the large language Q&A model sequence corresponding to the question to be answered, and the Q&A container sends the question to be answered to the large language Q&A model sequence corresponding to the question to be answered to generate Q&A content;
[0043] Step 6: The Q&A container returns the Q&A content to the Q&A device to complete the intelligent Q&A.
[0044] The cloud model call server obtains the answering efficiency of large language models in the corresponding fields based on the answering speed and accuracy of different large language models in different fields, including:
[0045] Answering efficiency = γ × Answering speed + δ × Accuracy
[0046] Among them, γ represents the weight coefficient of the answering speed, δ represents the weight coefficient of the accuracy.
[0047] Obtaining the call priorities of each large language model in the large language model sequence of the corresponding field according to the resource occupancy of the large language models in the large language model sequence of the corresponding field, and obtaining the large language model Q&A priority according to the communication delay between the Q&A device and each large language model in the large language model sequence of the corresponding field and the large language model call priority, including:
[0048] The resource occupancy of the large language model includes computing power occupancy, answer load, and answer efficiency. The call priority of the large language model is obtained according to the resource occupancy, and the following formula is used:
[0049] Call priority = α × Computing power occupancy + β × Answer load + Answer efficiency
[0050] Where, α Represents the computing power occupancy weight coefficient, β Represents the answer load weight coefficient;
[0051] According to the communication delay between the Q&A device and each large language model in the large language model sequence of the corresponding field and the large language model call priority, the large language model Q&A priority is obtained:
[0052] Large language model Q&A priority = α × Computing power occupancy + β × Answer load + Answer efficiency + Communication delay between the Q&A device and the large language model.
[0053] The communication delay between the Q&A device and each large language model in the large language model sequence of the corresponding field includes:
[0054] The Q&A device sends a test request with a timestamp to the large language model, records the timestamps of the responses of each large language model, and determines the communication delay by calculating the time difference between the request sending and response receiving.
[0055] The Q&A container sends the question to be answered to the large language Q&A model sequence corresponding to the question to be answered to generate Q&A content, including:
[0056] The Q&A container decomposes the question to be answered into sub-questions in the corresponding field, obtains the sub-question sequence in the corresponding field according to the order of the fields for answering the questions, sends the sub-questions in the corresponding field to the large language Q&A model sequence corresponding to the question to be answered according to the sub-question sequence in the corresponding field, and sends the Q&A result returned by the large language Q&A model and the next sub-question in the corresponding field to the next large language Q&A model until the Q&A of all sub-questions in the corresponding field is completed to obtain the Q&A content.
[0057] Such as Figure 2As shown in the figure, an intelligent Q&A device based on cognitive reasoning, applying the intelligent Q&A method based on cognitive reasoning, includes a communication device, a latency test module, a data processing module, and an information input module; the communication device, the latency test module, and the information input module are respectively connected to the data processing module.
[0058] As Figure 3 shown in the figure, an intelligent Q&A system based on cognitive reasoning, characterized by applying the intelligent Q&A device based on cognitive reasoning, includes a cloud model call server and a communication device;
[0059] The intelligent Q&A device based on cognitive reasoning is communicatively connected to the cloud model call server through the communication device.
[0060] Specifically, the present invention proposes an intelligent Q&A method based on cognitive reasoning. Through the collaborative work of components such as the cloud model call server and the Q&A container, it realizes the rapid and accurate answering of questions in different fields, while effectively reducing resource consumption and communication latency. The specific steps are as follows:
[0061] Step 1: Construct a large language model sequence
[0062] The cloud model call server first calculates the answering efficiency of the large language model in the corresponding field according to the answering speed and accuracy rate of different large language models in different fields. The calculation formula for the answering efficiency is: Answering efficiency = γ × Answering speed + δ × Accuracy rate, where γ and δ are the weight coefficients of the answering speed and accuracy rate respectively, and can be adjusted according to the actual situation. Based on the answering efficiency, the server generates a large language model sequence for the corresponding field for subsequent calls.
[0063] When the Q&A device sends the information of the question to be answered, the cloud model call server receives it and generates a Q&A container for subsequent question processing.
[0064] Step 2: Question domain identification
[0065] The Q&A container uses natural language processing technology to identify the domain of the question according to the information of the question to be answered. If the question belongs to a single domain, it directly enters Step 3; if the question involves multiple domains, it enters Step 4 for special processing.
[0066] Step 3: Single domain question processing
[0067] For single domain questions, the Q&A container selects the large language model with the highest answering efficiency from the large language model sequence generated in Step 1 to answer the question. This can not only ensure the accuracy of the answer but also improve the response speed.
[0068] Step 4: Multi-domain cross question processing
[0069] For multi-domain intersection problems, the Q&A container first obtains the sequence of large language models in the corresponding domains according to the domains involved in the question. Then, considering factors such as the resource occupancy of the large language models (including computing power occupancy, answer load, and answer efficiency), communication latency, etc., the Q&A priorities of each large language model are calculated. The specific calculation method is: call priority = α × computing power occupancy + β × answer load + answer efficiency, where α and β are the weight coefficients of computing power occupancy and answer load respectively; combined with the communication latency, the Q&A priority of the large language model is obtained: Q&A priority of the large language model = α × computing power occupancy + β × answer load + answer efficiency + communication latency between the Q&A device and the large language model.
[0070] The measurement of communication latency is carried out by the Q&A device sending a test request with a timestamp to the large language model, recording the timestamps of the responses of each large language model, and determining the time difference between the request sending and response receiving.
[0071] Step Five: Generate the sequence of large language Q&A models
[0072] According to the Q&A priorities of the large language models calculated in Step Four and the order of each domain in the question to be answered, the Q&A container generates the sequence of large language Q&A models for the corresponding question to be answered. This sequence ensures that the question can be answered in the optimal order and manner.
[0073] Step Six: Question Answering and Return
[0074] The Q&A container decomposes the question to be answered into sub-questions in the corresponding domains, and sequentially sends the sub-questions to the corresponding large language models for answering according to the sequence of large language Q&A models. The Q&A results returned by each large language model will be used as the input for the next large language model to process until all sub-questions are answered. Finally, the Q&A container returns the complete Q&A content to the Q&A device to complete the intelligent Q&A process.
[0075] Example 1: Example of single-domain problem processing
[0076] Suppose the user submits a question about "the principle of quantum computing" through the Q&A device, and this question belongs to a single domain (i.e., the quantum computing domain in physics). According to the intelligent Q&A method based on cognitive reasoning proposed by the present invention, the processing flow is as follows:
[0077] Construct the sequence of large language models:
[0078] The cloud model call server has pre-evaluated the response speeds and accuracies of multiple large language models in the field of physics and calculated their response efficiencies. Suppose there are three large language models A, B, and C, and their response efficiencies in the field of physics are 0.85, 0.90, and 0.75 respectively (according to the formula: response efficiency = γ × response speed + δ × accuracy rate, where γ and δ have been set according to actual requirements).
[0079] Based on these response efficiencies, the server has generated a sequence of large language models in the field of physics: B (highest efficiency) > A > C.
[0080] Problem domain identification:
[0081] After the Q&A container receives the question "Principle of Quantum Computing" submitted by the user, it uses natural language processing technology to identify that this question belongs to the field of physics, specifically the sub-field of quantum computing.
[0082] Single-domain problem processing:
[0083] According to the sequence of large language models in the field of physics generated in step one, the Q&A container directly selects the large language model B with the highest response efficiency to process this question.
[0084] After receiving the question, large language model B conducts in-depth understanding and reasoning, and then generates a detailed answer about the "Principle of Quantum Computing".
[0085] Question answering and returning:
[0086] Large language model B returns the answer content to the Q&A container.
[0087] The Q&A container returns the complete Q&A content to the Q&A device, and the user can then see the accurate answer about the "Principle of Quantum Computing".
[0088] In this embodiment, by selecting the large language model B with the highest response efficiency to process single-domain problems, not only the accuracy of the answer is guaranteed, but also the response speed is improved, and resource consumption and communication latency are reduced.
[0089] Embodiment 2: Example of multi-domain cross-problem processing
[0090] Suppose the user submits a question about "Applications of Artificial Intelligence in the Medical Field and Its Impact on Ethics" through the Q&A device, and this question involves three fields: artificial intelligence, medicine, and ethics. According to the intelligent Q&A method based on cognitive reasoning proposed by the present invention, the processing flow is as follows:
[0091] Constructing a sequence of large language models:
[0092] The cloud model call server has pre-evaluated the response speeds and accuracies of multiple large language models in the fields of artificial intelligence, medicine, and ethics, and calculated their response efficiencies in their respective fields.
[0093] Based on these response efficiencies, the server generated sequences of large language models for the corresponding fields. Assume that in the field of artificial intelligence, A1 > A2 > A3; in the medical field, M1 > M2; and in the ethics field, E1 > E2.
[0094] Problem domain identification:
[0095] After the Q&A container receives the question submitted by the user, it uses natural language processing technology to identify that the question involves the three fields of artificial intelligence, medicine, and ethics.
[0096] Multi-domain cross-question processing:
[0097] The Q&A container obtains the sequences of large language models for the corresponding fields according to the fields involved in the question.
[0098] Taking into account factors such as the resource occupancy of large language models (including computing power occupancy, response load, and response efficiency), communication latency, etc., the Q&A priorities of each large language model are calculated. Assume that after calculation, the order of the Q&A priorities of large language models is: M1 (medical field) > A1 (artificial intelligence field) > E1 (ethics field).
[0099] Generate a sequence of large language Q&A models:
[0100] According to the calculated Q&A priorities of large language models and the order of each field in the question to be answered (assuming the user hopes to first understand the applications of the medical field, then artificial intelligence technology, and finally explore ethical implications), the Q&A container generates a sequence of large language Q&A models for the question to be answered: M1 -> A1 -> E1.
[0101] Question answering and return:
[0102] The Q&A container decomposes the question to be answered into sub-questions for the corresponding fields and sequentially sends the sub-questions to the corresponding large language models for answering according to the sequence of large language Q&A models.
[0103] First, the M1 large language model answers the sub-question about "the application of artificial intelligence in the medical field" and returns the result to the Q&A container.
[0104] Next, the Q&A container sends the answer result of M1 and the sub-question about "artificial intelligence technology" to the A1 large language model for answering.
[0105] After the large language model A1 provides an answer, the result is returned to the Q&A container, which then sends the answer from A1 and the sub-question about "ethical implications" to the large language model E1 for answering.
[0106] After the large language model E1 provides an answer, the final result is returned to the Q&A container.
[0107] The Q&A container returns the complete Q&A content (including applications in the medical field, introduction to artificial intelligence technology, and analysis of ethical implications) to the Q&A device, and the user can then see a comprehensive and accurate answer about "the application of artificial intelligence in the medical field and its impact on ethics".
[0108] In this embodiment, by comprehensively considering factors such as resource occupancy and communication latency of the large language model, the Q&A priority of the large language model is calculated, and a sequence of large language Q&A models corresponding to the questions to be answered is generated, ensuring that cross-domain questions can be comprehensively and accurately answered in the optimal order and manner.
Claims
1. An intelligent question-answering method based on cognitive reasoning, characterized in that: The steps include: Step 1: The cloud model calls the server to obtain the answer efficiency of the large language model in the corresponding field according to the answer speed and accuracy of different large language models in different fields, and obtains the large language model sequence in the corresponding field according to the answer efficiency of the large language model in the corresponding field. The cloud model calls the server to obtain the question information to be answered sent by the question-answering device and generate a question-answering container; Step 2: The question-answering container obtains the domain of the question to be answered according to the question information to be answered sent by the question-answering device. If the domain of the question to be answered is a single-domain question, the process proceeds to step 3; otherwise, the process proceeds to step 4. Step 3: According to the field of the question to be answered, a sequence of large language models in the corresponding field is obtained. According to the answer efficiency of each large language model, the question-answering container sends the question information to be answered to the large prediction model with the highest answer efficiency to generate question-answering content, and then proceeds to step 6. Step 4: According to the fields included in the questions to be answered, a large language model sequence for each corresponding field is obtained, and according to the resource occupation of the large language model in the large language model sequence for the corresponding field, the calling priority of each large language model in the large language model sequence for the corresponding field is obtained respectively, and according to the communication delay between the question-answering device and each large language model in the large language model sequence for the corresponding field and the large language model calling priority, the large language model question-answering priority is obtained, and the large language model question-answering sequence for the corresponding field is obtained according to the question-answering priority; Step 5: According to the large language model with the highest question-answering priority in the large language model question-answering sequence of each corresponding field and the order of the fields of the questions to be answered, a large language question-answering model sequence corresponding to the questions to be answered is obtained, and the question-answering container sends the questions to be answered to the large language question-answering model sequence corresponding to the questions to be answered to generate question-answering content; Step 6: The question-and-answer container returns the question-and-answer content to the question-and-answer device to complete the intelligent question-and-answer process. The cloud model calling server obtains the answer efficiency of the large language model in the corresponding field according to the answer speed and accuracy of different large language models in different fields, including: Answer efficiency = γ*response speed+δ *Accuracy Among them γ represents the weight coefficient of answer speed, δ The weight coefficient representing the accuracy rate.
2. The intelligent question-answering method based on cognitive reasoning according to claim 1, characterized in that: The method of obtaining the calling priority of each large language model in the large language model sequence of the corresponding field according to the resource occupation of the large language model in the large language model sequence of the corresponding field, and obtaining the large language model question-answering priority according to the communication delay between the question-answering device and each large language model in the large language model sequence of the corresponding field and the large language model calling priority, includes: The resource usage of the large language model includes computing power usage, answer load and answer efficiency. The calling priority of the large language model is obtained according to the resource usage, using the following formula: Call Priority = α*computing power usage+β *Answer Load + Answer Efficiency in, α Indicates the weight coefficient of computing power, β represents the answer load weight coefficient; According to the communication delay between the question-answering device and the large language model sequence in the corresponding field and the large language model call priority, the large language model question-answering priority is obtained: Large language model question answering priority = α*computing power usage+β *Answer load + answer efficiency + communication delay between the question-answering device and the large language model.
3. The intelligent question-answering method based on cognitive reasoning according to claim 2, characterized in that: The communication delay between the question-answering device and each language model in the large language model sequence of the corresponding field includes: The question-and-answer device sends a test request with a timestamp to the large language model, records the timestamps of the responses from each language model, and determines the communication delay by calculating the time difference between sending the request and receiving the response.
4. The intelligent question-answering method based on cognitive reasoning according to claim 3 is characterized in that: The question-and-answer container sends the question to be answered to the large language question-and-answer model sequence corresponding to the question to be answered, and generates question-and-answer content, including: The question-and-answer container decomposes the question to be answered into sub-questions in the corresponding fields, and obtains a sequence of sub-questions in the corresponding fields according to the order of the fields in which the questions are answered. According to the sequence of sub-questions in the corresponding fields, the sub-questions in the corresponding fields are sent to the large language question-and-answer model sequence of the corresponding questions to be answered. According to the question-and-answer results returned by the large language question-and-answer model and the next sub-question in the corresponding field, they are sent to the next large language question-and-answer model until the question-and-answer of all sub-questions in the corresponding fields is completed and the question-and-answer content is obtained.
5. An intelligent question-answering device based on cognitive reasoning, characterized in that: The intelligent question-answering method based on cognitive reasoning described in any one of claims 1-4 includes a communication device, a delay test module, a data processing module and an information input module; the communication device, delay test module and information input module are respectively connected to the data processing module.
6. Intelligent question-answering system based on cognitive reasoning, characterized by: The intelligent question-answering device based on cognitive reasoning as claimed in claim 5 includes a cloud model calling server and a communication device; The intelligent question-answering device based on cognitive reasoning is connected to the cloud model calling server through a communication device.
Citation Information
Patent Citations
Intelligent question answering system and method based on large model and knowledge graph
CN118673113A
Systems and Methods for Programmatic Labeling of Training Data for Machine Learning Models via Clustering and Language Model Prompting
US20240160900A1