Intelligent Question-Answering Method, Device, Computer Equipment and Program Product
By using the combination of large language models and reinforcement learning models in the intelligent question-answer system, cognitive correction strategies are constructed and action planning is optimized, the problem of logical understanding deviation in traditional intelligent question-and-answer systems is solved, and higher accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202410963627.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-07-17
AI Technical Summary
In traditional intelligent question-and-answer systems, the agent's analysis logic dimension of user problems is relatively single, resulting in comprehension deviation and low accuracy.
By obtaining multimodal problem data, using the large language model in the proxy agent for deconstruction and decomposition, combining reinforcement learning models and knowledge data to build cognitive correction strategies, optimizing the action planning of the large language model, continuously learning user needs, and improving logical reasoning capabilities.
It improves the logical reasoning ability of the intelligent question-and-answer system in the professional field and the accuracy of target results, adapts to user needs, and provides high-quality response results.
Smart Images

Figure CN118820436B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an intelligent question-answering method, device, computer device, and computer program product. Background Art
[0002] With the development of artificial intelligence and large model technology, large models can be used as kernel agents to build intelligent agents for knowledge data in specific fields to adapt to intelligent question-answering in specific fields.
[0003] In traditional technologies, first, a knowledge base is built based on business data or knowledge data in a specific field. Then, a large language model is trained according to the pre-built knowledge base to obtain a large language model that has learned the knowledge data in the knowledge base, and then an intelligent agent is obtained. The user analyzes the question through the intelligent agent in combination with the content of the knowledge base to obtain a target result, and feeds the target result back to the user interface.
[0004] However, in traditional technologies, since the dimension of reasoning of the intelligent agent for the questions raised by the user is relatively single, it is easy for the logic of the intelligent agent to analyze the questions raised by the user to have a misunderstanding, resulting in a low accuracy rate of intelligent question-answering by the intelligent agent. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide an intelligent question-answering method, device, computer device, computer-readable storage medium, and computer program product.
[0006] In a first aspect, this application provides an intelligent question-answering method, including:
[0007] Obtain multi-modal question data input by the user;
[0008] Based on the question data, determine the knowledge data corresponding to the question data, and decompose the question data according to the large language model in the proxy intelligent agent to obtain the first context data corresponding to the question data;
[0009] Determine the second context data according to the initial result of the large language model in the proxy intelligent agent for the question data;
[0010] Based on the reinforcement learning model, the knowledge data, the first context data, and the second context data in the proxy intelligent agent, construct a cognitive correction strategy;
[0011] Based on the cognitive correction strategy, perform action planning on the large language model to determine a task planning strategy, and obtain a target result based on the task planning strategy and the large language model.
[0012] In one embodiment, after performing action planning on the large language model based on the cognitive correction strategy to determine a task planning strategy and obtaining a target result based on the task planning strategy and the large language model, the method further includes:
[0013] Obtaining feedback information on the target result fed back by the user;
[0014] Constructing a historical output sequence corresponding to the user based on the feedback information, the task planning strategy, and the multimodal problem data;
[0015] Determining user habit type information and action feedback information of the user according to the historical output sequence, and training the reinforcement learning model based on the user habit type and the action feedback information to obtain a trained reinforcement learning model.
[0016] In one embodiment, constructing the cognitive correction strategy based on the reinforcement learning model, the knowledge data, the first context data, and the second context data in the agent includes:
[0017] Constructing an initial cognitive correction strategy for the initial result according to the knowledge data, the first context data, and the second context data;
[0018] Parsing and adjusting the inference logic in the initial cognitive correction strategy based on the reinforcement learning model in the agent to obtain an inference-calibrated cognitive correction strategy.
[0019] In one embodiment, performing action planning on the large language model based on the cognitive correction strategy to determine a task planning strategy includes:
[0020] Analyzing the cognitive correction strategy based on the large language model to obtain an initial task planning strategy;
[0021] Verifying the initial task planning strategy according to the reinforcement learning model in the agent to obtain verification results of each task in the initial task planning strategy;
[0022] Determining a task to be corrected based on the verification results, and correcting the task to be corrected according to the large language model to obtain a task planning strategy corresponding to the problem data.
[0023] In one embodiment, the knowledge data includes long-term knowledge data and short-term knowledge data; determining the knowledge data corresponding to the problem data based on the problem data includes:
[0024] Determine the long-term knowledge data corresponding to the problem data based on the similarity between the problem data and the knowledge data in the preset knowledge base in the agent intelligent body;
[0025] Retrieve and determine the short-term knowledge data corresponding to the problem data in the dialogue context and other task contexts of the agent intelligent body.
[0026] In one embodiment, the deconstructing and decomposing the problem data to obtain the first context data corresponding to the problem data includes:
[0027] Expand the problem data according to the large language model in the agent intelligent body to generate the detailed information of the problem data;
[0028] Based on the detailed information, structurally decompose the problem data to obtain the first context data corresponding to the problem data.
[0029] In one embodiment, the determining the second context data according to the initial result of the problem data by the large language model in the agent intelligent body includes:
[0030] Generate the initial result corresponding to the problem data according to the large language model in the agent intelligent body;
[0031] Based on the large language model in the agent intelligent body, decompose the initial result to obtain the elements included in the initial result;
[0032] Based on the elements included in the initial result and the large language model, perform relevant information retrieval on the initial result to obtain the second context data.
[0033] In a second aspect, the present application also provides an intelligent question answering device, including:
[0034] The first acquisition module is used to acquire the multimodal problem data input by the user;
[0035] The deconstruction module is used to determine the knowledge data corresponding to the problem data based on the problem data, and deconstruct and decompose the problem data according to the large language model in the agent intelligent body to obtain the first context data corresponding to the problem data;
[0036] The determination module is used to determine the second context data according to the initial result of the problem data by the large language model in the agent intelligent body;
[0037] The first construction module is used to construct a cognitive correction strategy based on the reinforcement learning model, the knowledge data, the first context data, and the second context data in the agent intelligent body;
[0038] A planning module, configured to perform action planning on the large language model based on the cognitive correction strategy, determine a task planning strategy, and obtain a target result based on the task planning strategy and the large language model.
[0039] In one embodiment, the apparatus further includes:
[0040] A second acquisition module, configured to acquire feedback information on the target result fed back by a user;
[0041] A second construction module, configured to construct a historical output sequence corresponding to the user based on the feedback information, the task planning strategy, and the multimodal problem data;
[0042] A training module, configured to determine user habit type information and action feedback information of the user according to the historical output sequence, and train the reinforcement learning model based on the user habit type and the action feedback information to obtain a trained reinforcement learning model.
[0043] In one embodiment, the first construction module is specifically configured to construct an initial cognitive correction strategy for the initial result according to the knowledge data, the first context data, and the second context data;
[0044] Parse and adjust the inference logic in the initial cognitive correction strategy based on the reinforcement learning model in the agent to obtain a cognitively corrected strategy with calibrated inference.
[0045] In one embodiment, the planning module is specifically configured to analyze the cognitive correction strategy based on the large language model to obtain an initial task planning strategy;
[0046] Verify the initial task planning strategy according to the reinforcement learning model in the agent to obtain verification results of each task in the initial task planning strategy;
[0047] Determine a task to be corrected based on the verification result, and correct the task to be corrected according to the large language model to obtain a task planning strategy corresponding to the problem data.
[0048] In one embodiment, the knowledge data includes long-term knowledge data and short-term knowledge data; the deconstruction module is specifically configured to determine the long-term knowledge data corresponding to the problem data based on the similarity between the problem data and the knowledge data in a preset knowledge base in the agent;
[0049] Retrieve and determine the short-term knowledge data corresponding to the problem data in the dialogue context and other task contexts of the agent.
[0050] In one embodiment, the deconstruction module is specifically configured to expand the problem data according to the large language model in the agent intelligent body to generate detailed information of the problem data;
[0051] Based on the detailed information, perform a structural decomposition on the problem data to obtain first context data corresponding to the problem data.
[0052] In one embodiment, the determination module is specifically configured to generate an initial result corresponding to the problem data according to the large language model in the agent intelligent body;
[0053] Based on the large language model in the agent intelligent body, decompose the initial result to obtain the elements included in the initial result;
[0054] Based on the elements included in the initial result and the large language model, perform relevant information retrieval on the initial result to obtain second context data.
[0055] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0056] Obtain multi-modal problem data input by a user;
[0057] Based on the problem data, determine knowledge data corresponding to the problem data, and according to the large language model in the agent intelligent body, perform deconstruction decomposition on the problem data to obtain first context data corresponding to the problem data;
[0058] Determine second context data according to the initial result of the problem data by the large language model in the agent intelligent body;
[0059] Based on the reinforcement learning model, the knowledge data, the first context data, and the second context data in the agent intelligent body, construct a cognitive correction strategy;
[0060] Based on the cognitive correction strategy, perform action planning on the large language model to determine a task planning strategy, and based on the task planning strategy and the large language model, obtain a target result.
[0061] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0062] Obtain multi-modal problem data input by a user;
[0063] Determine the knowledge data corresponding to the problem data based on the problem data, and decompose the problem data according to the large language model in the agent to obtain the first context data corresponding to the problem data;
[0064] Determine the second context data according to the initial result of the large language model in the agent for the problem data;
[0065] Construct a cognitive correction strategy based on the reinforcement learning model, the knowledge data, the first context data, and the second context data in the agent;
[0066] Perform action planning on the large language model based on the cognitive correction strategy to determine a task planning strategy, and obtain a target result based on the task planning strategy and the large language model.
[0067] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:
[0068] Obtain multi-modal problem data input by a user;
[0069] Determine the knowledge data corresponding to the problem data based on the problem data, and decompose the problem data according to the large language model in the agent to obtain the first context data corresponding to the problem data;
[0070] Determine the second context data according to the initial result of the large language model in the agent for the problem data;
[0071] Construct a cognitive correction strategy based on the reinforcement learning model, the knowledge data, the first context data, and the second context data in the agent;
[0072] Perform action planning on the large language model based on the cognitive correction strategy to determine a task planning strategy, and obtain a target result based on the task planning strategy and the large language model.
[0073] The above intelligent question-answering method, device, computer device, and computer program product can determine knowledge data through question data, and can provide high-quality logical reasoning ability in professional fields. By performing data reasoning on the question data and knowledge data through the reinforcement learning model in the agent intelligent body, the first context data and the second context data are obtained. The reinforcement learning model can continuously learn the correlation between the target result and the user's needs in the application of the large language model, continuously correct the large language model, and construct a cognitive correction strategy based on the first context data and the second context data. It can verify the action plan output by the large language model in the agent intelligent body, obtain an action plan that is more suitable for the question data, and then execute the action plan, which can improve the accuracy of the target result. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0075] Figure 1 It is a schematic flowchart of the intelligent question-answering method in an embodiment;
[0076] Figure 2 It is a schematic flowchart of continuously training the agent intelligent body through the reinforcement learning model in an embodiment;
[0077] Figure 3 It is a schematic diagram of the reinforcement learning model for learning in an embodiment;
[0078] Figure 4 It is a schematic flowchart of constructing a cognitive correction strategy in another embodiment;
[0079] Figure 5 It is a schematic diagram of the principle of the cognitive correction strategy in an embodiment;
[0080] Figure 6 It is a schematic flowchart of the steps of constructing a task planning strategy in an embodiment;
[0081] Figure 7 It is a schematic diagram of the architecture of the agent intelligent body in an embodiment;
[0082] Figure 8 It is a schematic flowchart of the steps of determining knowledge data in an embodiment;
[0083] Figure 9 It is a schematic flowchart of the steps of retrieving the first context data in an embodiment;
[0084] Figure 10 Flow diagram of the step of retrieving second context data in an embodiment;
[0085] Figure 11 Block diagram of the structure of an intelligent question - answering device in an embodiment;
[0086] Figure 12 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0087] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0088] In one embodiment, as Figure 1 shown, an intelligent question - answering method is provided. In this embodiment, it is exemplified that the method is applied to a terminal. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is realized through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0089] Step 102, obtain multimodal question data input by the user.
[0090] In the embodiment of the present application, the terminal obtains the question data input by the user. The question data can be multimodal data. For example, the question data can be question text, voice, image or video data in a certain field. The terminal converts the multimodal question data input by the user through the perception structure in the proxy agent, that is, converts the multimodal question data into a format that the large - language model can understand and process, and inputs the question data after format conversion into the large - language model (LLM, Large Language Model) in the proxy agent, and further processes the question data through the large - language model.
[0091] For the voice data input by the user, the voice is converted into text format through the voice recognition structure included in the proxy agent, so that the large - language model can process it; for the image data, the proxy agent needs to perform image processing and feature extraction, and convert the image into text describing the image content or a related data structure; for the video data, the proxy agent extracts key frames from the video or performs video content analysis, and finally converts it into a processable data form.
[0092] Step 104: Determine the knowledge data corresponding to the problem data based on the problem data, and decompose the problem data according to the large language model in the agent intelligent body to obtain the first context data corresponding to the problem data.
[0093] In the embodiment of the present application, the terminal first retrieves and expands relevant information for the problem data according to the large language model. First, the terminal queries, through the large language model, knowledge data related to the problem input by the user from its internal or external knowledge base, including definitions, rules, or other domain-specific information about a particular domain, such as the medical domain or the financial domain. The agent intelligent body uses natural language processing techniques to extract information directly related to the user's problem from the obtained knowledge base data, including techniques such as entity recognition, relationship extraction, and event extraction, to ensure that the extracted knowledge data is closely related to the user's problem data.
[0094] The terminal deeply analyzes the problem data through the large language model in the agent intelligent body and decomposes the structure of the problem data. First, the large language model determines the intent and purpose in the problem data, determines the task type or task theme involved in the problem data, and combines the knowledge data with the intent and purpose of the problem data based on the task type to generate an extension of relevant information in the dimension of the problem proposed by the user, that is, to obtain the first context data corresponding to the problem data.
[0095] In an exemplary embodiment, if the problem proposed by the user is: "What are the symptoms of COVID-19?" The agent intelligent body first accesses its internal or external knowledge base to find information related to the symptoms of COVID-19. The agent intelligent body uses natural language processing techniques to extract knowledge data considered directly related to the user's problem from the knowledge base, such as: "fever, cough, shortness of breath, fatigue, etc.". Then, the agent intelligent body deeply analyzes and structures the user's problem through the large language model, including understanding the semantics of the problem, such as determining that the user is asking for information about the symptoms of COVID-19, and then considering the context of the problem, such as the current global epidemic situation, and other context data that affect the user's specific needs and expectations. Combining the information extracted from the knowledge base and the analysis results of the large language model, the agent intelligent body generates the first context data of the problem. In this example, the first context data may include a detailed list containing the common symptoms of COVID-19 and corresponding diagnosis and treatment recommendations, etc.
[0096] Step 106: Determine the second context data according to the initial result of the large language model in the agent intelligent body for the problem data.
[0097] In the embodiments of the present application, first, the terminal preliminarily analyzes the problem input by the user through a large language model, including semantic understanding, key information identification, etc., obtains the problem theme for the problem data, and defines a specific initial task planning strategy according to the task theme. For example, specific information is extracted from the input data, calculations or analyses are performed, and the results are formatted and output.
[0098] Then, the terminal executes the initial task planning strategy obtained according to the task analysis through the proxy agent to obtain an initial result. For example, the initial task planning strategy can be to extract a text summary, calculate parameter values, and analyze data trends. The proxy agent combines the results of the execution of the initial task planning strategy to generate second context data. The second context data is the result of further processing the initial processing result to provide more detailed and specific information or responses to the large language model to enrich information such as the environment of the potential results corresponding to the current problem data. In addition, during the process of generating the second context data, the proxy agent can combine and analyze the previous interaction history, domain-specific background knowledge, or other relevant information to further optimize and improve the generated data to ensure its accuracy and relevance in practical applications.
[0099] In an alternative embodiment, the terminal can also construct a chain of thought (COT, Chain-of-thought) based on the semantic information of the problem data, and use the chain-of-thought technology to guide the large language model to perform "step-by-step thinking", so as to decompose difficult tasks into smaller and simpler steps by using more test time calculations. The chain of thought transforms a large task into multiple manageable tasks and explains the thinking process of the large language model.
[0100] Step 108, construct a cognitive correction strategy based on the reinforcement learning model, knowledge data, first context data, and second context data in the proxy agent.
[0101] Among them, the reinforcement learning model is used to determine the best decision-making strategy most relevant to different types of problem data according to the cognitive detection logic learned from the interaction with the environment to achieve a specific goal.
[0102] In the embodiments of the present application, the terminal integrates the first context data, the second context data, and the knowledge data through the proxy agent, and inputs the facts and rules extracted from the knowledge base, the key information and semantic analysis results extracted from the user question, and the analysis results or detailed summaries extracted from a large amount of information for the initial result into the large language model according to the pre-constructed inference template. The large language model comprehensively processes the relevant explanatory texts, extended texts, and environmental descriptions of the above-mentioned multiple problem data, and starts to construct a cognitive correction strategy in combination with the cognitive detection logic learned in continuous training. The reinforcement learning model optimizes the cognitive correction strategy by analyzing the environment (user question data and preliminary results) and executing actions (the initial task planning strategy of the response), and analyzes and adjusts the inference logic of the cognitive correction strategy to ensure the consistency and effectiveness of the cognitive correction strategy in different situations.
[0103] Step 110, perform action planning on the large language model based on the cognitive correction strategy to determine the task planning strategy, and obtain the target result based on the task planning strategy and the large language model.
[0104] In the embodiments of the present application, the proxy agent performs action planning on the large language model according to the cognitive correction strategy and under the guidance of the reinforcement learning model to determine how to effectively apply the large language model to solve specific tasks or problems. Among them, the action planning includes determining which specific large language model functions, algorithms, or operation steps should be called and executed to generate the best response result.
[0105] The terminal guides the proxy agent on how to call different parts of the large language model and how to organize and process the generated data and results during the actual execution process through the task planning strategy. The task planning strategy includes determining the data input method, executing specific configurations or parameter settings of the large language model, and how to process and interpret the output results of the large language model. For example, the target result is the response result to the user's question data, including a detailed analysis report, structured data output, or domain-specific suggestions and solutions.
[0106] In an exemplary embodiment, the purpose of the question data is to first extract a summary, calculate certain parameters in the text, and output them. The proxy agent calls a suitable text summary extraction algorithm through the large language model. After completing the summary extraction, the proxy agent selects an appropriate parameter calculation method according to the cognitive correction strategy, which involves mathematical calculations, statistical analysis, or other modeling techniques, to derive the required specific parameters from the summary. After completing the parameter calculation, the proxy agent formats the result into an understandable form and outputs it as the final target result.
[0107] In an optional embodiment, when the agent performs a specific task, it can call the required external resource tools, which can be various APIs, databases, search engines, or other software services. By calling the external resource tools, the agent can obtain the data required for each task in the task planning strategy or perform specific operations.
[0108] In the above intelligent question-answering method, knowledge data is determined through question data, which can provide high-quality logical reasoning ability in the professional field. Data reasoning is performed on the question data and knowledge data through the reinforcement learning model in the agent, and the first context data and the second context data are obtained. The reinforcement learning model can continuously learn the correlation between the target result and the user's needs in the application of the large language model, continuously correct the large language model, and construct a cognitive correction strategy based on the first context data and the second context data, which can verify the action plan output by the large language model in the agent, obtain an action plan that is more adapted to the question data, and then execute the action plan, which can improve the accuracy of the target result.
[0109] In an exemplary embodiment, after a complete question-answering session, the terminal can also train the reinforcement learning model according to the user feedback information to continuously adjust the logical reasoning ability of the agent for intelligent question-answering. For example, Figure 2 as shown, after step 110, the method further includes steps 202 to 206. Among them:
[0110] Step 202, obtain the feedback information on the target result fed back by the user.
[0111] In the embodiment of the present application, the reinforcement learning model can adopt the Asynchronous Advantage Actor-Critic (A3C, an asynchronous concurrent reinforcement learning framework) framework. After the user ends the question behavior with the agent, the user can feedback the evaluation of the target result output by the agent. The agent determines whether each reply content conforms to the environment of the question data by the evaluation fed back by the user and extracting the satisfaction degree of the user with respect to each reply content of the agent in the text information of each user question during the question-answering process, and then obtains all the feedback information of the user on the target result.
[0112] Step 204, construct a historical output sequence corresponding to the user based on the feedback information, the task planning strategy, and the multimodal question data.
[0113] In the embodiments of the present application, the proxy agent records the specific tasks executed each time and the selected large language model functions. Additionally, if the question data is multimodal (e.g., including text, images, or audio), the proxy agent will record various forms and inputs of the user's question to better understand and process the user's needs.
[0114] The proxy agent integrates the collected user feedback information, task planning strategies, and question data into a historical output sequence for each user, which is used to reflect the comprehensive history of the user's interaction with the proxy agent, including the characteristics of the question, the proxy's response, user satisfaction, and the specific actions executed by the proxy agent.
[0115] Step 206: Determine the user habit type information and action feedback information of the user based on the historical output sequence, and train the reinforcement learning model based on the user habit type and action feedback information to obtain a trained reinforcement learning model.
[0116] In the embodiments of the present application, as Figure 3 shown, the proxy agent analyzes the historical output sequence to identify and determine the habit type information of each user, including the service type, response style, and behavior preferences in specific situations preferred by the user. The reinforcement learning model provides the proxy agent with the ability of dynamic memory and adjustment to environmental feedback to improve the reasoning ability. Among them, the reinforcement learning model adopts standard reinforcement learning settings, in which the reward mechanism in reinforcement learning provides simple binary rewards (0 / 1), and the actions follow the settings of ReAct. For the task planning strategies in the historical output sequence, a heuristic function is calculated, and according to the result of the heuristic function, it is determined whether each task meets the environment of the current question data. Using the user habit type and action feedback information of the current user determined in the historical output sequence, the proxy agent trains the already constructed reinforcement learning model, updates the policy network and value network of the model, so as to more accurately guide the selection and correction of task planning strategies in different situations, and complete the construction of personalized proxy agents for different users.
[0117] In this embodiment, the proxy agent continuously learns and optimizes from user feedback through the reinforcement learning model, determines the user habit types of different users, constructs personalized cognitive correction strategies under different user types according to the user habit types, ensures that the analysis and reasoning environment of the large language model in the proxy agent is adapted to the user's needs, and improves the accuracy of the proxy agent's output target results.
[0118] In an exemplary embodiment, as Figure 4 shown, step 108 includes steps 402 to 404. Among them:
[0119] Step 402: Construct an initial cognitive correction strategy for the preliminary test results based on the knowledge data, the first context data, and the second context data.
[0120] In the embodiments of the present application, the agent combines the knowledge data, the first context data, and the second context data, and based on the existing information and analysis results, preliminarily constructs an initial cognitive correction strategy for the large language model to further process and respond to the questions or requirements of the user.
[0121] Step 404: Parse and adjust the inference logic in the initial cognitive correction strategy based on the reinforcement learning model in the agent to obtain a cognitively corrected strategy with calibrated inference.
[0122] In the embodiments of the present application, the terminal analyzes and parses the inference logic in the initial cognitive correction strategy through the reinforcement learning model, and evaluates the effectiveness and adaptability of the existing cognitive correction strategy in the actual situation. According to the feedback and guidance of the reinforcement learning model, the terminal adjusts and calibrates the inference of the task planning strategy, including modifying the execution order, conditional judgment, or response action in the task planning strategy, to ensure the consistency and optimization effect of the task planning strategy in different situations.
[0123] In an exemplary embodiment, as Figure 5 shown, an example of the cognitive correction strategy is a behavior trajectory composed of action, thought, and observation. Through the cognitive correction strategy, self-reflection on the inference of the large language model is realized, and circular verification is carried out in combination with the environment of the problem data, the inference of the large language model, and the task planning strategy until the environment, inference, and task planning strategy all meet the preset matching degree threshold, and the cognitive correction of the inference of the large language model is completed.
[0124] In this embodiment, the agent can start from the initial information processing and analysis, gradually optimize its cognitive correction strategy, improve the intelligence level of the agent, and also ensure the provision of consistent and efficient services and responses in a dynamic and changing environment, and improve the accuracy of the output target results of the agent.
[0125] In an exemplary embodiment, as Figure 6 shown, step 110 includes steps 602 to 606. Among them:
[0126] Step 602: Analyze the cognitive correction strategy based on the large language model to obtain an initial task planning strategy.
[0127] In the embodiments of the present application, the agent intelligent body analyzes the current cognitive correction strategy through a large language model, understands the logical flow, conditional judgment, and data processing steps in the cognitive correction strategy, and based on the analysis of the cognitive correction strategy, the agent intelligent body generates an initial task planning strategy, which describes the specific tasks and operation steps that should be taken during the execution process.
[0128] Step 604: Verify the initial task planning strategy according to the reinforcement learning model in the agent intelligent body to obtain the verification results of each task in the initial task planning strategy.
[0129] In the embodiments of the present application, after each action of the reinforcement learning model, the agent calculates a heuristic function and selectively decides whether to reset the environment to start a new trial according to the result of the heuristic function, so as to verify the initial task planning strategy. The agent intelligent body analyzes the execution order and resource allocation of each task in the initial task planning strategy through the reinforcement learning model, and verifies the execution results and effects of each task in the initial task planning strategy to obtain the verification results of each task.
[0130] Step 606: Determine the task to be corrected based on the verification results, and perform correction processing on the task to be corrected according to the large language model to obtain the task planning strategy corresponding to the problem data.
[0131] In the embodiments of the present application, the agent intelligent body determines the specific tasks that need to be adjusted or corrected in the initial task planning strategy according to the verification results. The tasks to be corrected may include parts with low execution efficiency, inaccurate results, or non-matching with user requirements. The agent intelligent body uses the large language model to process and correct the tasks to be corrected. For example, reset the task execution order, optimize the data processing method, improve the user interaction method, etc., to improve the quality and effect of the task planning strategy. After the correction processing, the agent intelligent body generates the final task planning strategy corresponding to the problem data. This task planning strategy has been optimized and adjusted and can better meet the user's needs and expected results.
[0132] In a specific embodiment, as Figure 7 shown, Figure 7 is the system architecture of the agent intelligent body. In the task planning strategy, corresponding external resource tools can also be called according to task requirements, such as calculators, search engines, and calendars, etc., to perform calculation processing on data based on external tools to meet the data processing requirements in subsequent actions and inferences.
[0133] In this embodiment, the initial task planning policy is verified and corrected by a reinforcement learning model to obtain the task planning policy corresponding to the problem data, enabling the agent to effectively obtain the task planning policy, ensuring that in practical applications, the task planning policy can provide accurate and personalized service responses, and improving the accuracy of the target results.
[0134] In an exemplary embodiment, the knowledge data includes long-term knowledge data and short-term knowledge data. As Figure 8 shown, step 104 includes steps 802 to 804. Among them:
[0135] Step 802, determining the long-term knowledge data corresponding to the problem data based on the similarity between the problem data and the knowledge data in the preset knowledge base in the agent.
[0136] In the embodiment of the present application, the long-term knowledge data is stored by storing the knowledge data for a very long time, ranging from several days to several decades, and has an almost infinite storage capacity. The long-term knowledge data is the corresponding long-term knowledge data in the external vector storage or database storage that can be retrieved by the agent when querying the problem data and can be accessed through fast retrieval.
[0137] The agent first compares and analyzes the problem data with the knowledge data in the preset knowledge base, including semantic similarity calculation, keyword matching, or other natural language processing techniques, to evaluate the correlation and similarity between them, and based on the results of the similarity analysis, determines the knowledge data with a similarity greater than the preset threshold as the long-term knowledge data corresponding to the problem data.
[0138] Step 804, retrieving and determining the short-term knowledge data corresponding to the problem data in the dialogue context and other task contexts of the agent.
[0139] In the embodiment of the present application, the short-term knowledge data can be a context window similar to the Transformer (a self-attention based model architecture) architecture, which stores the input information in the current session or task. This information is immediately available when the agent processes the problem data and helps the agent understand the context of the current dialogue. The agent utilizes the current dialogue history and the context information of other relevant tasks. This information includes the recent dialogue content, the user's latest query, the agent's response history, etc., which helps to understand the specific context and user intent of the current problem data. Through comprehensive analysis of the dialogue and task contexts, the agent determines and collects the short-term knowledge data directly related to the problem data, providing the most relevant and effective information support for the agent in the current environment.
[0140] In this embodiment, the agent intelligent body can utilize both long-term knowledge data and short-term knowledge data simultaneously to provide comprehensive and real-time support and services for users. By leveraging knowledge data at different time scales, more accurate auxiliary data can be provided for the problem data, thereby improving the response efficiency and decision-making accuracy of the agent intelligent body.
[0141] In an exemplary embodiment, as Figure 9 shown, step 104 includes steps 902 to 904. Among them:
[0142] Step 902, expand the problem data according to the large language model in the agent intelligent body to generate detailed information of the problem data.
[0143] In the embodiment of the present application, first, through the semantic understanding and generation capabilities of the large language model, the terminal can understand the problem data from multiple perspectives and generate relevant detailed content. The large language model supplements more background information, explains different aspects of the problem, or provides relevant historical backgrounds, etc., according to the semantics and background of the problem, to generate detailed information of the problem data, enriching the description of the problem and helping to provide more accurate and comprehensive solutions.
[0144] Step 904, decompose the problem data structurally based on the detailed information to obtain the first context data corresponding to the problem data.
[0145] In the embodiment of the present application, the terminal analyzes the input problem data through the agent intelligent body, understands its requirements and target tasks, decomposes the target tasks into smaller subtasks, and decides how to execute these subtasks. Specifically, the agent intelligent body performs a structural analysis of the problem data according to the detailed information of the problem data, identifies the key elements of the problem, analyzes the structure and logic of the problem, understands the semantic meaning of the problem, etc., determines the theme and background information of the problem data, etc., to obtain the first context data of the problem data.
[0146] In this embodiment, by extracting more detailed and structured information from the problem data, the understanding degree of the large language model for the problem data and the accuracy of understanding the user's needs are improved, providing an information basis for the construction of subsequent cognitive correction strategies and the planning of task planning strategies, improving the accuracy of the large language model's reasoning for the problem data, and further improving the accuracy of the target result.
[0147] In an exemplary embodiment, as Figure 10 shown, step 106 includes steps 1002 to 1006. Among them:
[0148] Step 1002, generate an initial result corresponding to the problem data according to the large language model in the agent intelligent body.
[0149] In the embodiments of the present application, based on the understanding and analysis of problem data, the large language model helps the agent to generate an initial result, which can be a text response or response data representing an initial task planning strategy.
[0150] Step 1004: Decompose the initial result based on the large language model in the agent to obtain the elements included in the initial result.
[0151] In the embodiments of the present application, the agent decomposes the initial result based on the large language model in the agent to identify and extract the elements included therein. Specifically, the agent conducts a detailed analysis of the generated initial result to identify and extract information fragments, key data, important context, etc. in the initial result.
[0152] Step 1006: Perform relevant information retrieval on the initial result based on the elements included in the initial result and the large language model to obtain second context data.
[0153] In the embodiments of the present application, the agent utilizes the large language model again to perform information retrieval according to the elements in the initial result. Among them, the large language model can understand the deep and context relationships of the user's needs based on this element, and then obtain the second context data through information retrieval. Among them, information retrieval includes retrievals respectively performed on long-term knowledge data and short-term knowledge data, as well as real-time data retrieval in a search engine based on the elements included in the initial result to improve the environment of the problem data and complete context expansion in the dimension of the result of the problem data.
[0154] In this embodiment, through the decomposition and information retrieval of the initial result, second context data more relevant to the environment of the problem data is generated, further improving the data richness when the large language model conducts analysis, enhancing the accuracy of the large language model in semantically understanding and demand understanding of the problem data, and thus improving the accuracy of the target result.
[0155] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0156] Based on the same inventive concept, an embodiment of the present application further provides an intelligent question - answering device for implementing the intelligent question - answering method involved above. The solution provided by this device for solving problems is similar to the solution described in the above - mentioned method. Therefore, the specific limitations in one or more of the following embodiments of the intelligent question - answering device can refer to the limitations on the intelligent question - answering method in the foregoing text, and will not be elaborated here.
[0157] In an exemplary embodiment, as Figure 11 shown, an intelligent question - answering device 1100 is provided, including: a first acquisition module 1101, a deconstruction module 1102, a determination module 1103, a first construction module 1104, and a planning module 1105, where:
[0158] The first acquisition module 1101 is configured to acquire multimodal question data input by a user;
[0159] The deconstruction module 1102 is configured to determine knowledge data corresponding to the question data based on the question data, and deconstruct and decompose the question data according to the large - language model in the proxy agent to obtain first context data corresponding to the question data;
[0160] The determination module 1103 is configured to determine second context data according to the initial result of the question data by the large - language model in the proxy agent;
[0161] The first construction module 1104 is configured to construct a cognitive correction strategy based on the reinforcement learning model, knowledge data, first context data, and second context data in the proxy agent;
[0162] The planning module 1105 is configured to perform action planning on the large - language model based on the cognitive correction strategy, determine a task planning strategy, and obtain a target result based on the task planning strategy and the large - language model.
[0163] In one of the embodiments, the device 1100 further includes:
[0164] A second acquisition module, configured to acquire feedback information on the target result fed back by the user;
[0165] A second construction module, configured to construct a historical output sequence corresponding to the user based on the feedback information, the task planning strategy, and the multimodal question data;
[0166] A training module, configured to determine the user's user habit type information and action feedback information according to the historical output sequence, and train the reinforcement learning model based on the user habit type and the action feedback information to obtain a trained reinforcement learning model.
[0167] In one embodiment, the first construction module 1104 is specifically configured to construct an initial cognitive correction strategy for the initial result according to the knowledge data, the first context data, and the second context data;
[0168] Parse and adjust the inference logic in the initial cognitive correction strategy based on the reinforcement learning model in the agent to obtain a cognitively corrected strategy with calibrated inference.
[0169] In one embodiment, the planning module 1105 is specifically configured to analyze the cognitive correction strategy based on a large language model to obtain an initial task planning strategy;
[0170] Verify the initial task planning strategy according to the reinforcement learning model in the agent to obtain the verification results of each task in the initial task planning strategy;
[0171] Determine the task to be corrected based on the verification results, and correct the task to be corrected according to the large language model to obtain a task planning strategy corresponding to the problem data.
[0172] In one embodiment, the knowledge data includes long-term knowledge data and short-term knowledge data; the deconstruction module 1102 is specifically configured to determine the long-term knowledge data corresponding to the problem data based on the similarity between the problem data and the knowledge data in the preset knowledge base in the agent;
[0173] Retrieve and determine the short-term knowledge data corresponding to the problem data in the dialogue context and other task contexts of the agent.
[0174] In one embodiment, the deconstruction module 1102 is specifically configured to expand the problem data according to the large language model in the agent to generate detailed information about the problem data;
[0175] Perform structural decomposition on the problem data based on the detailed information to obtain the first context data corresponding to the problem data.
[0176] In one embodiment, the determination module 1103 is specifically configured to generate an initial result corresponding to the problem data according to the large language model in the agent;
[0177] Decompose the initial result based on the large language model in the agent to obtain the elements included in the initial result;
[0178] Perform relevant information retrieval on the initial result based on the elements included in the initial result and the large language model to obtain the second context data.
[0179] Each module in the above intelligent question-answering device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0180] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structural diagram can be as Figure 12 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store knowledge data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. The computer program, when executed by the processor, implements an intelligent question-answering method.
[0181] Those skilled in the art can understand that Figure 12 the structure shown in
[0182] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0183] Obtain multimodal question data input by the user;
[0184] Based on the question data, determine the knowledge data corresponding to the question data, and deconstruct and decompose the question data according to the large language model in the proxy agent to obtain the first context data corresponding to the question data;
[0185] Determine the second context data according to the initial result of the large language model in the proxy agent for the question data;
[0186] Construct a cognitive correction strategy based on the reinforcement learning model, knowledge data, first context data, and second context data in the proxy agent;
[0187] Based on the cognitive correction strategy, perform action planning on the large language model to determine a task planning strategy, and obtain a target result based on the task planning strategy and the large language model.
[0188] In one embodiment, when the processor executes the computer program, the following steps are also implemented:
[0189] Obtain feedback information on the target result provided by the user;
[0190] Based on the feedback information, task planning strategy, and multi-modal problem data, construct a historical output sequence corresponding to the user;
[0191] Determine the user's user habit type information and action feedback information based on the historical output sequence, and train the reinforcement learning model based on the user habit type and action feedback information to obtain a trained reinforcement learning model.
[0192] In one embodiment, when the processor executes the computer program, the following steps are also implemented:
[0193] Construct an initial cognitive correction strategy for the initial result according to the knowledge data, first context data, and second context data;
[0194] Parse and adjust the inference logic in the initial cognitive correction strategy based on the reinforcement learning model in the proxy agent to obtain a cognitively corrected strategy with calibrated inference.
[0195] In one embodiment, when the processor executes the computer program, the following steps are also implemented:
[0196] Analyze the cognitive correction strategy based on the large language model to obtain an initial task planning strategy;
[0197] Verify the initial task planning strategy according to the reinforcement learning model in the proxy agent to obtain the verification results of each task in the initial task planning strategy;
[0198] Determine the task to be corrected based on the verification results, and correct the task to be corrected according to the large language model to obtain a task planning strategy corresponding to the problem data.
[0199] In one embodiment, when the processor executes the computer program, the following steps are also implemented:
[0200] Determine the long-term knowledge data corresponding to the problem data based on the similarity between the problem data and the knowledge data in the preset knowledge base in the proxy agent;
[0201] Retrieve and determine short-term knowledge data corresponding to the problem data in the dialogue context and other task contexts of the agent intelligent agent.
[0202] In one embodiment, when the processor executes the computer program, the following steps are also implemented:
[0203] Expand the problem data according to the large language model in the agent intelligent agent to generate detailed information of the problem data;
[0204] Based on the detailed information, decompose the problem data structurally to obtain the first context data corresponding to the problem data.
[0205] In one embodiment, when the processor executes the computer program, the following steps are also implemented:
[0206] Generate an initial result corresponding to the problem data according to the large language model in the agent intelligent agent;
[0207] Based on the large language model in the agent intelligent agent, decompose the initial result to obtain the elements included in the initial result;
[0208] Based on the elements included in the initial result and the large language model, retrieve relevant information for the initial result to obtain the second context data.
[0209] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0210] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0211] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0212] Those of ordinary skill in the art can understand that all or part of the processes in the above-described method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-described method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0213] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0214] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. An intelligent question-answering method, characterized in that, The method includes: Obtaining multi-modal problem data input by the user; Determining knowledge data corresponding to the problem data based on the problem data, and performing structural decomposition on the problem data according to the large language model in the agent to obtain first context data corresponding to the problem data; Determining second context data according to the initial result of the large language model in the agent for the problem data; Constructing a cognitive correction strategy based on the reinforcement learning model, the knowledge data, the first context data, and the second context data in the agent; Performing action planning on the large language model based on the cognitive correction strategy to determine a task planning strategy, and obtaining a target result based on the task planning strategy and the large language model; The performing structural decomposition on the problem data to obtain first context data corresponding to the problem data includes: Expanding the problem data according to the large language model in the agent to generate detailed information of the problem data; Performing structural decomposition on the problem data based on the detailed information to obtain first context data corresponding to the problem data; The determining second context data according to the initial result of the large language model in the agent for the problem data includes: Generating an initial result corresponding to the problem data according to the large language model in the agent; Decomposing the initial result based on the large language model in the agent to obtain elements included in the initial result; Performing relevant information retrieval on the initial result based on the elements included in the initial result and the large language model to obtain second context data; The performing action planning on the large language model based on the cognitive correction strategy to determine a task planning strategy includes: Analyzing the cognitive correction strategy based on the large language model to obtain an initial task planning strategy; Verifying the initial task planning strategy according to the reinforcement learning model in the agent to obtain verification results of each task in the initial task planning strategy; Determining a task to be corrected based on the verification results, and performing correction processing on the task to be corrected according to the large language model to obtain a task planning strategy corresponding to the problem data.
2. The method according to claim 1, wherein After performing action planning on the large language model based on the cognitive correction strategy to determine a task planning strategy, and obtaining a target result based on the task planning strategy and the large language model, the method further includes: Obtaining feedback information of the user on the target result; Constructing a historical output sequence corresponding to the user based on the feedback information, the task planning strategy, and the multi-modal problem data; Determining user habit type information and action feedback information of the user according to the historical output sequence, and training the reinforcement learning model based on the user habit type and the action feedback information to obtain a trained reinforcement learning model.
3. The method according to claim 1, wherein The constructing a cognitive correction strategy based on the reinforcement learning model, the knowledge data, the first context data, and the second context data in the agent includes: Construct an initial cognitive correction strategy for the initial result based on the knowledge data, the first context data, and the second context data; Parse and adjust the inference logic in the initial cognitive correction strategy based on the reinforcement learning model in the agent to obtain a cognitively corrected strategy with calibrated inference.
4. The method according to claim 1, wherein The knowledge data includes long-term knowledge data and short-term knowledge data; Determining the knowledge data corresponding to the problem data based on the problem data includes: Determine the long-term knowledge data corresponding to the problem data based on the similarity between the problem data and the knowledge data in the preset knowledge base in the agent; Retrieve and determine the short-term knowledge data corresponding to the problem data in the dialogue context and other task contexts of the agent.
5. An intelligent question answering device, characterized in that, The device includes: A first acquisition module for acquiring multimodal problem data input by the user; A deconstruction module for determining the knowledge data corresponding to the problem data based on the problem data and structurally decomposing the problem data according to the large language model in the agent to obtain the first context data corresponding to the problem data; A determination module for determining the second context data according to the initial result of the problem data by the large language model in the agent; A first construction module for constructing a cognitive correction strategy based on the reinforcement learning model, the knowledge data, the first context data, and the second context data in the agent; A planning module for performing action planning on the large language model based on the cognitive correction strategy to determine a task planning strategy, and obtaining a target result based on the task planning strategy and the large language model; The deconstruction module is specifically used to expand the problem data according to the large language model in the agent to generate detailed information of the problem data; Structurally decompose the problem data based on the detailed information to obtain the first context data corresponding to the problem data; The determination module is specifically used to generate an initial result corresponding to the problem data according to the large language model in the agent; Decompose the initial result according to the large language model in the agent to obtain the elements included in the initial result; Retrieve relevant information for the initial result based on the elements included in the initial result and the large language model to obtain the second context data; The planning module is specifically used to analyze the cognitive correction strategy based on the large language model to obtain an initial task planning strategy; Verify the initial task planning strategy according to the reinforcement learning model in the agent to obtain the verification results of each task in the initial task planning strategy; Determine the task to be corrected based on the verification results, and correct the task to be corrected according to the large language model to obtain the task planning strategy corresponding to the problem data.
6. The device according to claim 5, wherein The device further includes: A second acquisition module for acquiring feedback information on the target result from the user; A second construction module for constructing a historical output sequence corresponding to the user based on the feedback information, the task planning strategy, and the multimodal problem data; A training module, configured to determine user habit type information and action feedback information of the user according to the historical output sequence, and train the reinforcement learning model based on the user habit type and the action feedback information to obtain a trained reinforcement learning model.
7. The device according to claim 5, characterized in that, The first construction module is specifically configured to construct an initial cognitive correction strategy for the initial result according to the knowledge data, the first context data, and the second context data; Parse and adjust the inference logic in the initial cognitive correction strategy based on the reinforcement learning model in the agent to obtain a cognitively corrected strategy with calibrated inference.
8. The device according to claim 5, characterized in that, The knowledge data includes long-term knowledge data and short-term knowledge data; the deconstruction module is specifically configured to determine the long-term knowledge data corresponding to the problem data based on the similarity between the problem data and the knowledge data in the preset knowledge base in the agent; Retrieve and determine the short-term knowledge data corresponding to the problem data in the dialogue context and other task contexts of the agent.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Method and device for training generative large language model based on knowledge base feedback
CN117009490A
Method, device, and system for providing inquiry and response to output data of artificial intelligence model-based space technology
KR102641137B1