Dialogue reasoning method and device, medium and product

By generating inference paths that reflect the diversity and coherence of reasoning and constructing the logical relationship of agents, the problem of insufficient inference diversity and coherence in the existing technology is solved, and the accuracy and diversity of dialogue reasoning are significantly improved.

CN120197714AActive Publication Date: 2025-06-24SHANG HAI JIE YUE XING CHEN ZHI NENG KE JI YOU XIAN GONG SI
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510685724.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

When existing dialogue reasoning techniques deal with multitasking or complex tasks, it is difficult to balance the diversity of reasoning and the consistency of reasoning, resulting in limited model reasoning capabilities.

Method used

The reasoning task is obtained through the inference model, and an inference path can be generated that reflects the diversity and coherence of reasoning, and reason based on this path to obtain the final result. This method includes parsing the inference task, generating inference paths, building logical relationships between agents, and verifying task results.

Benefits of technology

It significantly improves the accuracy and diversity of inference results, avoids the problem of strategy fixation and task conflict in monologue reasoning, and improves the application effect in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197714A_ABST
    Figure CN120197714A_ABST
Patent Text Reader

Abstract

The invention discloses a dialogue reasoning method and device, a medium and a product, and belongs to the field of artificial intelligence. The dialogue reasoning method comprises the steps of obtaining a reasoning task through a reasoning model, generating a reasoning path according to the reasoning task, and reasoning based on the reasoning path to obtain a reasoning result. The reasoning path capable of reflecting reasoning diversity and reasoning coherence is generated by analyzing the reasoning task, so that the reasoning model analyzes and processes the reasoning task based on the reasoning path, and the accuracy of a reasoning result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a dialogue reasoning method, device, medium and product. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) have demonstrated powerful capabilities in the field of dialogue reasoning. As an important research direction in artificial intelligence, dialogue reasoning aims to build an intelligent system that can understand the questions input by users and draw reasonable conclusions through a series of logical reasoning steps. However, current dialogue reasoning technologies still face many challenges, especially the limited reasoning ability in dealing with complex tasks and multi-task scenarios.

[0003] Currently, the reasoning methods based on large language models mainly adopt the monologue form, that is, a single model completes the entire reasoning process from the first-person perspective. In monologue reasoning, natural language is first processed for intention recognition and process planning, then process planning and process execution are carried out, and finally the execution results are organized and the reasoning conclusion is output.

[0004] However, the prior art still has the following deficiencies: First, the existing monologue reasoning methods are difficult to balance reasoning diversity and reasoning coherence when dealing with multi-tasks or complex tasks, which limits the reasoning ability of the model. The monologue model is prone to falling into a single-path dependence and is difficult to flexibly switch strategies, resulting in insufficient reasoning diversity. Second, the existing reasoning models lack a clear role division and management mechanism during the reasoning process, leading to frequent attention transfer and insufficient coherence. Finally, the prior art lacks an effective task allocation and scheduling mechanism, resulting in unreasonable resource allocation and reasoning path conflicts during the reasoning process. These problems seriously affect the application effect of the dialogue reasoning system in complex scenarios. Summary of the Invention

[0005] Aiming at the problem that the existing monologue reasoning methods are difficult to balance reasoning diversity and reasoning coherence when dealing with multi-tasks or complex tasks, the present application provides a dialogue reasoning method, device, medium and product aiming to significantly improve reasoning diversity and coherence.

[0006] To achieve the above object, some embodiments of the present application provide the following aspects: In a first aspect, some embodiments of the present application further provide a dialogue reasoning method, including: The reasoning model obtains a reasoning task; Based on the reasoning task, a reasoning path is generated, where the reasoning path includes at least one reasoning node, and each reasoning node is associated with a task type and corresponding scenario information; Perform reasoning based on the reasoning path to obtain a reasoning result.

[0007] Optionally, generating the reasoning path based on the reasoning task includes: Parse the reasoning task to obtain scenario information and at least one subtask, and each subtask is associated with a task type; Generate the reasoning path based on the subtask, the task type, and the scenario information, and each subtask is associated with a reasoning node.

[0008] Optionally, performing reasoning based on the reasoning path to obtain a reasoning result includes: The reasoning model generates agents according to the reasoning nodes in the reasoning path, and each reasoning node corresponds to an agent; Obtain the logical relationship between the agents according to the association relationship between the reasoning nodes in the reasoning path; Construct the logical connection between the agents based on the logical relationship; Use the agent to process the subtask associated with the corresponding reasoning node to obtain the corresponding task result; Obtain the reasoning result based on the logical connection between the agents and the task result.

[0009] Optionally, obtaining the reasoning result based on the logical connection between the agents and the task result includes: Each subtask corresponds to a verification strategy, and verify the task result associated with the subtask according to the verification strategy to obtain the verification result associated with the subtask; Generate the reasoning result based on the logical connection and the verification result.

[0010] Optionally, before the reasoning model obtains the reasoning task, it further includes: Train the initial reasoning model to confirm the reasoning model.

[0011] Optionally, training the initial reasoning model to confirm the reasoning model includes: The initial reasoning model obtains a training task set; Parse the training tasks in the training task set to obtain the task type and scenario information; Based on the task type and the scenario information, obtain scenario information and at least one subtraining task, and each subtraining task is associated with a task type; Generate an inference path based on the sub-training task, the task type, and the scenario information, where the inference path includes at least one inference node, and each inference node is associated with a task type and corresponding scenario information; The initial inference model generates agents according to the inference nodes in the inference path, and each of the inference nodes corresponds to an agent; Obtain the logical relationship between the agents according to the association relationship between the inference nodes in the inference path; Construct the logical connection between the agents based on the logical relationship; Use the agent to process the sub-training task associated with the corresponding inference node to obtain the corresponding task result; Obtain the training inference result based on the logical connection between the agents and the task result; Adjust the initial inference model according to the training inference result and the standard result corresponding to the training task.

[0012] Optionally, the adjusting the initial inference model according to the training inference result and the standard result corresponding to the training task includes: Use the PPO strategy to adjust the initial inference model based on the training inference result and the standard result corresponding to the training task.

[0013] In a second aspect, some embodiments of the present application further provide an electronic device, where the electronic device includes: one or more processors; and a memory storing computer program instructions, and the computer program instructions, when executed, cause the processors to execute the steps of the method described above.

[0014] In a third aspect, some embodiments of the present application further provide a computer-readable medium, on which computer program instructions are stored, and the computer program instructions can be executed by a processor to implement the method described above.

[0015] In a fourth aspect, some embodiments of the present application further provide a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method described above are implemented.

[0016] Compared with the related art, in the solution provided by the embodiments of the present application, the dialogue inference method obtains an inference task through an inference model, generates an inference path according to the inference task, and performs inference based on the inference path to obtain an inference result. By analyzing the inference task, an inference path that can reflect inference diversity and inference coherence is generated, so that the inference model can analyze and process the inference task based on the inference path, thereby improving the accuracy of the inference result. Description of the Drawings

[0017] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated. The drawings in the figures do not constitute a scale limitation.

[0018] Figure 1 It is a flowchart of a method of an embodiment of the dialogue reasoning method described in this application; Figure 2 It is a flowchart of a method of an embodiment of generating an inference path in this application; Figure 3 It is a flowchart of a method of an embodiment of obtaining an inference result in this application; Figure 4 It is a flowchart of a method of an embodiment of training an initial inference model in this application; Figure 5 It is an exemplary structural diagram of an electronic device in this application. Detailed implementation manners

[0019] The advantages of this application are further elaborated below in conjunction with the drawings and specific embodiments.

[0020] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0021] The terms used in this disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. The singular forms "a", "the" and "said" used in this disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0022] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0023] In the description of the present application, it should be understood that the numerical labels before the steps do not indicate the order of execution of the steps, but are only used to facilitate the description of the present application and to distinguish each step. Therefore, it should not be construed as a limitation to the present application.

[0024] Embodiment 1 To address the defect that existing monologue-based reasoning methods have difficulty in balancing reasoning diversity and reasoning coherence when dealing with multi-task or complex tasks, the present application proposes a dialogue reasoning method. Refer to Figure 1 , which is a flowchart of a dialogue reasoning method according to a preferred embodiment of the present application. As can be seen from the figure, a dialogue reasoning method provided in this embodiment includes the following steps: S1. The reasoning model obtains a reasoning task; S2. Generate a reasoning path based on the reasoning task, where the reasoning path includes at least one reasoning node, and each reasoning node is associated with a task type and corresponding scenario information; S3. Perform reasoning based on the reasoning path to obtain a reasoning result.

[0025] In this embodiment, the dialogue reasoning method obtains a reasoning task through a reasoning model, generates a reasoning path according to the reasoning task, performs reasoning based on the reasoning path, and obtains a reasoning result. By analyzing the reasoning task, a reasoning path that can reflect reasoning diversity and reasoning coherence is generated, so that the reasoning model can analyze and process the reasoning task based on this reasoning path, thereby improving the accuracy of the reasoning result.

[0026] Embodiment 2 Embodiment 2 of the present application relates to a dialogue reasoning method. Embodiment 2 is an improvement based on Embodiment 1. Refer to Figure 2 The specific improvement lies in: Step S2 may include the following steps: S21. Analyze the reasoning task to obtain scenario information and at least one subtask, and each subtask is associated with a task type; S22. Generate the reasoning path based on the subtask, the task type, and the scenario information, and each subtask is associated with one of the reasoning nodes.

[0027] In this embodiment, the inference model can be a dialogue inference system built based on a large language model. This system can receive the inference tasks input by the user and generate reasonable inference results through a series of inference steps. The inference tasks can be questions, commands, or requirements proposed by the user, such as "analyze the recent financial situation of a certain company", "evaluate the market prospect of a certain product", etc. Combining the system prompt and the analysis of the inference task, clarify the identities and functions of the roles corresponding to each subtask (such as: question decomposer, strategy proposer, solution verifier, etc.). Each subtask has clear responsibilities and inference tasks to ensure the effective combination of multiple perspectives and multiple strategies among the subtasks.

[0028] After the inference model obtains the inference task, it first analyzes the inference task to obtain the scenario information and at least one subtask, and each of the subtasks is associated with a task type. By way of example and not limitation, the scenario information includes Socratic dialogue scenarios, classroom scenarios, teacher-student scenarios, discussion scenarios, etc. For example, for the inference task of "analyze the recent financial situation of a certain company", the system may parse out the scenario information of "financial analysis", as well as subtasks such as "collect financial data", "calculate key financial indicators", and "analyze financial trends", and these subtasks are respectively associated with task types such as "data collection", "data calculation", and "trend analysis".

[0029] Based on the subtasks, the task types, and the scenario information, generate the inference path, and each of the subtasks is associated with one of the inference nodes. The inference path is a directed graph structure, where each node represents an inference step, and the connection between the nodes represents the logical order of the inference. In the above example, the inference path may include three inference nodes, corresponding to the three subtasks of "collect financial data", "calculate key financial indicators", and "analyze financial trends", respectively. These nodes are connected in logical order to form a complete inference path.

[0030] Embodiment Three Embodiment Three of this application relates to a dialogue inference method. Embodiment Three is an improvement based on Embodiment One. Refer to Figure 3 The specific improvement lies in: Step S3 may include: S31. The inference model generates agents according to the inference nodes in the inference path, and each of the inference nodes corresponds to an agent; In this embodiment, the agent is a functional module capable of performing specific tasks and has independent reasoning capabilities. Based on the scenario information corresponding to each reasoning node, the interaction environment can be clarified, and agents can exchange information, propose strategies, and verify viewpoints with each other in a structured dialogue form. The reasoning model has functions such as task recording, progress management, and dynamic feedback to support the stable progress of the reasoning process. For example, the agent corresponding to the "collect financial data" node has the ability to search and organize financial data, the agent corresponding to the "calculate key financial indicators" node has the ability of mathematical calculation and financial analysis, and the agent corresponding to the "analyze financial trends" node has the ability of trend recognition and prediction.

[0031] S32. According to the association relationship between the reasoning nodes in the reasoning path, obtain the logical relationship between the agents; In the above example, there is a sequential dependence relationship between the "collect financial data" node and the "calculate key financial indicators" node because only after collecting financial data can financial indicators be calculated; similarly, there is also a dependence relationship between the "calculate key financial indicators" node and the "analyze financial trends" node. These association relationships determine the logical relationship between the corresponding agents.

[0032] S33. Based on the logical relationship, construct the logical connection between the agents; The logical connection refers to the data transfer and collaboration mechanism between agents. The logical connection is reflected in the interaction between agents to achieve consensus building, view negotiation, and strategy supplementation; through the supervision of agents, the dynamic adjustment of task objectives and the update of role responsibilities are realized, and the reasoning process is continuously refined and optimized in the form of multi-round dialogues. For example, after the "collect financial data" agent completes the task, it will transfer the collected financial data to the "calculate key financial indicators" agent; after the latter completes the calculation, it will transfer the calculation results to the "analyze financial trends" agent. This data transfer and collaboration mechanism ensures the coherence and effectiveness of the entire reasoning process.

[0033] S34. Use the agent to process the subtask associated with the corresponding reasoning node to obtain the corresponding task result; Each agent executes the corresponding subtask according to its own capabilities and the received data, and generates a task result.

[0034] For example, the "collect financial data" agent may generate a dataset containing the company's recent financial statements, the "calculate key financial indicators" agent may generate an analysis report containing indicators such as profit margin and return on assets, and the "analyze financial trends" agent may generate a prediction result on the trend of the company's financial condition.

[0035] S35. Obtain the inference result based on the logical connection between the agents and the task result.

[0036] Specifically, step S35 may include the following steps: Each of the subtasks corresponds to a verification strategy. Verify the task result associated with the subtask according to the verification strategy to obtain the verification result associated with the subtask; Generate the inference result based on the logical connection and the verification result.

[0037] In this embodiment, the verification strategy is a mechanism to ensure the accuracy and reliability of the task result. For example, for the "collect financial data" subtask, the verification strategy may include checking the integrity, consistency, and timeliness of the data; For the "calculate key financial indicators" subtask, the verification strategy may include checking the correctness of the calculation process and the reasonableness of the result; For the "analyze financial trends" subtask, the verification strategy may include checking the reasonableness of the analysis logic and the credibility of the prediction result.

[0038] The inference result is the final output of the entire inference process. It synthesizes the task results of each agent and considers the logical connection between the agents. In the above example, the inference result may be a report comprehensively analyzing the company's recent financial situation, including financial data, key indicators, and trend predictions.

[0039] In this embodiment, the design space of subtask inference can be clearly defined through the system prompt, including the structure of the inference task, role configuration, environmental characteristics, and interaction mechanism, etc. Through scientific definition and flexible setting, the inference model can adaptively construct a suitable dialogue inference framework in different task scenarios. The inference model performs task analysis according to the system prompt and automatically completes the configuration of subtasks, including role assignment, dialogue task decomposition, inference chain construction, and result generation. By dynamically adjusting the inference path, ensure the flexible application of multi-strategy and multi-role collaborative inference in complex tasks. Evaluate the accuracy and reasonableness of the inference result through the verification strategy. Through preset rules and multi-dimensional evaluation indicators, analyze the inference result layer by layer, generate a comprehensive score, and quantify the dialogue inference effect. Feed the evaluation score back to the inference model as a reward, record data according to the performance of the inference task, and activate the reinforcement learning training process. The inference model continuously adjusts the inference strategy and role interaction method based on these reward signals, gradually enhancing the dialogue body inference ability and improving the accuracy and efficiency of task processing.

[0040] Compared with the current monologue-based reasoning model, this application introduces agent role division and environmental deduction, and uses the reinforcement learning algorithm to stimulate the reasoning model's ability to dynamically set and simulate "agent-agent interaction" and "agent-environment interaction", so as to achieve dynamic policy adjustment and multi-perspective analysis during the reasoning process, and effectively respond to the diverse needs of complex tasks. Through dialogue-based reasoning, the diversity and coherence of reasoning can be significantly improved, and problems such as fixed strategies, scattered attention, and task conflicts in monologue-based reasoning can be avoided.

[0041] Taking the reasoning task as a complex mathematical and physical problem as an example: The problems input into the reasoning model are complex compound problems that mix mathematics and physics, including: calculating the eigenvalues of the quantum Hamiltonian operator, verifying the high-energy regularization of physical theories, analyzing the structure of chemical reaction substances, and calculating the gyromagnetic moment, etc. The reasoning model first analyzes the task characteristics, decomposes the problem into multiple subtasks, and constructs a dialogue environment according to the characteristics of different disciplines. In the "quantum energy" environment, the reasoning model creates three subtasks: the subtask of dealing with theoretical physics, the subtask of mathematics, and the subtask of dealing with vector analysis. Each subtask is associated with a corresponding agent, and the agents derive and calculate the eigenvalues of the quantum operator through dialogue, gradually verifying the correctness of the answer. In the reasoning path, the agents respectively propose formula calculations, eigenvalue analyses, and physical meaning interpretations, and finally reach a consensus to obtain the correct answer to the first question. When the task is switched, the reasoning environment is adjusted to "physical theory" based on the reasoning path, and a subtask for dealing with physics is introduced to discuss the high-energy regularization problem of physical theories. Through agent dialogue, integrating physical model assumptions and experimental verifications, a scientific and reasonable conclusion is finally formed. After completing the subtask of physics, the environment is adjusted again based on the reasoning path to enter the "chemical laboratory", and the agent for dealing with chemistry deduces the changes in the structure of substances in chemical reactions and the calculations of related physical quantities to verify the correctness of the gyromagnetic moment. Through the above agent role division and dialogue scenario switching, the reasoning model can flexibly adjust the reasoning strategy when facing multi-disciplinary cross tasks, avoiding the problems of fixed strategies and task interference in the traditional monologue reasoning mode, and significantly improving the diversity, coherence, and accuracy of reasoning.

[0042] Taking the reasoning task as an open-ended problem as an example: The reasoning model generates symbolic agents, connection agents, and reasoning paths by analyzing the reasoning task based on the task type of the reasoning task. The symbolic agent interprets the concept of symbolism, represents knowledge using explicit symbols and logical rules, and emphasizes rule-driven reasoning methods. The connection agent conducts knowledge modeling through neural networks and distributed learning, and uses data to train the model to autonomously form representations. According to the reasoning path, the information output by the symbolic agent and the connection agent is discussed. For example, the word "bank" can refer to either a riverbank or a bank. The symbolic agent explicitly differentiates through context rules, while the connection agent relies on weighted connections to automatically deduce the meanings of polysemous words. By comparison, it is pointed out that the symbolic agent relies on artificially defined rules, while the connection agent learns from data. In terms of reasoning stability, the symbolic agent performs stably within the rules but lacks scalability; the connection agent performs poorly in distribution transfer and zero-shot tasks and is difficult to maintain robustness. The discussion finally reaches a consensus: the symbolic agent and the connection agent each have their advantages and disadvantages, and the combination of the two can achieve a balance between performance and interpretability. For example, through the integration of neural perception and symbolic reasoning, a complementary artificial intelligence model can be formed. This case demonstrates the advantages of conversational reasoning in the analysis of open-ended questions: through the expression of multi-role viewpoints and dialogue interactions, it can promote the comprehensive analysis and in-depth interpretation of academic issues.

[0043] Embodiment 4 Embodiment 4 of this application relates to a dialogue reasoning method. Embodiment 4 is an improvement based on Embodiment 1. The specific improvement lies in: Before performing step S1, it may further include: Training the initial reasoning model to confirm the reasoning model.

[0044] Furthermore, referring to Figure 4 The training of the initial reasoning model to confirm the reasoning model may include the following steps: A1. The initial reasoning model obtains a training task set; The training task set is a set of samples for training the model, and each sample contains a reasoning task and the corresponding standard result. For example, the training task set may contain multiple reasoning tasks in different fields and of different complexities, such as financial analysis, market prediction, scientific research, etc.

[0045] A2. Analyze the training tasks in the training task set to obtain the task type and scenario information; The initial reasoning model needs to learn how to extract key information from the training tasks, including the task type (such as analysis, prediction, evaluation, etc.) and scenario information (such as finance, market, science, etc.).

[0046] A3. Based on the task type and the scenario information, obtain the scenario information and at least one sub-training task, with each sub-training task associated with a task type; The initial inference model needs to learn how to break down complex training tasks into multiple sub-training tasks and assign appropriate task types to each sub-training task.

[0047] A4. Based on the sub-training tasks, the task types, and the scenario information, generate an inference path, where the inference path includes at least one inference node, and each inference node is associated with a task type and the corresponding scenario information; The initial inference model needs to learn how to construct an effective inference path, ensure the reasonable logical relationship between each inference node, and be able to effectively solve the training tasks.

[0048] A5. The initial inference model generates agents according to the inference nodes in the inference path, with each inference node corresponding to an agent; The initial inference model needs to learn how to generate appropriate agents for different types of inference nodes to ensure that each agent has the ability to execute the corresponding sub-training task.

[0049] A6. According to the association relationship between each inference node in the inference path, obtain the logical relationship between the agents; The initial inference model needs to learn how to identify the association relationship between inference nodes and convert these relationships into the logical relationship between agents.

[0050] A7. Based on the logical relationship, construct the logical connection between the agents; The initial inference model needs to learn how to establish the data transfer and cooperation mechanism between agents to ensure the coherence and effectiveness of the entire inference process.

[0051] A8. Use the agents to process the sub-training tasks associated with the corresponding inference nodes and obtain the corresponding task results; The initial inference model needs to learn how to coordinate each agent to execute the corresponding sub-training task and generate the task results.

[0052] A9. Based on the logical connection between the agents and the task results, obtain the training inference results; The initial inference model needs to learn how to synthesize the task results of each agent and consider the logical connection between agents to generate the final training inference results.

[0053] A10. Adjust the initial inference model according to the training inference results and the standard results corresponding to the training tasks.

[0054] The initial inference model needs to continuously adjust its parameters and strategies by comparing the differences between the training inference results and the standard results, so as to improve its inference ability.

[0055] Further, step A10 may include: adjusting the initial inference model based on the training inference results and the standard results corresponding to the training task by adopting the PPO (Proximal Policy Optimization) strategy.

[0056] PPO (Proximal Policy Optimization) is a reinforcement learning algorithm that maximizes the expected return by optimizing the policy function. In this embodiment, the PPO strategy is used to adjust the parameters of the initial inference model to make the generated training inference results closer to the standard results. Specifically, the PPO strategy calculates the differences between the training inference results and the standard results to generate gradient information, and then uses this gradient information to update the model parameters. This adjustment process will be repeated continuously until the model performance reaches the expected level.

[0057] Embodiment Five A dialogue inference method may include the following steps: The inference model obtains an inference task; Generate an inference path based on the inference task, where the inference path includes at least one inference node, and each inference node is associated with a task type and corresponding scenario information; Perform inference based on the inference path to obtain an inference result.

[0058] In this embodiment, the inference model may be a dialogue inference system constructed based on a neural network, which can receive the inference tasks input by users and generate reasonable inference results through a series of inference steps. The inference tasks may be questions, commands or requirements proposed by users, such as "analyze the market competitiveness of a certain product", "evaluate the development prospects of a certain technology", etc.

[0059] After the inference model obtains the inference task, it first parses the inference task to obtain scenario information and at least one subtask, and each of the subtasks is associated with a task type. For example, for the inference task of "analyze the market competitiveness of a certain product", the system may parse out the scenario information of "market analysis", as well as subtasks such as "collect market data", "analyze the situation of competing products", "evaluate product advantages", etc., and these subtasks are respectively associated with task types such as "data collection", "competitor analysis", "advantage evaluation", etc.

[0060] Generate the inference path based on the subtasks, the task type, and the scenario information. Each subtask is associated with an inference node. The inference path is a directed graph structure, where each node represents an inference step, and the connections between nodes represent the logical order of the inferences. In the above example, the inference path may contain three inference nodes, corresponding to the three subtasks of "collecting market data", "analyzing competitor situations", and "evaluating product advantages" respectively. These nodes are connected in logical order to form a complete inference path.

[0061] The inference model generates agents according to the inference nodes in the inference path. Each inference node corresponds to an agent. An agent is a functional module capable of performing specific tasks and has independent inference capabilities. For example, the agent corresponding to the "collecting market data" node has the ability to search for and organize market data, the agent corresponding to the "analyzing competitor situations" node has the ability to analyze and compare competitors, and the agent corresponding to the "evaluating product advantages" node has the ability to identify and evaluate advantages.

[0062] Obtain the logical relationships between the agents according to the association relationships between the various inference nodes in the inference path. In the above example, there is a sequential dependence relationship between the "collecting market data" node and the "analyzing competitor situations" node because only after collecting market data can competitor situations be analyzed; similarly, there is also a dependence relationship between the "analyzing competitor situations" node and the "evaluating product advantages" node. These association relationships determine the logical relationships between the corresponding agents.

[0063] Construct the logical connections between the agents based on the logical relationships. The logical connection refers to the data transfer and cooperation mechanism between the agents. For example, after the "collecting market data" agent completes the task, it will transfer the collected market data to the "analyzing competitor situations" agent; after the latter completes the analysis, it will transfer the analysis results to the "evaluating product advantages" agent. This data transfer and cooperation mechanism ensures the coherence and effectiveness of the entire inference process.

[0064] Use the agents to process the subtasks associated with the corresponding inference nodes and obtain the corresponding task results. Each agent executes the corresponding subtask according to its own capabilities and the received data and generates a task result. For example, the "collecting market data" agent may generate a dataset containing information such as market size and growth rate, the "analyzing competitor situations" agent may generate an analysis report containing information such as competitor characteristics and market share, and the "evaluating product advantages" agent may generate an evaluation result regarding the product's competitive advantages.

[0065] Each of the sub-tasks corresponds to a verification strategy. According to the verification strategy, the task result associated with the sub-task is verified to obtain the verification result associated with the sub-task. The verification strategy is a mechanism to ensure the accuracy and reliability of the task result. For example, for the sub-task of "collecting market data", the verification strategy may include checking the integrity, consistency, and timeliness of the data; for the sub-task of "analyzing competitor situations", the verification strategy may include checking the rationality of the analysis logic and the objectivity of the result; for the sub-task of "evaluating product advantages", the verification strategy may include checking the rationality of the evaluation criteria and the credibility of the result.

[0066] Based on the logical connection between the agents and the task result, the inference result is obtained. Alternatively, based on the logical connection and the verification result, the inference result is generated. The inference result is the final output of the entire inference process. It synthesizes the task results of each agent and takes into account the logical connection between the agents. In the above example, the inference result may be a comprehensive report analyzing the market competitiveness of the product, including market data, competitor situations, and product advantages, etc.

[0067] Before the inference model obtains the inference task, it is also necessary to train the initial inference model to confirm the inference model. The training process includes the following steps: The initial inference model obtains a training task set. The training task set is a set of samples used to train the model, and each sample contains an inference task and the corresponding standard result. For example, the training task set may contain multiple inference tasks in different fields and of different complexities, such as market analysis, technology evaluation, product R & D, etc.

[0068] Parse the training tasks in the training task set to obtain the task type and scenario information. The initial inference model needs to learn how to extract key information from the training tasks, including the task type (such as analysis, evaluation, prediction, etc.) and the scenario information (such as market, technology, product, etc.).

[0069] Based on the task type and the scenario information, obtain the scenario information and at least one sub-training task, and each sub-training task is associated with a task type. The initial inference model needs to learn how to decompose complex training tasks into multiple sub-training tasks and assign appropriate task types to each sub-training task.

[0070] Based on the sub-training tasks, the task type, and the scenario information, generate an inference path. The inference path includes at least one inference node, and each inference node is associated with a task type and the corresponding scenario information. The initial inference model needs to learn how to construct an effective inference path to ensure that the logical relationship between each inference node is reasonable and can effectively solve the training task.

[0071] The initial inference model generates agents according to the inference nodes in the inference path, and each inference node corresponds to an agent. The initial inference model needs to learn how to generate appropriate agents for different types of inference nodes to ensure that each agent has the ability to execute the corresponding sub-training task.

[0072] According to the association relationships between the inference nodes in the inference path, obtain the logical relationships between the agents. The initial inference model needs to learn how to identify the association relationships between inference nodes and transform these relationships into the logical relationships between agents.

[0073] Based on the logical relationships, construct the logical connections between the agents. The initial inference model needs to learn how to establish the data transfer and cooperation mechanisms between agents to ensure the coherence and effectiveness of the entire inference process.

[0074] Use the agents to process the sub-training tasks associated with the corresponding inference nodes and obtain the corresponding task results. The initial inference model needs to learn how to coordinate each agent to execute the corresponding sub-training task and generate task results.

[0075] Based on the logical connections between the agents and the task results, obtain the training inference results. The initial inference model needs to learn how to integrate the task results of each agent and consider the logical connections between agents to generate the final training inference results.

[0076] According to the training inference results and the standard results corresponding to the training tasks, adjust the initial inference model. The initial inference model needs to continuously adjust its own parameters and strategies by comparing the differences between the training inference results and the standard results to improve the inference ability.

[0077] Use the A2C strategy to adjust the initial inference model based on the training inference results and the standard results corresponding to the training tasks. A2C (Advantage Actor-Critic) is a reinforcement learning algorithm that combines the advantages of policy gradient and value function approximation. In this embodiment, the A2C strategy is used to adjust the parameters of the initial inference model to make the generated training inference results closer to the standard results. Specifically, the A2C strategy calculates the differences between the training inference results and the standard results to generate gradient information, and then uses this gradient information to update the model parameters. This adjustment process will be repeated continuously until the model performance reaches the expected level.

[0078] It should be noted that Embodiment 1, Embodiment 2, Embodiment 3, Embodiment 4, and Embodiment 5 are all types of dialogue inference methods.

[0079] In addition, some embodiments of the present application also provide an electronic device. The electronic device may be various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and so on. The electronic device may also be various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0080] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processors to perform the steps of the method provided in any one or more of the above embodiments. Figure 5 An exemplary structural diagram of the electronic device is disclosed. As Figure 5 shown, the electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and may be mounted on a common motherboard or otherwise installed as needed. The processor may process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if needed, multiple processors and / or multiple buses may be used together with multiple memories and multiple memories. Similarly, multiple electronic devices may be connected, with each device providing part of the necessary operations (such as, as a server array, a set of blade servers, or a multi-processor system). Among them, the components, their connections and relationships, and their functions shown herein are only examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0081] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 may be connected by a bus or other means, Figure 5 taking connection by bus as an example.

[0082] The input device 1103 can receive input digital or character information and generate key signal inputs related to the user settings and function controls of the electronic device, such as input devices like touchscreens, keypads, mice, trackpads, touchpads, pointing sticks, one or more mouse buttons, trackballs, joysticks, etc. The output device 1104 can include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors), etc. The display device can include, but is not limited to, liquid crystal displays (LCDs), light-emitting diode (LED) displays, and plasma displays. In some embodiments, the display device can be a touchscreen.

[0083] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide inputs to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback); and inputs from the user can be received in any form (including acoustic input, voice input, or haptic input).

[0084] In the embodiments of the present application, computer programs / instructions are stored on a computer-readable medium. When the computer programs / instructions are executed by a processor, the steps of the methods provided in any one or more of the above embodiments are implemented. The computer-readable medium can be included in the electronic device described in the above embodiments; or it can exist separately without being assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.

[0085] The memory 1102 can be used as a non-transitory computer-readable storage medium for storing non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, so as to implement the program instructions / modules corresponding to the methods provided in any one or more of the above embodiments of the present application.

[0086] The memory 1102 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may optionally include a memory remotely disposed relative to the processor 1101, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0087] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0088] The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0089] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0090] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. For example, an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device can be used. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, for example, a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. In addition, some steps or functions of this application can be implemented by hardware, for example, as a circuit that cooperates with a processor to execute each step or function.

[0091] The computer program product provided by the embodiments of this application includes one or more computer programs / instructions. When the computer programs / instructions are executed by a processor, they entirely or partially generate the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available media can be magnetic media (for example, floppy disks, hard disks, magnetic tapes), optical media (for example, DVDs), or semiconductor media (for example, solid state disks (SSDs)), etc.

[0092] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0093] The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference numerals in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. The multiple elements or devices recited in the apparatus claims may also be implemented by one element or device through software or hardware. The terms "first", "second", etc. are only used for descriptive distinction and do not represent any specific order, nor can they be construed as indicating or implying relative importance.

[0094] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily mention changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A dialogue reasoning method, characterized in that, Including: The inference model obtains an inference task; Generate an inference path based on the inference task, where the inference path includes at least one inference node, and each inference node is associated with a task type and corresponding scenario information; Perform inference based on the inference path to obtain an inference result.

2. The dialogue reasoning method according to claim 1, wherein The generating the inference path based on the inference task includes: Parse the inference task to obtain scenario information and at least one subtask, and each subtask is associated with a task type; Generate the inference path based on the subtask, the task type, and the scenario information, and each subtask is associated with one of the inference nodes.

3. The dialogue reasoning method according to claim 2, characterized in that, Performing inference based on the inference path to obtain an inference result includes: The inference model generates agents according to the inference nodes in the inference path, and each inference node corresponds to an agent; Obtain the logical relationship between the agents according to the association relationship between the inference nodes in the inference path; Construct the logical connection between the agents based on the logical relationship; Use the agent to process the subtask associated with the corresponding inference node to obtain the corresponding task result; Obtain the inference result based on the logical connection between the agents and the task result.

4. The dialogue reasoning method according to claim 3, characterized in that The obtaining the inference result based on the logical connection between the agents and the task result includes: Each subtask corresponds to a verification strategy, verify the task result associated with the subtask according to the verification strategy to obtain the verification result associated with the subtask; Generate the inference result based on the logical connection and the verification result.

5. The dialogue reasoning method according to claim 1, wherein Before the inference model obtains the inference task, it further includes: Train the initial inference model to confirm the inference model.

6. The dialogue reasoning method according to claim 5, wherein The training the initial inference model to confirm the inference model includes: The initial inference model obtains a training task set; Parse the training tasks in the training task set to obtain the task type and scenario information; Based on the task type and the scenario information, obtain scenario information and at least one subtraining task, and each subtraining task is associated with a task type; Generate an inference path based on the subtraining task, the task type, and the scenario information, where the inference path includes at least one inference node, and each inference node is associated with a task type and corresponding scenario information; The initial inference model generates agents according to the inference nodes in the inference path, and each inference node corresponds to an agent; Obtain the logical relationship between the agents according to the association relationship between the inference nodes in the inference path; Construct the logical connection between the agents based on the logical relationship; Use the agent to process the subtraining task associated with the corresponding inference node to obtain the corresponding task result; Obtain the training inference result based on the logical connection between the agents and the task result; Adjust the initial inference model according to the training inference result and the standard result corresponding to the training task.

7. The dialogue inference method according to claim 6, wherein Adjusting the initial inference model according to the training inference result and the standard result corresponding to the training task includes: Adjusting the initial inference model by using the PPO strategy based on the training inference result and the standard result corresponding to the training task.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which when executed cause the processor to execute the steps of the method according to any one of claims 1 to 7.

9. A computer-readable medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-agent knowledge reasoning method and system based on deep reinforcement learning

    CN115526317A

  • Knowledge graph-based reasoning method and system

    CN116720584A

  • Multi-granularity time sequence knowledge graph question and answer method and device and storage medium

    CN118013003A

  • Intelligent question and answer method based on cooperation of large language model and knowledge graph

    CN118797017A

  • Online reasoning method and device, electronic equipment and storage medium

    CN119149221A