Target-oriented agent task guiding reasoning method, device and equipment

Through the goal-oriented agent task guidance reasoning method, the problem of inconsistency between the execution actions and business goals in the agent task planning is solved, and efficient and accurate task planning and user experience improvement are achieved.

CN120012922APending Publication Date: 2025-05-16BEIJING GANYI INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510049476.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When planning tasks, existing agents perform actions that are inconsistent with business goals, resulting in inefficient task planning and affecting user experience.

Method used

A goal-oriented agent task-guided reasoning method is provided, by obtaining user input information, determining the starting state and target state, obtaining action state maps, and generating and executing target paths based on this information to ensure that the execution action is consistent with the business goal.

Benefits of technology

It improves the efficiency and accuracy of the agent task planning, makes the execution path consistent with the business goals, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012922A_ABST
    Figure CN120012922A_ABST
Patent Text Reader

Abstract

The invention provides a target-oriented agent task guiding reasoning method, device and equipment. The method comprises the following steps: acquiring user input information; determining an initial state and a target state of the user according to the user input information; acquiring an action state map; generating a target path according to the initial state, the target state and the action state map; and executing the action included in the target path, obtaining and outputting an execution result, and waiting for obtaining the user input information again. The task of the intelligent agent is efficiently planned, the consistency of the execution action and the business target is guaranteed, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence, and specifically relates to a goal-oriented intelligent agent task-guided reasoning method, device and equipment. Background Art

[0002] In the field of artificial intelligence, especially the design and application of AI agents, existing agent reasoning mainly relies on large models for problem decomposition and task planning. Although large models have strong general capabilities, they often lack a deep understanding of task objectives in specific industry scenarios, resulting in the agent's execution actions being inconsistent with business objectives, resulting in inefficient task planning and affecting user experience. Summary of the invention

[0003] This application provides a goal-oriented agent task guidance reasoning method, device and equipment to solve the problem that when the current agent performs task planning, the execution action is inconsistent with the business goal, which makes the task planning efficiency low and affects the user experience. It can achieve efficient planning of the agent's tasks, ensure the consistency of the execution action and the business goal, and improve the user experience.

[0004] This application provides a goal-oriented agent task-guided reasoning method, including: Get user input information; Determine the user's starting state and target state according to the user input information; Get the action state graph; Generate a target path according to the starting state, the target state and the action state graph; Execute the actions included in the target path, obtain and output the execution result, and wait for obtaining the user input information again.

[0005] According to the goal-oriented intelligent agent task-guided reasoning method provided in the present application, before obtaining the action state graph, the method also includes: obtaining a business document; parsing the business document to obtain multiple subtasks; converting each of the multiple subtasks into at least one executable instruction; establishing a logical dependency relationship between the at least one executable instruction to obtain an instruction sequence; and generating an action state graph according to the instruction sequence corresponding to each subtask.

[0006] According to the goal-oriented intelligent agent task-guided reasoning method provided in the present application, the target path is generated according to the starting state, the target state and the action state graph, including: obtaining a list of disabled actions; traversing all feasible actions in the action state graph according to the disabled action list and the target state; recursively generating a sub-plan that satisfies the preconditions of each feasible action according to the starting state, the sub-plan including the preconditions, actions and execution results; generating multiple paths according to the sub-plan and the feasible actions; and selecting a target path from the multiple paths.

[0007] According to the goal-oriented intelligent agent task-guided reasoning method provided in the present application, the selecting of the target path from the multiple paths includes: calculating the total execution cost of each path in the multiple paths, wherein the total execution cost is related to the execution cost and action effectiveness of the actions included in each path; and determining the path with the lowest execution cost as the target path.

[0008] According to the goal-oriented intelligent agent task-guided reasoning method provided in the present application, before traversing all feasible actions in the action state graph according to the target state, the method also includes: checking the validity of the target state; when the target state is valid, determining whether the target state has been satisfied in the starting state; if it has been satisfied, determining that the target path is empty.

[0009] According to the goal-oriented intelligent agent task-guided reasoning method provided in the present application, before outputting the execution result, the method also includes: determining the current state based on the execution result; determining whether the current state is consistent with the expected state of the action; if inconsistent, adding the action to the disabled action list, and regenerating the target path based on the updated disabled action list.

[0010] According to the goal-oriented intelligent agent task-guided reasoning method provided by the present application, the method also includes: obtaining interaction records; obtaining the state to be updated from the action state map according to the interaction records; determining all the actions to be updated corresponding to the state to be updated, and all the actions to be updated can reach the expected state from the state to be updated through different paths; simulating the process from the state to be updated to the expected state according to all the actions to be updated to obtain the action simulation validity of each action to be updated; determining the actual action validity of each action to be updated according to the interaction records; obtaining the overall validity of each action to be updated according to the action simulation validity and the actual action validity; updating the overall validity of each action to be updated, and the higher the overall validity, the greater the probability that the path where the corresponding action to be updated is located is selected when generating the target path.

[0011] The present application also provides a goal-oriented agent task-guided reasoning device, comprising: A first acquisition unit, used to acquire user input information; A determination unit, configured to determine a starting state and a target state of the user according to the user input information; A second acquisition unit, used for acquiring an action state graph; A generating unit, configured to generate a target path according to the starting state, the target state and the action state graph; The execution unit is used to execute the actions included in the target path, obtain and output the execution result, and wait for obtaining the user input information again.

[0012] The present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements any of the above-mentioned goal-oriented intelligent agent task-guided reasoning methods.

[0013] The present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned goal-oriented intelligent agent task-guided reasoning methods.

[0014] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned goal-oriented intelligent agent task-guided reasoning methods.

[0015] The goal-oriented agent task-guided reasoning method, device and equipment provided by the present application first determine the starting state and target state based on user input information, then obtain the action state map, and generate the target path based on the starting state, target state and action state map, and execute the actions included in the target path, obtain and output the execution result, and then wait for the user input information to be obtained again. In this way, the determined execution path can be consistent with the current business goal, and at the same time, the response efficiency and execution accuracy of the agent can be improved. It enables the agent to make smarter decisions that are more in line with industry characteristics. Enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1This is one of the flowcharts of the goal-oriented intelligent agent task-guided reasoning method provided in this application.

[0018] Figure 2 This is the second flowchart of the goal-oriented intelligent agent task-guided reasoning method provided in this application.

[0019] Figure 3 This is the third flowchart of the goal-oriented intelligent agent task-guided reasoning method provided in this application.

[0020] Figure 4 This is the fourth flowchart of the goal-oriented intelligent agent task-guided reasoning method provided in this application.

[0021] Figure 5 It is a structural diagram of the goal-oriented intelligent agent task-guided reasoning device provided in this application.

[0022] Figure 6 It is a structural schematic diagram of the electronic device provided by this application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0025] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0026] Currently, when intelligent agents perform task planning, they reason based on general processes. In specific implementation scenarios, there are inconsistencies between the planned execution actions and business goals, resulting in low task planning efficiency and poor user experience.

[0027] In response to the above problems, an embodiment of the present application provides a goal-oriented intelligent agent task guided reasoning method, device and equipment. The embodiment of the present application is described in detail below in conjunction with the accompanying drawings.

[0028] See also Figure 1 , Figure 1 This is one of the flowcharts of the goal-oriented intelligent agent task guidance reasoning method provided by the present application. The goal-oriented intelligent agent task guidance recommendation method includes the following steps.

[0029] S101, obtaining user input information.

[0030] The input information may refer to the content of the user's questions to the intelligent agent, or the intelligent agent's reply content, as well as the user's objective attributes, such as the user's gender, age group, whether a certain assessment has been completed, whether a certain condition has been met, and all other content that is helpful in determining the user's current status.

[0031] S102: Determine a starting state and a target state according to the user input information.

[0032] Among them, semantic analysis of user needs can be performed based on user input information, and application scenarios and business goals can be obtained based on the results of semantic analysis. At the same time, the target state corresponding to the business goal is determined, and the current state of the user is analyzed to obtain the current state corresponding to the business goal. When obtaining the starting state, in addition to determining it based on user input information, it can also include querying related databases and interfaces to obtain the starting state. In the specific implementation, the state in this solution refers to the specific links of the user journey or work tasks, such as "user has opened an account", "user deposited successfully", etc. The user journey refers to expressing the contact points (touchpoints) with the user in the form of a timeline from the user's perspective according to the process of business development, and analyzing what happens at each touchpoint, the user's feelings, benefits, costs and other information. Suitable for experience optimization, product design, etc.

[0033] S103, obtaining an action state graph.

[0034] The action state graph is associated with the business document to indicate the relationship between task-state-action. The business document may include Standard Operating Procedure (SOP) documents and Quality Assurance (QA) data. In particular, the action state graph may also be associated with the business goal description. That is, the action state graph may be generated from the business document and the business goal description.

[0035] S104, generating a target path according to the starting state, the target state and the action state graph.

[0036] There may be multiple actions to achieve the target state from the starting state, that is, multiple paths may be determined based on the action state graph, so it is necessary to select a target path from these multiple paths. The target path may be the path with the shortest time or the path with the shortest cost.

[0037] S105, executing the actions included in the target path, acquiring and outputting the execution results, and waiting to acquire the user input information again.

[0038] The execution result includes the answer reply and / or guidance content for the user input information. Figure 2 ,like Figure 2 As shown, the system obtains user input information, that is, the content of the user's question or the user's historical question content, and then performs a target check based on the user input information, determines the starting state and the target state based on the target check, and then performs path planning based on the starting state, the target state and the action state map to obtain the target path. Then, an action search is performed based on the action in the target path, and the action is executed. Then, the execution result is output to the user, and the execution result includes an answer reply and / or guidance content. At the same time, after the system executes the action in the target path, the user status can also be updated, and it can be used as a reference for the next path planning. When the system outputs the execution result, it will wait to obtain the user input information. If the user enters a question again, the target check, path planning, action search and action execution steps are repeated, and the execution result is output again, and this cycle is repeated until the user no longer asks questions.

[0039] It can be seen that in this embodiment, the starting state and the target state are first determined based on the user input information, and then the action state map is obtained, and the target path is generated based on the starting state, the target state and the action state map, and the actions included in the target path are executed, and the execution result is output. In this way, the determined execution path can be consistent with the current business goal, and the response efficiency and execution accuracy of the intelligent agent can be improved, so that the intelligent agent can make smarter decisions that are more in line with industry characteristics. Enhance the user experience.

[0040] In a possible embodiment, before obtaining the action state graph, the method also includes: obtaining a business document; parsing the business document to obtain multiple subtasks; converting each of the multiple subtasks into at least one executable instruction; establishing a logical dependency relationship between the at least one executable instruction to obtain an instruction sequence; and generating an action state graph according to the instruction sequence corresponding to each subtask.

[0041] Among them, for SOP documents, the logical relationship between tasks, actions and states can be extracted based on the document content. For QA data, it can be summarized and classified based on the question intent and the answer content. At the same time, the business goals can be decomposed, and then multiple subtasks corresponding to a business goal can be obtained based on the parsed business documents. Each subtask consists of a precondition, an action and a completion effect. The precondition is the state that needs to be achieved before executing the current subtask. For example, the precondition for providing product recommendations is to complete the risk assessment. In particular, before generating the action state map, the integrity and consistency of the instruction sequence can be verified first, and then the action state map can be generated based on the instruction sequence after the verification.

[0042] For example Figure 3 As shown in the figure, when generating the action state graph, document processing is first performed, which includes document parsing and classification of standard operating procedure documents, quality assurance data, and business goal descriptions. Then, based on the parsed and classified content, the target is disassembled to obtain multiple subtasks. Each subtask is then converted into a specific executable instruction, and a logical dependency relationship between instructions is established. Then, the action state graph is constructed based on the task-state-action relationship according to the instruction sequence.

[0043] It can be seen that in this embodiment, an action state graph is constructed based on business documents, so that the intelligent agent can be consistent with the business goals when performing tasks, thereby improving the accuracy of the intelligent agent.

[0044] In a possible embodiment, generating a target path according to the starting state, the target state and the action state graph includes: obtaining a list of disabled actions; traversing all available actions in the action state graph according to the list of disabled actions and the target state; recursively generating a sub-plan that satisfies the prerequisites of each available action according to the starting state, the sub-plan including the prerequisites, actions and execution results; generating multiple paths according to the sub-plans and the available actions; and selecting a target path from the multiple paths.

[0045] In the working process, the intelligent agent will output answer replies and / or guidance content based on the dialogue history and the current state of the user. Therefore, when generating the target path, you can first determine all the actions that can reach the target state. Of course, if the traversed action is in the disabled action list, the action is excluded. After traversing all the possible actions, recursively generate a sub-plan that meets its preconditions for each possible action. Each sub-plan corresponds to a precondition, an action, and an execution result. The precondition is the execution result of the next recursive sub-plan. Until the precondition of a recursively determined sub-plan is the starting state.

[0046] In a specific implementation, when multiple paths are generated based on sub-plans and actionable actions, after generating the sub-plans, the circular dependencies between actions can be processed, and then the sub-plans and current actions can be merged to obtain multiple paths, so as to ensure that the action sequence is non-repetitive and in the correct order. In a specific implementation, if the action corresponding to the sub-plan is also in the disabled action list, the path corresponding to the sub-plan is deleted. In particular, the disabled action lists corresponding to different business scenarios can be different, that is, the same action may exist in the disabled action list of a certain business scenario, but in another business scenario, the disabled action list does not include the action.

[0047] It can be seen that in this embodiment, path planning based on the two basic elements of state and action can make the guidance process clearer and improve the efficiency and accuracy of the intelligent agent's task execution.

[0048] In a possible embodiment, selecting a target path from the multiple paths includes: calculating a total execution cost of each of the multiple paths, wherein the total execution cost is related to the execution cost and action effectiveness of the actions included in each path; and determining the path with the lowest execution cost as the target path.

[0049] When calculating the total execution cost, the cost of each action included in each path can be calculated first, and then the total cost of the path can be calculated based on the costs of all actions in the path. In a specific implementation, the quotient of the action cost of each action and the action effectiveness of the action is the cost of the action.

[0050] It can be seen that in this embodiment, the optimal path is found through cost calculation based on action effectiveness, which can improve the intelligence of path selection and the efficiency of selecting the target path.

[0051] In a possible embodiment, before traversing all feasible actions in the action state graph according to the target state, the method also includes: checking the validity of the target state; if the target state is valid, determining whether the target state has been satisfied in the starting state; if so, determining that the target path is empty.

[0052] The validity check includes checking whether there is a conflict between the starting state and the target state. For example, if the starting state gender is female and the target state gender is male, there is a conflict. The validity check can also verify whether all states are in the state set. Finally, it is determined whether the target state has been met. That is, if the target state is included in the starting state, there is no need to execute the relevant actions. At this time, the actions in the target path are empty.

[0053] It can be seen that in this embodiment, before generating the target path, performing validity check and verifying whether the target state has been satisfied can improve the efficiency of business guidance reasoning and save computing resources.

[0054] In a possible embodiment, before outputting the execution result, the method further includes: determining a current state based on the execution result; determining whether the current state is consistent with an expected state of the action; if not, adding the action to a disabled action list, and regenerating a target path based on the updated disabled action list.

[0055] Among them, after obtaining the target path, you can also verify the feasibility of the path first to check whether there are infeasible actions. The infeasible action may refer to an action with too high an action cost or too low an action effectiveness. The action cost refers to the cost required to complete the action, and the action effectiveness refers to the probability that the action can achieve the expected state. When the action in the target path is an actionable action, execute the action and obtain the execution result. Then determine the current state based on the execution result, and output the execution result. If the current state is inconsistent with the expected state, add the current action to the disabled action list. Then regenerate the target path, and execute the actions in the newly generated target path until the current state is the target state.

[0056] It can be seen that in this embodiment, by monitoring the status after each action is executed, the disabled list and the target path can be updated in time, thereby improving the efficiency and accuracy of business guidance reasoning.

[0057] In a possible embodiment, the method also includes: obtaining interaction records; obtaining the state to be updated from the action state map according to the interaction records; determining all actions to be updated corresponding to the state to be updated, and all the actions to be updated can reach the expected state from the state to be updated through different paths; simulating the process from the state to be updated to the expected state according to all the actions to be updated to obtain the action simulation validity of each action to be updated; determining the actual action validity of each action to be updated according to the interaction records; obtaining the overall validity of each action to be updated according to the action simulation validity and the actual action validity; updating the overall validity of each action to be updated, and the higher the overall validity, the greater the probability that the path where the corresponding action to be updated is located will be selected when generating the target path.

[0058] When generating a target path, the agent can update the effectiveness of each action based on the execution results, thereby selecting the most effective action in the current state and generating the target path. In the specific implementation, feedback learning can be performed based on Monte Carlo Tree Search (MCTS) to optimize the response process.

[0059] For the specific iterative optimization process, please refer to Figure 4 After feedback learning begins, the state selection is performed first. This step is used to determine the state nodes to be updated. The selection method can be based on the upper confidence bound for trees (UCT) to select the state to be updated. This is to balance the development of the current high-win rate nodes and the exploration of unknown nodes with less statistical data. The specific formula is as follows: Among them, Q(v) represents the number of times action v is valid, N(V) represents the number of times action v is accessed, and N(v p ) indicates the preceding action v p The number of visits, c represents the exploration parameter, which is a constant.

[0060] After selecting the state to be updated, continue to explore downward with the state to be updated as the starting point, and obtain multiple actions corresponding to the state to be updated. These actions can all reach the desired state through different action paths. Then when calculating the effectiveness of each action, it can be divided into two stages. One stage is the cold start stage. At this time, there is no user data, so the large model intelligent system can be used to simulate user dialogues and calculate the effectiveness of the action simulation of each action. In particular, the Dual-Agent model can be used to simulate user behavior and calculate the effectiveness of the action. In the online stage, there is real-time feedback from users at this time, so the success rate of each attempted action can be calculated through the dialogue log data, and the actual effectiveness of each action can be calculated. Then the overall effectiveness of the action is calculated based on the effectiveness of the action simulation and the actual effectiveness of the action, and the effectiveness of the action is updated based on the overall effectiveness of the action. The effectiveness of the action simulation and the actual effectiveness of the action can both be calculated by the following formula: in, Indicates that the next step attempts to perform the action. Indicates that after attempting the action, the current state changes to match the expected state.

[0061] When calculating the overall effectiveness, it can be calculated by the following formula: Overall effectiveness = γ action simulation effectiveness + (1-γ) action actual effectiveness, Among them, γ represents the degree of exploration. When there is less user data, γ is relatively small, and vice versa. This can balance the cold start data and the actual user interaction data, thereby achieving a smooth transition and reducing the randomness of user clicks and completions when there is less online feedback data.

[0062] In particular, after each simulation, the following is updated for each node v along the visited path: Increase the number of visits: N(n)←N(n)+1; Update node: Q(v)←Q(v)+Q(v final ).

[0063] In the specific implementation, the interaction data includes simulation data and user feedback data, and the user feedback data includes session data, state data, and action selection data. Session data refers to the user's questions and the system's responses. State data can be the state at the time of the question and the state at the end of the guidance. Action selection data refers to the candidate items and the final selection of the action. For example, in the guidance question scenario, the user will click on the guidance question recommended by the system to participate in the action selection.

[0064] It can be seen that in this embodiment, the effectiveness of each action is continuously updated through the back-propagation mechanism, so as to screen out the most effective action, which can improve the intelligence of the system and user satisfaction.

[0065] The goal-oriented intelligent agent task-guided reasoning device provided in the present application is described below. The goal-oriented intelligent agent task-guided reasoning device described below and the goal-oriented intelligent agent task-guided reasoning method described above can be referenced to each other.

[0066] See also Figure 5 , Figure 5 It is a structural diagram of the goal-oriented intelligent agent task-guided reasoning device provided by the present application. The goal-oriented intelligent agent task-guided reasoning device 500 includes: a first acquisition unit 501, used to acquire user input information; a determination unit 502, used to determine the user's starting state and target state according to the user input information; a second acquisition unit 503, used to acquire an action state map; a generation unit 504, used to generate a target path according to the starting state, the target state and the action state map; an execution unit 505, used to execute the actions included in the target path, acquire and output the execution result, and wait for acquiring the user input information again.

[0067] In a possible embodiment, before obtaining the action state graph, the second acquisition unit 503 is also used to: obtain a business document; parse the business document to obtain multiple subtasks; convert each of the multiple subtasks into at least one executable instruction; establish a logical dependency relationship between the at least one executable instruction to obtain an instruction sequence; and generate an action state graph according to the instruction sequence corresponding to each subtask.

[0068] In a possible embodiment, in terms of generating a target path according to the starting state, the target state and the action state graph, the generating unit 504 is specifically used to: obtain a list of disabled actions; traverse all enabled actions in the action state graph according to the list of disabled actions and the target state; recursively generate a sub-plan that satisfies the prerequisites of each enabled action according to the starting state, the sub-plan including the prerequisites, actions and execution results; generate multiple paths according to the sub-plan and the enabled actions; and select a target path from the multiple paths.

[0069] In a possible embodiment, in terms of selecting the target path from the multiple paths, the generation unit 504 is specifically used to: calculate the total execution cost of each path in the multiple paths, the total execution cost is related to the execution cost and action effectiveness of the actions included in each path; determine the path with the lowest execution cost as the target path.

[0070] In a possible embodiment, before traversing all feasible actions in the action state graph according to the target state, the generation unit 504 is further used to: perform a validity check on the target state; if the target state is valid, determine whether the target state has been satisfied in the starting state; if so, determine that the target path is empty.

[0071] In a possible embodiment, before outputting the execution result, the execution unit 505 is also used to: determine the current state according to the execution result; determine whether the current state is consistent with the expected state of the action; if not, add the action to the disabled action list, and regenerate the target path based on the updated disabled action list.

[0072] In a possible embodiment, the goal-oriented intelligent agent task-guided reasoning device 500 also includes an updating unit, which is specifically used to: obtain interaction records; obtain the state to be updated from the action state map according to the interaction records; determine all actions to be updated corresponding to the state to be updated, and all the actions to be updated can reach the expected state from the state to be updated through different paths; simulate the process from the state to be updated to the expected state according to all the actions to be updated to obtain the action simulation validity of each action to be updated; determine the actual action validity of each action to be updated according to the interaction records; obtain the overall validity of each action to be updated according to the action simulation validity and the actual action validity; update the overall validity of each action to be updated, and the higher the overall validity, the greater the probability that the path where the corresponding action to be updated is located will be selected when generating the target path.

[0073] See also Figure 6 , Figure 6 Schematic diagram of the structure of the electronic device provided by this application. Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communication interface 620 and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute a goal-oriented agent task-guided reasoning method, the method comprising: obtaining user input information; determining the user's starting state and target state according to the user input information; obtaining an action state map; generating a target path according to the starting state, the target state and the action state map; executing the actions included in the target path, obtaining and outputting the execution result, and waiting to obtain the user input information again.

[0074] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0075] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the goal-oriented intelligent agent task-guided reasoning method provided by the above-mentioned methods, the method comprising obtaining user input information; determining the starting state and target state of the user based on the user input information; obtaining an action state map; generating a target path based on the starting state, the target state and the action state map; executing the actions included in the target path, obtaining and outputting the execution results, and waiting to obtain the user input information again.

[0076] On the other hand, the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements any one of the above-mentioned goal-oriented intelligent agent task-guided reasoning methods, the method including: obtaining user input information; determining the user's starting state and target state based on the user input information; obtaining an action state map; generating a target path based on the starting state, the target state and the action state map; executing the actions included in the target path, obtaining and outputting the execution results, and waiting to obtain the user input information again.

[0077] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0078] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A goal-oriented agent task-guided reasoning method, characterized in that: include: Get user input information; Determine the user's starting state and target state according to the user input information; Get the action state graph; Generate a target path according to the starting state, the target state and the action state graph; Execute the actions included in the target path, obtain and output the execution result, and wait for obtaining the user input information again.

2. The method according to claim 1, characterized in that: Before obtaining the action state graph, the method further includes: Obtain business documents; Parsing the business document to obtain multiple subtasks; converting each of the plurality of subtasks into at least one executable instruction; Establishing a logical dependency relationship between the at least one executable instruction to obtain an instruction sequence; An action state graph is generated according to the instruction sequence corresponding to each subtask.

3. The method according to claim 1, characterized in that The generating a target path according to the starting state, the target state and the action state graph includes: Get the list of disabled actions; Traversing all enabled actions in the action state graph according to the disabled action list and the target state; Recursively generate a sub-plan that satisfies the preconditions of each feasible action according to the initial state, wherein the sub-plan includes the preconditions, actions and execution results; generating a plurality of paths according to the sub-plans and the feasible actions; A target path is selected from the plurality of paths.

4. The method according to claim 3, characterized in that The selecting a target path from the multiple paths comprises: Calculating a total execution cost of each of the plurality of paths, wherein the total execution cost is related to an execution cost and an action effectiveness of actions included in each of the paths; The path with the lowest execution cost is determined as the target path.

5. The method according to claim 3, characterized in that: Before traversing all possible actions in the action state graph according to the target state, the method further includes: Performing validity check on the target state; If the target state is valid, determining whether the target state has been satisfied in the starting state; If it is satisfied, it is determined that the target path is empty.

6. The method according to claim 3, characterized in that Before outputting the execution result, the method further includes: Determine the current state according to the execution result; determining whether the current state is consistent with the expected state of the action; If they are inconsistent, the action is added to the disabled action list, and the target path is regenerated based on the updated disabled action list.

7. The method according to claim 1, characterized in that The method further comprises: Get interaction records; Acquire a state to be updated from the action state graph according to the interaction record; Determine all actions to be updated corresponding to the state to be updated, and all actions to be updated can reach the desired state from the state to be updated through different paths; Simulating the process from the state to be updated to the expected state according to all the actions to be updated, and obtaining the action simulation validity of each action to be updated; Determining the actual validity of each action to be updated according to the interaction record; Obtaining the overall validity of each action to be updated according to the action simulation validity and the action actual validity; The overall effectiveness of each action to be updated is updated, and the higher the overall effectiveness is, the greater the probability that the path where the corresponding action to be updated is located will be selected when generating the target path.

8. A goal-oriented agent task-guided reasoning device, characterized in that: include: A first acquisition unit, used for acquiring input information; A determination unit, configured to determine a starting state and a target state of the user according to the user input information; A second acquisition unit, used for acquiring an action state graph; A generating unit, configured to generate a target path according to the starting state, the target state and the action state graph; The execution unit is used to execute the actions included in the target path, obtain and output the execution result, and wait for obtaining the user input information again.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the goal-oriented agent task-guided reasoning method as described in any one of claims 1 to 7 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the goal-oriented agent task-guided reasoning method as described in any one of claims 1 to 7 are implemented.