Task planning reasoning device, method and equipment of intelligent agent, medium and product

Through the collaborative work of the agent task planning and reasoning device, the agent can independently learn, make decisions and execute diversified tasks in a changing and complex environment, solving the problem of insufficient general intelligence capabilities of the existing agent system and achieving a higher level of intelligence.

CN120146202APending Publication Date: 2025-06-13BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510624075.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing agent systems lack general intelligence capabilities and are unable to independently learn, make decisions and perform diversified tasks in a changing and complex environment, and lack the ability to independently learn, long-term memory and complex reasoning.

Method used

It provides a task planning and reasoning device for an agent, including an environment variable module, a belief update module, a path planning module and a task execution module. Through the coordinated work of these modules, the agent can independently drive based on value and capabilities, realize task continuation, and generate future paths through simulated skill library information.

Benefits of technology

It significantly improves the general intelligence capabilities of the agent, allowing it to make independent decisions based on value and ability, realize task continuation and generation of simulation skill library information, and achieve a higher level of intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146202A_ABST
    Figure CN120146202A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a task planning reasoning device of an agent, which can be applied to the technical field of artificial intelligence. The task planning reasoning device of the intelligent agent comprises an environment variable module, a belief updating module, a path planning module and a task execution module. The environment variable module is used for receiving environment variable information; the belief updating module is used for performing belief reasoning according to the environment variable information provided by the environment variable module to generate intention updating information; the path planning module is used for updating information according to the intention of the belief updating module and selecting an optimal task path by calling skill library information, and calling of action library information can be achieved through the skill library information; and the task execution module is used for executing the optimal task path of the path planning module, generating execution feedback information and sending the execution feedback information to the environment variable module as updating of the environment variable information. The embodiment of the invention further provides an agent task planning reasoning method and device, a storage medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a task planning and reasoning method, device, equipment, medium and product for an intelligent agent. Background Art

[0002] Artificial Intelligence (AI for short) is an important driving force for the new round of scientific and technological revolution and industrial transformation, and is a new key technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. As an important part of intelligent science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine (i.e., an intelligent agent) that can react in a way similar to human intelligence.

[0003] With the continuous development of artificial intelligence technology, although existing intelligent agent systems perform well in specific tasks, most systems are still limited to single functions and lack general intelligence like humans, that is, the ability to autonomously learn, make decisions and execute diverse tasks in a changing and complex environment. Although current technologies have made remarkable progress in specific fields (such as speech recognition, image recognition, etc.), they have not truly realized a general intelligent agent (AGI) with autonomous cognition and broad adaptability. For example, intelligent agents cannot make autonomous decisions based on intrinsic values and capabilities, still rely on rules or simple feedback to react, and are also unable to use the first-person perspective to achieve effective environmental perception and understanding, and follow up and react in a timely manner according to environmental changes. At the same time, they also lack autonomous learning ability, long-term memory ability and complex reasoning ability, etc., which greatly limits the general intelligent effect of existing intelligent agents. Summary of the Invention

[0004] In view of at least one of the above problems, embodiments of the present invention aim to provide a task planning and reasoning device, method, equipment, medium and product for an intelligent agent that can significantly improve general intelligent capabilities, thereby providing an intelligent agent that can be autonomously driven based on values and capabilities, effectively realizing task continuation and being able to simulate skill library information to generate future paths, and at the same time being able to support the implementation of simulation unification, so as to reach a higher level of intelligence.

[0005] An aspect of an embodiment of the present invention provides a task planning and reasoning device for an agent, which includes an environmental variable module, a belief update module, a path planning module, and a task execution module. The environmental variable module is used to receive environmental variable information; the belief update module is used to perform belief reasoning based on the environmental variable information provided by the environmental variable module to generate intention update information; the path planning module is used to select an optimal task path according to the intention update information of the belief update module by calling the skill library information, and the action library information can be called through the skill library information; the task execution module is used to execute the optimal task path of the path planning module, generate execution feedback information and send it to the environmental variable module as an update of the environmental variable information.

[0006] According to an embodiment of the present invention, the belief update module includes an environmental information unit, a voice data unit, and a mental language unit. The environmental information unit is used to receive the environmental observation information of the environmental variable information provided by the environmental variable module and generate belief understanding information; the voice data unit is used to receive the environmental voice information of the environmental variable information provided by the environmental variable module and generate voice extraction information; the mental language unit is used to receive the language information of the environmental variable information provided by the environmental variable module, the voice extraction information generated by the voice data unit, and the belief understanding information generated by the environmental information unit to generate language understanding information.

[0007] According to an embodiment of the present invention, the belief update module further includes a cancellation verification unit and a mental reasoning unit. The cancellation verification unit is used to receive the belief understanding information of the environmental information unit and generate an execution cancellation instruction for the current action of the agent; the mental reasoning unit is used to receive the belief understanding information of the environmental information unit and generate non-self cognitive information.

[0008] According to an embodiment of the present invention, the belief update module further includes a knowledge update unit. The knowledge update unit is used to receive the non-self cognitive information of the mental reasoning unit and the language understanding information of the mental language unit to generate updated knowledge information.

[0009] According to an embodiment of the present invention, the belief update module further includes a self-state unit, an action planning unit, and a task reflection unit. The self-state unit is used to receive the action knowledge information of the execution cancellation instruction of the cancellation verification unit and the updated knowledge information of the knowledge update unit to generate self-state update information; the action planning unit is used to receive the self-state update information of the self-state unit and generate action planning information; the task reflection unit is used to receive the self-state update information of the self-state unit and generate intention reflection information.

[0010] According to an embodiment of the present invention, the belief update module further includes an intention update unit. The intention update unit is configured to receive the action planning information of the action planning unit, the intention reflection information of the task reflection unit, and the updated knowledge information of the knowledge update unit, and generate intention update information.

[0011] According to an embodiment of the present invention, the path planning module includes a path planning unit and a reflection planning unit. The path planning unit is configured to, according to the intention update information, call the skill strategy information of the skill library information to generate at least one feasible task path, and select a feasible task path as the optimal task path based on the value score of each feasible task path in the at least one feasible task path; the reflection planning unit is configured to store the information of the optimal task path of the path planning unit.

[0012] Another aspect of the embodiments of the present invention provides a task planning and reasoning method for an intelligent agent, which is characterized by including: performing belief reasoning according to the received environmental variable information to generate intention update information; according to the intention update information, selecting an optimal task path by calling the skill library information, wherein the action library information can be called through the skill library information; and executing the optimal task path to generate execution feedback information, and the execution feedback information is used as an update of the environmental variable information.

[0013] Another aspect of the embodiments of the present invention provides an electronic device, including one or more processors and a memory, and the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned task planning and reasoning method of the intelligent agent.

[0014] Another aspect of the embodiments of the present invention provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned task planning and reasoning method of the intelligent agent.

[0015] Another aspect of the embodiments of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-mentioned task planning and reasoning method of the intelligent agent is implemented.

[0016] The task planning and reasoning device of the intelligent agent provided by the embodiments of the present invention can at least partially solve the problem of low general intelligence level in the task planning of intelligent agents in the related art, and thus can at least achieve one of the following technical effects: (1) UV drive: It can realize an intelligent agent autonomously driven based on value and ability. In particular, the skill library U will have various states, enabling its knowledge to be effectively utilized, and assuming the role of continuous planning; (2) Task continuation: Due to the discontinuity of the skill library, a task can be suspended and can be continued to be executed; (3) Simulation: The agent can generate future paths through the simulation skill library U to assist in implementing the execution decision of the agent; (4) Execution simulation unification: The knowledge of the same skill library U can support both execution and simulation simultaneously.

[0017] Therefore, the above-mentioned task planning and reasoning device of an agent according to an embodiment of the present invention can provide an integrated multi-module collaboration, achieve autonomous planning by integrating multi-module functions, and be able to explain or answer its own behaviors and other functions reflected by general agents based on feedback learning and replanning.

[0018] It should be understood that the above general description and the following specific embodiments are only exemplary and explanatory, and cannot limit the scope claimed by the present invention. Description of the Drawings

[0019] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings: Figure 1 A structural block diagram of a task planning and reasoning device of an agent according to an embodiment of the present invention is schematically shown; Figure 2 A structural block diagram of a specific application scenario of a task planning and reasoning device of an agent according to an embodiment of the present invention is schematically shown; Figure 3 An application scenario diagram of a task planning and reasoning method, device, equipment, medium, and program product of an agent according to an embodiment of the present invention is schematically shown; Figure 4 A flowchart of a task planning and reasoning method of an agent according to an embodiment of the present invention is schematically shown; and Figure 5 A block diagram of an electronic device suitable for implementing the task planning and reasoning method of an agent according to an embodiment of the present invention is schematically shown.

[0020] The above-mentioned drawings are part of the specification of the embodiments of the present invention, which illustrate the exemplary embodiments of the present invention. The accompanying drawings and the description of the specification are used together to explain the principles of the embodiments of the present invention. It should be understood that the above general description of the drawings and the following specific embodiments are only exemplary and explanatory, and cannot limit the scope claimed by the present invention. Detailed Embodiments

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clearly understood, the following will clearly explain the spirit of the content disclosed in the present invention with reference to the accompanying drawings and detailed descriptions. After any person skilled in the relevant technical field understands the embodiments of the content of the present invention, the techniques taught by the content of the present invention can be changed and modified without departing from the spirit and scope of the content of the present invention.

[0022] The schematic embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention. Additionally, elements / components using the same or similar reference numerals in the accompanying drawings and embodiments are used to represent the same or similar parts.

[0023] Regarding the use of "first", "second",... etc. in the present invention, it does not particularly refer to the order or sequence, nor is it used to limit the present invention. It is only used to distinguish elements or operations described with the same technical terms.

[0024] Regarding the directional terms used in the present invention, such as: up, down, left, right, front or back, etc., they are only references to the directions in the accompanying drawings. Therefore, the directional terms used are for explanation and not for limiting the creation.

[0025] Regarding the use of "comprising", "including", "having", "containing", etc. in the present invention, they are all open-ended terms, meaning including but not limited to.

[0026] Regarding the use of "and / or" in the present invention, it includes any one or all combinations of the described things.

[0027] Regarding "multiple" in the present invention, it includes "two" and "more than two"; regarding "multiple groups" in the present invention, it includes "two groups" and "more than two groups".

[0028] Regarding the terms "substantially", "about", etc. used in the present invention, they are used to modify any quantity or error that can vary slightly, but these slight variations or errors do not change their essence. Generally, the range of such slight variations or errors modified by such terms can be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments, or other values in some embodiments. Those skilled in the art should understand that the aforementioned values can be adjusted according to actual needs and are not limited thereto.

[0029] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used here should be interpreted as having meanings consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0030] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, or C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). Those skilled in the art should also understand that substantially any disjunctive conjunctions and / or phrases indicating two or more alternative items, whether in the specification, claims, or drawings, should be understood as giving the possibility of including one of these items, either side of these items, or both items. For example, the phrase "A or B" should be understood as including the possibility of "A" or "B", or "A and B".

[0031] In view of at least one of the above problems, embodiments of the present invention aim at a task planning and reasoning device, method, device, medium, and product of an intelligent agent capable of significantly improving general intelligent capabilities, thereby providing an intelligent agent capable of being autonomously driven based on value and ability, effectively realizing task continuation and being able to simulate skill library information to generate future paths, and at the same time being able to support the implementation of simulation unification, so as to achieve a higher level of intelligence.

[0032] The following will be based on Figures 1 to 2 A detailed description will be given of the task planning and reasoning device of the intelligent agent of the disclosed embodiments.

[0033] Figure 1 A structural block diagram of a task planning and reasoning device of an intelligent agent according to an embodiment of the present invention is schematically shown.

[0034] As Figure 1 shown, the task planning and reasoning device 100 of the intelligent agent of this embodiment includes an environmental variable module 110, a belief update module 120, a path planning module 130, and a task execution module 140.

[0035] The environmental variable module 110 is used to receive environmental variable information.

[0036] The belief update module 120 is used to perform belief reasoning based on the environmental variable information provided by the environmental variable module to generate intention update information.

[0037] The path planning module 130 is used to update the information of the intention according to the belief update module, and select the optimal task path by calling the information in the skill library, where the information in the action library can be called through the information in the skill library.

[0038] The task execution module 140 is used to execute the optimal task path of the path planning module, generate execution feedback information and send it to the environment variable module for updating the environment variable information.

[0039] The agent in the embodiment of the present invention can be a humanoid intelligent robot or other AI devices, and generally has an intelligent control-feedback-learning system composed of an actuator and the above-mentioned task planning and reasoning device in the embodiment of the present invention, and can complete specific action task planning and control execution. Among them, the task planning and reasoning device in the embodiment of the present invention can be implemented by a planning and reasoning engine driven by a UV (skill + value) architecture, and can realize the basic functions of the underlying skill library information, value evaluation information and action library information, and realize the cognitive loop of belief reasoning (belief update), planning (plan) and execution (execute), specifically as Figure 2 shown.

[0040] The environment variable module 110 (environment variable, abbreviated as ENV) can be used as the message interface module of the task planning and reasoning device to forward the received messages or the generated messages. Among them, the environment variable information can include the environmental interaction information generated during the interaction between the agent (such as through various sensors, actuators, etc.) and the outside world, such as the environmental visual information (information about people, objects, and the relationships between corresponding people, between people, and between objects, etc.) detected by a visual sensor (such as a camera, etc.) for the surrounding environment, and the environmental touch information generated by a dexterous hand touching an object. In addition, the environment variable information can also include the task planning information and environmental cognitive feedback information generated by the task planning and reasoning device itself, such as the planning information for a specific execution action, the reasoning cognitive feedback information, etc. In short, the environment variable information of the environment variable module can include the environmental interaction information between the agent and the external environment (environment) and the agent's own cognitive reasoning information, learning feedback information, etc.

[0041] The belief update module 120 is used to receive the environmental variable information from the environmental variable module, and perform cognitive completion (PG completion) of the environmental variable information based on the cognitive reasoning process for the environmental variable information. Finally, intention update information can be generated through the cognitive completion process. Among them, the intention update information can be cognitive reasoning information used to reflect the action intention of the agent itself, and can include inferred information about unobserved environmental information, inferred information about the future state of things, and mental state information (mental state) for updating the agent based on these inferred information. Therefore, the intention update information contains the agent's current cognition and focus on the environment (belief state), and can also reflect its execution intention for the target execution object.

[0042] The path planning module 130 (Plan) can usually implement task planning based on the UV architecture using the intention update information of the above-mentioned belief update module 120. Among them, in this task planning process, the path planning module 130 can call the skill library information (U) according to the inferred information of the intention update information to perform task path planning related to task planning. The task path can specifically be an action execution sequence, and the specific action execution tasks can be sorted in terms of execution in atomic-level actions to form this sequence.

[0043] The environmental variable information can include the agent's cognitive information and focus information at the current moment. In addition, it can also include the call information of the skill library information and the call information related to the value evaluation of task planning. Among them, the skill library information (U) describes the agent's skill library, such as the strategy of turning a Rubik's Cube and the possible results caused by different strategies, etc., regarding the execution strategies of different action paths and the possible results of different execution strategies (policy + world model). In this process, the action library information (Action) can be called by different skill libraries to form a part of the execution strategy of this skill library information. The action library information can include information related to the actions executed by the agent. In addition, the value library information (i.e., Value, abbreviated as V) can be used to reflect the value of the task to be executed by the agent, and the task to be executed can be evaluated for value according to this value library information.

[0044] Among them, the value library information (V) may include state -> scalar function mapping information. The skill library information (U) may include state -> state function mapping information, and at the same time support being interrupted at any time when an action is issued, and being able to continue receiving information at the interruption position to achieve continuous execution from the interruption position. Different skill library information can be called from each other, and can also achieve the call of different action library information (Action). Among them, the action library information (Action) can also include state->state mapping, corresponding to the environmental action, including the effect or influence of making an action on the agent state (state), and can have a certain degree of randomness.

[0045] Therefore, the path planning module 130 can select the skill library (including the call of action library information) and the relevant parameters of the skill library information, and the final output decision is the optimal task path corresponding to the current cognition and focus of the agent. The optimal task path can be the optimal action execution path that most conforms to the current cognitive state and action intention of the agent. Specifically, according to the current state and action intention of the agent, the path planning module 130 can use simulation to call different skill library information, and at the same time plan different task paths, and then evaluate these task paths respectively (such as V value scoring), so as to obtain the optimal task path according to the pros and cons or high and low of the evaluation results, as the final decision for the agent's subsequent task execution.

[0046] The specific evaluation process can be based on the simulation run of each task path and the possible state effects, perform value scoring based on the path and state, and select the task path with a significant increase or maximum value score as the planned path, that is, the optimal task path of the final decision.

[0047] The task execution module (Execute) 140 is mainly responsible for the execution of the final decision of the agent. At the same time, the task execution module will also send the final decision and the content generated during its execution as part of the new environmental variable information to the environmental variable module 110 for updating relevant information (beliefs), so that the updated environmental variable information can reflect the learning, reflection, decision-making task adjustment, etc. of the task execution module 140. At the same time, the task execution module 140 will output the action information (action) of the optimal task path to the agent itself, and the agent itself calls the corresponding actuator to implement the execution of the corresponding task action, that is, it can be understood as sending the corresponding planned action to the environment (Environment). During the execution of the task execution module 140, the agent can internally update its own state and relevant action models, etc.

[0048] In summary, in the process of continuous interaction with the external environment, the task planning and reasoning device of the intelligent agent provided by the embodiments of the present invention can generate the task planning path for the next moment in each round of iteration through the task loop iterative processing of the above-mentioned environment variable module 110, belief update module 120, path planning module 130, and task execution module 140, thus providing an intelligent agent system that can be based on real-time cognitive understanding and execution in a virtual environment. This system can be deployed in an interactive virtual home AI device. That is to say, the intelligent agent with this task planning and reasoning device has "embodiment", can "live" in the environment from the first-person perspective, can see, hear, speak, move, and do things like a human, and at the same time has the following key capabilities: (1) Real-time scene perception and understanding ability. For example, the intelligent agent can observe the environment from the first perspective and understand room items, states, tasks, etc.

[0049] (2) Decision-making ability based on value and simulation. For example, the intelligent agent can autonomously decide what to do and what to do more preferably according to internal goals or value functions through an internal world simulation mechanism.

[0050] (3) Ability to use multiple capabilities in combination. For example, the intelligent agent can achieve language dialogue, planning actions, executing actions, learning new concepts, explaining its own behavior, etc.

[0051] (4) Task execution ability. For example, the intelligent agent can complete daily tasks in the environment, such as cleaning up garbage and organizing items.

[0052] (5) Ability to execute language and actions in parallel. For example, the intelligent agent can complete tasks while speaking, rather than "speaking first and then doing".

[0053] (6) Adjustment ability based on feedback. For example, the intelligent agent can modify its behavior or plan according to user or system feedback.

[0054] (7) Learning ability based on communication and feedback. For example, the intelligent agent can continuously accumulate experience, update knowledge, and adapt to new tasks.

[0055] (8) Theory of Mind ability. For example, the intelligent agent can speculate on the thoughts, intentions, and feelings of others and adjust its dialogue or behavior accordingly, such as understanding whether the user is confused, whether they hope to obtain an explanation, or whether they have misunderstood the task.

[0056] Figure 2 The structural block diagram of a specific application scenario of the task planning and reasoning device of the intelligent agent according to the embodiments of the present invention is schematically shown.

[0057] Such as Figure 1 and Figure 2As shown, according to an embodiment of the present invention, the belief update module 120 includes an environmental information unit 201, a sound data unit 202, and a mentalese unit 203.

[0058] The environmental information unit 201 is configured to receive the environmental observation information of the environmental variable information provided by the environmental variable module 110 and generate belief understanding information. The sound data unit 202 is configured to receive the environmental sound information of the environmental variable information provided by the environmental variable module 110 and generate sound extraction information. The mentalese unit 203 is configured to receive the language information of the environmental variable information provided by the environmental variable module 110, the sound extraction information generated by the sound data unit 202, and the belief understanding information generated by the environmental information unit 201, and generate language understanding information.

[0059] The environmental information unit 201 (pg update) can serve as the environmental data interface of the belief update module 120. Among them, the environmental observation information (observation) can be the environmental detection information generated by the agent during the interaction with the outside world, such as the environmental visual information (cvpg) detected by the visual sensor to detect the surrounding environment, such as the environmental visual information of people, objects, and the relationships between corresponding people, between people and between objects. The environmental information unit 201 can obtain the corresponding belief understanding information (belief pg) by preliminarily processing the environmental observation information. The belief understanding information can be understood as the cognitive understanding information generated according to the environmental observation information (such as converting the environmental visual information cvpg into cognitive visual information, etc.). Specifically, it can be the cognitive data generated by extracting information from the environmental observation information (such as extracting the information of people, objects, and the relationships between things in the visual image, etc.), which is reflected as the understanding of the environmental observation information, such as gor a box update, and also includes recognized actions, physical events, etc.

[0060] The sound data unit 202 (sound data update) can serve as the sound data interface of the belief update module 120. Among them, the environmental sound information can be the sound data (sound) contained in the environmental variable information, such as specific voiceprint data like speaking voices, footsteps, motor rotation sounds, etc. By performing extraction operations on these sound data, corresponding sound extraction information can be obtained. These sound extraction information can be text information corresponding to these sound data, such as language information, etc. The mind language unit 203 (mind nlp update) can serve as the language data interface of the belief update module 120. Among them, the language information can be the natural language processing information (nature languageprocessing message, abbreviated as nlp message) contained in the environmental variable information, such as language text information of human speech, text information formed by image recognition and extraction, etc., which are language-related information that can be processed by a language model. In addition, the sound extraction information generated by the above-mentioned sound data unit 202 and the belief understanding information generated by the environmental information unit 201 can also serve as the input information of the mind language unit 203, enabling the mind language unit 203 to perform language understanding processing based on these language information, sound extraction information, and belief understanding information. For example, perform necessary grounding, augmentation, and retrieve language understanding operations to achieve language understanding of these information, thereby generating language understanding information. Among them, the language understanding information can be language-related cognitive reasoning information generated by performing language understanding on the input information of the above-mentioned mind language unit 203.

[0061] Therefore, in one round of iteration process, through the above-mentioned environmental information unit 201, sound data unit 202, and mind language unit 203, the ability to perceive and understand real-time scenarios can be achieved, and efficient and accurate information fusion and cognition in multiple aspects such as language dialogue, new knowledge learning, surrounding environment understanding, and self-behavior cognition can be realized.

[0062] It should be noted that the environmental information unit 201, the sound data unit 202, and the thinking language unit 203, as different data interfaces of the belief update module 120, can also be used for data access of the action execution information (ack information) in the environmental variable information. The action execution information can be the execution status information of the current agent's action or the action executed at the previous moment, such as execution completed, execution ended, execution aborted, etc. At the same time, it can also include the execution expectation information corresponding to these action execution statuses, such as the execution expectation of "no dirt on the tabletop" for the "wipe the table" action has been achieved, etc. Among them, the action execution information can be included in the corresponding environmental observation information, environmental sound information, and language information, and can be repeatedly called by subsequent module units to be used for intention judgment of the agent's action execution.

[0063] As Figure 1 and Figure 2 shown, according to an embodiment of the present invention, the belief update module 120 further includes a cancellation check unit 204 and a thinking reasoning unit 205.

[0064] The cancellation check unit 204 is used to receive the belief understanding information of the environmental information unit 201 and generate an execution cancellation instruction for the current action of the agent. The thinking reasoning unit 205 is used to receive the belief understanding information of the environmental information unit and generate non-self-cognitive information.

[0065] The cancellation check unit (cancel check) 204 can check the actions currently being executed and the actions to be executed by the current agent according to the action execution expectation information related to the action execution information (ack information) in the belief understanding information and the action information of the current agent in the state of being executed and expected to be executed. When these actions have met the expectations (for example, the outside can help achieve or has achieved the corresponding action expectations), it can provide quick decision-making information on whether to cancel the current action or the expected action, and these quick decision-making information can be sent in the form of an execution cancellation instruction. The execution cancellation instruction can be a computer-executable instruction for quickly canceling the current action or the expected action of the agent.

[0066] For example, when the user notifies the agent "Go and close the window in the living room, and then turn on the TV", the agent will check the current execution action (such as in the middle of handing a water cup to the user) and the expected execution actions (closing the window and turning on the TV) according to the execution content of the cancellation check unit 204. However, at this time, the window may have been closed by a gust of wind, or the user may not know that the window has been closed. Then, the agent needs to make an autonomous judgment and cognition based on the environmental detection information (observation), and then cancel the "window closing" action by generating an execution cancellation instruction. After completing the current action of "handing the water", directly operate the "turn on the TV" action.

[0067] The thinking and reasoning unit (mind pg update & mind action update) 205 can perform cognitive reasoning based on the content related to environmental detection information in the belief understanding information, and perform cognitive reasoning on the belief states and actions of non-self, such as the cognition of other people, other objects, and other thing relationships. For example, whether the agent can see something, hear a sound, or even recognize the value of other people (such as distinguishing customers and tourists in an exhibition hall). Thus, corresponding non-self cognitive information can be generated accordingly. During the generation process of this non-self cognitive information, the current actions of other people also need to be called and cognitively reasoned to understand the action purposes of other people and be able to predict the action expectations of these actions. For example, when seeing someone pointing at a display board with a finger, it can be directly recognized that this person may need a specific object shown on the display board. Therefore, the non-self cognitive information can be cognitive reasoning information about other people, objects, and thing relationships other than the agent itself, and can reflect the agent's cognitive level and understanding ability of the surrounding world.

[0068] Therefore, by means of the cancellation check unit 204 and the thinking and reasoning unit 205, the theory of mind ability of the agent can be greatly improved. For example, the agent can speculate on the thoughts, intentions, and feelings of others, and even adjust the conversation or behavior accordingly, and even be able to understand whether the user is confused or has misunderstood the task.

[0069] As Figure 1 and Figure 2 shown, according to an embodiment of the present invention, the belief update module 120 further includes a knowledge update unit 206.

[0070] The knowledge update unit 206 is used to receive the non-self cognitive information of the thinking and reasoning unit 205 and the language understanding information of the thinking language unit 203, and generate updated knowledge information.

[0071] The knowledge update unit (action precondition update & knowledge update) 206 can be used to receive the non-self-cognitive information of the thinking reasoning unit 205 and the language understanding information of the thinking language unit 203 to update its own knowledge reserve information. Specifically, it can update the relevant conceptual knowledge through the language understanding information. For example, it can learn and update the action precondition through the language understanding information. Therefore, the updated knowledge information can be the updated content of the knowledge update unit 206 relative to the existing knowledge. Specifically, the knowledge can be updated according to the processing of the environmental variable information received by the belief update module 120 in each iteration process (for example, after the processing of the above environmental information unit 201, sound data unit 202, thinking language unit 203, cancellation verification unit 204, and thinking reasoning unit 205).

[0072] Therefore, through the knowledge update unit 206, the intelligent agent can have the learning ability based on communication and feedback, such as continuously accumulating experience and updating knowledge, so that in the subsequent task planning process, it can use the existing knowledge to judge and understand the planning tasks, and thus the intelligent agent can have the adjustment ability based on feedback.

[0073] As Figure 1 and Figure 2 shown, according to an embodiment of the present invention, the belief update module 120 further includes a self-state unit 207, an action planning unit 208, and a task reflection unit 209.

[0074] The self-state unit 207 is used to receive the execution cancellation instruction of the cancellation verification unit 204 and the action knowledge information of the updated knowledge information of the knowledge update unit 206 to generate self-state update information; The action planning unit 208 is used to receive the self-state update information of the self-state unit 207 to generate action planning information; The task reflection unit 209 is used to receive the self-state update information of the self-state unit 207 to generate intention reflection information.

[0075] The self - status unit 207 (continue u update) can update the status of the action being executed by the agent itself during the current round of iteration based on the action knowledge information related to the execution experience of the agent that executes the cancellation instruction and updates the knowledge information, in combination with the action execution expectation. For example, if the execution result of the currently executed action meets the expectation of this action, the execution status of this action can be updated, that is, this action execution can continue. If the execution result of the currently executed action does not meet the action expectation, corresponding processing such as pausing or aborting the action can be performed, such as time - out processing. Specifically, the execution expectation of the action of "throwing garbage" is "the garbage is in the trash can", but it is found that "there is no trash can in the room", then the execution expectation of the "throwing garbage" action cannot be achieved. At this time, the execution status of this action can be updated to reflect the information that the action does not meet the expectation and cannot be executed. Therefore, the self - status update information can be the reflection information of the agent's own action execution, which can reflect the ability to speculate and recognize action execution in combination with the expectation of the action.

[0076] The action planning unit 208 (plan in action update) can plan the action expected to be executed during this round of iteration based on the self - status update information generated by the above - mentioned self - status unit 207. That is, it can plan subsequent actions according to the action completion information (such as ack information) of the current agent. Among them, the action planning information is the relevant information of the expected action execution generated by the action planning unit 208 through calling the corresponding action library information (Action), etc., in combination with the self - status update information of the agent itself for subsequent action planning, such as the action name, the sequence of action execution, etc.

[0077] The task reflection unit 209 (u proposal update) can reflect on the feedback information such as the task in which the agent is in the execution state or has completed the execution in combination with the self - status update information of the above - mentioned self - status unit 207, so as to generate intention reflection information. For example, it generates corresponding intention information for the reflection that the agent's own action cannot achieve the corresponding action execution expectation. For example, the user notifies the agent to "step on the stool and wipe the dust on the top of the bookcase". During the execution of the "step on the stool" action, the agent finds that "there is no stool in the surrounding environment". At this time, it causes the agent to be unable to achieve the action expectation of "the dust on the top of the bookcase is wiped clean". At this time, the agent can output the intention reflection information of "because there is no stool in the surrounding environment, the dust on the top of the bookcase cannot be wiped clean" through the task reflection unit 209. Therefore, this intention reflection information can be the reflection information generated by the agent for the current or expected action that cannot achieve the action expectation (i.e., the action intention intent), which can reflect the reason for being unable to achieve this action expectation.

[0078] Therefore, the combination of the above-mentioned self-state unit 207, action planning unit 208, and task reflection unit 209 can further reflect the theory of mind ability of the agent, and can provide the planning of the next action or the reflection information that the action cannot be carried out according to whether the action intention is achieved, and can realize the integrated use of multiple abilities, such as action planning, behavior explanation, etc.

[0079] As Figure 1 and Figure 2 shown, according to an embodiment of the present invention, the belief update module 120 further includes an intention update unit 210.

[0080] The intention update unit 210 is configured to receive the action planning information of the action planning unit 208, the intention reflection information of the task reflection unit 209, and the updated knowledge information of the knowledge update unit 206, and generate intention update information.

[0081] In the embodiment of the present invention, during each round of iteration, according to the above-mentioned action planning information, intention reflection information, and updated knowledge information, the intention update unit 210 (intent update) can regenerate multiple action execution intentions (intent) of the agent at the current moment, and evaluate these action execution intentions respectively. According to the evaluation scores of these action execution intentions, different technology library information can be selected, and the final intention update information can be expressed as different categories of labels (label). Therefore, the intention update information can be the operation intention information of the next task of the agent itself finally generated according to the cognition and understanding of the environmental variable information.

[0082] It can be seen that based on the intention update information, the agent can make decisions based on value and simulation. Specifically, the agent can determine what to do, what not to do, what to do first, or what to do last, etc. according to internal goals or values through an internal simulation mechanism. In addition, it can also be adjusted based on feedback, for example, modifying the behavior or plan according to user or system feedback.

[0083] As Figure 1 and Figure 2 shown, according to an embodiment of the present invention, the path planning module 130 includes a path planning unit 301 and a reflection plan unit 302.

[0084] The path planning unit 301 is configured to call the skill strategy information of the skill library information to generate at least one feasible task path according to the intention update information, and select a feasible task path as the optimal task path based on the value score of each feasible task path in the at least one feasible task path; The reflection plan unit 302 is configured to store the information of the optimal task path of the path planning unit 301.

[0085] The path planning unit 301 (random planner) can call the skill library information (U) based on the intention update information, and autonomously plan the feasible task path based on the skill policy information (policy) of the skill library information. Among them, the feasible task path is a set of possible required task actions when the task intention of the agent for the intention update information is large. These task actions can be combined in the execution order of the intentions achieved by the action execution to form a task action path. The skill policy information can be the policy rule information for calling relevant task actions by adapting different skills U according to the task intention.

[0086] Based on the above-mentioned feasible task path, the evaluation of the expected task path execution can be carried out according to the task intention for different feasible task paths. For example, the closer the execution expectation of a feasible task path is to the task intention, the higher the evaluation score; on the contrary, the lower the evaluation score. Therefore, these feasible task paths can be sorted based on the evaluation score, and the feasible task path with the highest evaluation score can be selected as the optimal task path. It should be noted that in the embodiments of the present invention, the action execution expectation and the task execution expectation can generally be predicted by the corresponding module unit through calling the corresponding skill library information to perform action execution or task execution simulation.

[0087] The feedback planning unit 302 (reflect plan) can store the above-generated optimal task path. In fact, the feedback planning unit 302 can also store the intention update information of the foregoing belief update module 120, which will not be elaborated here.

[0088] Therefore, based on the above path planning unit 301 and the reflection planning unit 302, the agent in the embodiments of the present invention can further achieve the decision-making ability based on value and simulation, autonomously decide what to do and what to do first through the inner world simulation mechanism. Moreover, it can complete the routine tasks in the environment, has the task execution ability, and can also maintain the parallel execution ability of language and action or multiple actions. At the same time, it can plan or modify behaviors based on feedback, reflecting the adjustment ability based on feedback.

[0089] In summary, combined with the task planning and reasoning device of the agent in the embodiments of the present invention, the following operation task planning of (1)-(8) can be realized: (1) Task generation based on intention In each iteration process, the belief update module 120 of the task planning and reasoning device of the agent can receive instructions or questions (such as observation and nlp message, etc.) from the external environment as environmental variable information, and the corresponding intent update unit 210 (intent update) can generate intent update information that obeys the instructions or answers the questions based on the scene and situation understanding. Further, when the path planning module 130 is planning, this intent update information can be processed in the speaking skill library to generate a feasible task path for the speaking content or action. If the feasible task path corresponding to a certain intent is evaluated and has a relatively high satisfaction score, this path will be selected as the optimal task path and the first action of this path will be executed... and so on in a loop until its own intent is satisfied or the intent is cancelled.

[0090] (2)Simulation-based task generation In each iteration process, based on its current state, the agent applies different skill libraries during planning to generate different paths and simulation states, and selects the optimal generated path as the final decision according to the agent value, and sends the first action of this path to the environment.

[0091] (3)Action expectation and visual information fusion, replanning (cancel check) In each iteration process, the agent continuously perceives new scenes and its own information while performing an action in the environment. The agent checks whether the goal of this action is achieved. If it is achieved, the action will no longer be performed (for example, during the process of closing the door, another agent closes the door), and the continuation state of the skill corresponding to this action is updated and used for subsequent planning. If the action is completed, then: 1) If the object related to the action goal is not within the field of view, it will be considered that the goal has been achieved, and the cognitive state continue u update will be updated; 2) If the action does not achieve the action goal, it will be retried, and the model of whether this action can be executed will be updated. Subsequently, if this action cannot be successfully executed, this action will no longer be planned, and other implementation paths or tasks will be selected; 3) If the action achieves the action goal, the continuation state of the corresponding skill instance will be updated, and the normal planning process will be followed.

[0092] (4)Performing actions and speaking in parallel The agent will perform planning in each round. If taking something was planned in the previous round, the action of taking something will be executed. If speaking is planned during the process of taking something, speaking will be performed simultaneously.

[0093] (5)Fast system implementation method Basically, each round of planning generates an execution plan. Then, the first action of the execution plan is taken and executed in the environment. When this action is completed, the previous execution plan is preferentially simulated based on the current scenario. If the simulation has no problems and can increase sufficient value, the next action of the plan will be immediately executed. This is the main process of fast thinking. Conversely, if the previous plan has unsatisfactory results or cannot be simulated after simulation, re-planning will be carried out, similar to the slow system.

[0094] (6)Implementation method of task suspension and switching During the execution of a task, if another task associated with a skill and its related path have a higher score, then the other task will be executed, and this task is equivalent to being suspended. However, as long as in a certain round, this task has a higher score after planning, then this task will be executed, equivalent to resuming this task. A typical example is doing laundry. After pressing the button, the laundry skill cannot perform subsequent planning because it has not waited for the washing machine to finish washing the clothes (the action of waiting for the laundry to complete is not finished, and no path can be planned). At this time, once other tasks can increase the score, other tasks will be done. If the clothes are washed, this event will trigger the completion of the action of waiting for the laundry to complete, and the laundry skill can then continue to plan subsequent actions, equivalent to the task being resumed.

[0095] (7)Reflection and re-planning In addition to the re-planning caused by the update of the action execution model due to the inability to execute a certain action described above, this architecture also supports re-planning based on environmental feedback. For example, when executing the action of picking up something, the environmental feedback that the thing cannot be reached is passed to the reflect module to construct the information that it cannot be reached, and suggested skills are provided for planning to help complete reaching the thing. At this time, it is very likely that operations such as moving a stool will be planned.

[0096] (8)Implementation method of communication learning Communication learning is that the user gives language feedback to help the intelligent agent obtain knowledge or skill updates. Currently, it is learning triggered by analyzing the user's intention, mainly including the addition of new concepts, the modification of existing skills, and the modification of the content in the prompt of the large model to achieve the purpose of learning.

[0097] In summary, the task planning and reasoning device of the intelligent agent in the embodiments of the present invention can achieve the goal of verifying and exploring the core mechanisms for realizing general artificial intelligence (AGI) by constructing a general intelligent agent prototype system. Specifically, this prototype system is not only for demonstrating the feasibility of current technologies, but more importantly, for laying a foundation for the future development of AGI. Combining the above-mentioned belief update module 120 and path planning module 130, the task planning and reasoning device can at least achieve the following corresponding levels of AGI capabilities in the following key technical fields: Value-driven and ability-driven decision-making mechanism: The intelligent agent can make autonomous decisions based on internal values and abilities, rather than relying solely on rules or simple feedback. This is a key ability for autonomously choosing behavioral paths.

[0098] Perception system: The system can understand the environment in real time, effectively perceive using the first-person perspective, identify environmental changes, and react timely.

[0099] Learning mechanism: Through autonomous learning, the intelligent agent can accumulate experience, expand the knowledge base, and continuously optimize decision-making strategies in a changing environment.

[0100] Memory system: The intelligent agent has long-term memory ability, can save and utilize historical information to improve the execution efficiency and accuracy of current tasks.

[0101] Reasoning and decision-making mechanism: Through complex reasoning, the intelligent agent can make reasonable decisions in unknown and uncertain environments, ensuring that the execution of tasks does not rely on preset rules.

[0102] Therefore, the task planning and reasoning device of the intelligent agent in the embodiments of the present invention can provide an integrated multi-module collaboration, achieve autonomous planning by integrating multi-module functions, and can, based on feedback learning and replanning, explain or answer its own behaviors and other functions demonstrated by general intelligent agents, and for the first time provides a complete prototype system aiming to realize general artificial intelligence.

[0103] According to an embodiment of the present invention, any multiple of the environmental variable module 110, belief update module 120, path planning module 130, and task execution module 140 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the environmental variable module 110, belief update module 120, path planning module 130, and task execution module 140 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), programmable logic array (PLA), system on chip, system on substrate, system on package, application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as hardware or firmware through circuit integration or packaging, or can be implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the environmental variable module 110, belief update module 120, path planning module 130, and task execution module 140 can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.

[0104] Based on the above task planning and reasoning device of the intelligent agent, the present invention also provides a task planning and reasoning method for the intelligent agent. The following will be combined with Figure 3 and Figure 4 to describe this method in detail.

[0105] An aspect of an embodiment of the present invention provides a task planning and reasoning method for an intelligent agent, which includes: performing belief reasoning based on the received environmental variable information to generate intention update information; according to the intention update information, selecting an optimal task path by calling the skill library information, where the action library information can be called through the skill library information; and executing the optimal task path to generate execution feedback information, and the execution feedback information is used as an update of the environmental variable information.

[0106] Figure 3 Schematically shows an application scenario diagram of the task planning and reasoning method, device, equipment, medium, and program product of the intelligent agent according to an embodiment of the present invention.

[0107] As Figure 3 shown, the application scenario 300 according to this embodiment may include terminal devices 301, 302, 303, network 304, and server 305. The network 304 is used to provide a medium for communication links between the terminal devices 301, 302, 303 and the server 305. The network 304 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0108] Users can use terminal devices 301, 302, 303 to interact with server 305 via network 304 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 301, 302, 303, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0109] Terminal devices 301, 302, 303 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, and so on.

[0110] Server 305 can be a server that provides various services, such as a background management server that supports the websites browsed by users using terminal devices 301, 302, 303 (for example only). The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0111] It should be noted that the task planning and reasoning method of the agent provided by the embodiments of the present invention can generally be executed by server 305. Correspondingly, the task planning and reasoning device of the agent provided by the embodiments of the present invention can generally be set in server 305. The task planning and reasoning method of the agent provided by the embodiments of the present invention can also be executed by a server or a server cluster different from server 305 and capable of communicating with terminal devices 301, 302, 303 and / or server 305. Correspondingly, the task planning and reasoning device of the agent provided by the embodiments of the present invention can also be set in a server or a server cluster different from server 305 and capable of communicating with terminal devices 301, 302, 303 and / or server 305.

[0112] It should be understood that Figure 3 the numbers of terminal devices, networks, and servers in

[0113] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 3 Based on the scenario described below Figure 1 , Figure 2 and Figure 4 the task planning and reasoning method of the agent in the disclosed embodiments will be described in detail.

[0114] As Figure 4 shown, one aspect of the embodiments of the present invention provides a task planning and reasoning method for an agent, which includes operations S401 to S403.

[0115] In operation S401, belief reasoning is performed based on the received environmental variable information to generate intention update information; in one embodiment, operation S401 can be implemented by the belief update module 120 described above, which will not be elaborated here.

[0116] In operation S402, according to the intention update information, the optimal task path is selected by calling the skill library information, and the action library information can be called through the skill library information; in one embodiment, operation S402 can be implemented by the path planning module 130 described above, which will not be elaborated here. And In operation S403, the optimal task path is executed to generate execution feedback information, and the execution feedback information is used as an update of the environmental variable information. In one embodiment, operation S403 can be implemented by the task execution module 140 described above, which will not be elaborated here.

[0117] It should be noted that the technical effects achievable by the task planning and reasoning method of the intelligent agent in the embodiments of the present invention are exactly realized based on the aforementioned task planning and reasoning device of the intelligent agent, which will not be elaborated here.

[0118] Figure 5 A block diagram of an electronic device suitable for implementing the task planning and reasoning method of an intelligent agent according to an embodiment of the present invention is schematically shown.

[0119] The above-mentioned electronic device provided by the embodiment of the present invention includes one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned task planning and reasoning method of the intelligent agent.

[0120] As Figure 5 shown, the electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. The processor 501 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 501 can also include on-board memory for caching purposes. The processor 501 can include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present invention.

[0121] In the RAM 503, various programs and data required for the operation of the electronic device 500 are stored. The processor 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. The processor 501 performs various operations of the method flow according to the embodiments of the present invention by executing programs in the ROM 502 and / or the RAM 503. It should be noted that the programs may also be stored in one or more memories other than the ROM 502 and the RAM 503. The processor 501 may also perform various operations of the method flow according to the embodiments of the present invention by executing programs stored in the one or more memories.

[0122] According to an embodiment of the present invention, the electronic device 500 may further include an input / output (I / O) interface 505, and the input / output (I / O) interface 505 is also connected to the bus 504. The electronic device 500 may further include one or more of the following components connected to the I / O interface 505: an input portion 506 including a keyboard, a mouse, etc.; an output portion 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 508 including a hard disk, etc.; and a communication portion 509 including a network interface card such as a LAN card, a modem, etc. The communication portion 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read from it can be installed into the storage portion 508 as needed.

[0123] The present invention also provides a computer-readable storage medium having executable instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the task planning and reasoning method of the above-mentioned agent.

[0124] Among them, the computer-readable storage medium may be included in the device / apparatus / system described in the above embodiments; or it may exist alone and not be assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0125] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or apparatus. For example, according to an embodiment of the present invention, the computer-readable storage medium may include one or more memories other than the above-described ROM 502 and / or RAM 503 and / or ROM 502 and RAM 503.

[0126] An embodiment of the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the above-mentioned task planning and reasoning method of the intelligent agent.

[0127] Among them, the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, this program code is used to enable the computer system to implement the method provided by the embodiment of the present invention.

[0128] When the computer program is executed by the processor 501, it executes the above-mentioned functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0129] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 509, and / or be installed from the removable medium 511. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0130] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or be installed from the removable medium 511. When the computer program is executed by the processor 501, it executes the above-mentioned functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0131] In accordance with embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0133] In addition, all actions of obtaining information, signals, or data in the present invention are carried out on the premise of complying with the corresponding data protection laws, regulations, and policies of the country where it is located, and with the authorization given by the owner of the corresponding device.

[0134] Those skilled in the art can understand that the features described in the various embodiments and / or claims of the present invention can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments and / or claims of the present invention can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0135] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present invention.

Claims

1. A task planning reasoning device for an intelligent agent, characterized in that: include: Environment variable module, used to receive environment variable information; A belief updating module, used for performing belief reasoning based on the environment variable information provided by the environment variable module to generate intention update information; A path planning module, used to update the information according to the intention of the belief updating module, and select the optimal task path by calling the skill library information, wherein the calling of the action library information can be realized through the skill library information; as well as The task execution module is used to execute the optimal task path of the path planning module, generate execution feedback information and send it to the environment variable module as an update of the environment variable information.

2. The device according to claim 1, characterized in that The belief updating module includes: An environmental information unit, configured to receive environmental observation information of the environmental variable information provided by the environmental variable module, and generate belief understanding information; A sound data unit, used for receiving the environmental sound information of the environmental variable information provided by the environmental variable module, and generating sound extraction information; The thinking language unit is used to receive the language information of the environmental variable information provided by the environmental variable module, the sound extraction information generated by the sound data unit and the belief understanding information generated by the environmental information unit, and generate language understanding information.

3. The device according to claim 2, characterized in that The belief updating module also includes: A cancellation checking unit, configured to receive the belief understanding information of the environment information unit and generate an execution cancellation instruction for the current action of the agent; The thinking and reasoning unit is used to receive the belief understanding information of the environmental information unit and generate non-self-cognitive information.

4. The device according to claim 2, characterized in that The belief updating module also includes: The knowledge updating unit is used to receive the non-self-cognitive information of the thinking and reasoning unit and the language comprehension information of the thinking and language unit, and generate updated knowledge information.

5. The device according to claim 4, characterized in that The belief updating module also includes: A self-state unit, used for receiving the execution cancellation instruction of the cancellation checking unit and the action knowledge information of the updated knowledge information of the knowledge updating unit, and generating self-state update information; An action planning unit, configured to receive the self-state update information of the self-state unit and generate action planning information; The task reflection unit is used to receive the self-state update information of the self-state unit and generate intention reflection information.

6. The device according to claim 5, characterized in that The belief updating module also includes: The intention updating unit is used to receive the action planning information of the action planning unit, the intention reflection information of the task reflection unit and the updated knowledge information of the knowledge updating unit, and generate intention updating information.

7. The device according to claim 1, characterized in that The path planning module includes: a path planning unit, configured to generate at least one feasible task path according to the intention update information, call the skill strategy information of the skill library information, and select a feasible task path as the optimal task path based on the value score of each feasible task path in the at least one feasible task path; The reflection planning unit is used to store information about the optimal task path of the path planning unit.

8. A task planning reasoning method for an intelligent agent, characterized in that: include: Perform belief reasoning based on the received environmental variable information to generate intention update information; According to the intention update information, the optimal task path is selected by calling the skill library information, wherein the action library information can be called through the skill library information; as well as The optimal task path is executed to generate execution feedback information, where the execution feedback information is used as an update of the environment variable information.

9. An electronic device, comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are caused to perform the method of claim 8.

10. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the method of claim 8.

11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to claim 8 is implemented.

Citation Information

Patent Citations

  • Planning method based on feedback and value driving, computer equipment and storage medium

    CN119377387A

  • Control method and device for intelligent agent with body and readable storage medium

    CN119416881A

  • Information generation method and device, electronic equipment and storage medium

    CN119670863A

  • Environment space relation reasoning method based on multi-agent debate, medium and equipment

    CN119692470A

  • Apparatus for Recognizing Human Behavior Patterns using Affordance and Belief-Desire-Intension based Agent Model

    KR1020150003561A