Intelligent agent skill autonomous learning method and system, storage medium and electronic equipment

Through the agent extracting scene information, value evaluation and task planning from the observation data, the agent's difficulty in independently identifying skills is solved, and the agent's ability to independently learn new skills is realized, and its ability to independently learn and adapt to the new environment is improved.

CN120106166AInactive Publication Date: 2025-06-06BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE

Patent Information

Application Number
CN202510591602.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult for an agent to actively identify their shortcomings in skill learning and generate new learning tasks, which limits the agent's independent learning ability and flexibility of the learning mechanism.

Method used

Scene information is extracted through observation data, candidate goals are proposed, value evaluation is performed, targets are selected for task planning, task sequence is executed, skill database supports task execution, confirm learning new skills, train skill models and update skill database.

Benefits of technology

The motivation for the agent to independently generate new skills has been realized, and the agent's independent learning ability has been improved, so that it can actively identify insufficient skills learning and generate new learning tasks, which has enhanced the ability to adapt to the new environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106166A_ABST
    Figure CN120106166A_ABST
Patent Text Reader

Abstract

The invention provides an agent skill autonomous learning method and system, a storage medium and electronic equipment, and the method comprises the steps: proposing a plurality of candidate targets based on scene information; any candidate target is selected for value evaluation; when the value evaluation result accords with the expectation, taking the candidate target of which the value evaluation result accords with the expectation as a selected target, and performing task planning according to the selected target to obtain a new task sequence; executing tasks in the task sequence based on the skill library; when the skill library does not support task execution, new skill learning is confirmed, a skill model is trained for the new skill, and the skill library is updated. Based on the value-driven normal form, the agent is enabled to autonomously generate a motivation for learning new skills, so that the autonomous evolution of the skills is realized, the adaptive capacity of the agent in coping with a new environment is enhanced, the agent can actively recognize the defects of the agent in the aspect of skill learning, a new learning task is generated, and the autonomous learning capacity of the agent is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent agent skill autonomous learning method, system, storage medium and electronic equipment. Background Art

[0002] At present, the learning of embodied agents is a key research issue in the field of embodied intelligence, covering how to enable agents to independently learn and adapt to the tasks they face through interaction with the environment. The current mainstream learning paradigm usually relies on designing specific task training environments for embodied agents and providing specialized skill training support for agents through these environments. For example, in the training tasks of manipulators grasping objects or robots opening cabinet doors, researchers will first set the task goals and help agents improve their skills in these specific tasks through traditional learning methods such as reinforcement learning. These methods are usually based on manually designed training tasks and use certain learning algorithms to achieve the task completion ability of the agent. However, although this method can help agents complete specified tasks, it still has limitations in terms of the autonomy of the agent because its learning process depends on manually set tasks and goals.

[0003] In recent years, researchers have begun to explore how to achieve autonomous learning of intelligent agents and improve their autonomy. Specifically, the current technical solutions explore the learning methods of intelligent agents under autonomous drive, which can generally be divided into two levels. The first level is: after setting the task goals, the robot can autonomously plan tasks and perform these tasks without relying on human intervention. At this level, although the behavior and task execution of the intelligent agent are autonomous, the goal itself is still set externally. The second level is more complex. The intelligent agent can not only perform tasks, but also actively define its own goals based on its own state and external environmental information, and drive behavior generation. This ability to define goals autonomously significantly improves the autonomy of the intelligent agent, because the intelligent agent can not only perceive changes in the external environment, but also make decisions and actions based on internal needs based on its own value system.

[0004] However, although the existing technology has made some progress, there is still an important challenge that has not been solved: how to enable the intelligent agent to actively identify its own deficiencies in skill learning and generate new learning tasks. In other words, current research focuses more on how to enable the intelligent agent to generate behaviors and achieve goals, but ignores the decision-making process of how the intelligent agent identifies which new skills need to be learned in this process. In order to further improve the autonomous learning ability of the intelligent agent, future research needs to make innovations in this direction, especially how to enable the intelligent agent to actively perceive and decide on the skills that need to be further learned through interaction with the environment, so as to achieve a more flexible and general learning mechanism. Summary of the invention

[0005] One of the purposes of the present invention is to provide a method for autonomous learning of intelligent agent skills to solve the above-mentioned technical problems.

[0006] An embodiment of the present invention provides an agent skill autonomous learning method, which controls the agent to perform operations including: S1, extract effective scene information based on observation data; S2, based on the scene information, propose multiple candidate targets; S3. Select any candidate target for value assessment; S4. When the value assessment result meets expectations, execute S5; otherwise, return and re-execute S2; S5. The candidate targets whose value assessment results meet expectations are selected as the targets, and task planning is performed according to the selected targets to obtain a new task sequence; S6. Execute tasks in the task sequence based on the skill library; S7. When the skill library supports task execution, execute S8; otherwise, execute S9; S8. When the task execution status is updated to all tasks completed, return to re-execute S1; otherwise, return to re-execute S6; S9, when it is confirmed that the new skill is to be learned, execute S10; otherwise, return and re-execute S5; S10: Train the skill model for the new skill and update the skill library, and return to re-execute S6 based on the updated skill library.

[0007] Optionally, the step S1, extracting effective scene information according to the observed data, includes: Extracting key object information from the observation data; wherein the key object information includes at least: semantic information, position and posture information, and attributes of the object; A graphical model is used to represent key object information and obtain scene information.

[0008] Optionally, the step S3, selecting any candidate target for value assessment, includes: Based on the pre-trained value calculation model, the selected candidate targets are evaluated for value.

[0009] Optionally, performing task planning according to the selected target in S5 to obtain a new task sequence includes: Take the selected target as the final expected goal to be achieved, decompose and plan the tasks, and obtain a new task sequence that closes the loop from the task logic.

[0010] Optionally, the steps to determine whether the skill library supports task execution are as follows: When one or more skills in the skill library are applied to the scenario of the task, if the scenario state of the task changes from the initial state to the target state, it is judged that the skill library supports the task execution; otherwise, it is judged that it does not support it.

[0011] Optionally, the steps for confirming learning of the new skill are as follows: The missing skills for acquiring the task's scene state from the initial state to the target state; Use missing skills as confirmation of new skills learned.

[0012] Optionally, in S10, training a skill model for a new skill includes: Build task training environments for new skills; Build a large number of scenarios in the mission training environment; Based on reinforcement learning, the target state is expected, and training is performed according to a large number of scenarios in a task training environment to obtain a skill model.

[0013] An embodiment of the present invention provides an agent skill autonomous learning system, comprising: The control module is used to control the agent to perform operations including: S1, extract effective scene information based on observation data; S2, based on the scene information, propose multiple candidate targets; S3. Select any candidate target for value assessment; S4. When the value assessment result meets expectations, execute S5; otherwise, return and re-execute S2; S5. The candidate targets whose value assessment results meet expectations are selected as the targets, and task planning is performed according to the selected targets to obtain a new task sequence; S6. Execute tasks in the task sequence based on the skill library; S7. When the skill library supports task execution, execute S8; otherwise, execute S9; S8. When the task execution status is updated to all tasks completed, return to re-execute S1; otherwise, return to re-execute S6; S9, when it is confirmed that the new skill is to be learned, execute S10; otherwise, return and re-execute S5; S10: Train the skill model for the new skill and update the skill library, and return to re-execute S6 based on the updated skill library.

[0014] Optionally, the step S1, extracting effective scene information according to the observed data, includes: Extracting key object information from the observation data; wherein the key object information includes at least: semantic information, position and posture information, and attributes of the object; A graphical model is used to represent key object information and obtain scene information.

[0015] Optionally, the step S3, selecting any candidate target for value assessment, includes: Based on the pre-trained value calculation model, the selected candidate targets are evaluated for value.

[0016] Optionally, performing task planning according to the selected target in S5 to obtain a new task sequence includes: Take the selected target as the final expected goal to be achieved, decompose and plan the tasks, and obtain a new task sequence that closes the loop from the task logic.

[0017] Optionally, the steps to determine whether the skill library supports task execution are as follows: When one or more skills in the skill library are applied to the scenario of the task, if the scenario state of the task changes from the initial state to the target state, it is judged that the skill library supports the task execution; otherwise, it is judged that it does not support it.

[0018] Optionally, the steps for confirming learning of the new skill are as follows: The missing skills for acquiring the task's scene state from the initial state to the target state; Use missing skills as confirmation of new skills learned.

[0019] Optionally, in S10, training a skill model for a new skill includes: Build task training environments for new skills; Build a large number of scenarios in the mission training environment; Based on reinforcement learning, the target state is expected, and training is performed according to a large number of scenarios in a task training environment to obtain a skill model.

[0020] A computer-readable storage medium provided by an embodiment of the present invention is characterized in that a computer program is stored on the computer-readable storage medium, and a processor executes the computer program to implement any of the methods described above.

[0021] An embodiment of the present invention provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any one of the methods described above.

[0022] This application has achieved the following beneficial effects: Based on the value-driven paradigm, the intelligent agent is motivated to learn new skills, thereby realizing autonomous evolution of skills, enhancing the adaptability of the intelligent agent to new environments, enabling the intelligent agent to actively identify its own deficiencies in skill learning and generate new learning tasks, greatly improving the intelligent agent's autonomous learning ability and realizing a more flexible and universal learning mechanism.

[0023] A value-driven intelligent agent is realized in an embodied interactive virtual environment. The intelligent agent can autonomously generate task goals and task plans driven by its own values. During task execution, it can make subjective judgments on whether to learn new skills based on the task execution status, and can use the programming interface of software tools such as Unreal Engine to set up a new skill training environment until new skills can be trained and added to the original skill library to solve problems that could not be solved before.

[0024] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.

[0025] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 Schematic diagram of an agent's autonomous skill learning method in an embodiment of the present invention; Figure 2 Schematic diagram of a specific implementation module of the method for autonomous learning of skills of an intelligent agent in an embodiment of the present invention; Figure 3 It is a schematic diagram of a specific implementation model of the method for autonomous learning of intelligent agent skills in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0028] The embodiment of the present invention provides a method for autonomous learning of intelligent agent skills, such as Figure 1 As shown, controlling the agent to perform operations includes: S1, extract effective scene information based on observation data; S2, based on the scene information, propose multiple candidate targets; S3. Select any candidate target for value assessment; S4. When the value assessment result meets expectations, execute S5; otherwise, return and re-execute S2; S5. The candidate targets whose value assessment results meet expectations are selected as the targets, and task planning is performed according to the selected targets to obtain a new task sequence; S6. Execute tasks in the task sequence based on the skill library; S7. When the skill library supports task execution, execute S8; otherwise, execute S9; S8. When the task execution status is updated to all tasks completed, return to re-execute S1; otherwise, return to re-execute S6; S9, when it is confirmed that the new skill is to be learned, execute S10; otherwise, return and re-execute S5; S10: Train the skill model for the new skill and update the skill library, and return to re-execute S6 based on the updated skill library.

[0029] The working principle and beneficial effects of the above technical solution are: like Figure 2 As shown, in the specific implementation, the following multiple modules are used to perform operations: Scene information acquisition module: extracts scene structure information from the input scene information.

[0030] Task target sampling module based on scene information: performs target sampling based on scene structure information to generate several candidate task targets.

[0031] Task planning module: Plan tasks according to the selected objectives and generate task sequences to achieve task objectives.

[0032] Task execution module based on skill library: Based on the skill information in the skill library, the tasks in the task sequence are executed one by one. If they can be executed, the action is output. If the existing skills cannot meet the task, the conclusion of missing skills is output.

[0033] Task status update module: applies action execution to the environment, updates the task status, and obtains the result of whether the task is completed.

[0034] Value decision module: Calculate the value of candidate targets sampled from tasks. At the same time, after receiving reports on skill loss, it can also make decisions on re-planning or activating skill learning.

[0035] Skill training module: For confirmed learned skills, the corresponding task training environment is constructed, and a reasonable model structure and training method are called for automatic training to obtain new model parameters that can solve the corresponding tasks, that is, new skills are acquired.

[0036] Skill library management module: manages skill information and can handle the addition of new skills.

[0037] like Figure 3 As shown, the specific implementation process is as follows: First, start the virtual environment and start the agent in the virtual environment to make the agent enter the normal operation state. At this time, the agent can observe the virtual environment. According to the observed information, extract the key object information in the environment (such as the semantic information, position and posture information, various attributes, etc. of the object), and use the graph model (such as various probabilistic graph models) to represent it. Take the scene information represented by the graph model as input data and enter the task target sampling module. The task target sampling module will calculate according to the scene information and output multiple candidate targets, each candidate target is represented as a different scene state. In other words, the task target sampled here is actually the scene target state caused by the task execution. The candidate targets generated in the previous step are input into the value decision module for value evaluation. If it meets expectations (the value evaluation score exceeds a certain set threshold), the target that meets expectations is selected as the selected target and input into the task planning module; if it does not meet expectations, re-sample the task targets and re-evaluate the value until the target that meets expectations is sampled. The task planning module analyzes and plans the selected targets and outputs the task sequence. The agent further performs each task in the task sequence based on the skill library. If the current task can find the corresponding skill in the skill library and it is determined that the current task can be completed in sequence, the action required to complete the task is output according to the skill information; if the current task cannot be completed by a single skill or a combination of skills in the skill library, the missing skills (that is, the required skills corresponding to the current task that cannot be completed) are reported to the value decision module. The value decision module determines whether to learn new skills. The scene is updated according to the action output by the task execution module, and the task status is updated at the same time. If the task status is that all tasks have been completed, return to the scene information extraction module and continue to observe new scene information; if the task status is that the task has not been completed, a signal is returned to the execution module to notify it to execute the next task in the task sequence.

[0038] This application is based on the value-driven paradigm, which enables the intelligent agent to autonomously generate motivation to learn new skills, thereby realizing autonomous evolution of skills, enhancing the adaptability of the intelligent agent to cope with new environments, and enabling the intelligent agent to actively identify its own deficiencies in skill learning and generate new learning tasks, which greatly improves the autonomous learning ability of the intelligent agent and realizes a more flexible and universal learning mechanism. An intelligent agent driven by value in an embodied interactive virtual environment is realized, which can autonomously generate task goals and generate task plans under the drive of its own value. During task execution, a subjective judgment can be made on whether new skills should be learned based on the task execution status, and the programming interface of software tools such as Unreal Engine can be used to set up a training environment for new skills until new skills can be trained and added to the original skill library to solve problems that could not be solved before.

[0039] In one embodiment, the step S1, extracting effective scene information according to the observed data, includes: Extracting key object information from the observation data; wherein the key object information includes at least: semantic information, position and posture information, and attributes of the object; A graphical model is used to represent key object information and obtain scene information.

[0040] This step implements the input of scene information and the output of task objectives. The scene information is represented by a graph model. Each object in the scene is represented by a "node" in the graph model, and the relationship between objects is represented by an "edge". In this way, all objects in the entire scene and the relationship between objects can be represented. Based on this representation method, data collection can be carried out, that is, a large number of simulation environments are constructed, and the graph model representation of the scene state and the corresponding value are annotated to form a training data set, and then a neural network model can be trained to achieve the ability to score any scene state. The scene state can use the Markov Chain Monte Carlo (MCMC) method to sample new states, and each new state can be scored using the trained neural network model, thereby generating multiple task objectives and scores. Each task objective corresponds to a scene state, so scoring the scene state is equivalent to scoring the task objective, so that different task objectives (i.e., the scene normality corresponding to the objective) can be quantitatively evaluated and compared, preparing for further value decision-making based on the objective score.

[0041] In one embodiment, the step S3, selecting any candidate target for value assessment, includes: Based on the pre-trained value calculation model, the selected candidate targets are evaluated for value.

[0042] The task goal (i.e., the target state of the scenario) and the score generated in the previous step are input into the value decision module. The value decision module can model multiple value dimensions and take the weighted sum of the value results of different value dimensions as the total value output. Each value dimension has a numerical result, and then the results of each dimension are weighted and summed. The weights can use empirical values ​​predefined by technicians, such as average weighting or specially adjusting the weights of certain value dimensions. The value decision module takes different task target states and the corresponding original scores as module input, enters the pre-trained value calculation model for value calculation, and outputs different value calculation results for different goals.

[0043] The implementation steps of the value calculation model are as follows: First, a number of text descriptions are implemented by manual writing or AI model generation methods. The text content contains status description information of various scenarios. Then, for each text description of the scene status, its corresponding value label is marked (for example, it contains several different dimensions and the weight information of each dimension, and ensures that the sum of the weights of multiple dimensions is 1) to form a training data set. Finally, a neural network model based on basic structures such as Transformer can be selected and trained on the produced data set to obtain a value calculation model.

[0044] In addition, some pre-trained large models can be directly used as value calculation models, such as OpenAI GPT4, DeepSeek R1, Tongyi Qianwen large model, etc. These pre-trained large models themselves have certain value judgment strategies built in and can be directly used as value calculation models.

[0045] Finally, if the highest value exceeds the preset value threshold, the target corresponding to the highest value is selected as the selected target; otherwise, this batch of targets is abandoned.

[0046] In one embodiment, the step S5 performs task planning according to the selected target to obtain a new task sequence, including: Take the selected target as the final expected goal to be achieved, decompose and plan the tasks, and obtain a new task sequence that closes the loop from the task logic.

[0047] After the selected goal is determined, the target state is used as the final expected goal to be achieved, and the task is decomposed and planned to close the loop from the task logic. For example, if the initial scene state is "the key is on the ground", and the target state is "the key is in the drawer", then a possible task planning result is: Step 1, the agent picks up the key on the ground; Step 2, the agent takes the key to the table with a drawer; Step 3, the agent opens the drawer; Step 4, the agent puts the key in the drawer; Step 5, the agent closes the drawer. The above example reflects a task planning result that can be logically closed, and forms a task sequence consisting of 5 steps.

[0048] In one embodiment, the steps for determining whether the skill library supports task execution are as follows: When one or more skills in the skill library are applied to the scenario of the task, if the scenario state of the task changes from the initial state to the target state, it is judged that the skill library supports the task execution; otherwise, it is judged that it does not support it.

[0049] According to the planned task sequence, each task in the sequence is analyzed and executed one by one. Each task is the change of the scene state from the initial state to the target state. By applying one or more skills to the scene, the scene state may change. If the skills in the skill library can reach the target state after combination, it means that the task can be completed based on the existing skills, then the corresponding actions of the skills involved should be performed in the scene; if the skills in the skill library cannot reach the above target state after combination, it means that the existing skills are not enough to solve the current task. The conclusion of missing skills is notified to the value decision module, and the value decision module determines whether to learn new skills or re-plan a new task sequence. For the value decision module, if the number of re-planning has not exceeded a certain preset threshold, it is considered that there is no need to learn new skills for the time being, and the task is directly planned again; otherwise, new skills need to be learned.

[0050] In one embodiment, the steps of confirming and learning the new skill are as follows: The missing skills for acquiring the task's scene state from the initial state to the target state; Use missing skills as confirmation of new skills learned.

[0051] If the value decision module confirms to learn new skills, it means that it confirms to learn how to transform the scene from a certain initial state to the target state. For example, if the "agent opens the drawer" task cannot be solved by relying on the existing skill library, it means that there is no skill that can change the drawer from the "closed" state to the "open" state, and the missing skill is the "open drawer" skill.

[0052] In one embodiment, in S10, training a skill model for a new skill includes: Build task training environments for new skills; Build a large number of scenarios in the mission training environment; Based on reinforcement learning, the target state is expected, and training is performed according to a large number of scenarios in a task training environment to obtain a skill model.

[0053] In order to build a task training environment for missing skills, it is necessary to construct a large number of various scenarios (for example, various drawers in closed states, and the scenarios support the pushing and pulling of drawers), and provide an initial control model (such as a policy network) for the agent's arm control, take the expected target state as the agent's operation target (for example, take the open state of the drawer as the target state), and train the control model based on reinforcement learning methods. Finally, the new control model trained is expected to solve the target task to a certain extent, and it can be added to the skill library as a newly learned skill.

[0054] After learning new skills, the skills in the skill library can support tasks that were previously impossible to perform, and produce a series of corresponding actions to be applied to the scene.

[0055] An embodiment of the present invention provides an agent skill autonomous learning system, comprising: The control module is used to control the agent to perform operations including: S1, extract effective scene information based on observation data; S2, based on the scene information, propose multiple candidate targets; S3. Select any candidate target for value assessment; S4. When the value assessment result meets expectations, execute S5; otherwise, return and re-execute S2; S5. The candidate targets whose value assessment results meet expectations are selected as the targets, and task planning is performed according to the selected targets to obtain a new task sequence; S6. Execute tasks in the task sequence based on the skill library; S7. When the skill library supports task execution, execute S8; otherwise, execute S9; S8. When the task execution status is updated to all tasks completed, return to re-execute S1; otherwise, return to re-execute S6; S9, when it is confirmed that the new skill is to be learned, execute S10; otherwise, return and re-execute S5; S10: Train the skill model for the new skill and update the skill library, and return to re-execute S6 based on the updated skill library.

[0056] S1, extracting effective scene information according to the observation data, comprises: Extracting key object information from the observation data; wherein the key object information includes at least: semantic information, position and posture information, and attributes of the object; A graphical model is used to represent key object information and obtain scene information.

[0057] S3, selecting any candidate target for value assessment, includes: Based on the pre-trained value calculation model, the selected candidate targets are evaluated for value.

[0058] In S5, task planning is performed according to the selected target to obtain a new task sequence, including: Take the selected target as the final expected goal to be achieved, decompose and plan the tasks, and obtain a new task sequence that closes the loop from the task logic.

[0059] The steps to determine whether the skill library supports task execution are as follows: When one or more skills in the skill library are applied to the scenario of the task, if the scenario state of the task changes from the initial state to the target state, it is judged that the skill library supports the task execution; otherwise, it is judged that it does not support it.

[0060] The steps for confirming learning of the new skill are as follows: The missing skills for acquiring the task's scene state from the initial state to the target state; Use missing skills as confirmation of new skills learned.

[0061] In S10, training a skill model for a new skill includes: Build task training environments for new skills; Build a large number of scenarios in the mission training environment; Based on reinforcement learning, the target state is expected, and training is performed according to a large number of scenarios in a task training environment to obtain a skill model.

[0062] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. A processor executes the computer program to implement any of the above methods.

[0063] An embodiment of the present invention provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any one of the methods described above.

[0064] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A method for autonomous learning of intelligent agent skills, characterized in that: Controlling the agent to perform operations includes: S1, extract effective scene information based on observation data; S2, based on the scene information, propose multiple candidate targets; S3. Select any candidate target for value assessment; S4. When the value assessment result meets expectations, execute S5; otherwise, return and re-execute S2; S5. The candidate targets whose value assessment results meet expectations are selected as the targets, and task planning is performed according to the selected targets to obtain a new task sequence; S6. Execute tasks in the task sequence based on the skill library; S7. When the skill library supports task execution, execute S8; otherwise, execute S9; S8. When the task execution status is updated to all tasks completed, return to re-execute S1; otherwise, return to re-execute S6; S9, when it is confirmed that the new skill is to be learned, execute S10; otherwise, return and re-execute S5; S10: Train the skill model for the new skill and update the skill library, and return to re-execute S6 based on the updated skill library.

2. The method for autonomous learning of intelligent agent skills as claimed in claim 1, characterized in that: S1, extracting effective scene information according to the observation data, comprises: Extracting key object information from the observation data; wherein the key object information includes at least: semantic information, position and posture information, and attributes of the object; A graphical model is used to represent key object information and obtain scene information.

3. The method for autonomous learning of intelligent agent skills as claimed in claim 1, characterized in that: S3, selecting any candidate target for value assessment, includes: Based on the pre-trained value calculation model, the selected candidate targets are evaluated for value.

4. The method for autonomous learning of intelligent agent skills as claimed in claim 1, characterized in that: In S5, task planning is performed according to the selected target to obtain a new task sequence, including: Take the selected target as the final expected goal to be achieved, decompose and plan the tasks, and obtain a new task sequence that closes the loop from the task logic.

5. The method for autonomous learning of intelligent agent skills as claimed in claim 1, characterized in that: The steps to determine whether the skill library supports task execution are as follows: When one or more skills in the skill library are applied to the scenario of the task, if the scenario state of the task changes from the initial state to the target state, it is judged that the skill library supports the task execution; otherwise, it is judged that it does not support it.

6. The method for autonomous learning of intelligent agent skills as claimed in claim 1, characterized in that: The steps for confirming learning of the new skill are as follows: The missing skills for acquiring the task's scene state from the initial state to the target state; Use missing skills as confirmation of new skills learned.

7. The method for autonomous learning of intelligent agent skills as claimed in claim 1, characterized in that: In S10, training a skill model for a new skill includes: Build task training environments for new skills; Build a large number of scenarios in the mission training environment; Based on reinforcement learning, the target state is expected, and training is performed according to a large number of scenarios in a task training environment to obtain a skill model.

8. An agent skill autonomous learning system, characterized in that: include: The control module is used to control the agent to perform operations including: S1, extract effective scene information based on observation data; S2, based on the scene information, propose multiple candidate targets; S3. Select any candidate target for value assessment; S4. When the value assessment result meets expectations, execute S5; otherwise, return and re-execute S2; S5. The candidate targets whose value assessment results meet expectations are selected as the targets, and task planning is performed according to the selected targets to obtain a new task sequence; S6. Execute tasks in the task sequence based on the skill library; S7. When the skill library supports task execution, execute S8; otherwise, execute S9; S8. When the task execution status is updated to all tasks completed, return to re-execute S1; otherwise, return to re-execute S6; S9, when it is confirmed that the new skill is to be learned, execute S10; otherwise, return and re-execute S5; S10: Train the skill model for the new skill and update the skill library, and return to re-execute S6 based on the updated skill library.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the processor executes the computer program to implement the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Robot autonomous learning method, device and equipment and storage medium

    CN114529010A

  • Planning method based on feedback and value driving, computer equipment and storage medium

    CN119377387A

  • Closed-loop online self-learning framework applied to autonomous vehicle

    US20240086776A1

Cited By

  • Space intelligent reasoning method and system

    CN120471182A

  • Intelligent agent skill generation method and electronic equipment

    CN121562660A