A task-driven universal robot intelligent control method and system
By building a task-driven general-purpose robot intelligent control system, the problem of insufficient autonomous working ability of robots in unstructured scenarios is solved, efficient and low-cost skill generation and reuse in complex tasks are achieved, and the robot's task completion ability in generalized scenarios is improved.
Patent Information
- Application Number
- CN202510912917.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-03
AI Technical Summary
Existing technologies lack the ability for robots to work autonomously in unstructured general scenarios, traditional models cannot be effectively expanded, data requirements are large and prone to "data explosion", making it difficult to maintain high performance in complex tasks.
Build a task-driven general-purpose robot intelligent control system, including a skill library, task parser, skill generator, and skill arbitrator. Decompose tasks through the robot perception-language-planning model, dynamically combine skill elements, generate and arbitrate new skills, and achieve cross-scenario reuse.
It reduces data requirements, improves task success rates, maintains high performance, and enables robots to complete complex tasks in generalized scenarios with better precision and versatility.
Smart Images

Figure CN120395912B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of robot control technology and computer control, and specifically relates to a task-driven universal robot intelligent control method and system, aiming to improve the robot's working ability in generalized scenarios. Background Art
[0002] The application of robots in general scenarios is showing a diversified development trend, covering multiple fields such as industry, home services, medical care, logistics, education, etc., and is still in the transition stage from structured scenarios to unstructured scenarios. Large-scale application has been achieved in some fields, while the robot's ability to work autonomously in generalized scenarios is still in the early stages of exploration.
[0003] Numerous achievements have been made in applying robots to unstructured, general-purpose scenarios. These include approaches based on classical control theory, training methods using end-to-end data collection, and embodied manipulation models. These methods offer diverse solutions for robots operating in complex scenarios. However, traditional models are only applicable to situations where the robot is operating in a single environment. Training methods using end-to-end data collection and embodied manipulation models are data-intensive and cannot scale to new tasks. Furthermore, as task complexity increases, the amount of data must be increased to maintain satisfactory performance, leading to the risk of "data explosion." Summary of the Invention
[0004] In response to the shortcomings of the existing technology, the present invention provides a task-driven universal robot intelligent control method and system, which enables the robot to autonomously adapt to environmental changes based on environmental information, complete skill combination and generalization, and perform complex tasks in generalized scenarios.
[0005] The method and system of the present invention are of great significance for robot control in generalized scenarios. By constructing a skill library using skill elements, the present invention eliminates the need for robots to train themselves to perform tasks with large datasets. Instead, they only need to train on a small amount of data to form skill elements. This helps reduce data requirements, improve task success rates, and significantly reduce data costs while maintaining high performance. Furthermore, the inclusion of a skill generator and skill arbitrator gives these skill elements greater precision and versatility, allowing them to be reused across different scenarios and tasks, enabling robots to complete complex tasks in generalized scenarios.
[0006] To achieve the above purpose, the specific technical solutions adopted by the present invention are as follows:
[0007] In one aspect, the present invention provides a task-driven universal robot intelligent control method, comprising the following steps:
[0008] Step 1: Build an initial skill library for the robot. The skill elements included in the initial skill library are core capabilities of the robot that are highly universal and frequently used, while avoiding excessive complexity.
[0009] Step 2: The robot decomposes the complex task into subtasks through the task decoder. The robot perception-language-planning model is used as the task parser.
[0010] Step 3: When a subtask is obtained, the robot combines the initial skill library to complete the subtask. The robot realizes modular execution of complex tasks by dynamically combining and parameterizing predefined skill elements according to task requirements.
[0011] Step 4: When there is a missing skill element in the skill library and the subtask cannot be completed, a new skill element is generated through the skill generator.
[0012] Step 5: Complete the execution of the subtask by combining the generated new skill element, and at the same time determine whether the skill element should be added to the skill library through the skill arbitrator.
[0013] Step 6: If the newly generated skill element passes the arbitration of the skill arbitrator, it will be added to the skill library. If the newly generated skill element fails to pass the arbitration of the skill arbitrator, it will be forgotten.
[0014] Furthermore, the skill elements in the robot's initial skill library described in step 1 are: Based on the learned skill actions, the actions are divided into discrete actions that can represent the completion of a task. Similar discrete actions are then grouped and defined as skill elements. (For example, in the grasping skill, the robot can grasp an apple, a doorknob, or other objects; these discrete actions are combined into a single skill element labeled "grasp object").
[0015] The skill meta-element in the initial skill library is constructed in two ways: First, when constructing a skill meta-element that only involves single-object interaction (for example, grabbing an apple), skill meta-element learning is completed through data-driven simulation end-to-end training.
[0016] The second type: When it is necessary to build a skill element involving multi-object interaction (such as washing dishes), the skill element is learned through two stages: human prior and autonomous environment interaction.
[0017] First, the human prior stage is carried out: the expert teaching trajectory database is collected through motion capture teleoperation and VR teleoperation, and the preliminary strategy network is trained (the preliminary strategy network is Define, where is the entire network training space, is the observation state, It's action. is the initial state distribution, are unknown and possibly random transition probabilities that depend on the system dynamics, Encoding tasks for reward functions simultaneously), establishing basic action behavior patterns.
[0018] Secondly, it enters the autonomous environment interaction stage: the robot optimizes its strategy in environmental interaction, improves its performance through reward functions, continuously collects data generated in the actual environment, and adds it to the training data, so that it can adapt to the state distribution in the actual environment, optimize basic action behavior patterns, and finally complete the learning of skill elements.
[0019] Furthermore, the complex task described in step 2 includes environmental information perception and natural language understanding. The robot perception-language-planning model comprises a multimodal perception layer, a language understanding layer, and a planning and decision-making layer. The multimodal perception layer is used to perform the environmental information perception task within the complex task, while the language understanding layer is used to perform the natural language understanding task within the complex task. The outputs of these two layers are generated through a multimodal Transformer architecture to generate cognitive features that are fed into the planning and decision-making layer, ultimately outputting the subtask.
[0020] Multimodal Perception Layer: This layer is responsible for acquiring the environmental perception information required for complex tasks. It uses corresponding encoders for different environmental information inputs (preferably, convolutional neural networks (CNNs) for image inputs, recurrent neural networks (RNNs) for audio inputs, and Kalman filters for digital information inputs such as LiDAR). All encoder outputs are mapped into a unified multimodal vector space.
[0021] Language Understanding Layer: This layer is responsible for acquiring the natural language understanding information required for complex tasks. It uses a large language model (preferably GPT-4) to convert the natural language in complex tasks into structured target parsing instructions.
[0022] The unified encoding data in the multimodal vector space output by the multimodal perception layer and the structured target parsing instructions output by the language understanding layer are used to generate cognitive features through the multimodal Transformers architecture.
[0023] Planning and decision-making layer: Cognitive features and noise actions are passed as input tokens to the Transformer module. The step information output by the Transformer module is added to the cognitive features through sinusoidal position encoding. The cognitive features with added step information are combined with the noise actions and input into the Transformer module again. Action sequence planning is gradually completed, and the related action sequences are connected into subtasks.
[0024] Furthermore, the robot skill generator described in step 4 can generate skill meta-elements in two ways: combinatorial generation and generalized generation. Combinatorial generation involves building on existing skill meta-elements, optimizing the skill call sequence using the PPO and SAC algorithms. Using existing skill meta-elements as nodes, the state transition diagram is used to search for optimal paths and generate new skill meta-elements through combination (for example, completing the skill meta "entering a room" by following the steps "open door → walk → close door"). Generalized generation involves two implementation paths for learning new skill meta-elements: First, a complementary combination of an autoregressive model and a diffusion model is used to generate new skill meta-elements. The autoregressive model sequentially predicts the discrete actions required for the new task, while the diffusion model gradually denoises the discrete actions generated by the autoregressive model to generate an action sequence, forming the new skill meta-elements. Second, reinforcement learning is used to design reward terms based on energy efficiency and contact force stability. A physics-guided reward function is used to guide the generation of physically intuitive action sequences as new skill meta-elements.
[0025] Furthermore, the robot skill arbitrator described in step 5 evaluates the acquired new skill element through basic performance analysis based on actual task completion; adopts a comprehensive scoring model and sets a score threshold to determine whether the new skill element is added to the initial skill library.
[0026] Furthermore, when building a skill library based on skill elements, symbolic predicate logic is first used to define the prerequisites and effects of skill elements, abstracting the skill elements. Each skill element contains adjustable parameters that can be adjusted to adapt to different scenarios. The skill element storage structure adopts a three-dimensional mesh storage method. First, skill elements are hierarchically classified by function and scenario labels, forming a single-layer mesh structure with similar functions and scenarios. Second, the single-layer mesh structure is connected through the knowledge graph to form a three-dimensional structure, realizing the association of skill elements.
[0027] Preferably, the robot skill library implements skill library capacity management through the following operations, including: (1) removing skill elements that have not been called within a set period; (2) replacing old skill elements with new skill elements in the skill arbitrator that have passed the evaluation, and clustering and modularizing similar skill elements.
[0028] On the other hand, the present invention provides a task-driven universal robot intelligent control system, including a skill library, a task parser, a skill generator and a skill arbitrator.
[0029] The skill library contains skill elements that are core capabilities of the robot that are highly universal and frequently used, while avoiding excessive complexity.
[0030] The task parser utilizes a robotic perception-language-planning model, comprising a multimodal perception layer, a language understanding layer, and a planning and decision-making layer. The multimodal perception layer is used to perform environmental information perception tasks within complex tasks, while the language understanding layer is used to perform natural language understanding tasks within complex tasks. The outputs of these two layers are fed into a multimodal Transformer architecture to generate cognitive features that are fed into the planning and decision-making layer, ultimately outputting subtasks.
[0031] The skill generator is used to generate new skill elements when there are missing skill elements in the skill library and the subtask cannot be completed.
[0032] The skill arbitrator is used to determine whether the generated new skill element is added to the skill library.
[0033] The beneficial effects of the present invention are as follows:
[0034] (1) To address the problem of data exponential explosion caused by the increasing complexity of robot tasks, the present invention proposes to build a skill library by learning skill elements, so that when training robots to perform tasks, they do not need to face a huge data set. Instead, they only need to train a small amount of data to form skill elements. This helps to reduce data requirements, improve task success rates, maintain high performance, and significantly reduce data costs.
[0035] (2) In order to solve the problem that robots have difficulty in skill transfer in generalized scenario control work, the present invention proposes to add a skill generator and a skill arbitrator to make the skill element have better precision and versatility, allowing reuse across different scenarios and tasks, and enabling the robot to complete complex tasks in generalized scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Overall workflow diagram of the task-driven general robot intelligent control framework.
[0037] Figure 2 Build a flowchart for your skills library. DETAILED DESCRIPTION
[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings and the robot workflow. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by ordinary technicians in this field based on the embodiments of the present invention without inventive efforts are within the scope of protection of the present invention.
[0039] In one aspect, this embodiment provides a task-driven universal robot intelligent control method, comprising the following steps:
[0040] Step 1: Build an initial skill library for the robot. The skill elements included in the initial skill library are core capabilities of the robot that are highly universal and frequently used (for example, skill elements related to basic movement and navigation, object manipulation, and interaction), while avoiding excessive complexity of the skill elements.
[0041] The skill elements in the robot's initial skill library are: Based on the learned skill actions, the actions are divided into discrete movements that can represent the completion of a task. Similar discrete movements are then grouped and defined as skill elements. (For example, in the grasping skill, a robot can grasp an apple, a doorknob, or other objects; these discrete movements are combined into a single skill element labeled "grasp object").
[0042] The skill meta-levels in the initial skill library are constructed using two methods. First, when constructing skill meta-levels involving single-object interactions (e.g., grasping an apple), skill meta-level learning is accomplished through end-to-end training using data-driven simulation. Specifically, a digital twin environment that closely replicates the real world is constructed using a high-fidelity simulation platform. Reinforcement learning is performed within this digital twin environment, and the resulting learning strategies are then distilled into skill meta-levels.
[0043] The second type: When it is necessary to build a skill element involving multi-object interaction (such as washing dishes), the skill element is learned through two stages: human prior and autonomous environment interaction. First, the human prior stage is carried out: through motion capture teleoperation and VR teleoperation, the expert teaching trajectory database is collected, and the preliminary policy network is trained (the preliminary policy network is Define, where is the entire network training space, is the observation state, It's action. is the initial state distribution, are unknown and possibly random transition probabilities that depend on the system dynamics, The reward function also encodes the task), establishing basic action patterns. Next, the autonomous environment interaction phase begins: optimizing strategies during these interactions, improving performance through the reward function, and continuously collecting data from the real environment and incorporating it into the training data. This allows the system to adapt to the state distribution in the real environment, optimize basic action patterns, and ultimately complete the learning of skill meta-data.
[0044] Step 2: The robot decomposes the complex task into subtasks through the task decoder. The robot perception-language-planning model is used as the task parser.
[0045] Complex robotic tasks include environmental perception and natural language understanding. The robot's perception-language-planning model comprises a multimodal perception layer, a language understanding layer, and a planning and decision-making layer. The multimodal perception layer is used to perform the environmental perception required for complex tasks, while the language understanding layer is used to perform the natural language understanding required for these tasks. The outputs of these two layers are generated through a multimodal Transformer architecture to generate cognitive features that are fed into the planning and decision-making layer, ultimately delivering subtask outputs.
[0046] Multimodal Perception Layer: This layer is responsible for acquiring the environmental perception information required for complex tasks. It uses encoders tailored to the specific environmental inputs: for image inputs, a convolutional neural network (CNN); for audio inputs, a recurrent neural network (RNN); and for digital inputs such as LiDAR, a Kalman filter. All encoder outputs are mapped into a unified multimodal vector space.
[0047] Language Understanding Layer: This layer is responsible for acquiring the natural language understanding information required for complex tasks. Using a large language model (GPT-4), it converts the natural language in complex tasks into structured target parsing instructions.
[0048] The unified encoding data in the multimodal vector space output by the multimodal perception layer and the structured target parsing instructions output by the language understanding layer are used to generate cognitive features through the multimodal Transformers architecture.
[0049] Planning and decision-making layer: Cognitive features and noise actions are passed as input tokens to the Transformer module. Step information is added to the cognitive features through sinusoidal positional encoding. This cognitive feature is combined with the noise action and input into the Transformer module again to gradually complete the planned action sequence. At the same time, the related action sequences are connected into subtasks.
[0050] For example, during a robot's operation, a task (verbal or text) is given: "Open the top drawer, retrieve the cup, and pay attention to the vase." The multimodal perception layer first acquires the environmental perception information required for this complex task. In this example, the shape and spatial position of objects and components in the environment, such as the cup, vase, and drawer, are unified in a multimodal vector space. The language understanding layer then converts the natural language into structured target parsing instructions: 1. Grab the handle of the top drawer; 2. Move the handle outward; 3. Move the robot away from the vase; 4. Open the drawer and search for the cup; 5. Remove the cup. These unified data and structured target parsing instructions in the multimodal vector space output by these two layers are then used to generate cognitive features using a multimodal Transformer architecture. In this example, these cognitive features include "the vase is on the top of the drawer" and "the drawer has three layers." Finally, the planning decision layer passes the generated cognitive features and noise actions as input words to the Transformer module. The step information is added to the cognitive features through sinusoidal position encoding. The cognitive features are combined with the noise actions and input into the Transformer module again to gradually complete the planned action sequence. At the same time, the related action sequences are connected into subtasks. In this embodiment, they are subtask 1: "The drawer is 100 cm to the right, and you should first move to 30 cm away from the drawer"; subtask 2: "Find the top layer of the drawer and grab the handle"; subtask 3: "Pay attention to the shaking amplitude of the vase on the top of the drawer, pull the handle and move outward and adjust the speed at any time"; subtask 4: "Find the cup in the drawer and take it out."
[0051] Step 3: When a subtask is obtained, the robot combines the skill library to complete the subtask. The robot realizes modular execution of complex tasks by dynamically combining and parameterizing predefined skill elements according to task requirements.
[0052] The specific operations are as follows: First, the robot retrieves and matches skill meta-elements, performs semantic similarity matching based on the semantic alignment between the skill meta-description and the task requirements, and searches for the corresponding skill meta-elements in the skill library; second, a symbolic planner is used to generate an action sequence that meets physical and logical constraints; finally, the Q-learn network is trained to evaluate the expected benefits of candidate skill combinations and select the optimal solution.
[0053] Step 4: When the robot cannot complete the subtask, it derives the missing skill element through the knowledge graph, and the skill element generator generates new skill elements to complete the subtask.
[0054] Robot skill metagenerators generate skill metagenerators in two ways: compositional generation and generalization generation. Combination generation generates new skill metagenerators by using existing skill metagenerators as a foundation, optimizing the skill metagenerator call sequence using the PPO and SAC algorithms. Using the skill metagenerators as nodes, the optimal path is searched for through a state transition diagram to combine and generate new skill metagenerators. (For example, learning the new skill metagenerator "entering a room" from the three existing skill metagenerators "opening a door → walking → closing a door"). Generalization generation can be implemented in two ways: First, a complementary combination of an autoregressive model and a diffusion model is used to generate new skill metagenerators. The autoregressive model sequentially predicts the discrete actions required for the new task, while the diffusion model gradually denoises the discrete actions generated by the autoregressive model to generate an action sequence to form a new skill metagenerator. (For example, this approach is used when a robot learns skill metagenerators with similarities between "unscrewing a bottle cap" and "tightening a screw.") Second: Using reinforcement learning, we design reward items based on energy efficiency and contact force stability, and adopt a physics-guided reward function to guide the generation of action sequences that conform to physical intuition as new skill elements (for example, the robot has learned the two skill elements of "walking" and "running", and in this way it can learn the new skill element of "passing through road sections in different states (muddy roads, gravel roads)").
[0055] When choosing a skill element generation method, combinatorial generation is more efficient than generalized generation, but the resulting skill elements are more limited. Therefore, combinatorial generation is preferred for generating new skill elements. If the required skill element cannot be generated, generalized generation should be used to generate new skill elements.
[0056] Step 5: Complete the execution of the subtask by combining the generated new skill element, and at the same time use the robot skill arbitrator to determine whether the skill element should be added to the skill library.
[0057] The robot skill arbitrator is divided into two steps: first, by analyzing the basic performance of the new skill element in completing actual tasks, the basic performance evaluation score of the new skill element is obtained; second, a comprehensive scoring model is used to set the score threshold to determine whether it should be included in the database.
[0058] Different basic performance indicators are selected based on the task type to evaluate the new skill element. These include task completion rate, execution efficiency, physical feasibility, safety, robustness, dangerous action detection, anti-interference capability, failure recovery mechanism, generalization capability assessment, and resource consumption analysis.
[0059] The comprehensive scoring model adopts a pre-trained neural network model based on activation contribution. The model uses forward propagation analysis to adjust the activation value in the model in real time according to the ongoing task type, calculates the weight of each basic performance through variance contribution, and obtains the score by weighting the basic performance evaluation score of the new skill element and the weight of the basic performance. Skill elements with a score greater than the set score threshold are added to the skill library.
[0060] For example: When a robot is faced with explosion-proof leakage handling in a chemical plant, the task tolerance rate is extremely low. It is necessary to completely eliminate the source of danger to avoid residual risks. More emphasis is placed on the task completion rate, safety, and robustness in basic performance, and repeated verification of operation results. When the robot is in logistics and warehousing scenarios, order throughput directly affects commercial profits. It is necessary to maximize the task volume per unit time, and more emphasis is placed on testing the robot's execution efficiency, generalization ability assessment, and resource consumption analysis.
[0061] On the other hand, the present invention provides a task-driven universal robot intelligent control system, including a skill library, a task parser, a skill generator and a skill arbitrator.
[0062] The skill elements included in the skill library are core capabilities of the robot that are highly universal and frequently used, while avoiding excessive complexity.
[0063] The skill elements in the robot's initial skill library are: Based on the learned skill actions, the actions are divided into discrete movements that can represent the completion of a task. Similar discrete movements are then grouped and defined as skill elements. (For example, in the grasping aspect, a robot can grasp an apple, a doorknob, or other objects; these discrete movements are combined into a single skill element labeled "grasp object").
[0064] The skill meta-element in the initial skill library is constructed in two ways: First, when constructing a skill meta-element that only involves single-object interaction (for example, grabbing an apple), skill meta-element learning is completed through data-driven simulation end-to-end training.
[0065] The second type: When it is necessary to build a skill element involving multi-object interaction (such as washing dishes), the skill element is learned through two stages: human prior and autonomous environment interaction.
[0066] First, the human prior stage is carried out: the expert teaching trajectory database is collected through motion capture teleoperation and VR teleoperation, and the preliminary strategy network is trained (the preliminary strategy network is Define, where is the entire network training space, is the observation state, It's action. is the initial state distribution, are unknown and possibly random transition probabilities that depend on the system dynamics, Encoding tasks for reward functions simultaneously), establishing basic action behavior patterns.
[0067] Secondly, it enters the autonomous environment interaction stage: the robot optimizes its strategy in environmental interaction, improves its performance through reward functions, continuously collects data generated in the actual environment, and adds it to the training data, so that it can adapt to the state distribution in the actual environment, optimize basic action behavior patterns, and finally complete the learning of skill elements.
[0068] When building a skill library based on skill elements, symbolic predicate logic is first used to define the prerequisites and effects of skill elements, abstracting them. Each skill element contains adjustable parameters that can be adjusted to adapt to different scenarios. The skill element storage structure uses a three-dimensional mesh storage method. First, skill elements are hierarchically classified by function and scenario labels, forming a single-layer mesh structure with similar functions and scenarios. Second, the single-layer mesh structure is connected through a knowledge graph to form a three-dimensional structure, realizing the association of skill elements.
[0069] The robot skill library achieves skill library capacity management through the following operations, including: (1) removing skill elements that have not been called within a set period; (2) replacing old skill elements with new skill elements from the skill arbitrator that have passed the evaluation, and clustering and modularizing similar skill elements.
[0070] The task parser utilizes a robotic perception-language-planning model, comprising a multimodal perception layer, a language understanding layer, and a planning and decision-making layer. The multimodal perception layer is used to perform environmental information perception tasks within complex tasks, while the language understanding layer is used to perform natural language understanding tasks within complex tasks. The outputs of these two layers are fed into a multimodal Transformer architecture to generate cognitive features that are fed into the planning and decision-making layer, ultimately outputting subtasks.
[0071] Multimodal Perception Layer: This layer is responsible for acquiring the environmental perception information required for complex tasks. It uses corresponding encoders for different environmental information inputs (preferably, convolutional neural networks (CNNs) for image inputs, recurrent neural networks (RNNs) for audio inputs, and Kalman filters for digital information inputs such as LiDAR). All encoder outputs are mapped into a unified multimodal vector space.
[0072] Language Understanding Layer: This layer is responsible for acquiring the natural language understanding information required for complex tasks. It uses a large language model (preferably GPT-4) to convert the natural language in complex tasks into structured target parsing instructions.
[0073] The unified encoding data in the multimodal vector space output by the multimodal perception layer and the structured target parsing instructions output by the language understanding layer are used to generate cognitive features through the multimodal Transformers architecture.
[0074] Planning and decision-making layer: Cognitive features and noise actions are passed as input tokens to the Transformer module. The step information output by the Transformer module is added to the cognitive features through sinusoidal position encoding. The cognitive features with added step information are combined with the noise actions and input into the Transformer module again. Action sequence planning is gradually completed, and the related action sequences are connected into subtasks.
[0075] The skill generator is used to generate new skill elements when there are missing skill elements in the skill library and the subtask cannot be completed.
[0076] There are two approaches to generating skill meta-elements in a robot skill generator: combinatorial generation and generalized generation. Combinatorial generation builds on existing skill meta-elements, optimizing the skill invocation sequence using the PPO and SAC algorithms. Using existing skill meta-elements as nodes, the state transition diagram is used to search for optimal paths and combine them to generate new skill meta-elements (for example, completing the skill meta-elements "enter a room" by following the steps "open door → walk → close door"). Generalized generation involves two approaches to learning new skill meta-elements. First, a complementary combination of autoregressive and diffusion models is used to generate new skill meta-elements. The autoregressive model sequentially predicts the discrete actions required for the new task, while the diffusion model gradually denoises the discrete actions generated by the autoregressive model to generate action sequences, forming the new skill meta-elements. Second, reinforcement learning is used to design rewards based on energy efficiency and contact force stability. A physics-inspired reward function is used to guide the generation of physically intuitive action sequences as new skill meta-elements.
[0077] The skill arbitrator is used to determine whether the generated new skill element is added to the skill library.
[0078] The robot skill arbitrator is divided into two steps: first, by analyzing the basic performance of the new skill element in completing actual tasks, the basic performance evaluation score of the new skill element is obtained; second, a comprehensive scoring model is used to set the score threshold to determine whether it should be included in the database.
[0079] Different basic performance indicators are selected based on the task type to evaluate the new skill element. These include task completion rate, execution efficiency, physical feasibility, safety, robustness, dangerous action detection, anti-interference capability, failure recovery mechanism, generalization capability assessment, and resource consumption analysis.
[0080] The comprehensive scoring model adopts a pre-trained neural network model based on activation contribution. The model uses forward propagation analysis to adjust the activation value in the model in real time according to the ongoing task type, calculates the weight of each basic performance through variance contribution, and obtains the score by weighting the basic performance evaluation score of the new skill element and the weight of the basic performance. Skill elements with a score greater than the set score threshold are added to the skill library.
[0081] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A task-driven general robot intelligent control method, characterized in that: The following steps are involved: Build the robot's initial skill library; The skill element in the robot's initial skill library is: based on the learned skill action, the action is divided into discrete actions that can represent the completion of a task, and then similar discrete actions are grouped and defined as skill elements; The skill elements in the initial skill library are constructed in two ways: The first method is to build a skill meta-learning model that only involves single-object interaction, using data-driven simulation end-to-end training to complete the skill meta-learning. The second method is to build a skill element involving multi-object interaction. This is done through two stages: human prior knowledge and autonomous environment interaction. The specific steps are as follows: First, the human prior stage is carried out: through motion capture teleoperation and VR teleoperation, a database of expert teaching trajectories is collected, a preliminary strategy network is trained, and basic action behavior patterns are established; The initial strategy network is implemented through the following Define, where is the entire network training space, is the observation state, It's action. is the initial state distribution, are unknown transition probabilities that depend on the system dynamics, Encode tasks simultaneously for the reward function; Next, the robot enters the autonomous environment interaction stage: it optimizes its strategy in environmental interaction, improves its performance through reward functions, continuously collects data generated in the real environment, and adds it to its training data. This allows it to adapt to the state distribution in the real environment, optimize basic action behavior patterns, and ultimately complete the learning of skill meta-data. The robot decomposes complex tasks into subtasks through a task decoder; the robot perception-language-planning model is used as a task parser; When a subtask is obtained, the robot combines the initial skill library to complete the subtask. The robot dynamically combines and parameterizes predefined skill elements according to task requirements to achieve modular execution of complex tasks. When there is a missing skill element in the skill library and the subtask cannot be completed, a new skill element is generated through the skill generator; The subtask is completed by combining the generated new skill element, and the skill arbitrator determines whether the skill element should be added to the skill library; If the newly generated skill element passes the arbitration of the skill arbitrator, it will be added to the skill library; if the newly generated skill element fails to pass the arbitration of the skill arbitrator, it will be forgotten.
2. A task-driven universal robot intelligent control method according to claim 1, characterized in that: The complex tasks include environmental information perception tasks and natural language understanding tasks; the robot perception-language-planning model includes a multimodal perception layer, a language understanding layer, and a planning and decision-making layer; the multimodal perception layer is used to complete the environmental information perception task in the complex tasks, and the language understanding layer is used to complete the natural language understanding task in the complex tasks; the outputs of the above two layers generate cognitive features through the multimodal Transformers architecture and enter the planning and decision-making layer, and finally output subtasks; at the same time, in the process of decomposing complex tasks into subtasks, the model determines the indicators based on the actual situation, and completes multi-objective optimization through the context-aware weight network.
3. The task-driven universal robot intelligent control method according to claim 1, characterized in that: The robot skill generator has two ways to generate skill elements, namely, combination generation and generalization generation. The combination generation method generates new skill elements, that is, based on existing skill elements, the skill call sequence is optimized through the PPO algorithm and the SAC algorithm, and the existing skill elements are used as nodes to search for the optimal path through the state transition diagram to generate new skill elements in combination. Learning new skill meta-elements through generalization generation includes two implementation paths: first, using the complementary combination of autoregressive model and diffusion model to generate new skill meta-elements; The autoregressive model is responsible for sequentially predicting the discrete actions required in the new task, and the diffusion model gradually denoises the discrete actions generated by the autoregressive model to generate an action sequence, forming a new skill element. Second: Using reinforcement learning, we design reward items based on energy efficiency and contact force stability, and adopt a physics-guided reward function to guide the generation of action sequences that conform to physical intuition as new skill elements.
4. The task-driven universal robot intelligent control method according to claim 1, characterized in that: The robot skill arbitrator evaluates the acquired new skill element by performing basic performance analysis based on actual task completion; and adopts a comprehensive scoring model to set a score threshold to determine whether the new skill element should be added to the initial skill library.
5. The task-driven universal robot intelligent control method according to claim 1, characterized in that: When building a skill library based on skill elements, we first use symbolic description predicate logic to define the prerequisites and effects of skill elements, abstracting the skill elements. Each skill element contains adjustable parameters that can be adjusted to adapt to different scenarios. The skill element storage structure adopts a three-dimensional mesh storage method. First, the skill elements are hierarchically classified according to function and scenario labels. Similar functions and scenarios form a single-layer mesh structure. Secondly, the single-layer network structure is connected through the knowledge graph to form a three-dimensional structure, realizing the association of skill elements.
6. The task-driven universal robot intelligent control method according to claim 1, characterized in that: The robot skill library achieves skill library capacity management through the following operations, including: (1) removing skill elements that have not been called within a set period; (2) replacing old skill elements with new skill elements from the skill arbitrator that have passed the evaluation, and clustering and modularizing similar skill elements.
7. A task-driven universal robot intelligent control system for implementing the method according to claim 1, characterized in that: Includes skill library, task parser, skill generator and skill arbitrator; The skill library includes skill elements; The task parser uses a robot perception-language-planning model, comprising a multimodal perception layer, a language understanding layer, and a planning and decision-making layer. The multimodal perception layer is used to perform environmental information perception tasks within complex tasks, while the language understanding layer is used to perform natural language understanding tasks within complex tasks. The outputs of these two layers are fed into the planning and decision-making layer through a multimodal Transformer architecture to generate cognitive features that ultimately output subtasks. The skill generator is used to generate new skill elements when there are missing skill elements in the skill library and the subtask cannot be completed; The skill arbitrator is used to determine whether the generated new skill element is added to the skill library.
Citation Information
Patent Citations
Batch techniques for handling unbalanced training data for chat robots
CN115485690A
Ai system
US20250061307A1