Robot skill learning method, system and terminal for text large model assisted reinforcement learning
Through the combination of predefined parameterized skills and text big models, the problem of high data volume and low efficiency of robot long sequence task planning in the prior art is solved, and more efficient reinforcement learning agent training and better task execution results are achieved.
Patent Information
- Application Number
- CN202411992120.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
Existing robot long-sequence task planning methods based on large models and reinforcement learning have high requirements for data volume and low efficiency.
Through predefined parameterized skills, visual object detection algorithms are used to identify object information in the robot's workspace, and task instructions are analyzed in combination with text big models, action planning results are output, and reinforcement learning agents are trained through jump reinforcement learning and conservative Q learning.
The complexity of the model is reduced, the training time of reinforcement learning agents is reduced, efficiency is improved, and the impact of extrapolation error is reduced later in the reinforcement learning agent training is improved, and the final performance is improved.
Smart Images

Figure CN120012819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotics technology, and in particular to a robot skill learning method, system, terminal and computer-readable storage medium using a large text model to assist in reinforcement learning. Background Art
[0002] When a robot faces long sequence tasks, it is difficult to obtain effective guidance experience from natural language. In this regard, existing solutions mainly focus on mapping task instructions into action sequences, but they usually require specially designed symbolic languages, and fixed behavior patterns cannot complete complex tasks.
[0003] In order to improve the flexibility of language instructions and the success rate of tasks, large model technology and reinforcement learning methods can be combined. The former is used to identify language instructions, and the latter is used to execute tasks. For example, the human-computer interactive assembly method and system based on multimodal large models and reinforcement learning in the prior art include: collecting visual, text and voice data, dividing the assembly task into multiple independent subtasks through a multimodal large model; each subtask is trained in advance through a reinforcement learning method, and a skill library is composed of multiple intelligent agents; the large model determines the calling order of the intelligent agents in the current task. However, the above scheme requires that the number of skill libraries must be large enough to be applicable to complex tasks, and the training time for each skill is long. The introduction of large models does not improve the exploration efficiency.
[0004] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0005] The main purpose of the present invention is to provide a robot skill learning method, system, terminal and computer-readable storage medium assisted by text large model reinforcement learning, aiming to solve the problem that the existing robot long sequence task planning method based on large model and reinforcement learning has high data volume requirements and low efficiency.
[0006] To achieve the above-mentioned object of the invention, the present invention provides a robot skill learning method using a large text model to assist reinforcement learning, and the robot skill learning method using a large text model to assist reinforcement learning comprises:
[0007] Predefined parameterized skills, wherein the parameterized skills include skills and parameters;
[0008] Use visual target detection algorithms to identify object information in the robot's workspace and output the semantic labels of objects and the positional relationships between objects in the form of scene graphs;
[0009] The scene graph and the task instruction are input into a text big model, the text big model parses the task instruction through a prompt word engineering, and outputs an action planning result corresponding to the task instruction, wherein the action planning result includes an observation object and an action sequence, the action sequence includes an action, and the action is the parameterized skill;
[0010] The text big model is used as the guiding agent, the reinforcement learning agent is used as the exploring agent, and the jump reinforcement learning and conservative Q learning are combined. The reinforcement learning agent is trained using the data set generated by the interaction between the guiding agent and the exploring agent and the environment to obtain a trained reinforcement learning agent;
[0011] Obtaining a task instruction to be executed, inputting the scene graph and the task instruction to be executed into the text big model, the text big model parsing the task instruction to be executed through a prompt word engineering, and outputting a target observation object corresponding to the task instruction to be executed;
[0012] The target observation object and the to-be-executed task instructions are input into the trained reinforcement learning agent, and the trained reinforcement learning agent outputs a target action sequence.
[0013] Optionally, the predefined parameterized skills, the parameterized skills including skills and parameters, specifically include:
[0014] Predefine parameterized skills, wherein the parameterized skills include skills and parameters, and the skills are composed of atomic skills;
[0015] Among them, the skills include reaching, grasping, pushing and opening the gripper; the atomic skills include horizontal movement, vertical movement, closing the end effector and opening the end effector.
[0016] Optionally, the method of using a visual target detection algorithm to identify object information in the robot workspace and outputting semantic labels of objects and positional relationships between objects in the form of a scene graph is specifically as follows:
[0017] Use visual target detection algorithms to identify object information in the robot's workspace and output the semantic labels of objects and the positional relationships between objects in the form of scene graphs;
[0018] The scene graph includes node features and edge features, wherein the node features are used to describe objects existing in the workspace and the physical properties of the objects, and the edge features are used to describe the spatial positional relationship between objects.
[0019] Optionally, the step of inputting the scene graph and the task instruction into a large text model, the large text model parsing the task instruction through a prompt word engineering, and outputting an action planning result corresponding to the task instruction specifically includes:
[0020] Inputting the scene graph and the task instruction into a textual macro model, wherein the scene graph is used to describe the environment observation and the task instruction is used to describe the task goal;
[0021] The text big model parses the task instruction through prompt word engineering, selects the observation object corresponding to the task goal from the environmental observation and plans the skill sequence required to complete the task goal;
[0022] Using the position of the observed object as a parameter corresponding to a skill in the skill sequence to obtain an action sequence;
[0023] Taking the observed object and the action sequence as the action planning result corresponding to the task instruction, and outputting the action planning result;
[0024] The skill sequence includes several skills, the action sequence includes several actions, and the actions are the parameterized skills.
[0025] Optionally, the text big model is used as a guiding agent, the reinforcement learning agent is used as an exploring agent, and jump reinforcement learning and conservative Q learning are combined. The reinforcement learning agent is trained using a data set generated by the interaction between the guiding agent and the exploring agent and the environment to obtain a trained reinforcement learning agent, which specifically includes:
[0026] Using the large text model as a guiding agent and the reinforcement learning agent as an exploring agent;
[0027] In the early stage of the reinforcement learning agent training, the guiding agent interacts with the environment to generate first training data, and the first training data is stored in an experience replay pool;
[0028] Sampling the first training data from the experience replay pool, and training the exploration agent using the sampled first training data;
[0029] In the middle of the training of the reinforcement learning agent, an exploration step is gradually added, the guiding agent is used to interact with the environment to generate first training data, and the exploration agent is used to interact with the environment to generate second training data, and the first training data and the second training data are stored in an experience replay pool;
[0030] Sampling the first training data and the second training data from the experience replay pool, and training the exploration agent using the sampled first training data and the second training data;
[0031] In the later stage of the training of the reinforcement learning agent, gradually reducing the guiding steps until only the exploration agent interacts with the environment to generate second training data, and storing the second training data in the experience replay pool;
[0032] Sampling the second training data from the experience replay pool, and training the exploration agent using the sampled second training data to obtain a trained exploration agent, wherein the trained exploration agent is a trained reinforcement learning agent;
[0033] Among them, the use of the guiding agent to interact with the environment to generate the first training data refers to executing the action planning result output by the guiding agent in a simulation environment to obtain the first training data; the use of the exploration agent to interact with the environment to generate the second training data refers to using the exploration agent to perform random exploration in a simulation environment to obtain the second training data; and the conservative Q learning loss function is used when training the exploration agent using the sampled first training data.
[0034] Optionally, the step of obtaining the task instruction to be executed, inputting the scene graph and the task instruction to be executed into the text big model, the text big model parsing the task instruction to be executed through prompt word engineering, and outputting the target observation object corresponding to the task instruction to be executed, specifically includes:
[0035] Obtaining a task instruction to be executed, and inputting the scene graph and the task instruction to be executed into the text macro model, wherein the scene graph is used to describe the environment observation, and the task instruction to be executed is used to describe the task target to be executed;
[0036] The text large model parses the to-be-executed task instructions through prompt word engineering, and selects a target observation object corresponding to the to-be-executed task target from the environmental observation.
[0037] Optionally, the inputting the target observation and the to-be-executed task instruction object into the trained reinforcement learning agent, and the trained reinforcement learning agent outputting the target action sequence, specifically includes:
[0038] Inputting the target observation object and the to-be-executed task instruction into the skill network and parameter network of the trained reinforcement learning agent respectively;
[0039] The skill network outputs the skill corresponding to the task instruction to be executed;
[0040] The parameter network outputs the parameters corresponding to the skill;
[0041] The trained reinforcement learning agent uses the skill and the parameters corresponding to the skill as a target action sequence, and outputs the target action sequence.
[0042] To achieve the above-mentioned purpose of the invention, the present invention also provides a robot skill learning system assisted by reinforcement learning using a large text model, and the robot skill learning system assisted by reinforcement learning using a large text model comprises:
[0043] Predefined module: used to predefine parameterized skills, wherein the parameterized skills include skills and parameters;
[0044] Visual detection module: used to identify object information in the robot workspace using visual target detection algorithms, and output semantic labels of objects and positional relationships between objects in the form of scene graphs;
[0045] The text big model action planning module is used to input the scene graph and task instructions into the text big model, the text big model parses the task instructions through prompt word engineering, and outputs the action planning result corresponding to the task instructions, wherein the action planning result includes an observation object and an action sequence, the action sequence includes an action, and the action is the parameterized skill;
[0046] Reinforcement learning agent training module: used to use the large text model as a guide agent, the reinforcement learning agent as an exploration agent, combine jump reinforcement learning and conservative Q learning, and use the data set generated by the interaction between the guide agent and the exploration agent and the environment to train the reinforcement learning agent to obtain a trained reinforcement learning agent;
[0047] Target observation object output module: used to obtain the task instruction to be executed, input the scene graph and the task instruction to be executed into the text big model, the text big model parses the task instruction to be executed through the prompt word engineering, and outputs the target observation object corresponding to the task instruction to be executed;
[0048] Target action sequence output module: used to input the target observation object and the task instructions to be executed into the trained reinforcement learning agent, and the trained reinforcement learning agent outputs the target action sequence.
[0049] To achieve the above-mentioned purpose of the invention, the present invention also provides a terminal, which includes: a memory, a processor, and a robot skill learning program of text large model assisted reinforcement learning stored in the memory and run on the processor, and when the robot skill learning program of text large model assisted reinforcement learning is executed by the processor, the steps of the robot skill learning method of text large model assisted reinforcement learning as described above are implemented.
[0050] To achieve the above-mentioned purpose of the invention, the present invention also provides a computer-readable storage medium, which stores a robot skill learning program assisted by text large model reinforcement learning. When the robot skill learning program assisted by text large model reinforcement learning is executed by a processor, the steps of the robot skill learning method assisted by text large model reinforcement learning as described above are implemented.
[0051] In the present invention, parameterized skills are predefined, and the parameterized skills include skills and parameters; a visual target detection algorithm is used to identify object information in the robot workspace, and semantic labels of objects and positional relationships between objects are output in the form of a scene graph; the scene graph and task instructions are input into a large text model, and the large text model parses the task instructions through a prompt word engineering, and outputs an action planning result corresponding to the task instruction, wherein the action planning result includes an observed object and an action sequence, and the action sequence includes an action, and the action is the parameterized skill; the large text model is used as a guiding agent, and a reinforcement learning agent is used as a probe. A search agent is provided, which combines jump reinforcement learning and conservative Q learning, and uses the data set generated by the interaction between the guiding agent and the exploring agent and the environment to train the reinforcement learning agent to obtain a trained reinforcement learning agent; the task instructions to be executed are obtained, and the scene graph and the task instructions to be executed are input into the text big model, and the text big model parses the task instructions to be executed through prompt word engineering, and outputs the target observation object corresponding to the task instructions to be executed; the target observation object and the task instructions to be executed are input into the trained reinforcement learning agent, and the trained reinforcement learning agent outputs the target action sequence. The present invention uses a large text model to perform scene and instruction parsing, wherein the intermediate process does not rely on additional symbolic languages; visual observations are converted into scene graphs and input into a large text model, and with prompt words, the large text model also has the ability of scene parsing, wherein the process does not rely on a multimodal large model, thereby reducing the complexity of the model; parameterized skills are introduced as a bridge between a task planning method based on a large text model and a reinforcement learning method, thereby enabling the large text model to accelerate the training process of a reinforcement learning agent, thereby reducing training time and improving efficiency; a conservative Q-learning method is combined when using data collected by a guiding agent to update the parameters of an exploration agent, thereby reducing the influence of extrapolation errors; in the later stage of reinforcement learning agent training, the guiding steps are reduced, and the data of an experience replay pool mainly comes from the exploration steps, thereby enabling the final performance of the reinforcement learning agent to be better than the guiding strategy based on the large text model; after the reinforcement learning agent training is completed, the large text model is used to select observation objects for the reinforcement learning agent, thereby realizing language-guided task execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1It is a flow chart of a preferred embodiment of the robot skill learning method of the present invention using a large text model to assist reinforcement learning;
[0053] Figure 2 It is a structural schematic diagram of a preferred embodiment of the robot skill learning method of the present invention using a large text model to assist reinforcement learning;
[0054] Figure 3 It is a flowchart of prompt word design of the present invention;
[0055] Figure 4 It is a structural diagram of a preferred embodiment of the robot skill learning system of the present invention using a large text model to assist in reinforcement learning;
[0056] Figure 5 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0058] A preferred embodiment of the robot skill learning method of the present invention using a large text model to assist reinforcement learning is as follows: Figure 1 As shown, specifically including:
[0059] S1. Predefine parameterized skills, where the parameterized skills include skills and parameters.
[0060] In an implementation of this embodiment, the predefined parameterized skills include skills and parameters, specifically including:
[0061] Predefined parameterized skills, wherein the parameterized skills include skills and parameters, the skills are composed of one or more atomic skills, and the parameterized skills are equivalent to the actions, that is, the parameterized skills and actions are synonymous, both referring to skills with parameters;
[0062] Among them, the skills include reaching, grasping, pushing and opening the gripper; the atomic skills include horizontal movement, vertical movement, closing the end effector and opening the end effector.
[0063] Specifically, in order to make the action planning results of the text model and the dataset generated by the interaction with the environment (i.e., the first training data) directly act on the training process of the reinforcement learning agent, the actions generated by the two must have the same paradigm. Therefore, predefined parameterized skills are introduced. Each parameterized skill includes high-level skills and their corresponding low-level parameters. One or more atomic skills constitute a skill, such as "reach", "grasp", "push", and "open the gripper". Atomic skills can be understood as the most basic control unit. Intuitively, executing an atomic skill means controlling the end of the robot to move to a specified position in space, and controlling the opening and closing of the end effector, including horizontal movement (movement along the x-axis), vertical movement (movement along the y-axis), movement along the z-axis, rotation along the z-axis, closing the end effector, and opening the end effector. The corresponding control parameters include x, y, z, yaw, and g, where x, y, and z represent the displacement distances in the x-axis, y-axis, and z-axis directions respectively; yaw represents the rotation angle in the z-axis direction; and g represents whether the end effector is closed. Skills plus control parameters (i.e., parameters) constitute parameterized skills. "Reach" means to move horizontally first, then rotate along the z-axis, and then move vertically. The corresponding control parameters are (x, y, z, yaw). "Grab" means to execute "Reach" first, then close the end effector. The corresponding control parameters are (x, y, z, yaw, 1). First, execute a "Reach" skill with control parameters (x, y, z, yaw). When "Reach" is completed, execute an "Close End Effector" atomic skill with control parameters (0, 0, 0, 0, 1). "Push" means to close the end effector first, then execute "Reach", and then move in the specified direction. The corresponding control parameters are (x, y, z, yaw, 1, dx, dy, dz). First, execute an "Close End Effector" atomic skill with control parameters (0, 0, 0, 0, 1), then execute an "Reach" skill with control parameters (x, y, z, yaw), and then execute an "Move in Specified Direction" atomic skill with control parameters (dx, dy, dz). "Open the gripper" means opening the end effector, and the corresponding control parameters are (0, 0, 0, 0, 0). The "Open End Effector" atomic skill with control parameters of (0, 0, 0, 0, 0) is only executed once, which can be understood as having no parameters.
[0064] The moderate use of "atomic skills" in the present invention can make the execution process of the task smoother. For the task planning method based on the large text model, a skill sequence is first obtained. Each skill uses the position of the object related to the task as the parameter of the skill to form an action sequence. Therefore, the result is an action sequence; for the reinforcement learning method, the hierarchical idea is used, and the result is a high-level skill category (i.e., skill) plus a low-level motion parameter (i.e., parameter), i.e., an action sequence. In this way, the two methods are unified in the form of action. In summary, the present invention introduces parameterized skills as a bridge between the task planning method based on the large text model and the reinforcement learning method, so that the former can accelerate the training process of the latter and reduce the training time of the reinforcement learning agent.
[0065] S2. Use visual target detection algorithms to identify object information in the robot workspace and output the semantic labels of objects and the positional relationships between objects in the form of scene graphs.
[0066] In one implementation of this embodiment, the visual target detection algorithm is used to identify object information in the robot workspace, and the semantic labels of the objects and the positional relationships between the objects are output in the form of a scene graph, specifically:
[0067] Use visual target detection algorithms to identify object information in the robot's workspace and output the semantic labels of objects and the positional relationships between objects in the form of scene graphs;
[0068] The scene graph includes node features and edge features, wherein the node features are used to describe objects existing in the workspace and the physical properties of the objects, and the edge features are used to describe the spatial positional relationship between objects.
[0069] S3. Input the scene graph and task instruction into the text big model, and the text big model parses the task instruction through prompt word engineering, and outputs the action planning result corresponding to the task instruction.
[0070] In an implementation of this embodiment, the scene graph and the task instruction are input into a large text model, the large text model parses the task instruction through a prompt word project, and outputs an action planning result corresponding to the task instruction, specifically including:
[0071] Inputting the scene graph and the task instruction into a textual macro model, wherein the scene graph is used to describe the environment observation and the task instruction is used to describe the task goal;
[0072] The text big model parses the task instruction through prompt word engineering, selects the observation object corresponding to the task goal from the environmental observation and plans the skill sequence required to complete the task goal;
[0073] Using the position of the observed object as a parameter corresponding to a skill in the skill sequence to obtain an action sequence;
[0074] Taking the observed object and the action sequence as the action planning result corresponding to the task instruction, and outputting the action planning result;
[0075] The skill sequence includes several skills, the action sequence includes several actions, and the actions are the parameterized skills.
[0076] Specifically, the present invention uses a visual target detection algorithm to identify object information in a workspace, and saves the semantic labels of objects and their positional relationships in the form of a scene graph. The scene graph and natural language instructions (i.e., task instructions) are used as inputs to the text model. The former describes environmental observations and the latter describes task objectives. The text model is assisted by prompt word engineering to parse the task objectives, and the operation objects involved in completing the task, i.e., the observation objects, and the action plan for completing this task are selected from environmental observations. It should be noted that skills represent a behavior mode / concept. For example, "reaching" means "horizontal displacement first, then rotation, and then vertical displacement"; parameters specify the motion parameters of the behavior represented by the skill. After the "skills" and "parameters" are determined, the interaction mode between the agent and the environment is finally determined. Using this "skill + parameter" to interact with the environment once is called "performing an action once."
[0077] S4. Use the large text model as the guiding agent and the reinforcement learning agent as the exploring agent, combine jump reinforcement learning and conservative Q learning, and use the data set generated by the interaction between the guiding agent and the exploring agent and the environment to train the reinforcement learning agent to obtain a trained reinforcement learning agent.
[0078] In an implementation of this embodiment, the text big model is used as a guiding agent, the reinforcement learning agent is used as an exploration agent, jump reinforcement learning and conservative Q learning are combined, and the reinforcement learning agent is trained using a data set generated by the interaction between the guiding agent and the exploration agent and the environment to obtain a trained reinforcement learning agent, which specifically includes:
[0079] Using the large text model as a guiding agent and the reinforcement learning agent as an exploring agent;
[0080] In the early stage of the reinforcement learning agent training, the guiding agent interacts with the environment to generate first training data, and the first training data is stored in an experience replay pool;
[0081] Sampling the first training data from the experience replay pool, and training (i.e., updating) the exploration agent using the sampled first training data;
[0082] In the middle of the training of the reinforcement learning agent, an exploration step is gradually added, the guiding agent is used to interact with the environment to generate first training data, and the exploration agent is used to interact with the environment to generate second training data, and the first training data and the second training data are stored in an experience replay pool;
[0083] Sampling the first training data and the second training data from the experience replay pool, and training (i.e., updating) the exploration agent using the sampled first training data and the sampled second training data;
[0084] In the later stage of the training of the reinforcement learning agent, gradually reducing the guiding steps until only the exploration agent interacts with the environment to generate second training data, and storing the second training data in the experience replay pool;
[0085] Sampling the second training data from the experience replay pool, and training (i.e., updating) the exploration agent using the sampled second training data to obtain a trained exploration agent, wherein the trained exploration agent is a trained reinforcement learning agent;
[0086] Among them, the use of the guiding agent to interact with the environment to generate the first training data refers to executing the action planning result output by the guiding agent (i.e., the text large model) in a simulation environment to obtain the first training data; the use of the exploration agent to interact with the environment to generate the second training data refers to using the exploration agent to perform random exploration in a simulation environment to obtain the second training data, i.e., collecting the second training data online; and the conservative Q learning loss function is used when training the exploration agent using the sampled first training data.
[0087] The present invention accelerates the training process of reinforcement learning through a task planning method based on a large text model, so that the robot can learn complex tasks from natural language. Specifically, the action planning result of the large text model is executed in a simulation environment to obtain the first training data, i.e., the four-tuple data {observation, action, reward, next state}, wherein observation refers to the observed object, and action refers to the action sequence, which is stored in the experience replay pool. At this time, the quality of the data in the experience replay pool is greater than the data collected by random exploration (exploration step), so in the early stage of the training of the reinforcement learning agent, by combining jump reinforcement learning and conservative Q learning, the reinforcement learning agent is quickly initialized. With the improvement of the performance of the reinforcement learning agent, the guiding steps of the large text model are gradually reduced, and the online collection of data, i.e., random exploration data (exploration step), is turned to make the final performance of the reinforcement learning agent exceed the planning results of the large text model.
[0088] It should be noted that Jump-Start RL (JSRL) has a pre-existing but not necessarily the best guiding (or “teacher”) agent (corresponding to the text model of the present invention), and an exploring agent (corresponding to the reinforcement learning agent that needs to be trained in the present invention). Under normal circumstances, reinforcement learning collects training data through random exploration in the early stage of training, and the reward value of the collected training data is very low at this time. However, in JSRL, the guiding agent first interacts with the environment to guide the robot to a state “closer to success”, and then the exploring agent interacts with the environment, so that the training data collected has a higher reward value than random exploration. In the later stage of training, as the performance of the exploring agent improves, the guiding steps of the guiding agent are gradually reduced, and the exploring agent is allowed to completely take over the process of interacting with the environment. Considering that there is an extrapolation error problem in the training data collected by the guiding agent and the exploring agent, when the training data collected by the guiding agent is used to update the exploring agent, the loss function of Conservative Q Learning (CQL) can be used to reduce the impact of the extrapolation error. In summary, the present invention uses the text big model as the guiding agent, and through JSRL+CQL, it can accelerate the training process of reinforcement learning. It should be noted that CQL can be replaced by independent Q learning (IQL) to achieve the same effect.
[0089] S5. Obtain the task instructions to be executed, input the scene graph and the task instructions to be executed into the text big model, the text big model parses the task instructions to be executed through prompt word engineering, and outputs the target observation object corresponding to the task instructions to be executed.
[0090] In an implementation of this embodiment, the step of obtaining the task instruction to be executed, inputting the scene graph and the task instruction to be executed into the text big model, the text big model parsing the task instruction to be executed through the prompt word engineering, and outputting the target observation object corresponding to the task instruction to be executed, specifically includes:
[0091] Obtaining a task instruction to be executed, and inputting the scene graph and the task instruction to be executed into the text macro model, wherein the scene graph is used to describe the environment observation, and the task instruction to be executed is used to describe the task target to be executed;
[0092] The text large model parses the to-be-executed task instructions through prompt word engineering, and selects a target observation object corresponding to the to-be-executed task target from the environmental observation.
[0093] S6. Input the target observation object and the to-be-executed task instructions into the trained reinforcement learning agent, and the trained reinforcement learning agent outputs a target action sequence.
[0094] In an implementation of this embodiment, the inputting of the target observation object and the to-be-executed task instruction into the trained reinforcement learning agent, and the trained reinforcement learning agent outputting the target action sequence, specifically includes:
[0095] Inputting the target observation object and the to-be-executed task instruction into the skill network and parameter network of the trained reinforcement learning agent respectively;
[0096] The skill network outputs the skill corresponding to the task instruction to be executed;
[0097] The parameter network outputs the parameters corresponding to the skill;
[0098] The trained reinforcement learning agent uses the skill and the parameters corresponding to the skill as a target action sequence, and outputs the target action sequence.
[0099] Specifically, since the text big model can select the correct observation target (i.e., observation object) according to the environmental description of the scene diagram and the natural language instructions, when the reinforcement learning agent is trained, the text big model can be used to select observation targets for it, so that the reinforcement learning agent can complete operations on specific objects according to the language in a specified scenario.
[0100] In another implementation of this embodiment, Figure 2 As shown, the task parsing module 1 is used to parse natural language instructions (i.e., task instructions), and select the operation object (i.e., observation object) that meets the description of the instruction from the environment according to the environmental information, and plan the specific skill execution steps (i.e., skill sequence) to complete the task described in the instruction, and then use the position of the operation object as the skill parameter to finally obtain the action sequence. First, the scene graph generator converts the image observation in the environment into a scene graph, which is a way to represent environmental observations in text. It consists of node features and edge features. Node features describe which objects are in the space and the physical properties of each object; edge features describe the spatial relationship between each node, such as object A is to the left of object B. The natural language instructions are combined with the scene graph. Figure 1The instructions are input into the LLM (Large Language Model) planner, and combined with the prompt words, it is asked to output the operation observation objects corresponding to the task, as well as the action execution steps corresponding to the task. For example, the natural language instruction is "stack the blue square on top of the red square". The instruction and the scene graph are input into the LLM planner, and it will output that the operation objects involved in the instruction are "blue square" and "red square", and the corresponding skill execution steps (i.e. skill sequence) are "reach -> grab -> reach -> open the gripper". The parameter of each skill is the position of the operation object, and finally the action sequence is "reach (x, y, z, yaw) -> grab (x, y, z, yaw, 1) -> reach (x, y, z, yaw) -> open the gripper". So far, Figure 2 The task parsing module 1 in the figure can be regarded as the guiding agent in JSRL, and the reinforcement learning module 2 is the exploration agent to be trained.
[0101] Figure 2 The reinforcement learning module 2 in is divided into two layers. The first layer determines the skill category (i.e., skill), and the second layer has multiple networks. Each network outputs the execution parameters (i.e., parameters) of the corresponding skill. After obtaining observations (i.e., observation objects) from the environment, they are sent to the upper and lower layers respectively. After the first layer determines the skill category, the output of the parameter network corresponding to the second layer is selected as the parameter of the skill. Under normal circumstances, the agent (i.e., reinforcement learning agent / exploration agent) is allowed to interact with the environment continuously and collect data. After a long enough time, the agent can also complete specific tasks, such as "stacking blocks". Here it is combined with JSRL. First, the data is completely collected by the guiding agent, so that the model (referring to the overall model, including the large text model and the reinforcement learning agent, and the reinforcement learning agent is actually trained) can collect data with higher reward values in the early stage of training and store it in the experience replay pool 3. After collecting a certain amount of data, samples are taken from it to update the exploration agent. Considering the influence of extrapolation error, the loss function of conservative Q learning is combined when backpropagating. As the performance of the exploration agent improves, an exploration step is added appropriately when collecting training data, and the data is stored in the experience replay pool 3, and the "collect data -> sampling update" process is repeated. As the exploration agent is continuously updated, the exploration step is gradually lengthened until the newly collected data is completely generated by the exploration agent, and the guiding agent no longer plays a role in the training process.
[0102] In addition, after the model (actually referring to the reinforcement learning agent) is trained, for a simple task such as "stacking blocks", considering that there are only two blocks on the table, this trained reinforcement learning agent can complete it. However, in actual situations, there may be multiple objects on the table, and which two blocks to stack can be specified by natural language. Therefore, unlike the guiding agent in JSRL, which only plays a role in the training process, the present invention uses the task planner here, that is, the guiding agent / text model has the ability to select observation objects, which can select the correct observation target for the trained reinforcement learning agent, so that it can complete specific tasks according to human language.
[0103] It should be noted that the "guidance" in the present invention refers to letting the guiding agent interact with the environment and collect data, specifically refers to executing the skill planning results of the text large model in the simulation environment to obtain quadruple data; "exploration" refers to letting the exploration agent interact with the environment and collect data, specifically refers to the reinforcement learning agent collecting data through random exploration, that is, collecting data online. Under normal circumstances, JSRL only has "guidance->exploration", and the present invention can explore one more time, and collect data in the way of "exploration->guidance->exploration", which can further reduce the impact of the extrapolation error problem and accelerate the training process of the reinforcement learning agent.
[0104] In another implementation of this embodiment, Figure 3 The following is a design flow chart of prompt words. Prompt words are used to assist Figure 2 The LLM planner in the program understands the meaning of the input natural language instructions and scene graphs, as well as the standard output format. In simple terms, the prompt word is a paragraph of text that is pre-input to the LLM planner, asking it to generate a certain output format for future inputs. Figure 3 What is described is how to design the pre-input paragraph. First, tell the LLM planner that it plays the role of a robot with a two-finger gripper and can complete the predefined parameterized skills "reach", "grasp", "push", and "open the gripper" (role-playing). It is required to understand the object information and spatial relationship in the environment based on the input scene graph, and understand natural language instructions, select appropriate observation targets (i.e., observation objects) from the environment, and plan the action sequence (task description) to complete the task. In order to make the output results accurate and logical, the LLM planner is required to output the thinking process (thinking chain). The thinking chain in the prompt word tells the LLM planner the output content and form of each step. Finally, through appropriate examples, the output form is further standardized.
[0105] The present invention designs prompt words, and the text large model can select the correct target from visual observation according to the text instructions and plan the action sequence; the action sequence planned by the former can reduce the exploration steps of reinforcement learning and speed up the training process of the model; when the reinforcement learning agent performs the task, its action sequence is smoother than the action sequence planned by the text large model to ensure that the machine will not be damaged; the target recognized by the text large model can be used as the observation of reinforcement learning, so as to realize the reinforcement learning agent to complete the task according to the task instruction; therefore, using the text large model to guide reinforcement learning to realize robot skill learning has many advantages. It should be noted that the skill learning here is consistent with the skill learning in the question, both of which refer to task learning, including skill and parameter learning.
[0106] In summary, the present invention uses big model technology to understand language instructions and generate feasible guidance strategies, thereby reducing random exploration steps and accelerating the training process of reinforcement learning; for the trained intelligent agent, the big model can select the correct operation object from language instructions and visual observations as the observation of the intelligent agent, so that the robot can complete actions according to human instructions.
[0107] In addition, based on the above-mentioned text large model assisted reinforcement learning robot skill learning method, the present invention also provides a text large model assisted reinforcement learning robot skill learning system, wherein the text large model assisted reinforcement learning robot skill learning system preferred embodiment, such as Figure 4 As shown, specifically including:
[0108] Predefined module 01: used to predefine parameterized skills, wherein the parameterized skills include skills and parameters;
[0109] Visual detection module 02: used to identify object information in the robot workspace using a visual target detection algorithm, and output the semantic labels of objects and the positional relationships between objects in the form of a scene graph;
[0110] Text big model action planning module 03: used for inputting the scene graph and task instruction into the text big model, the text big model parses the task instruction through prompt word engineering, and outputs the action planning result corresponding to the task instruction, wherein the action planning result includes an observation object and an action sequence, the action sequence includes an action, and the action is the parameterized skill;
[0111] Reinforcement learning agent training module 04: used to use the large text model as a guide agent, the reinforcement learning agent as an exploration agent, combine jump reinforcement learning and conservative Q learning, and use the data set generated by the interaction between the guide agent and the exploration agent and the environment to train the reinforcement learning agent to obtain a trained reinforcement learning agent;
[0112] Target observation object output module 05: used to obtain the task instruction to be executed, input the scene graph and the task instruction to be executed into the text big model, the text big model parses the task instruction to be executed through the prompt word engineering, and outputs the target observation object corresponding to the task instruction to be executed;
[0113] Target action sequence output module 06: used to input the target observation object and the task instruction to be executed into the trained reinforcement learning agent, and the trained reinforcement learning agent outputs the target action sequence.
[0114] In addition, based on the above-mentioned text large model assisted reinforcement learning robot skill learning method and system, the present invention also provides a terminal accordingly, wherein a preferred embodiment of the terminal is as follows: Figure 5 As shown, it specifically includes a processor 10, a memory 20 and a display 30. Figure 5 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0115] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, and a flash card (Flash Card) equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as program codes of the storage terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a robot skill learning program 40 for text large model assisted reinforcement learning is stored on the memory 20, and the robot skill learning program 40 for text large model assisted reinforcement learning can be executed by the processor 10, thereby realizing the steps of the robot skill learning method for text large model assisted reinforcement learning in this application.
[0116] In some embodiments, the processor 10 can be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing a robot skill learning program 40 assisted by text large model reinforcement learning.
[0117] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface.
[0118] In one embodiment, when the processor 10 executes the robot skill learning program 40 of text large model assisted reinforcement learning in the memory 20, the steps of the robot skill learning method of text large model assisted reinforcement learning as described above are implemented.
[0119] The present invention also provides a computer-readable storage medium accordingly, wherein the computer-readable storage medium stores a robot skill learning program of text large model assisted reinforcement learning, and when the robot skill learning program of text large model assisted reinforcement learning is executed by a processor, the steps of the robot skill learning method of text large model assisted reinforcement learning as described above are implemented.
[0120] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or terminal including the element.
[0121] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.
[0122] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A robot skill learning method assisted by reinforcement learning with a large text model, characterized in that: The robot skill learning method using a large text model to assist reinforcement learning includes: Predefined parameterized skills, wherein the parameterized skills include skills and parameters; Use visual target detection algorithms to identify object information in the robot's workspace and output the semantic labels of objects and the positional relationships between objects in the form of scene graphs; The scene graph and the task instruction are input into a text big model, the text big model parses the task instruction through a prompt word engineering, and outputs an action planning result corresponding to the task instruction, wherein the action planning result includes an observation object and an action sequence, the action sequence includes an action, and the action is the parameterized skill; The text big model is used as the guiding agent, the reinforcement learning agent is used as the exploring agent, and the jump reinforcement learning and conservative Q learning are combined. The reinforcement learning agent is trained using the data set generated by the interaction between the guiding agent and the exploring agent and the environment to obtain a trained reinforcement learning agent; Obtaining a task instruction to be executed, inputting the scene graph and the task instruction to be executed into the text big model, the text big model parsing the task instruction to be executed through a prompt word engineering, and outputting a target observation object corresponding to the task instruction to be executed; The target observation object and the to-be-executed task instructions are input into the trained reinforcement learning agent, and the trained reinforcement learning agent outputs a target action sequence.
2. The robot skill learning method of text large model assisted reinforcement learning according to claim 1 is characterized in that: The predefined parameterized skills include skills and parameters, specifically including: Predefine parameterized skills, wherein the parameterized skills include skills and parameters, and the skills are composed of atomic skills; Among them, the skills include reaching, grasping, pushing and opening the gripper; the atomic skills include horizontal movement, vertical movement, closing the end effector and opening the end effector.
3. The robot skill learning method of text large model assisted reinforcement learning according to claim 2 is characterized in that: The method of using a visual target detection algorithm to identify object information in the robot workspace and outputting semantic labels of objects and positional relationships between objects in the form of a scene graph is as follows: Use visual target detection algorithms to identify object information in the robot's workspace and output the semantic labels of objects and the positional relationships between objects in the form of scene graphs; The scene graph includes node features and edge features, wherein the node features are used to describe objects existing in the workspace and the physical properties of the objects, and the edge features are used to describe the spatial positional relationship between objects.
4. The robot skill learning method of text large model assisted reinforcement learning according to claim 3 is characterized in that: The scene graph and the task instruction are input into the text big model, the text big model parses the task instruction through the prompt word engineering, and outputs the action planning result corresponding to the task instruction, specifically including: Inputting the scene graph and the task instruction into a textual macro model, wherein the scene graph is used to describe the environment observation and the task instruction is used to describe the task goal; The text big model parses the task instruction through prompt word engineering, selects the observation object corresponding to the task goal from the environmental observation and plans the skill sequence required to complete the task goal; Using the position of the observed object as a parameter corresponding to a skill in the skill sequence to obtain an action sequence; Taking the observed object and the action sequence as the action planning result corresponding to the task instruction, and outputting the action planning result; The skill sequence includes several skills, the action sequence includes several actions, and the actions are the parameterized skills.
5. The robot skill learning method of text large model assisted reinforcement learning according to claim 1 is characterized in that: The method uses the large text model as a guiding agent, uses a reinforcement learning agent as an exploring agent, combines jump reinforcement learning and conservative Q learning, and uses a data set generated by the interaction between the guiding agent and the exploring agent and the environment to train the reinforcement learning agent to obtain a trained reinforcement learning agent, specifically including: Using the large text model as a guiding agent and the reinforcement learning agent as an exploring agent; In the early stage of the reinforcement learning agent training, the guiding agent interacts with the environment to generate first training data, and the first training data is stored in an experience replay pool; Sampling the first training data from the experience replay pool, and training the exploration agent using the sampled first training data; In the middle of the training of the reinforcement learning agent, an exploration step is gradually added, the guiding agent is used to interact with the environment to generate first training data, and the exploration agent is used to interact with the environment to generate second training data, and the first training data and the second training data are stored in an experience replay pool; Sampling the first training data and the second training data from the experience replay pool, and training the exploration agent using the sampled first training data and the second training data; In the later stage of the training of the reinforcement learning agent, gradually reducing the guiding steps until only the exploration agent interacts with the environment to generate second training data, and storing the second training data in the experience replay pool; Sampling the second training data from the experience replay pool, and training the exploration agent using the sampled second training data to obtain a trained exploration agent, wherein the trained exploration agent is a trained reinforcement learning agent; Among them, the use of the guiding agent to interact with the environment to generate the first training data refers to executing the action planning result output by the guiding agent in a simulation environment to obtain the first training data; the use of the exploration agent to interact with the environment to generate the second training data refers to using the exploration agent to perform random exploration in a simulation environment to obtain the second training data; and the conservative Q learning loss function is used when training the exploration agent using the sampled first training data.
6. The robot skill learning method of text large model assisted reinforcement learning according to claim 1 is characterized in that: The step of obtaining the task instruction to be executed, inputting the scene graph and the task instruction to be executed into the text big model, the text big model parsing the task instruction to be executed through the prompt word engineering, and outputting the target observation object corresponding to the task instruction to be executed specifically includes: Obtaining a task instruction to be executed, and inputting the scene graph and the task instruction to be executed into the text macro model, wherein the scene graph is used to describe the environment observation, and the task instruction to be executed is used to describe the task target to be executed; The text large model parses the to-be-executed task instructions through prompt word engineering, and selects a target observation object corresponding to the to-be-executed task target from the environmental observation.
7. The robot skill learning method of text large model assisted reinforcement learning according to claim 2 is characterized in that: The inputting the target observation and the to-be-executed task instruction object into the trained reinforcement learning agent, and the trained reinforcement learning agent outputting the target action sequence, specifically includes: Inputting the target observation object and the to-be-executed task instruction into the skill network and parameter network of the trained reinforcement learning agent respectively; The skill network outputs the skill corresponding to the task instruction to be executed; The parameter network outputs the parameters corresponding to the skill; The trained reinforcement learning agent uses the skill and the parameters corresponding to the skill as a target action sequence, and outputs the target action sequence.
8. A robot skill learning system assisted by reinforcement learning with a large text model, characterized in that: The robot skill learning system using a large text model to assist in reinforcement learning includes: Predefined module: used to predefine parameterized skills, wherein the parameterized skills include skills and parameters; Visual detection module: used to identify object information in the robot workspace using visual target detection algorithms, and output semantic labels of objects and positional relationships between objects in the form of scene graphs; The text big model action planning module is used to input the scene graph and task instructions into the text big model, the text big model parses the task instructions through prompt word engineering, and outputs the action planning result corresponding to the task instructions, wherein the action planning result includes an observation object and an action sequence, the action sequence includes an action, and the action is the parameterized skill; Reinforcement learning agent training module: used to use the large text model as a guide agent, the reinforcement learning agent as an exploration agent, combine jump reinforcement learning and conservative Q learning, and use the data set generated by the interaction between the guide agent and the exploration agent and the environment to train the reinforcement learning agent to obtain a trained reinforcement learning agent; Target observation object output module: used to obtain the task instruction to be executed, input the scene graph and the task instruction to be executed into the text big model, the text big model parses the task instruction to be executed through the prompt word engineering, and outputs the target observation object corresponding to the task instruction to be executed; Target action sequence output module: used to input the target observation object and the task instructions to be executed into the trained reinforcement learning agent, and the trained reinforcement learning agent outputs the target action sequence.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a robot skill learning program for text large model assisted reinforcement learning stored in the memory and executable on the processor. When the robot skill learning program for text large model assisted reinforcement learning is executed by the processor, the steps of the robot skill learning method for text large model assisted reinforcement learning as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a robot skill learning program using text large model assisted reinforcement learning, and when the robot skill learning program using text large model assisted reinforcement learning is executed by a processor, the steps of the robot skill learning method using text large model assisted reinforcement learning as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Humanoid robot control method, system, equipment and medium
CN120816481A
Intelligent agent training method, data processing method and question answering method
CN121117622A
Humanoid robot indoor action planning and action control method and system and robot
CN121187142A
Humanoid robot indoor motion planning and motion control method, system and robot
CN121187142B