Robot control model training method, robot control model control method, robot control model training device, robot control model control device and electronic equipment

By obtaining preset action data sets and action planning models and generating action strategy information, the versatility problem of robot control methods in dynamic unknown scenarios is solved, and more precise robot control is achieved.

CN120620237AActive Publication Date: 2025-09-12BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD

Patent Information

Application Number
CN202511128836.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-09-12
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing robot control methods lack a general mechanism in dynamic unknown scenarios and cannot effectively handle unknown tasks. Multi-strategy dependence leads to task specificity, and the lack of effective text quantification methods leads to insufficient control accuracy and generalization ability.

Method used

By obtaining a preset action dataset, the preset large language model in the robot control model is used to generate action sequence text and predict target direction. Combined with the initial text encoder and action planning model, action strategy information is generated, and the model generalization ability is improved through parameter training.

Benefits of technology

The generalization ability of the robot control model is improved, making the control of the robot more precise and adaptable to different scenarios and environmental changes, reducing the need for task-specific simulation environments, and improving control efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120620237A_ABST
    Figure CN120620237A_ABST
Patent Text Reader

Abstract

The invention provides a training method and device for a robot control model, a control method and device and electronic equipment, and relates to the technical field of equipment control, and the method comprises the steps: obtaining a preset motion data set, and employing a preset large language model in the robot control model to obtain a preset motion data set; generating an action sequence text and a prediction target direction corresponding to the sample natural language command based on the sample natural language command; and encoding the action sequence text by adopting an initial text encoder in the robot control model, determining a target action code corresponding to the action code from a preset action codebook, and performing action planning according to the target action code and a predicted target direction by adopting an initial action planning model in the robot control model. Generating action strategy information corresponding to the sample natural language command; and parameter adjustment training is carried out on the initial motion planning model in the robot control model. The generalization ability of the robot control model is improved, and control over the robot is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of equipment control technology, and in particular to a training method, control method, device and electronic equipment for a robot control model. Background Art

[0002] The development of humanoid robot control technology relies on the deep integration of reinforcement learning and imitation learning, aiming to achieve autonomous execution of complex tasks. In recent years, generative adversarial imitation learning (GAL) has provided an effective framework for policy training by using a discriminator to distinguish between expert trajectories and policy-generated trajectories. However, it can only generate actions for specific tasks and has difficulty handling dynamic, unknown scenarios.

[0003] Traditional approaches, such as generative adversarial imitation learning, use policy and discriminator training to generate actions that mimic expert behavior. Alternatively, adversarial skill embedding utilizes a hierarchical model to enable low-level policies to learn reusable skills, and then train high-level policies to solve tasks. Adversarial motion priors learn to imitate different skills from reference data to achieve high-level tasks.

[0004] However, there is a contradiction between the randomness of large language model output and the accuracy of instruction execution, there is a lack of effective text quantification methods, and the task specificity caused by multi-strategy dependence makes it impossible to handle unknown tasks through a single strategy. In motion control, there is a lack of a general mechanism that takes into account both action imitation and direction control. Summary of the Invention

[0005] The purpose of this application is to address the deficiencies in the above-mentioned prior art and to provide a training method, control method, device and electronic equipment for a robot control model to improve the generalization ability of the robot control model and make the control of the robot more precise.

[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of the present application are as follows: In a first aspect, an embodiment of the present application provides a method for training a robot control model, the method comprising: Acquire a preset action data set, wherein the action data set includes: sample natural language commands and preset reference actions corresponding to the sample natural language commands; Using a preset large language model in the robot control model, based on the sample natural language command, generate an action sequence text and a predicted target direction corresponding to the sample natural language command; Encoding the action sequence text using an initial text encoder in the robot control model to obtain an action code corresponding to the sample natural language command; According to the action code, determining a target action code corresponding to the action code from a preset action code book, wherein the preset action code book pre-stores a plurality of preset action codes; Using the initial motion planning model in the robot control model, performing motion planning according to the target motion code and the predicted target direction, and generating motion strategy information corresponding to the sample natural language command; According to the preset reference action and the action state information of the target robot under the action strategy information, the initial action planning model in the robot control model is adjusted and trained.

[0007] Optionally, the initial action planning model includes: a skill encoder, an initial policy network, and a discriminator; The initial motion planning model in the robot control model is used to perform motion planning according to the target motion code and the predicted target direction, and to generate motion strategy information corresponding to the sample natural language command, including: Performing skill encoding on the target action code using the skill encoder to obtain an action vector corresponding to the sample natural language command; Using the initial strategy network, generating action strategy information corresponding to the sample natural language command according to the predicted target direction and the action vector; The step of training the initial motion planning model in the robot control model according to the preset reference motion and the motion state information of the target robot under the motion strategy information includes: Using the discriminator, based on the preset reference action and the action state information of the target robot under the action strategy information, calculate the action reward parameter under the sample natural language command; The initial strategy network is trained based on the action reward parameters.

[0008] Optionally, the action dataset further includes: a preset target direction corresponding to the sample natural language command; and before adjusting and training the initial policy network according to the action reward parameter, the method further includes: Obtaining the actual motion direction of the target robot under the motion strategy information; Calculating a task reward parameter under the sample natural language command according to the preset target direction and the actual action direction; The adjusting and training of the initial strategy network according to the action reward parameter includes: The initial strategy network is trained based on the action reward parameter and the task reward parameter.

[0009] Optionally, the method further includes: Calculating a coding loss parameter according to the action code corresponding to the sample natural language command and the corresponding target action code; The initial text encoder is trained based on the encoding loss parameter.

[0010] Optionally, determining, according to the action code, a target action code corresponding to the action code from a preset action codebook includes: Calculating, based on the action code, the similarity between the action code and each preset action code in the preset action codebook; According to the similarity, a preset action code closest to the action code is determined from the preset action codebook as the target action code.

[0011] Optionally, the using of a preset large language model in the robot control model to generate an action sequence text and a predicted target direction corresponding to the sample natural language command based on the sample natural language command includes: Generate a prompt word based on the sample natural language command, the description information of the target robot, the description information of the scene where the target robot is located, and the basic skill information of the target robot; The preset large language model is used to generate the action sequence text and predict the target direction based on the prompt word.

[0012] In a second aspect, another embodiment of the present application provides a robot control method, the method comprising: Get input natural language commands; Using a preset large language model in the robot control model, based on the natural language command, an action sequence text and a predicted target direction corresponding to the natural language command are generated; Encoding the action sequence text using an initial text encoder in the robot control model to obtain an action code corresponding to the natural language command; According to the action code, determining a target action code corresponding to the action code from a preset action code book, wherein the preset action code book pre-stores a plurality of preset action codes; Using the target action planning model in the robot control model, action planning is performed according to the target action code and the predicted target direction, and action strategy information corresponding to the sample natural language command is generated; wherein the robot control model is a model trained using the training method of any robot control model described in the first aspect above; The robot is controlled according to the predicted action strategy information.

[0013] In a second aspect, another embodiment of the present application provides a training device for a robot control model, the device comprising: An acquisition module is configured to acquire a preset action data set, wherein the action data set includes: a sample natural language command and a preset reference action corresponding to the sample natural language command; a generation module, configured to use a preset large language model in the robot control model to generate an action sequence text and a predicted target direction corresponding to the sample natural language command based on the sample natural language command; An encoding module, configured to encode the action sequence text using an initial text encoder in the robot control model to obtain an action code corresponding to the sample natural language command; a determination module, configured to determine, based on the action code, a target action code corresponding to the action code from a preset action code book, wherein the preset action code book pre-stores a plurality of preset action codes; a planning module, configured to adopt an initial motion planning model in the robot control model, perform motion planning according to the target motion code and the predicted target direction, and generate motion strategy information corresponding to the sample natural language command; The training module is used to adjust the parameters of the initial motion planning model in the robot control model according to the preset reference action and the action state information of the target robot under the action strategy information.

[0014] In the third aspect, another embodiment of the present application provides an electronic device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and the processor executes the machine-readable instructions to perform the steps of the method described in any one of the first and second aspects above.

[0015] In a fourth aspect, another embodiment of the present application provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described in any one of the first and second aspects are executed.

[0016] The beneficial effects of this application are: The present application provides a training method, control method, device and electronic device for a robot control model, which obtains a preset action data set, adopts a preset large language model in the robot control model, and generates an action sequence text and a predicted target direction corresponding to the sample natural language command based on the sample natural language command; adopts an initial text encoder in the robot control model to encode the action sequence text to obtain an action code corresponding to the sample natural language command; according to the action code, determines a target action code corresponding to the action code from a preset action code book, wherein the preset action code book pre-stores a plurality of preset action codes; adopts an initial action planning model in the robot control model to perform action planning according to the target action code and the predicted target direction, and generates action strategy information corresponding to the sample natural language command; according to the preset reference action and the action state information of the target robot under the action strategy information, adjusts the parameters of the initial action planning model in the robot control model for training. In the present application, based on the preset action data set, a preset large language model is adopted to generate the corresponding action sequence text and the predicted target direction, and the preset large language model is used to understand the corresponding task, so as to adapt to different scenarios and improve the understanding of natural language. The initial text encoder in the robot control model is used to encode the action sequence text to obtain the target action code, which can standardize the action sequence text and adapt it to various tasks. The initial action planning model in the robot control model is used to generate action strategy information, ensure the accurate generation of action strategy information, dynamically adapt to the environment and robot status, and adjust the parameters of the initial action planning model in the robot control model according to the preset reference action and the action state information of the target robot under the action strategy information. The generalization ability of the robot control model can be improved. The method of the present application can ensure the accuracy of the generated action strategy information, thereby improving the generalization ability of the robot control model and making the control of the robot more precise. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 A flowchart of a method for training a robot control model provided in an embodiment of the present application; Figure 2 A schematic diagram of the process of initial strategy network training in a robot model training method provided in an embodiment of the present application; Figure 3A schematic diagram of the process of initial strategy network training in another robot model training method provided in an embodiment of the present application; Figure 4 A schematic diagram of a process for adjusting parameters of an initial text encoder in a robot model training method provided in an embodiment of the present application; Figure 5 A schematic diagram of a flow chart for encoding target actions in a training method for a robot control model provided in an embodiment of the present application; Figure 6 A process for determining text and direction in a training method for a robot control model provided in an embodiment of the present application; Figure 7 A schematic diagram of the structure of a robot control model provided in an embodiment of the present application; Figure 8 A scene diagram of the robot control method provided in an embodiment of the present application; Figure 9 A schematic flow chart of a robot control method based on a robot control model provided in an embodiment of the present application; Figure 10 A schematic diagram of the structure of a training device for a robot control model provided in an embodiment of the present application; Figure 11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.

[0020] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0021] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.

[0022] During robot operation, the operating environment is often complex and changeable, with many uncertainties. For example, in dynamic scenarios, robots may encounter different terrains, interference, and real-time adjustments to task requirements. Therefore, robot control needs to be highly adaptable, able to quickly perceive environmental changes and make corresponding adjustments. At the same time, it must also have efficient control strategies to accurately and smoothly perform tasks under complex working conditions. Currently, robots are controlled through reinforcement learning and imitation learning. However, existing methods require the creation of simulated environments and rewards for specific tasks, resulting in the need for multiple strategies and an inability to handle complex and unknown tasks. To this end, the present application provides a training method for a robot control model, which obtains a preset action data set, adopts a preset large language model in the robot control model, and generates an action sequence text and a predicted target direction corresponding to the sample natural language command based on the sample natural language command; adopts an initial text encoder in the robot control model to encode the action sequence text to obtain the action code corresponding to the sample natural language command; according to the action code, determines the target action code corresponding to the action code from a preset action code book, wherein the preset action code book pre-stores a plurality of preset action codes; adopts an initial action planning model in the robot control model, performs action planning according to the target action code and the predicted target direction, and generates action strategy information corresponding to the sample natural language command; according to the preset reference action and the action state information of the target robot under the action strategy information, adjusts the parameters of the initial action planning model in the robot control model for training. The robot control model trained by the training method of the robot control model in the present application can control the robot, and the robot is controlled based on the predicted action strategy information obtained by the robot control model. The robot control model in this application is a general model that can control the robot based solely on predicted action strategy information. There is no need to create a corresponding simulation environment for specific tasks, which improves the control efficiency of the robot and reduces the complexity of the robot control.

[0023] The following is an explanation of the training method of the robot control model provided in an embodiment of the present application in conjunction with the accompanying drawings. The method can be executed by an electronic device, which includes a memory and a processor. The processor can be a local controller of the robot or a cloud server. The processor is used to execute the robot motion control model training method and the robot control method. Figure 1 A flow chart of a training method for a robot control model provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method includes: Step 101: Obtain a preset action data set.

[0024] The action dataset includes sample natural language commands and preset reference actions corresponding to the sample natural language commands. Sample natural language commands are control instructions for the robot, such as "Please go to the kitchen door." The preset reference actions corresponding to the sample natural language commands can be reference actions obtained from a professional database. Specifically, they can include action sequences of specific actions. For example, when the characteristic action is leg lifting, the preset reference actions correspond to a series of leg lifting action sequences.

[0025] Optionally, the preset action data set can be data obtained from a preset model training data set, or it can be a data set created by itself according to the requirements of the robot control model. This application does not impose any restrictions on this.

[0026] Step 102: Using a preset large language model in the robot control model, generate an action sequence text and a predicted target direction corresponding to the sample natural language command based on the sample natural language command.

[0027] The action sequence text includes multiple actions and their corresponding predicted target directions. For example, the action sequence includes: turn right, run forward, left, walk forward, turn right, and walk forward. The predicted target direction can be a corresponding three-dimensional coordinate, for example, [0, 1, 2]. The default large language model is the Large Language Model (LLM).

[0028] Optionally, a preset large language model in the robot control model is used to generate action sequence text and predicted target direction corresponding to the sample natural language command based on preset prompt words and according to the sample natural language command and the current scene context information.

[0029] For example, when the sample natural language command is "Please go to the kitchen door", the preset large language model in the robot control model is used to generate the action sequence text and predicted target direction of "turn right [0, 1, 2], run forward [0, 1, 2], left [0, 1, 5], walk forward [0, 1, 5], turn right [0, 6, 5] and walk forward [0, 6, 5]" based on "Please go to the kitchen door".

[0030] Step 103: Encode the action sequence text using the initial text encoder in the robot control model to obtain the action code corresponding to the sample natural language command.

[0031] The initial text encoder may be a Contrastive Language-Image Pretraining (CLIP) text encoder. Specifically, the initial text encoder converts the action sequence text into a 512-dimensional vector.

[0032] Optionally, an initial text encoder is used to encode each action in the action sequence text to obtain a 512-dimensional vector for each action. The set of 512-dimensional vectors for each action is the action encoding corresponding to the sample natural language command.

[0033] Step 104: Determine a target action code corresponding to the action code from a preset action codebook according to the action code.

[0034] The preset action codebook contains multiple preset action codes. The preset action codebook can be a 512-dimensional vector corresponding to the text labels in the motion dataset used during training. The motion dataset is the action-related data during robot training, and each action has a corresponding text label.

[0035] Optionally, according to the action code corresponding to each action, the Euclidean distance between each action code and multiple preset action codes in the preset action codebook is calculated respectively, and the preset action code with the closest distance to the action code in the preset action codebook is obtained as the target action code of the action.

[0036] For example, the action in the action sequence text is "turn right", and the initial text encoder is used to obtain the corresponding 512-dimensional vector. The Euclidean distance between the 512-dimensional vector corresponding to "turn right" and multiple preset action codes in the preset action codebook is calculated, and the preset action code with the closest distance is "turn right" in the motion data set, and the initial text encoder is used to obtain the corresponding 512-dimensional vector.

[0037] Step 105: Using the initial motion planning model in the robot control model, perform motion planning according to the target motion encoding and the predicted target direction, and generate motion strategy information corresponding to the sample natural language command.

[0038] Among them, the initial motion planning model is used to generate motion strategy information based on the target motion encoding and predicted target direction, and the motion strategy information is used to control the motion of the target robot.

[0039] Step 106: Based on the preset reference action and the action state information of the target robot under the action strategy information, the initial action planning model in the robot control model is adjusted and trained.

[0040] Optionally, based on the preset reference action and the action state information of the target robot under the action strategy information, the difference between the action state information of the target robot under the action strategy information and the preset reference action is determined, and the initial action planning model in the robot control model is adjusted and trained through adversarial imitation learning.

[0041] In an embodiment of the present application, a preset action data set is obtained, and a preset large language model in a robot control model is used to generate an action sequence text and a predicted target direction corresponding to the sample natural language command based on the sample natural language command; an initial text encoder in the robot control model is used to encode the action sequence text to obtain an action code corresponding to the sample natural language command; based on the action code, a target action code corresponding to the action code is determined from a preset action code book, wherein the preset action code book is pre-stored with multiple preset action codes; an initial action planning model in the robot control model is used to perform action planning based on the target action code and the predicted target direction to generate action strategy information corresponding to the sample natural language command; and the initial action planning model in the robot control model is trained based on the preset reference action and the action state information of the target robot under the action strategy information. In this application, based on the preset action data set, a preset large language model is used to generate the corresponding action sequence text and the predicted target direction. By understanding the corresponding task through the preset large language model, it can adapt to different scenarios and improve the understanding of natural language. The initial text encoder in the robot control model is used to encode the action sequence text to obtain the target action code, which can standardize the action sequence text and thus adapt to various tasks. The initial motion planning model in the robot control model is used to generate motion strategy information, ensuring accurate generation of motion strategy information and dynamically adapting to the environment and robot state. Based on preset reference motions and the target robot's motion state information under the motion strategy information, the initial motion planning model in the robot control model is trained to improve the generalization ability of the robot control model. The method of the present application can ensure the accuracy of generated motion strategy information, thereby improving the generalization ability of the robot control model and making the control of the robot more precise.

[0042] Based on the above embodiment, the initial action planning model includes: a skill encoder, an initial policy network and a discriminator. This application also provides a process for training the initial policy network in a training method for a robot model. Figure 2 A schematic diagram of the process of initial strategy network training in a robot model training method provided in an embodiment of the present application, such as Figure 2As shown, in the above step 105, the initial motion planning model in the robot control model is used to perform motion planning according to the target motion encoding and the predicted target direction, and generate motion strategy information corresponding to the sample natural language command, including: Step 201: Use a skill encoder to perform skill encoding on the target action encoding to obtain an action vector corresponding to the sample natural language command.

[0043] The skill encoder is used to reduce the high-dimensional text features in target action encoding into low-dimensional action vectors. Action vectors are used to represent the semantics of actions.

[0044] Step 202: Using the initial strategy network, generate action strategy information corresponding to the sample natural language command based on the predicted target direction and action vector.

[0045] Optionally, an initial strategy network is used to convert the predicted target direction and action vector into action strategy information that can be executed by the robot based on the predicted target direction, action vector and current scene context information. The current scene context information may include the motion conditions of each joint of the robot and environmental information. The motion conditions of each joint of the robot may be information such as the height of the robot when it is turned off, the position of each limb, the speed of each limb, the current direction, etc. The environmental information may include information such as whether the ground is flat and whether there are obstacles in the current environment. The embodiment of the present application does not limit the specific content of the current scene context information.

[0046] In the above step 106, the initial motion planning model in the robot control model is trained based on the preset reference motion and the motion state information of the target robot under the motion strategy information, including: Step 203: Using a discriminator, calculate the action reward parameters under the sample natural language command based on the preset reference action and the action state information of the target robot under the action strategy information.

[0047] Among them, the preset reference action includes the preset joint state and preset motion trajectory of the robot; the action state information of the target robot under the action strategy information includes the joint state and motion trajectory of the robot under the action strategy information.

[0048] Optionally, according to the state of the target robot not executing the action strategy information st , and the state after executing the action strategy information st +1 Determine the actual movement direction of the target robot and determine the status of the unexecuted action strategy information in the preset action data set based on the preset action data set st And the status after executing the action strategy information stThe preset target direction corresponding to +1 is used to distinguish the state transition distribution between the preset target direction and the actual action direction through the discriminator objective function.

[0049] Optionally, the objective function is:

[0050] in, is the discriminator, It is used to find the discriminator parameters that minimize the objective function, and E[•] is used to find the mathematical expectation. is the state transition distribution corresponding to the preset target direction, is the state transition distribution corresponding to the actual action direction, log(•) is the logarithmic function, Z is the action strategy information corresponding to the sample natural language command, is the gradient penalty weight, is the input variable of the discriminator , For the discriminator to input variables The gradient vector of .

[0051] Optionally, the discriminator's action reward parameter is determined based on the objective function If the difference between the preset target direction and the actual action direction is large, then The closer it is to 0, the smaller the difference between the preset target direction and the actual action direction. The closer to 1.

[0052] Step 204: Adjust and train the initial strategy network according to the action reward parameters.

[0053] Optionally, a feedback signal of the discriminator is generated according to the action reward parameter, and the initial policy network is trained and adjusted so that the action state information corresponding to the action policy information obtained by the adjusted policy network is closer to the preset reference action.

[0054] In this embodiment, a skill encoder is used to obtain the action vector corresponding to a sample natural language command based on the target action encoding; an initial policy network is used to generate action policy information corresponding to the sample natural language command; a discriminator is used to calculate the action reward parameters under the sample natural language command; and the initial policy network is trained based on the action reward parameters. This application enables efficient training and resource optimization, improving the accuracy of the action policy information generated by the policy network.

[0055] Based on the above embodiment, the action data set also includes: a preset target direction corresponding to a sample natural language command. This application also provides another process for training the initial strategy network in a training method for a robot model. Figure 3A schematic diagram of the process of initial strategy network training in another robot model training method provided in an embodiment of the present application is shown as follows: Figure 3 As shown, before the initial policy network is trained in step 204, the method further includes: Step 301: Obtain the actual motion direction of the target robot under the motion strategy information.

[0056] The target robot's actual motion direction under the motion strategy information includes the direction vectors of the robot's pelvis, left hip, and right hip. If the target robot is holding another object, the target robot's actual motion direction under the motion strategy information also includes parameters such as the position, velocity, rotation relative to the pelvis, and angular velocity of the other object. The other object can be any object, such as a handkerchief or shield, and is not limited in this embodiment.

[0057] Optionally, inertial sensors are set according to the pelvis, left hip joint, and right hip joint of the target robot to obtain the actual movement direction of the target robot under the movement strategy information.

[0058] Step 302: Calculate the task reward parameters under the sample natural language command based on the preset target direction and the actual action direction.

[0059] The preset target direction includes the preset direction vectors of the robot's pelvis, left hip, and right hip. If the target robot is holding another object, the target robot's actual motion direction based on the motion strategy information also includes parameters such as the position, speed, rotation relative to the pelvis, and angular velocity of the other object. The other object can be any object, such as a handkerchief or shield, and this embodiment of the application does not impose any restrictions on this.

[0060] Optionally, a pelvic reward function is calculated based on the actual robot's pelvic direction vector, the preset robot's pelvic direction vector, and a preset pelvic weight coefficient. A left hip reward function is calculated based on the actual robot's left hip direction vector, the preset robot's left hip direction vector, and a preset left hip weight coefficient. A right hip reward function is calculated based on the actual robot's right hip direction vector, the preset robot's right hip direction vector, and a preset right hip weight coefficient. The sum of the preset pelvic weight coefficient, the left hip weight coefficient, and the right hip weight coefficient is 1.

[0061] Optionally, the task reward parameter is determined based on the pelvic reward function, the left hip reward function, and the right hip reward function. Specifically, the sum of the pelvic reward function, the left hip reward function, and the right hip reward function is used as the task reward parameter.

[0062] In step 204, the initial policy network is trained based on the action reward parameters, including: Step 303: Adjust and train the initial policy network according to the action reward parameters and the task reward parameters.

[0063] Optionally, a reward function of the initial policy network is determined based on the action reward parameters and the action reward parameter weights corresponding to the action reward parameters, and the task reward parameters and the task reward parameter weights corresponding to the task reward parameters, and the initial policy network is trained based on the reward function of the initial policy network. The sum of the action reward parameter weights and the task reward parameter weights is 1.

[0064] In this embodiment, the initial policy network is trained based on the action reward parameters and task reward parameters to obtain a policy network. This application integrates imitation and physical direction control as optimization goals, which can not only maintain the action accuracy of imitation learning, but also provide an efficient and robust action strategy for humanoid robot control in complex scenarios through the physical guidance of task rewards.

[0065] Based on the above embodiment, the present application also provides a process for adjusting the parameters of the initial text encoder in a training method for a robot control model. Figure 4 A schematic diagram of a process for adjusting parameters of an initial text encoder in a training method for a robot model provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, based on the above steps 101 to 106, the method further includes: Step 401: Calculate a coding loss parameter based on the action code corresponding to the sample natural language command and the corresponding target action code.

[0066] Optionally, the Euclidean distance in the 512-dimensional space is calculated according to the action code corresponding to the sample natural language command and the corresponding target action code, thereby obtaining the coding loss parameter.

[0067] For example, according to the sample natural language command, a 512-dimensional vector of the action code corresponding to the sample natural language command is output as "kick forward quickly", and the target action code corresponding to the encoder is a 512-dimensional vector corresponding to "kick", thereby obtaining the encoding loss parameter according to the action code corresponding to the sample natural language command and the corresponding target action code.

[0068] For example, a sample natural language command is "quickly kick forward" and a 512-dimensional vector of the action encoding corresponding to the sample natural language command is output. The label corresponding to the sample natural language command is passed through an encoder to obtain a 512-dimensional vector corresponding to the label. The encoding loss function is calculated based on the 512-dimensional vector of the action encoding and the 512-dimensional vector corresponding to the label. The label corresponding to the sample natural language command can be a label corresponding to a sample natural language command pre-stored in a preset action dataset.

[0069] Step 402: Perform parameter training on the initial text encoder based on the encoding loss parameters.

[0070] Optionally, according to the encoding loss function, the parameters of the initial text encoder are adjusted by back propagation, thereby reducing the loss function of the initial text encoder.

[0071] In this embodiment of the present application, based on the action encoding corresponding to the sample natural language command and the corresponding target action encoding, the encoding loss parameter is calculated and the initial text encoder is trained. This application can improve the text encoder's understanding of natural language, thereby ensuring the accuracy of the action encoding corresponding to the sample natural language command.

[0072] Based on the above embodiment, the present application also provides a process for determining target action encoding in a training method for a robot control model. Figure 5 A schematic diagram of a flow chart for determining target action encoding in a training method for a robot control model provided in an embodiment of the present application, such as Figure 5 As shown, in the above step 104, according to the action code, the target action code corresponding to the action code is determined from the preset action code book, including: Step 501: Calculate the similarity between the action code and each preset action code in the preset action codebook.

[0073] Optionally, the geometric distances between the action code and each preset action code in the preset action codebook in the multidimensional space are calculated respectively, thereby obtaining the similarity between the action code and each preset action code in the preset action codebook.

[0074] Step 502: According to the similarity, determine the preset action code closest to the action code from the preset action codebook as the target action code.

[0075] Optionally, the preset action code with the greatest similarity to the action code in the preset action codebook is determined as the target action code, wherein the preset action code with the greatest similarity is the preset action code with the smallest geometric distance to the preset action code in the preset action codebook.

[0076] In this embodiment of the present application, the similarity between the action code and each preset action code in the preset action codebook is calculated. Based on the similarity, the preset action code closest to the action code in the preset action codebook is determined as the target action code. In this embodiment of the present application, determining the target action code based on the similarity can reduce computational complexity and improve training and inference efficiency.

[0077] Based on the above embodiments, the present application also provides a process for determining text and direction in a training method for a robot control model. Figure 6The process of determining text and direction in a training method for a robot control model provided in an embodiment of the present application is as follows: Figure 6 As shown, in the above step 101, a preset large language model in the robot control model is used to generate an action sequence text and a predicted target direction corresponding to the sample natural language command based on the sample natural language command, including: Step 601: Generate prompt words based on the sample natural language command, the description information of the target robot, the description information of the scene where the target robot is located, and the basic skill information of the target robot.

[0078] The description information of the target robot may be the basic physical properties of the target robot, for example, it may include: body structure and shape, equipment and tools. The description information of the target robot may be "a humanoid robot with two arms and two legs, a simple gripper on the right hand, and no equipment on the left hand".

[0079] The scene description information of the target robot may include information such as environmental details, spatial structure, etc. For example, the scene description information of the target robot may be “kitchen scene: dishes on the table are messy”.

[0080] The target robot's basic skill information may include motor skill limitations, such as movement capability boundaries and mastered basic skills. For example, the target robot's basic skill information may include "walking forward and backward, turning left and right, and grasping small objects with its right gripper."

[0081] The prompt word can be an action and the coordinates corresponding to the action.

[0082] Optionally, the description information of the target robot, the description information of the scene in which the target robot is located, and the basic skill information of the target robot are used to determine the actions that the target robot can perform, and generate corresponding prompt words according to the sample natural language commands.

[0083] Step 602: Using a preset large language model, generate action sequence text and predict target direction based on the prompt words.

[0084] In the embodiment of the present application, a preset large language model is used to generate action sequence text and predicted target direction based on the prompt word. This application can ensure the accuracy of the generated action sequence text and predicted target direction, and can ensure that the action sequence text and predicted target direction are actions that the target robot can perform, thereby improving the control efficiency of the target robot.

[0085] The training method of the robot control model provided in the embodiment of the present application is described below in conjunction with the robot control model. Figure 7 A schematic diagram of a robot control model provided in an embodiment of the present application is shown in FIG. Figure 7As shown, the robot control model includes: a preset large language model, an initial text encoder, and an initial motion planning model. The initial motion planning model includes: a skill encoder, an initial policy network, and a discriminator. Specifically, the preset large language model in the robot control model is used to generate the action sequence text and predicted target direction corresponding to the sample natural language command based on the sample natural language command. The initial text encoder and the preset action codebook in the robot control model are used to encode the action sequence text to obtain the target action code corresponding to the sample natural language command. The skill encoder is used to skill encode the target action code to obtain the action vector corresponding to the sample natural language command. The initial policy network is used to generate the action strategy information corresponding to the sample natural language command based on the predicted target direction and action vector. The discriminator is used to calculate the action reward parameters under the sample natural language command based on the preset reference action and the action state information of the target robot under the action strategy information. The initial policy network is trained based on the action reward parameters.

[0086] On the basis of the above embodiments, the present application provides a robot control method based on a robot control model. Figure 8 A scene diagram of the robot control method provided in an embodiment of the present application, Figure 9 A flow chart of a robot control method based on a robot control model provided in an embodiment of the present application is shown as follows: Figure 8 As shown, this application takes a household robot as an example. The robot is a humanoid robot with two arms and two legs. The right hand is equipped with a simple gripper, and the left hand is unequipped. The current scene is a kitchen scene, including a table with a mess of dishes on it. The robot can walk forward and backward, turn left and right, and can grab small objects with the gripper on its right hand. Figure 9 As shown, the method includes: Step 901: Obtain input natural language commands.

[0087] The natural language command may be manually input, obtained through voice recognition, or determined by selecting a preset function.

[0088] Step 902: Using the preset large language model in the robot control model, generate the action sequence text and predicted target direction corresponding to the natural language command based on the natural language command.

[0089] Optionally, a preset large language model in the robot control model is used to generate action sequence text and predicted target direction corresponding to the sample natural language command based on the sample natural language command and the current scene context information.

[0090] Step 903: Encode the action sequence text using the initial text encoder in the robot control model to obtain the action code corresponding to the natural language command.

[0091] Optionally, an initial text encoder is used to encode each action in the action sequence text to obtain a 512-dimensional vector for each action. The set of 512-dimensional vectors for each action is the action encoding corresponding to the sample natural language command.

[0092] Step 904: Determine a target action code corresponding to the action code from a preset action codebook according to the action code.

[0093] The preset action codebook pre-stores a plurality of preset action codes.

[0094] Optionally, according to the action code corresponding to each action, the Euclidean distance between each action code and multiple preset action codes in the preset action codebook is calculated respectively, and the preset action code with the shortest distance to the action code is obtained as the target action code of the action.

[0095] Step 905: Use the target action planning model in the robot control model to perform action planning based on the target action coding and the predicted target direction, and generate action strategy information corresponding to the sample natural language command.

[0096] The robot control model is a model trained using any robot control model training method.

[0097] Step 906: Control the robot according to the predicted action strategy information.

[0098] Optionally, the predicted motion strategy information is used to control the robot so that the robot moves based on the predicted motion strategy.

[0099] In an embodiment of the present application, an input natural language command is obtained; a preset large language model in a robot control model is used to generate an action sequence text and a predicted target direction corresponding to the natural language command based on the natural language command; an initial text encoder in the robot control model is used to encode the action sequence text to obtain an action code corresponding to the natural language command; based on the action code, a target action code corresponding to the action code is determined from a preset action code book, a target action planning model in the robot control model is used to perform action planning based on the target action code and the predicted target direction, and action strategy information corresponding to the sample natural language command is generated; based on the predicted action strategy information, the robot is controlled. The present application can ensure the efficiency of generating action sequence text and predicting target directions by using a preset large language model, and can ensure that the robot executes corresponding actions based on a sequence by determining the action strategy information through the target action planning model, thereby ensuring the accuracy and efficiency of the robot control.

[0100] Based on the same inventive concept, the embodiments of the present application also provide a training device for a robot control model corresponding to the training method for the robot control model. Since the principle of solving the problem by the device in the embodiments of the present application is similar to the training method for the robot control model in the above-mentioned embodiments of the present application, the implementation of the device can refer to the implementation of the method.

[0101] Figure 10 A schematic diagram of a training device for a robot control model provided in an embodiment of the present application is shown in FIG. Figure 10 As shown, the device includes: an acquisition module 1001, a generation module 1002, an encoding module 1003, a determination module 1004, a planning module 1005, and a training module 1006; wherein the acquisition module 1001 is used to obtain a preset action data set, and the action data set includes: sample natural language commands and preset reference actions corresponding to the sample natural language commands; A generation module 1002 is configured to use a preset large language model in the robot control model to generate an action sequence text and a predicted target direction corresponding to the sample natural language command based on the sample natural language command; The encoding module 1003 is used to encode the action sequence text using the initial text encoder in the robot control model to obtain the action code corresponding to the sample natural language command; A determination module 1004 is configured to determine, based on the action code, a target action code corresponding to the action code from a preset action code book, wherein the preset action code book pre-stores a plurality of preset action codes; Planning module 1005, for using the initial motion planning model in the robot control model to perform motion planning based on the target motion code and the predicted target direction, and generating motion strategy information corresponding to the sample natural language command; The training module 1006 is used to adjust the parameters of the initial motion planning model in the robot control model according to the preset reference action and the motion state information of the target robot under the motion strategy information.

[0102] Optionally, the initial action planning model includes: a skill encoder, an initial policy network, and a discriminator; the planning module 1005 is specifically configured to: skill encode the target action encoding using the skill encoder to obtain an action vector corresponding to the sample natural language command; Using the initial policy network, the action policy information corresponding to the sample natural language command is generated based on the predicted target direction and action vector; The training module 1006 is specifically configured to: use a discriminator to calculate action reward parameters under sample natural language commands based on preset reference actions and action state information of the target robot under action strategy information; According to the action reward parameters, the initial policy network is trained.

[0103] Optionally, the action data set further includes: a preset target direction corresponding to the sample natural language command; the training module 1006 is further configured to: obtain the actual action direction of the target robot under the action strategy information; Calculate the task reward parameters under the sample natural language command based on the preset target direction and the actual action direction; The training module 1006 is specifically used to adjust and train the initial strategy network according to the action reward parameters and the task reward parameters.

[0104] Optionally, the device further includes: a calculation module, the calculation module being specifically configured to calculate a coding loss parameter according to the action code corresponding to the sample natural language command and the corresponding target action code; According to the encoding loss parameters, the initial text encoder is trained.

[0105] Optionally, the encoding module 1005 is specifically configured to: calculate, based on the action code, the similarity between the action code and each preset action code in the preset action codebook; According to the similarity, a preset action code closest to the action code is determined from the preset action codebook as the target action code.

[0106] Optionally, the generating module 1002 is specifically configured to generate prompt words based on the sample natural language command, the description information of the target robot, the description information of the scene where the target robot is located, and the basic skill information of the target robot; Using a preset large language model, it generates action sequence text and predicts target direction based on prompt words.

[0107] For descriptions of the processing flow of each module in the device and the interaction flow between each module, reference can be made to the relevant descriptions in the above method embodiment, which will not be described in detail here.

[0108] The embodiment of the present application also provides an electronic device, Figure 11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 11 As shown, the electronic device includes: a processor 1101, a memory 1102, and optionally, a bus 1103. The memory 1102 stores machine-readable instructions executable by the processor 1101. When the electronic device is running, the processor 1101 communicates with the memory 1102 via the bus 1103, and the machine-readable instructions are used by the processor 1101 to execute the steps of the above method.

[0109] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are executed.

[0110] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0111] In addition, the functional units in the various embodiments of the present application can be integrated into a single processing unit, each unit can exist physically separately, or two or more units can be integrated into a single unit. If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0112] The above is only a specific implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the protection scope of the present application.

Claims

1. A training method for a robot control model, characterized in that: The method comprises: Acquire a preset action data set, wherein the action data set includes: a sample natural language command and a preset reference action corresponding to the sample natural language command; Using a preset large language model in the robot control model, based on the sample natural language command, generate an action sequence text and a predicted target direction corresponding to the sample natural language command; Encoding the action sequence text using an initial text encoder in the robot control model to obtain an action code corresponding to the sample natural language command; Determining, according to the action code, a target action code corresponding to the action code from a preset action codebook, wherein the preset action codebook pre-stores a plurality of preset action codes; Using the initial motion planning model in the robot control model, performing motion planning according to the target motion code and the predicted target direction, and generating motion strategy information corresponding to the sample natural language command; According to the preset reference action and the action state information of the target robot under the action strategy information, the initial action planning model in the robot control model is adjusted and trained.

2. The method according to claim 1, characterized in that The initial action planning model includes: a skill encoder, an initial strategy network and a discriminator; The initial motion planning model in the robot control model is used to perform motion planning according to the target motion code and the predicted target direction, and to generate motion strategy information corresponding to the sample natural language command, including: Performing skill encoding on the target action code using the skill encoder to obtain an action vector corresponding to the sample natural language command; Using the initial strategy network, generating action strategy information corresponding to the sample natural language command according to the predicted target direction and the action vector; The parameter adjustment and training of the initial motion planning model in the robot control model according to the preset reference motion and the motion state information of the target robot under the motion strategy information includes: Using the discriminator, based on the preset reference action and the action state information of the target robot under the action strategy information, calculate the action reward parameter under the sample natural language command; The initial strategy network is trained based on the action reward parameters.

3. The method according to claim 2, characterized in that The action dataset further includes: a preset target direction corresponding to the sample natural language command; and before training the initial policy network according to the action reward parameter, the method further includes: Obtaining the actual motion direction of the target robot under the motion strategy information; Calculating a task reward parameter under the sample natural language command according to the preset target direction and the actual action direction; The adjusting and training of the initial strategy network according to the action reward parameter includes: The initial strategy network is trained based on the action reward parameter and the task reward parameter.

4. The method according to claim 1, wherein The method further comprises: Calculating a coding loss parameter according to the action code corresponding to the sample natural language command and the corresponding target action code; The initial text encoder is trained based on the encoding loss parameter.

5. The method according to claim 1, characterized in that The step of determining, according to the action code, a target action code corresponding to the action code from a preset action codebook includes: Calculating, based on the action code, the similarity between the action code and each preset action code in the preset action codebook; According to the similarity, a preset action code closest to the action code is determined from the preset action codebook as the target action code.

6. The method according to claim 1, wherein The method of using a preset large language model in the robot control model to generate an action sequence text and a predicted target direction corresponding to the sample natural language command based on the sample natural language command includes: Generate a prompt word based on the sample natural language command, the description information of the target robot, the description information of the scene where the target robot is located, and the basic skill information of the target robot; The preset large language model is used to generate the action sequence text and predict the target direction based on the prompt word.

7. A robot control method, characterized in that: The method comprises: Get input natural language commands; Using a preset large language model in the robot control model, based on the natural language command, an action sequence text and a predicted target direction corresponding to the natural language command are generated; Encoding the action sequence text using an initial text encoder in the robot control model to obtain an action code corresponding to the natural language command; Determining, according to the action code, a target action code corresponding to the action code from a preset action codebook, wherein the preset action codebook pre-stores a plurality of preset action codes; Using the target action planning model in the robot control model, performing action planning according to the target action code and the predicted target direction, and generating action strategy information corresponding to the sample natural language command; wherein the robot control model is a model trained using the training method of the robot control model described in any one of claims 1 to 6 above; The robot is controlled according to the predicted action strategy information.

8. A training device for a robot control model, characterized in that: The device comprises: An acquisition module is configured to acquire a preset action data set, wherein the action data set includes: a sample natural language command and a preset reference action corresponding to the sample natural language command; a generation module, configured to use a preset large language model in the robot control model to generate an action sequence text and a predicted target direction corresponding to the sample natural language command based on the sample natural language command; An encoding module, configured to encode the action sequence text using an initial text encoder in the robot control model to obtain an action code corresponding to the sample natural language command; a determination module, configured to determine, based on the action code, a target action code corresponding to the action code from a preset action code book, wherein the preset action code book pre-stores a plurality of preset action codes; a planning module, configured to adopt an initial motion planning model in the robot control model, perform motion planning according to the target motion code and the predicted target direction, and generate motion strategy information corresponding to the sample natural language command; The training module is used to adjust the parameters of the initial motion planning model in the robot control model according to the preset reference action and the action state information of the target robot under the action strategy information.

9. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor executes the machine-readable instructions to perform the steps of any one of the methods according to claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Training and / or utilizing machine learning models for use in natural language-based robot control

    CN115551681A

  • Robot logistics distribution method based on natural language and logistics distribution robot

    CN117993422A

  • Natural language control method for humanoid robot

    CN119610090A

  • Humanoid robot control method and system and humanoid robot

    CN119974029A

  • Robot action generation method and device, robot, medium and program product

    CN120002669A

Cited By

  • Three-stage training method and device for robot action strategy model and storage medium

    CN121290411A

  • Three-stage training method, device and storage medium of robot action policy model

    CN121290411B

  • Model training method and device for robots of various configurations and storage medium

    CN121340258A

  • Model training method, device and storage medium for multiple-configuration robot

    CN121340258B