Humanoid robot motion training method and device based on real scene
By building world models and online learning technology, the problem of failure of humanoid robots in real scenes is solved, and more stable motion control and higher task execution capabilities are achieved.
Patent Information
- Application Number
- CN202510616606.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-14
AI Technical Summary
When existing humanoid robots move from simulated training environment to real scenes, motion control strategies often fail due to environmental differences, resulting in unstable motion control in real scenes and inability to perform tasks effectively.
Using a world model and online learning method, a world model is built to simulate the physical characteristics of the real world and the motion laws of the robot, and the initial motion control strategy is trained in the simulated environment, and the strategy is quickly fine-tuned by collecting data in real time in real scenes.
Through this method, humanoid robots can more stably adapt to changes in real scenes, improve the stability and accuracy of motion control, and enhance the ability to perform tasks in real scenes.
Smart Images

Figure CN120116237A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of robot control, and particularly to a method and device for training the movement of a humanoid robot based on a real scenario. Background Art
[0002] With the continuous development of humanoid robot technology, the application fields of humanoid robots have gradually expanded and the complexity has been increasing. For example, rescue tasks in complex environments, industrial production tasks with fine operations, etc. all require humanoid robots to be able to stably and efficiently execute motion control instructions in real scenarios.
[0003] However, the main problem currently faced is that there is a large gap between the simulation training environment (simulation environment) and the real scenario, that is, the sim2real gap problem. For example, in a rescue scenario, a humanoid robot needs to move flexibly in a disaster site full of uncertainties, avoid obstacles and accurately reach the target position. However, for existing robots based on simulation training, after entering the real scenario, due to the differences between the environment and the simulation training environment, the motion control strategies of humanoid robots are often less effective or even ineffective, and they cannot execute tasks stably. Secondly, in an industrial production scenario, a humanoid robot needs to accurately complete operations such as the assembly of components. The strategies trained in the simulation environment are difficult to ensure stable motion accuracy in the real production line due to factors such as equipment vibration and light changes.
[0004] In summary, when the current humanoid robots are put into new real scenarios to execute tasks, the differences between the real scenario and the simulation training scenario lead to a reduction in the effect of the motion strategy, thereby reducing the performance of the humanoid robot in the real scenario. Summary of the Invention
[0005] The present application discloses a method and device for training the movement of a humanoid robot based on a real scenario, which is used to improve the performance of the humanoid robot in the real scenario.
[0006] The specific task is to design a technical solution for training the motion control of a humanoid robot that can solve the sim2real gap. Technically, it is required to use a world model and online learning to quickly fine-tune the motion control strategy on a real machine, so that the motion control of the robot in the real scenario is more stable and meets the requirements of the actual business scenario.
[0007] The first aspect of the present application discloses a control method for a humanoid robot to walk with straight knees, including: Construct a world model for training the humanoid robot, and the world model is used to simulate the physical characteristics of the real world and the motion laws of the humanoid robot; Generate an initial motion control strategy for the humanoid robot according to the target training task and the structural parameters of the humanoid robot; Train the initial motion control strategy of a humanoid robot in a simulation environment; Cause the humanoid robot to enter the real world to perform real-world scenario motions; Collect real-world scenario motion data of the humanoid robot in real time during real-world scenario motions through sensors; Input the real-world scenario motion data and the real-time motion actions determined by the initial motion control strategy into the world model for interactive training in real time to generate the simulated state of the robot at the next moment; Perform error analysis on the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real-world scenario motion data to generate a total prediction error; Update the initial motion control strategy according to the total prediction error.
[0008] Optionally, the steps of training the initial motion control strategy of a humanoid robot in a simulation environment include: Obtain the state information of the humanoid robot at the first moment and the operating environment information at the first moment; Determine the control instruction for the next moment from the initial motion control strategy according to the state information at the first moment and the operating environment information at the first moment; Control the humanoid robot to move through the control instruction; Generate motion evaluation information for the humanoid robot after the motion is completed; Adjust the initial motion control strategy according to the motion evaluation information and a preset reward function.
[0009] Optionally, the reward function includes a navigation task reward function and a balance posture control task reward function; The steps of adjusting the initial motion control strategy according to the motion evaluation information and a preset reward function include: Calculate the navigation task reward according to the distance data, obstacle handling data, and action stability data in the motion evaluation information and the preset navigation task reward function; Calculate the balance posture control task reward according to the vertical posture data, angular velocity data, and action stability data in the motion evaluation information and the preset balance posture control task reward function; Optimize the initial motion control strategy according to the navigation task reward and the balance posture control task reward.
[0010] Optionally, the steps of performing error analysis on the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real-world scenario motion data to generate a total prediction error include: Determine the speed data of each key joint and the pose data of each end of the skeleton from the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real-world scenario motion data; Generate joint error weights for each key joint of the humanoid robot according to the target training task; Determine the adjacent association weights of each key joint according to the real-time motion actions; Generate key joint data weights for the speed data of each key joint according to the joint error weights and the adjacent association weights; According to the target training task, determine the pose weights of the end of each skeleton of the humanoid robot, and bind the pose weights to the corresponding skeleton end pose data of each skeleton end; Perform error analysis based on the key joint data weights, pose weights, the simulated state of the robot at the next moment, and the real state of the robot at the corresponding moment in the real scene motion data to generate the total prediction error.
[0011] Optionally, the steps of updating the initial motion control strategy according to the total prediction error include: Determine the key joint errors from the total prediction error; Update the corresponding strategy weights of the key joints in the initial motion control strategy according to the key joint errors; Update all the strategy weights in the initial motion control strategy according to the total prediction error.
[0012] Optionally, the steps of constructing a world model for training a humanoid robot include: Collect the empirical motion data of the humanoid robot in the simulation world and the real world; Construct the empirical motion data for training the world model of the humanoid robot.
[0013] Optionally, after the step of collecting the real scene motion data of the humanoid robot in the real world scene in real time through the sensor, and before the step of inputting the real scene motion data into the world model in real time for interactive training to generate the simulated state of the robot at the next moment, the humanoid robot motion training method further includes: Denoise the collected real scene motion data; Map the numerical values of the real scene motion data collected by different sensors to a unified interval.
[0014] The second aspect of the present application discloses a humanoid robot motion training device based on a real scene, including: A construction unit for constructing a world model for training a humanoid robot, and the world model is used to simulate the physical characteristics of the real world and the motion laws of the humanoid robot; A first generation unit for generating an initial motion control strategy for the humanoid robot according to the target training task and the structural parameters of the humanoid robot; A training unit for training an initial motion control strategy of a humanoid robot in a simulation environment; An execution unit for making the humanoid robot enter the real world to execute real-world scenario motions; An acquisition unit for collecting real-scenario motion data of the humanoid robot in real time during the real-world scenario motion through sensors; A second generation unit for interactively training the real-scenario motion data and the real-time motion actions determined by the initial motion control strategy by inputting them into the world model in real time to generate the simulated state of the robot at the next moment; A third generation unit for performing error analysis on the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real-scenario motion data to generate a total prediction error; An update unit for updating the initial motion control strategy according to the total prediction error.
[0015] Optionally, the training unit includes: An acquisition module for acquiring the state information of the humanoid robot at the first moment and the operation environment information at the first moment; A determination module for determining the control instruction at the next moment from the initial motion control strategy according to the state information at the first moment and the operation environment information at the first moment; Control the humanoid robot to move through the control instruction; A generation module for generating motion evaluation information for the humanoid robot after the motion is completed; An adjustment module for adjusting the initial motion control strategy according to the motion evaluation information and a preset reward function.
[0016] Optionally, the reward function includes a navigation task reward function and a balance attitude control task reward function; The adjustment module includes: Calculate the navigation task reward according to the distance data, obstacle handling data, and action stability data in the motion evaluation information and the preset navigation task reward function; Calculate the balance attitude control task reward according to the vertical pose data, angular velocity data, and action stability data in the motion evaluation information and the preset balance attitude control task reward function; Optimize the initial motion control strategy according to the navigation task reward and the balance attitude control task reward.
[0017] Optionally, the third generation unit includes: Determine the speed data of each key joint and the pose data of the end of each skeleton from the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real-scenario motion data; Generate joint error weights for each key joint of the humanoid robot according to the target training task; Determine the adjacent association weights of each key joint according to the real-time motion actions; Generate key joint data weights for the velocity data of each key joint according to the joint error weights and the adjacent association weights; According to the pose weights of the end of each skeleton of the humanoid robot for the target training task, bind the pose weights to the corresponding skeleton end pose data of each skeleton end; Perform error analysis based on the key joint data weights, pose weights, the simulated state of the robot at the next moment, and the real state of the robot at the corresponding moment in the real-scene motion data to generate the total prediction error.
[0018] Optionally, the update unit includes: Determine the key joint errors from the total prediction error; Update the corresponding policy weights of the key joints in the initial motion control strategy according to the key joint errors; Update all the policy weights in the initial motion control strategy according to the total prediction error.
[0019] Optionally, the construction unit includes: Collect the empirical motion data of the humanoid robot in the simulation world and the real world; Construct the empirical motion data for training the world model of the humanoid robot.
[0020] Optionally, after the acquisition unit and before the second generation unit, the humanoid robot motion training device further includes: A denoising unit for denoising the collected real-scene motion data; A mapping unit for mapping the values of the real-scene motion data collected by different sensors into a unified interval.
[0021] The third aspect of the present application provides a humanoid robot motion training device based on a real scene, including: A processor, a memory, an input / output unit, and a bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, and the processor calls the program to execute the humanoid robot motion training method as described in the first aspect and any optional aspects of the first aspect.
[0022] The fourth aspect of the present application provides a computer-readable storage medium, on which a program is stored, and when the program is executed on a computer, it executes the humanoid robot motion training method as described in the first aspect and any optional aspects of the first aspect.
[0023] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages: In the present application, first, a world model for training a humanoid robot is constructed. The world model is used to simulate the physical characteristics of the real world and the motion laws of the humanoid robot. An initial motion control strategy is generated for the humanoid robot according to the target training task and the structural parameters of the humanoid robot. The initial motion control strategy of the humanoid robot is trained in a simulated environment. The humanoid robot is made to enter the real world to execute real-world scenario motions. Real-world scenario motion data of the humanoid robot during the real-world scenario motion is collected in real time through sensors. The real-world scenario motion data and the real-time motion actions determined by the initial motion control strategy are input into the world model in real time for interactive training to generate the simulated state of the robot at the next moment. An error analysis is performed on the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real-world scenario motion data to generate the total prediction error. The initial motion control strategy is updated according to the total prediction error.
[0024] The business requirements of the real scenario are realized by constructing a world model and using the method of online learning. First, a world model that can accurately simulate the physical characteristics of the real world and the motion laws of the robot is constructed, which is the basic support for the entire solution. In the simulated environment, the initial motion control strategy of the humanoid robot is trained based on the world model. Then, when the humanoid robot enters the real scenario, the online learning means are used to collect the motion data of the robot in the real environment in real time. These data are fed back to the world model to quickly fine-tune the initial strategy. The world model and online learning cooperate with each other to continuously optimize the strategy, enabling the humanoid robot to adapt to various changes in the real scenario, thereby realizing the stable operation of the motion control strategy in the real scenario and improving the performance of the humanoid robot in the real scenario operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a schematic diagram of an embodiment of the method for training the motion of a humanoid robot based on a real scenario in the present application; Figure 2 It is a schematic diagram of an embodiment of the method for training the initial motion control strategy of a humanoid robot in the present application; Figure 3 It is a schematic diagram of an embodiment of the method for adjusting the initial motion control strategy in the present application; Figure 4Schematic diagram of an embodiment of the method for generating the total prediction error for this application; Figure 5 Schematic diagram of an embodiment of the method for updating the initial motion control strategy for this application; Figure 6 Schematic diagram of an embodiment of the method for constructing a world model for training a humanoid robot for this application; Figure 7 Schematic diagram of an embodiment of the method for processing real - world scenario motion data for this application; Figure 8 Schematic diagram of an embodiment of the real - world - based humanoid robot motion training device for this application; Figure 9 Schematic diagram of another embodiment of the real - world - based humanoid robot motion training device for this application; Figure 10 Schematic diagram of the overall framework process for training a humanoid robot's motion based on a real - world scenario for this application; Figure 11 Schematic diagram of the framework process for fine - tuning and learning of the humanoid robot of this application in the real world. Detailed implementation manners
[0027] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well - known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.
[0028] It should be understood that when used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0029] It should also be understood that the term " / and" as used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0030] As used in the specification of this application and the appended claims, the term "if" may be construed as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be construed as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.
[0031] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0032] The reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0033] In the prior art, the main problem currently faced is that there is a large gap between the simulation training environment (simulation environment) and the real scenario, that is, the sim2real gap problem. For example, in a rescue scenario, a humanoid robot needs to move flexibly in a disaster scene full of uncertainties, avoid obstacles and accurately reach the target position. However, for existing robots based on simulation training, after entering the real scenario, due to the differences between the environment and the simulation training environment, the motion control strategies of humanoid robots are often less effective or even ineffective, and they cannot stably execute tasks. Secondly, in an industrial production scenario, a humanoid robot needs to accurately complete operations such as assembling parts. The strategies trained in the simulation environment are difficult to ensure stable motion accuracy in a real production line due to factors such as equipment vibration and light changes.
[0034] In summary, when the current humanoid robots are put into new real scenarios to execute tasks, the differences between the real scenario and the simulation training scenario lead to a reduction in the effectiveness of the motion strategy, and thus reduce the performance of humanoid robots in the real scenario.
[0035] Based on this, this application discloses a method and device for training the motion of a humanoid robot based on a real scenario, which is used to improve the performance of the humanoid robot in the real scenario.
[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0037] The method of the present application can be applied to a server, a device, a terminal, or other devices with logical processing capabilities. In this regard, the present application makes no limitation. For the convenience of description, the following will take the execution entity as a terminal as an example for description.
[0038] Please refer to Figure 1 , an embodiment of a method for training a humanoid robot's motion based on a real scenario is provided in the present application, including: 101. Construct a world model for training the humanoid robot, where the world model is used to simulate the physical characteristics of the real world and the motion laws of the humanoid robot; In this embodiment, the terminal first constructs a world model according to the humanoid robot and the tasks that the humanoid robot needs to execute, and generates a world model for simulating the physical characteristics of the real world and the motion laws of the humanoid robot. Specifically, the terminal first collects simulation data, collects various data related to the motion of the humanoid robot in the simulation environment, including the motion states of the robot under different scenario settings, environmental parameters and other information, providing a data basis for constructing the world model. Then, the world model is constructed, and machine learning algorithms are used to analyze and process the collected simulation data to construct a world model that can simulate the physical characteristics of the real world and the motion laws of the robot.
[0039] 102. Generate an initial motion control strategy for the humanoid robot according to the target training task and the structural parameters of the humanoid robot; 103. Train the initial motion control strategy of the humanoid robot in the simulation environment; In this embodiment, the terminal generates an initial motion control strategy for the humanoid robot according to the target training task and the structural parameters of the humanoid robot. Then, the initial strategy is trained in the simulation environment based on the world model. The terminal uses the constructed world model to train the initial motion control strategy of the humanoid robot in the simulation environment with a suitable algorithm, enabling the humanoid robot to learn a preliminary effective motion method.
[0040] 104. Let the humanoid robot enter the real world to execute the real world scenario motion; The humanoid robot enters the real world: Deploy the robot trained in the simulation environment to the real scenario to execute tasks.
[0041] 105. Real-time scene motion data of the humanoid robot during its movement in the real-world scenario is collected in real time through sensors; In this embodiment, the terminal collects real-time scene motion data of the humanoid robot during its movement in the real-world scenario. Specifically, the sensors collect real-time motion data, and the humanoid robot collects motion data in real time through the sensors equipped on itself, such as information on joint angles, positions, speeds, etc.
[0042] The real-time scene motion data is the current robot state st: This is a multi-dimensional vector that contains various information describing the state of the robot at time t. The specific content includes joint angles, joint speeds, pose data of the robot, and humanoid robot sensor data or environmental sensor data.
[0043] Joint angle: The degree of bending of each joint, reflecting the posture of the robot's limbs.
[0044] Joint speed: The rate of change of the joint angle, reflecting the speed of the robot's limb movement.
[0045] Position and pose of the robot: Information such as coordinates and orientation in three-dimensional space.
[0046] Sensor data: Such as the force condition detected by the force sensor, environmental image features obtained by the vision sensor, etc. (if there are relevant sensors).
[0047] The current action at is the real-time motion action determined by the initial motion control strategy, representing the action instruction executed by the robot at time t, and is usually also a vector. The real-time motion action includes torque instructions, speed instructions, etc. of the joints, such as the force to be applied to each joint or the motion speed to be achieved.
[0048] 106. The real-time scene motion data and the real-time motion action determined by the initial motion control strategy are input into the world model for interactive training in real time to generate the simulated state of the robot at the next moment; 107. Error analysis is performed on the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real-time scene motion data to generate the total prediction error; The terminal inputs the real-time scene motion data and the real-time motion action determined by the initial motion control strategy into the world model for interactive training in real time to generate the simulated state of the robot at the next moment.
[0049] Specifically, the humanoid robot state s^t+1 predicted by the terminal (the robot simulation state) is also a multi-dimensional vector, with the same dimension as the real-scene motion data st, and contains the state information of the robot predicted by the world model at time t+1, such as the predicted joint angles, velocities, positions, and postures. This prediction result will be compared with the real s^t+1 to calculate the corresponding error loss, so as to optimize the parameters of the world model and the initial motion control strategy.
[0050] In this embodiment, the terminal performs error analysis on the robot simulation state at the next moment and the real state of the robot at the corresponding moment in the real-scene motion data to generate the total prediction error. Specifically, it is necessary to compare the collected data with the prediction result of the world model to calculate the error, input the collected real motion data into the world model, compare it with the prediction result of the model, and calculate the error between the two.
[0051] Specifically, in this embodiment, in the world model, the main goal is to predict the robot state s^t+1 at the next moment based on the current robot state st and the current action at. Suppose there are N samples in total, and each sample contains the current state, action, and the corresponding real state at the next moment.
[0052] The MSE loss function L_MSE is used to measure the error between the predicted next moment state s^t+1 and the real next moment state st+1, and its formula is:
[0053] Among them, N is the number of samples, that is, the number of state-action pairs at different moments used in the training process. d is the dimension of the state vector of the humanoid robot. The humanoid robot state includes information such as joint angles, velocities, and positions, and d is the length of the vector composed of this information. st+1,ij represents the j-th element of the real next moment state vector in the i-th sample. s^t+1,ij represents the j-th element of the next moment state vector predicted by the model in the i-th sample.
[0054] When the humanoid robot enters the real world to perform tasks, its motion data, including joint states, positions, velocities, etc., is collected in real time through the sensors on the humanoid robot and the authenticity of the environment.
[0055] That is, after the humanoid robot enters the real world to perform real tasks, various sensors carried by it (such as joint position sensors, accelerometers, gyroscopes, vision sensors, etc.) start to work. These sensors collect the motion data of the humanoid robot in real time, including joint angles, angular velocities, linear velocities, position coordinates, attitude information, and surrounding environment information, etc. For example, in a complex indoor environment, the vision sensor collects information such as the room layout and the position of obstacles, and the joint position sensor records the real-time angles of each joint of the robot. Finally, these data are input into the stored world model, compared with the prediction results of the world model, and the error is calculated. Based on this error, using the online learning algorithm and stochastic gradient descent, the initial motion control strategy is quickly fine-tuned. This process also needs to be continuously repeated at all times during the on-site task execution. As the humanoid robot continuously collects on-site data in the real world, each piece of data will optimize the motion control strategy of the humanoid robot, gradually improving its stability, which enhances the usage effect in the real scenario.
[0056] The loss function of the world model is used for error calculation here. The difference is that the true value is taken from the simulation and becomes the true value in the real world.
[0057] The preprocessed data is input into the constructed world model. The world model predicts the state of the humanoid robot at the next moment based on the current input state and actions of the humanoid robot. The prediction result is compared with the real state at the next moment collected by the sensor, and the error between the two is calculated. The error calculation method used in this embodiment is the mean square error. For example, if the world model predicts a joint angle of 30 degrees, while the actual angle measured by the sensor is 32 degrees, the error of the joint angle prediction is calculated through the MSE formula, and the world model is trained using gradient descent.
[0058] In addition to the above mean square error method, a method of task objective weighted mean square error is also created, which will be described in detail in subsequent embodiments. The method of task objective weighted mean square error is used to analyze the error by combining the important joints and specific end poses of the humanoid robot, so as to accurately calculate the error between the real action and the predicted action of the humanoid robot in a specific task objective.
[0059] 108. Update the initial motion control strategy according to the total prediction error.
[0060] Finally, the terminal uses the online learning algorithm and the total prediction error to fine-tune the initial motion control strategy. Based on the calculated total prediction error, and using the online learning algorithm to quickly adjust and optimize the initial motion control strategy. The optimized initial motion control strategy is fed back to the humanoid robot to make its motion control in the real world more stable, and this process loops continuously to continuously improve the motion performance of the humanoid robot in the real world.
[0061] Please refer to Figure 10 and Figure 11 , Figure 10 is a schematic diagram of the overall framework process for training the motion of a humanoid robot based on a real scenario, Figure 11 is a schematic diagram of the framework process for fine-tuning and learning of a humanoid robot in the real world. The two figures respectively introduce the entire training scheme and the real-world fine-tuning and learning scheme. In this embodiment, the terminal uses online learning algorithms (stochastic gradient descent SGD and its variants Adagrad, Adadelta). First, the initial motion control strategy is adjusted according to the calculated error. Taking SGD as an example, the gradient of the error with respect to the parameters of the policy network is calculated, and then the parameters of the policy network are updated in the opposite direction of the gradient, so that the policy network can be closer to the real situation in subsequent predictions. If the gradient of a weight parameter w in the initial policy network is g and the learning rate is alpha, then the updated weight is (w' = w - alpha * g).
[0062] Next, the terminal performs cyclic optimization in the real-world scenario, that is, continuously repeats the above steps of data collection, preprocessing, interaction with the world model, and policy update. As the running time of the humanoid robot in the real world increases, more and more data is accumulated, the policy network is continuously optimized, gradually adapts to the real-world environment, and the motion control ability is improved.
[0063] In this embodiment, first, a world model for training a humanoid robot is constructed. The world model is used to simulate the physical characteristics of the real world and the motion laws of the humanoid robot. An initial motion control strategy is generated for the humanoid robot according to the target training task and the structural parameters of the humanoid robot. The initial motion control strategy of the humanoid robot is trained in a simulated environment. The humanoid robot is made to enter the real world to execute real-world scenario motions. Real-world scenario motion data of the humanoid robot during the real-world scenario motion process is collected in real time through sensors. The real-world scenario motion data and the real-time motion actions determined by the initial motion control strategy are input into the world model for interactive training in real time to generate the simulated state of the robot at the next moment. The error analysis is performed between the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real-world scenario motion data to generate the total prediction error. The initial motion control strategy is updated according to the total prediction error.
[0064] The real - world business requirements are realized by constructing a world model and using online learning. First, a world model that can accurately simulate the physical characteristics of the real world and the motion laws of the robot is constructed, which is the basic support for the whole solution. In the simulation environment, the initial motion control strategy of the humanoid robot is trained based on the world model. Then, when the humanoid robot enters the real scenario, online learning means are used to collect the motion data of the robot in the real environment in real - time. These data are fed back to the world model to quickly fine - tune the initial strategy. The world model and online learning cooperate with each other to continuously optimize the strategy, enabling the humanoid robot to adapt to various changes in the real scenario, thus realizing the stable operation of the motion control strategy in the real scenario and improving the performance of the humanoid robot in the real - scenario operation.
[0065] Secondly, the motion control training of the real - world humanoid robot is realized with the help of the world model and online learning. The main function of this training method is to make the motion control of the humanoid robot more stable in the real environment and effectively solve the problem of the gap between simulation and reality. For example, in an industrial production scenario, after the robot is trained in the simulation environment and enters the real production line, it should be able to stably perform tasks such as grasping and assembling without making action mistakes due to environmental differences. It constructs a world model to simulate the physical characteristics of the real world and the motion laws of the robot. Online learning uses the data collected in real - time by the robot in the real environment to quickly fine - tune the strategy trained based on the world model. In this way, the robot can quickly adapt to the real environment, making the motion control strategy more stable and reliable and accurately completing various tasks.
[0066] Secondly, this embodiment also has the beneficial effect of quick fine - tuning, efficiently solving the sim2real gap: through online learning, the initial strategy trained based on the world model is quickly fine - tuned using the data collected in the real world, greatly reducing the gap between simulation and reality and enabling the robot to quickly adapt to the real environment.
[0067] Secondly, this embodiment also has the beneficial effect of enhancing the stability of motion control. By adjusting the strategy based on the world - model prediction and real - time data feedback, the stability of the motion control of the humanoid robot in the real scenario is effectively improved, the action error rate is reduced, and the accuracy of task execution is guaranteed.
[0068] Secondly, in this embodiment, relying on simulation to construct the world model has a lower cost. Only simulation data is needed to construct the world model, avoiding the high cost and potential risks of large - scale data collection in the real scenario. At the same time, by using the repeatability and controllability of the simulation environment, the model training efficiency is improved.
[0069] Secondly, in this embodiment, real - device iteration optimization is also used. During the operation of the real device, data is continuously collected, and the motion control strategy is continuously iteratively optimized. As time goes by, the motion control performance of the robot gets better and better, and it can better adapt to complex and changeable real scenarios.
[0070] Please refer to Figure 2 , an embodiment of a method for training an initial motion control strategy of a humanoid robot is provided in this application, including: 201. Obtain the state information of the humanoid robot at the first moment and the operation environment information at the first moment; 202. Determine the control instruction for the next moment from the initial motion control strategy according to the state information at the first moment and the operation environment information at the first moment; 203. Control the humanoid robot to move through the control instruction; 204. Generate motion evaluation information for the humanoid robot after the motion is completed; 205. Adjust the initial motion control strategy according to the motion evaluation information and the preset reward function.
[0071] In this embodiment, in the simulated operation environment, with the help of the constructed world model, the initial motion control strategy of the humanoid robot is trained using the Proximal Policy Optimization (PPO) reinforcement learning algorithm. Taking the humanoid robot to complete a specific task (such as walking to a specified location in a simulated obstacle environment) as the goal, a corresponding reward mechanism (reward function) is set. Then, according to the results predicted by the world model, the humanoid robot continuously adjusts its own actions and gradually learns an effective motion strategy.
[0072] In this embodiment, the terminal obtains the state information of the humanoid robot at the first moment and the operation environment information at the first moment. The operation environment information at the first moment is the input (obs - observation value), specifically mainly including the current state information of the humanoid robot. Specifically, it covers the angles, angular velocities, and angular accelerations of each joint of the humanoid robot, which are used to accurately describe the posture and motion trend of the humanoid robot's limbs. It also includes the position coordinates of the humanoid robot in space (such as the x, y, z coordinates in the Cartesian coordinate system) and the posture (such as the orientation represented by Euler angles), which determine the three - dimensional position and direction of the humanoid robot in the virtual environment. In addition, it also includes the data collected by various sensors, such as the ground reaction force sensed by the pressure sensor, the contact situation with surrounding objects detected by the collision sensor, and the environmental feature information (such as the position, shape, distance, etc. of obstacles) extracted by the vision sensor. These information combined provide a comprehensive description of the environment where the robot is currently located and its own state for the reinforcement learning algorithm, enabling the algorithm to make appropriate decisions based on this.
[0073] Based on the state information at the first moment and the operating environment information at the first moment, the terminal determines the control instruction for the next moment from the initial motion control strategy. The output of the initial motion control strategy is a motion instruction (action), which is used to drive the motion of the humanoid robot. This includes control signals for each joint, specifying the target angle, target angular velocity, or target torque of each joint, so as to precisely adjust the actions of the robot's limbs. At the same time, it also involves the overall motion instruction of the humanoid robot, such as setting the speed and direction of forward, backward, and turning, enabling the humanoid robot to move as expected in the environment and complete specific tasks.
[0074] After obtaining the motion instruction, the terminal controls the humanoid robot to move through the control instruction, and then generates motion evaluation information for the humanoid robot after the motion is completed. The motion evaluation information refers to the respective motion data of the humanoid robot collected from the start to the end of this motion instruction. Finally, the terminal adjusts the initial motion control strategy according to the motion evaluation information and the preset reward function to achieve the effect of adjusting the initial motion control strategy.
[0075] Please refer to Figure 3 , an embodiment of a method for adjusting an initial motion control strategy provided by this application includes: 301. Calculate the navigation task reward according to the distance data, obstacle handling data, and action stability data in the motion evaluation information and the preset navigation task reward function; 302. Calculate the balance attitude control task reward according to the vertical pose data, angular velocity data, and action stability data in the motion evaluation information and the preset balance attitude control task reward function; 303. Optimize the initial motion control strategy according to the navigation task reward and the balance attitude control task reward.
[0076] The reward function is the key to guiding the humanoid robot to learn effective strategies. In the initial stage of training, positive rewards are given for positive behaviors such as the humanoid robot moving towards the target position, avoiding collisions with obstacles, and maintaining body balance. In this embodiment, every time the humanoid robot approaches the target by a certain distance, the reward value increases. If no collision occurs within a certain period of time, an additional reward is given. Successfully maintaining balance in complex terrains also receives corresponding rewards. On the contrary, negative rewards are given for negative behaviors such as deviating from the target, colliding, and losing balance. For example, every time the robot moves away from the target by a certain distance, the reward value decreases; a large amount of rewards are deducted for a collision; a serious negative reward is given for losing balance and falling. By reasonably designing the reward function, the robot can gradually learn optimized motion strategies through continuous attempts and develop in the direction of completing tasks and obtaining the maximum reward. The reward function in this embodiment is mainly based on maintaining balance while performing the navigation task, so it is set as the navigation task reward and the balance attitude control task reward.
[0077] In this embodiment, the terminal first calculates the navigation task reward according to the distance data, obstacle handling data, movement stability data in the motion evaluation information and a preset navigation task reward function. Specifically, in the navigation task, the goal is to move the robot from the starting point to the specified target point. The reward function should encourage the robot to approach the target while avoiding collisions with obstacles.
[0078]
[0079] Among them, R is the reward value of the navigation task reward function. dt is the distance between the robot and the target point at time t. is a very small positive number to prevent the denominator from being zero, usually taking a value of 10^-6. collision is the collision flag, which is 1 when a collision occurs and 0 otherwise. fall is the fall flag, which is 1 when the humanoid robot falls and 0 otherwise. 、 、 are weight coefficients used to adjust the importance of each item in the reward function. For example, it can be set that = 10, = 20, = 50.
[0080] Balance and attitude control task reward. In the balance and attitude control task, the focus is on keeping the robot in a good attitude and balance.
[0081]
[0082] Among them: Here, R is the reward value of the balance attitude control task reward function, is the angle between the humanoid robot's body and the vertical direction. The smaller the angle, the more balanced the attitude. is the angular velocity of the robot's body. The smaller the angular velocity, the more stable it is. 、 、 are weight coefficients. For example, in this embodiment, = 15, = 5, = 30.
[0083] Please refer to Figure 4 , this application provides an embodiment of a method for generating a total prediction error, including: 401. Determine the speed data of each key joint and the pose data of each end of the skeleton from the robot simulation state at the next moment and the real state of the robot at the corresponding moment in the real scene motion data; In this embodiment, the speed data of each key joint is determined from the robot simulation state at the next moment, and the speed data of each key joint in the real-scene motion data at the next moment. The two sets of speed data are in one-to-one correspondence through the key joints. Next, the pose data of each skeleton end is determined from the robot simulation state at the next moment, and the pose data of each skeleton end in the real-scene motion data at the next moment.
[0084] 402. Generate joint error weights for each key joint of the humanoid robot according to the target training task; The terminal determines the running accuracy importance for each key joint at the next moment according to different training tasks of the humanoid robot. Since there are multiple core actions in different target training tasks, and each action has different functions, this makes the importance of key joints different in each action, and the importance of the key joints used in each key action is also different. For example, several key joints that are core in the squatting action and several key joints that are core in the lifting action. In different actions, the same key joint plays different roles, so the possible deformation amount brought by this key joint is also different, resulting in different importance of the key joint in different actions. The joint error weight is a value greater than or equal to 1.
[0085] 403. Determine the adjacent association weight of each key joint according to the real-time motion action; In this embodiment, since different key joints have different actions in one action, the actions of some key joints can affect several adjacent key joints. This makes the deformation of the action on this key joint drive the deformation of other adjacent key joints. At this time, we give a part of these key joints an adjacent association weight, so that when this key joint deforms, its corresponding error can be increased through this association weight, thereby increasing the total error. The adjacent association weight is a value greater than 1 or equal to 1.
[0086] 404. Generate key joint data weights for the speed data of each key joint according to the joint error weights and the adjacent association weights; In this embodiment, the terminal generates key joint data weights for the velocity data of each key joint according to the joint error weight and the adjacent connection weight. In this embodiment, the joint error weight and the adjacent connection weight are multiplied to obtain the key joint data weights. The key joint data weights are inserted into the MSE loss function in step 107. After taking the difference between each state in the robot simulation state and the true state of the robot at the corresponding moment in the real-scene motion data and then squaring it, if it belongs to a key joint with key joint data weights, it needs to be multiplied by the key joint data weights and then divided by N after squaring. After superimposing all the errors, the errors of specific key joints are amplified.
[0087] 405. According to the pose weights of the end of each skeleton of the humanoid robot for the target training task, bind the pose weights to the corresponding skeleton end pose data of each skeleton end; In this embodiment, the terminal binds the pose weights of the end of each skeleton of the humanoid robot to the corresponding skeleton end pose data according to the target training task. Since the pose of the skeleton end represents the overall completion degree of the action, and the end of the humanoid robot has different pose accuracy requirements according to different tasks, in a specific target training task, when different skeleton ends complete the action, the required accuracies are different, such as grasping actions and gesture actions. The pose weights of each skeleton end, and the pose weights of the skeleton end are a value greater than 1. Subsequently, they are inserted into the MSE loss function in step 107, and the method is the same as that of the key joint data weights in step 404, which will not be elaborated here.
[0088] 406. Perform error analysis according to the key joint data weights, pose weights, the robot simulation state at the next moment, and the true state of the robot at the corresponding moment in the real-scene motion data to generate the total prediction error.
[0089] Finally, the terminal performs error analysis according to the key joint data weights, pose weights, the robot simulation state at the next moment, and the true state of the robot at the corresponding moment in the real-scene motion data to generate the total prediction error. In this embodiment, by setting corresponding deformation influence weights for different key joints and skeleton ends, if detectable deformation errors occur in the parameters of these positions during the movement process, these errors will be amplified according to the magnitude of their influence, increasing their importance and making them easier to notice. The generated total prediction error can increase the accuracy of subsequent adjustment of the initial motion control strategy.
[0090] Please refer to Figure 5 , this application provides an embodiment of a method for updating an initial motion control strategy, including: 501. Determine the key joint errors from the total prediction error; 502. Update the policy weights corresponding to the key joints in the initial motion control strategy according to the key joint error; 503. Update all the policy weights in the initial motion control strategy according to the total prediction error.
[0091] In this embodiment, the terminal not only needs to adjust the entire initial motion control strategy through the total prediction error, but also needs to separately adjust the control strategies of different key joints, making it more targeted during adjustment. Specifically, in this embodiment, after determining the key joint error from the total prediction error, determine each policy weight related to the key joint from the initial motion control strategy, and use the key joint error to update this part of the policy weights separately. Such an update strategy can specifically adjust the action strategies of specific key joints for the error of each action, increasing the accuracy of the initial motion control strategy.
[0092] Please refer to Figure 6 , an embodiment of a method for constructing a world model for training a humanoid robot provided by this application includes: 601. Collect the empirical motion data of the humanoid robot in the simulation world and the real world; 602. Construct a world model for training the humanoid robot from the empirical motion data.
[0093] In this embodiment, the terminal first collects a large amount of data related to the motion of the humanoid robot in the simulation world or the real world. Specifically, it is necessary to cover the joint angles, speeds, accelerations, and force conditions of the robot during motion under various environmental parameters such as task terrains, lighting conditions, and temperature environments, as well as the positions, shapes, and materials of surrounding obstacles and targets. Then, use machine learning algorithms and deep neural networks to analyze and process these collected data to construct a world model that can accurately simulate the physical characteristics of the real world and the motion laws of the robot. The function of this world model is to predict the operating state at the next moment based on the current state and action input of the humanoid robot.
[0094] Please refer to Figure 7 , an embodiment of a method for processing real-scene motion data provided by this application includes: 701. Denoise the collected real-scene motion data; 702. Map the numerical values of the real-scene motion data collected by different sensors to a unified interval.
[0095] The terminal processes the real - world motion data collected. Noise data is removed. Specifically, the terminal removes random noise during sensor measurement through a filtering algorithm (such as Kalman filtering), and then normalizes the data, mapping the values collected by different sensors to a unified interval, such as [0, 1] or [-1, 1], so that the neural network can better process the data and improve the training efficiency and stability.
[0096] Please refer to Figure 8 , this application provides an embodiment of a humanoid robot motion training device based on a real - world scenario, including: A construction unit 801, configured to construct a world model for training a humanoid robot, where the world model is used to simulate the physical characteristics of the real world and the motion laws of the humanoid robot; Optionally, the construction unit 801 includes: collecting the empirical motion data of the humanoid robot in the simulation world and the real world; constructing a world model for training the humanoid robot based on the empirical motion data.
[0097] A first generation unit 802, configured to generate an initial motion control strategy for the humanoid robot according to the target training task and the structural parameters of the humanoid robot; A training unit 803, configured to train the initial motion control strategy of the humanoid robot in a simulation environment; Optionally, the training unit 803 includes: An acquisition module, configured to acquire the state information of the humanoid robot at the first moment and the operation environment information at the first moment; A determination module, configured to determine the control instruction for the next moment from the initial motion control strategy according to the state information at the first moment and the operation environment information at the first moment; Control the humanoid robot to move through the control instruction; A generation module, configured to generate motion evaluation information for the humanoid robot after the motion is completed; An adjustment module, configured to adjust the initial motion control strategy according to the motion evaluation information and a preset reward function.
[0098] Optionally, the reward function includes a navigation task reward function and a balance attitude control task reward function; The adjustment module includes: calculating the navigation task reward according to the distance data, obstacle handling data, and action stability data in the motion evaluation information and the preset navigation task reward function; calculating the balance attitude control task reward according to the vertical pose data, angular velocity data, and action stability data in the motion evaluation information and the preset balance attitude control task reward function; optimizing the initial motion control strategy according to the navigation task reward and the balance attitude control task reward.
[0099] The execution unit 804 is configured to make the humanoid robot enter the real world to execute real-world scenario motions; The acquisition unit 805 is configured to collect real-scenario motion data of the humanoid robot in the process of real-world scenario motions in real time through sensors; The denoising unit 806 is configured to perform denoising processing on the collected real-scenario motion data; The mapping unit 807 is configured to map the numerical values of real-scenario motion data collected by different sensors into a unified interval; The second generation unit 808 is configured to interactively train the real-scenario motion data and the real-time motion actions determined by the initial motion control strategy in the world model in real time to generate the simulated state of the robot at the next moment; The third generation unit 809 is configured to perform error analysis on the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real-scenario motion data to generate the total prediction error; Optionally, the third generation unit 809 includes: determining the speed data of each key joint and the pose data of the end of each skeleton from the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real-scenario motion data; generating joint error weights for each key joint of the humanoid robot according to the target training task; determining the adjacent association weights of each key joint according to the real-time motion actions; generating key joint data weights for the speed data of each key joint according to the joint error weights and the adjacent association weights; generating pose weights for the pose of the end of each skeleton of the humanoid robot according to the target training task, and binding the pose weights to the pose data of the end of each corresponding skeleton; performing error analysis according to the key joint data weights, the pose weights, the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real-scenario motion data to generate the total prediction error.
[0100] The update unit 810 is configured to update the initial motion control strategy according to the total prediction error.
[0101] Optionally, the update unit 810 includes: determining the key joint error from the total prediction error; updating the strategy weights corresponding to the key joints in the initial motion control strategy according to the key joint error; updating all the strategy weights in the initial motion control strategy according to the total prediction error.
[0102] Please refer to Figure 9 , this application provides a humanoid robot motion training device based on a real scenario, including: A processor 901, a memory 902, an input / output unit 903, and a bus 904.
[0103] The processor 901 is connected to the memory 902, the input / output unit 903, and the bus 904.
[0104] The memory 902 stores a program, and the processor 901 calls the program to execute the humanoid robot motion training method as Figure 1 , Figure 2 and Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 7 in.
[0105] This application provides a computer-readable storage medium, on which a program is stored. When the program is executed on a computer, it executes the humanoid robot motion training method as Figure 1 , Figure 2 and Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 7 in.
[0106] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0107] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0108] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0109] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-integrated units can be implemented in the form of hardware or in the form of software functional units.
[0110] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.
Claims
1. A humanoid robot motion training method based on real scenes, characterized in that: include: Constructing a world model for training a humanoid robot, wherein the world model is used to simulate real-world physical properties and motion laws of the humanoid robot; generating an initial motion control strategy for the humanoid robot according to a target training task and structural parameters of the humanoid robot; Training an initial motion control strategy of the humanoid robot in the simulation environment; Allowing the humanoid robot to enter the real world and perform real-world scene movements; Collecting real-world scene motion data of the humanoid robot in real time during the movement of the real-world scene through sensors; Inputting the real scene motion data and the real-time motion action determined by the initial motion control strategy into the world model in real time for interactive training to generate a robot simulation state at the next moment; Perform error analysis on the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real scene motion data to generate a total prediction error; The initial motion control strategy is updated according to the total prediction error.
2. The humanoid robot motion training method according to claim 1, characterized in that: The step of training the initial motion control strategy of the humanoid robot in a simulated environment comprises: Acquire the state information and operating environment information of the humanoid robot at the first moment; Determining a control instruction at a next moment from the initial motion control strategy according to the state information at the first moment and the operating environment information at the first moment; Controlling the humanoid robot to move by means of the control instructions; generating movement evaluation information for the humanoid robot after the movement is completed; The initial motion control strategy is adjusted according to the motion evaluation information and a preset reward function.
3. The humanoid robot motion training method according to claim 2, characterized in that: The reward function includes a navigation task reward function and a balance posture control task reward function; The step of adjusting the initial motion control strategy according to the motion evaluation information and the preset reward function comprises: Calculate the navigation task reward based on the distance data, obstacle handling data, action stability data and the preset navigation task reward function in the motion evaluation information; Calculate the balance posture control task reward according to the vertical posture data, angular velocity data, action stability data and the preset balance posture control task reward function in the motion evaluation information; The initial motion control strategy is optimized according to the navigation task reward and the balance posture control task reward.
4. The humanoid robot motion training method according to claim 1, characterized in that: The step of performing error analysis on the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real scene motion data to generate a total prediction error comprises: Determine the speed data of each key joint and the position data of each skeleton end from the robot simulation state at the next moment and the real state of the robot at the corresponding moment in the real scene motion data; generating a joint error weight for each key joint of the humanoid robot according to a target training task; Determine the adjacent association weight of each key joint according to the real-time motion action; Generate a key joint data weight for the velocity data of each key joint according to the joint error weight and the adjacent association weight; Binding the pose weights to the pose data of the skeleton end corresponding to each skeleton end according to the pose weights of each skeleton end of the humanoid robot as a target training task; An error analysis is performed based on the key joint data weights, the posture weights, the robot simulation state at the next moment, and the real state of the robot at the corresponding moment in the real scene motion data to generate a total prediction error.
5. The humanoid robot movement training method according to claim 4, characterized in that: The step of updating the initial motion control strategy according to the total prediction error comprises: determining a critical joint error from the total prediction error; Updating the strategy weights corresponding to the key joints in the initial motion control strategy according to the key joint errors; All strategy weights in the initial motion control strategy are updated according to the total prediction error.
6. The humanoid robot movement training method according to any one of claims 1 to 5, characterized in that: The step of constructing a world model for training a humanoid robot comprises: Collect empirical motion data of humanoid robots in simulation and the real world; The empirical motion data is used to construct a world model for training a humanoid robot.
7. The humanoid robot movement training method according to any one of claims 1 to 5, characterized in that: After the step of collecting the real scene motion data of the humanoid robot in the process of moving in the real world scene in real time through the sensor, and before the step of inputting the real scene motion data into the world model in real time for interactive training to generate the robot simulation state at the next moment, the humanoid robot motion training method further includes: Performing denoising processing on the collected real scene motion data; The numerical values of the real scene motion data collected by different sensors are mapped into a unified interval.
8. A humanoid robot motion training device based on real scenes, characterized in that: include: A construction unit, used to construct a world model for training a humanoid robot, wherein the world model is used to simulate the physical properties of the real world and the motion laws of the humanoid robot; A first generating unit, configured to generate an initial motion control strategy for the humanoid robot according to a target training task and structural parameters of the humanoid robot; A training unit, used for training an initial motion control strategy of the humanoid robot in the simulation environment; An execution unit, used to enable the humanoid robot to enter the real world and execute real-world scene movements; A collection unit, used for collecting real scene motion data of the humanoid robot in the process of moving in the real world scene in real time through a sensor; A second generating unit is used to input the real scene motion data and the real-time motion action determined by the initial motion control strategy into the world model in real time for interactive training, so as to generate a robot simulation state at the next moment; A third generating unit is used to perform error analysis between the simulated state of the robot at the next moment and the real state of the robot at the corresponding moment in the real scene motion data to generate a total prediction error; An updating unit is used to update the initial motion control strategy according to the total prediction error.
9. The humanoid robot motion training device according to claim 8, characterized in that: The training unit comprises: An acquisition module, used to acquire the state information of the humanoid robot at the first moment and the operating environment information at the first moment; A determination module, configured to determine a control instruction at a next moment from the initial motion control strategy according to the state information at the first moment and the operating environment information at the first moment; Controlling the humanoid robot to move through the control instructions; A generating module, used for generating motion evaluation information for the humanoid robot after the motion is completed; The adjustment module is used to adjust the initial motion control strategy according to the motion evaluation information and a preset reward function.
10. A humanoid robot motion training device based on real scenes, characterized in that: include: A processor, a memory, an input-output unit and a bus, wherein the processor is connected to the memory, the input-output unit and the bus, the memory stores a program, and the processor calls the program to execute the humanoid robot motion training method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Mitigating reality gaps by training simulated-to-real models using vision-based robotic task models
CN114585487A
Sim2Real model construction method and device based on reinforcement learning
CN119669952A
Operation control model training method and robot operation control method and system
CN119748447A
Mitigating reality gap through feature-level domain adaptation in training of vision-based robot action model
US20230154160A1
Cited By
Humanoid robot acquisition, training and evaluation integrated method and system
CN121018562A
Humanoid robot control system and control method based on world knowledge enhancement and related equipment
CN121340269A
Multi-modal perception humanoid robot behavior test method and related equipment
CN121492069A
Migratable humanoid robot control method based on world model and related equipment
CN121946475A