Robot high jump platform motion control model training method and device

By adding noise to the observation data of the inertial measurement device in the robot high jump motion control and using deep reinforcement learning to train the control model, the control instability problems caused by measurement deviation and noise interference of the inertial measurement device are solved, and the motion stability and control reliability are improved.

CN120065752AActive Publication Date: 2025-05-30SHENZHEN ZHUJI POWER TECH CO LTD

Patent Information

Application Number
CN202510527225.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

In robot high jump platform motion control, measurement deviations and noise interference of inertial measurement devices lead to control instability, affecting the decision-making process and increasing the risk of falling.

Method used

By receiving the current status data of the virtual robot, adding noise to the observation data of the inertial measurement device, simulating measurement deviations and noise interference in the real environment, and training the robot's high jump platform motion control model based on deep reinforcement learning.

Benefits of technology

This method can reduce the risk of control instability caused by measurement errors and improve the stability and control reliability of the robot in the high jump platform movement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120065752A_ABST
    Figure CN120065752A_ABST
Patent Text Reader

Abstract

The invention provides a robot high jump platform motion control model training method and device, and relates to the technical field of sensors and robots. The method comprises the steps of obtaining current state data of a virtual robot, including observation data of an inertial measurement device, then adding noise into the observation data, and training a robot high jump platform motion control model in a high platform terrain based on deep reinforcement learning in combination with the current state data added with the noise. The diversity and robustness of state data are enhanced by introducing noise, so that the robot high jump platform motion control model obtained by training has the capability of adapting to complex terrains and sensor interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of sensors and robotics, and relates to a method and device for training a robot high-jump motion control model. Background Art

[0002] In the field of robot motion control, inertial measurement units (IMUs) are key state observation sensors and are widely used in robot attitude estimation, motion prediction, and control strategy optimization. However, in the actual environment, IMU observation data is often affected by shocks, vibrations, and external noise interference, resulting in a decrease in observation accuracy and affecting the stability of robot motion control. Especially in the scenario of a robot jumping over a high platform, measurement deviations of the IMU may cause the robot to be unable to accurately obtain its own state information when performing jumping or landing actions, thus affecting the decision-making process and even causing the robot to lose balance and fall.

[0003] In addition, when a robot continuously performs jumping tasks, the impact noise of the IMU during landing will gradually accumulate, causing the observation data to be used for the next control decision before it has returned to the normal state, further exacerbating the motion deviation and making it difficult for the robot to maintain a stable jumping rhythm.

[0004] In the prior art, robot motion control strategies are usually trained in a simulation environment and adapted to the real environment through transfer learning. However, the IMU data in the simulation environment is usually in an ideal state. Therefore, when the trained control strategy is applied to the real environment, IMU observation errors may cause the robot to be unable to correctly perceive the state deviation at the moment of landing, thereby leading to control instability and even causing the robot to fall. Summary of the Invention

[0005] The present disclosure provides a method and device for training a robot high-jump motion control model, which is used to reduce control instability caused by measurement deviations of the IMU in the scenario of a robot jumping over a high platform.

[0006] Additional aspects and advantages of the present disclosure will be partially described below, and will partially become apparent from the description, or can be learned through the practice of the present disclosure.

[0007] According to a first aspect of the present disclosure, there is provided a method for training a robot high-jump motion control model, including: Receiving current state data of a virtual robot, where the current state data includes observation data of an inertial measurement unit; Adding first noise to the observation data of the inertial measurement unit; Based on the current state data after adding the first noise, a high-platform terrain is combined to train a robot high-platform jumping motion control model based on deep reinforcement learning.

[0008] In an exemplary embodiment of the present disclosure, the first noise includes impact noise; adding the first noise to the observation data of the inertial measurement device includes: When it is determined according to the current state data of the virtual robot that the virtual robot lands, impact noise is added to the observation data of the inertial measurement device.

[0009] In an exemplary embodiment of the present disclosure, adding the first noise to the observation data of the inertial measurement device further includes: When it is determined according to the current state data of the virtual robot that the virtual robot is in continuous jumping motion, impact noise is added after each jumping action.

[0010] In an exemplary embodiment of the present disclosure, the first noise further includes vibration noise; adding the first noise to the observation data of the inertial measurement device further includes: When it is determined according to the current state data of the virtual robot that the virtual robot is in continuous jumping motion, vibration noise is added to the observation data of the inertial measurement device.

[0011] In an exemplary embodiment of the present disclosure, the first noise includes random Gaussian noise, and the mean and variance of the random Gaussian noise are determined according to the data characteristics of the real inertial measurement device.

[0012] In an exemplary embodiment of the present disclosure, the method further includes: Dynamically adjusting the intensity of the first noise according to the jumping height of the virtual robot.

[0013] In an exemplary embodiment of the present disclosure, the first noise includes impact noise; dynamically adjusting the intensity of the first noise according to the jumping height of the virtual robot includes: According to: Adjust the intensity of the first noise; wherein, is the impact noise at time t, is the impact indicator function, is a Gaussian distribution with a mean of 0, is the variance, is the basic noise variance, is the linear amplification coefficient of the noise intensity with respect to the jumping height, and h is the jumping height.

[0014] In an exemplary embodiment of the present disclosure, the first noise includes vibration noise; dynamically adjusting the intensity of the first noise according to the jumping height of the virtual robot includes: According to: Adjust the intensity of the first noise; Wherein, is the vibration noise at time t, is the amplitude of the vibration noise, is the basic vibration amplitude, is the linear coefficient of the vibration amplitude increasing with the jump height, h is the jump height, is the vibration frequency, is the phase offset.

[0015] In an exemplary embodiment of the present disclosure, the first noise includes random Gaussian noise; dynamically adjusting the intensity of the first noise according to the jump height of the virtual robot includes: According to: Adjust the intensity of the first noise; Wherein, is the current noise standard deviation calculated according to the jump height h, which is used to construct the Gaussian distribution subsequently, is the basic noise standard deviation, is the amplification ratio coefficient of the noise changing with the jump height.

[0016] In an exemplary embodiment of the present disclosure, the intensity range of the first noise is determined based on the measurement error statistical distribution of the real inertial measurement device at the moment when the robot lands.

[0017] In an exemplary embodiment of the present disclosure, adding the first noise to the observation data of the inertial measurement device includes: Adding the first noise to one or more of the acceleration observation data and the angular velocity observation data of the inertial measurement device.

[0018] In an exemplary embodiment of the present disclosure, training a robot high-jump motion control model based on a high platform terrain through deep reinforcement learning includes: In the simulation environment of the high platform terrain, using the teacher-student model framework, training the policy network for outputting the action strategy for controlling the robot's motion to obtain the robot high-jump motion control model.

[0019] In an exemplary embodiment of the present disclosure, using the teacher-student model framework to train the policy network for outputting the action strategy for controlling the robot's motion includes: The historical motion state information of the robot itself is encoded by the student encoder to generate a first latent vector, and the privileged state information of the robot is encoded by the teacher encoder to generate a second latent vector; Select the first latent vector or the second latent vector according to a preset policy and input it into the policy network; Based on the received first latent vector or second latent vector and the current motion state information, the policy network outputs an action policy for controlling the motion of the robot; wherein, the current motion state information includes the observation data of the inertial measurement device; The value network uses the privileged state information to output an estimated value of the current state; Based on the difference between the first latent vector and the second latent vector, optimize the parameters of the student encoder; Based on the action policy and the value estimation, use the deep reinforcement learning algorithm to update the parameters of the policy network.

[0020] In an exemplary embodiment of the present disclosure, the value network uses the privileged state information to output an estimated value of the current state, and further includes: The value network uses the privileged state information and combines the first latent vector or the second latent vector selected by the preset policy to output an estimated value of the current state.

[0021] According to a second aspect of the present disclosure, there is provided a method for controlling a robot's high jump movement, including: Obtain the current motion state information of the robot; Based on the current motion state information and a pre-trained robot high jump motion control model, output an action policy for controlling the motion of the robot; Wherein, the robot high jump motion control model is obtained according to the robot high jump motion control model training method in the above embodiment.

[0022] In an exemplary embodiment of the present disclosure, the pre-trained robot high jump motion control model includes a student encoder and a policy network; Based on the current motion state information and a pre-trained robot high jump motion control model to output an action policy for controlling the motion of the robot, including: The student encoder encodes the historical motion state information of the robot itself to generate a first latent vector and inputs it into the policy network; Based on the received first latent vector and the current motion state information, the policy network outputs an action policy for controlling the motion of the robot.

[0023] According to a third aspect of the present disclosure, there is provided a device for training a robot high jump motion control model, including: A data acquisition module, configured to receive the current state data of the virtual robot, where the current state data includes the observation data of the inertial measurement device; A noise addition module, configured to add first noise to the observation data of the inertial measurement device; A model training module, configured to combine the current state data after adding the first noise, and based on the high platform terrain, train a robot high platform jumping motion control model through deep reinforcement learning.

[0024] According to a fourth aspect of the present disclosure, there is provided a robot high platform jumping motion control device, including: An information acquisition module, configured to acquire the current motion state information of the robot; A policy output module, configured to output an action policy for controlling the motion of the robot based on the current motion state information and a pre-trained robot high platform jumping motion control model; Wherein, the robot high platform jumping motion control model is obtained according to the robot high platform jumping motion control model training method in the above embodiment.

[0025] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: A processor; and A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods in the above embodiments are implemented.

[0026] According to a sixth aspect of the present disclosure, there is provided a robot, including: A processor; and A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods in the above embodiments are implemented.

[0027] In an exemplary embodiment of the present disclosure, the robot includes any one of a legged robot, a quadruped robot, a biped robot, a wheeled robot, a wheel-legged robot, a four-wheel-legged robot, a humanoid robot, a cleaning robot, a transportation robot, a mobile robot, and a robotic arm.

[0028] According to a seventh aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program code instructions are stored, and when the computer program code instructions are called by the processor of the robot, the robot is enabled to execute the methods in the above embodiments.

[0029] It can be seen from the above technical solutions that the present disclosure has at least one of the following advantages and positive effects: In the method for training a robot high-platform jumping motion control model in an exemplary embodiment of the present disclosure, when receiving the current state data of the robot, noise is added to the observation data to simulate the measurement deviation of the IMU in the real environment, so that the control model can adapt to the observation error caused by impact or vibration during the training phase. The prior art usually uses ideal observation data for training, resulting in difficulty for the control strategy to make correct judgments when encountering IMU deviation in the real environment. Especially at the moment when the robot lands or during continuous jumping, the error may gradually accumulate and affect the motion stability. By actively introducing noise into the observation data, the present disclosure enables the trained control strategy to better adapt to the deviation of the actual IMU data, thereby reducing the risk of decision-making instability caused by measurement errors. In addition, based on the observation data with added noise, the present disclosure performs deep reinforcement learning training in the high-platform terrain, enabling the control model to adapt to different height terrain conditions and IMU observation errors simultaneously. When the robot lands at different heights, the impact on the IMU is different. If the control strategy fails to fully consider this impact, it may still lead to attitude estimation deviation, which in turn affects the next jump. During the training process, the present disclosure enables the control strategy to maintain a more stable decision-making ability when facing the impact vibration of landing at different heights, thereby improving the reliability of motion control and further avoiding control instability. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0031] Figure 1 The system architecture diagrams showing the method for training a robot high-platform jumping motion control model and the method for controlling a robot high-platform jumping motion in the embodiments of the present disclosure can be applied.

[0032] Figure 2 The flowchart showing a method for training a robot high-platform jumping motion control model in the embodiments of the present disclosure.

[0033] Figure 3 The schematic diagram showing a two-stage training-inference framework of a policy network in the embodiments of the present disclosure.

[0034] Figure 4 The flowchart showing a process of pre-training a policy network in the embodiments of the present disclosure.

[0035] Figure 5 The schematic diagram showing another two-stage training-inference framework of a policy network in the embodiments of the present disclosure.

[0036] Figure 6 The schematic diagram of the principle of the robot high-platform jumping motion control model training method in the embodiments of the present disclosure is shown.

[0037] Figure 7 The schematic diagram of a scene of a robot jumping on a high platform in the embodiments of the present disclosure is shown.

[0038] Figure 8 The schematic flowchart of a robot high-platform jumping motion control method in the embodiments of the present disclosure is shown.

[0039] Figure 9 The block diagram of a robot high-platform jumping motion control model training device in the embodiments of the present disclosure is shown.

[0040] Figure 10 The block diagram of a robot high-platform jumping motion control device in the embodiments of the present disclosure is shown.

[0041] Figure 11 The schematic diagram of a robot in the embodiments of the present disclosure is shown.

[0042] Figure 12 The schematic diagram of another robot in the embodiments of the present disclosure is shown.

[0043] Figure 13 The schematic diagram of yet another robot in the embodiments of the present disclosure is shown.

[0044] Figure 14 The schematic diagram of still another robot in the embodiments of the present disclosure is shown.

[0045] Figure 15 The schematic structural diagram of an electronic device in the embodiments of the present disclosure is shown.

[0046] Figure 16 The schematic structural diagram of a program product in the embodiments of the present disclosure is shown. Detailed implementation manners

[0047] In the description of the present disclosure, the terms "first" and "second" are only used for description and do not indicate relative importance or imply the number of technical features. Therefore, the features of "first" and "second" may explicitly or implicitly include at least one of such features. The meaning of "a plurality" is at least two, unless otherwise clearly defined.

[0048] Figure 1 The system architecture diagram to which the robot high-platform jumping motion control model training method and the robot high-platform jumping motion control method in the embodiments of the present disclosure can be applied is shown.

[0049] As Figure 1As shown, the system architecture 100 may include a terminal device 101, a robot 102, a network 103, and a server 104. Among them, the terminal device 101 includes, but is not limited to, desktop computers, portable computers, smartphones, tablet computers, etc., which are used to provide a human-computer interaction interface, support users to set training parameters, monitor the training process or view training results, and can also issue task control instructions and receive status feedback from the server 104 or the robot 102.

[0050] The robot 102 is a virtual robot deployed in a simulation platform, with complete state observation capabilities and action execution capabilities, and is used to collect training data and verify strategies for jumping tasks in a high-platform terrain simulation environment. The current state data of the robot 102 includes the observation data of inertial measurement devices. To improve the robustness of the strategy and its adaptability to actual deployment, the server 104 injects noise into the observation data to simulate measurement errors or environmental interference existing in real sensors, thereby training a more fault-tolerant control strategy.

[0051] The network 103 is a medium used to provide a communication link between the terminal device 101, the robot 102, and the server 104. It can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc., to ensure stable and efficient data interaction between the terminal device 101, the robot 102, and the server 104, including uploading status data, synchronizing training results, and downloading model parameters, etc.

[0052] The server 104 is used to construct a high-platform terrain simulation environment, schedule the virtual robot to complete the high-jump action in different initial states, collect training samples, and execute a deep reinforcement learning training process based on the state data with added noise to continuously optimize the jumping strategy. After the final training is completed, the server 104 can deploy the high-jump control model to the robot 102 for verification and testing, and the terminal device 101 can also remotely view the strategy execution effect.

[0053] The architecture 100 realizes an integrated closed-loop process from simulation environment construction, noisy state modeling to robust control strategy training, effectively improving the robot's jumping ability in high-difference scenarios and its adaptability to sensor errors.

[0054] It should be understood that Figure 1 the number and types of terminal devices, robots, networks, and servers in

[0055] The exemplary embodiments of the present disclosure provide a method for training a robot high-jump motion control model. Referring to Figure 2 as shown, this method may include the following steps S210 to step S230: Step S210: Receive the current state data of the virtual robot, where the current state data includes the observation data of the inertial measurement device. Step S220: Add first noise to the observation data of the inertial measurement device. Step S230: Based on the current state data after adding the first noise, train a robot high-platform jumping motion control model through deep reinforcement learning based on a high-platform terrain.

[0056] Execute the robot high-platform jumping motion control model training method provided by the exemplary embodiment of the present disclosure. By actively introducing noise into the observation data, the trained control strategy can better adapt to the deviation of the actual IMU data, thereby reducing the risk of decision-making instability caused by measurement errors. In addition, based on the observation data with added noise, deep reinforcement learning training is carried out in a high-platform terrain, enabling the control model to adapt to different height terrain conditions and IMU observation errors simultaneously. Moreover, during the training process, the control strategy can maintain a more stable decision-making ability when facing the landing impact vibration at different heights, thereby improving the reliability of motion control and further avoiding control instability.

[0057] Next, the robot high-platform jumping motion control model training method in this exemplary embodiment will be described in detail.

[0058] In step S210, receive the current state data of the virtual robot, where the current state data includes the observation data of the inertial measurement device.

[0059] Among them, the virtual robot refers to a robot model constructed in a simulation environment, which can simulate the motion, perception, and control processes of a real robot. For example, the virtual robot can output virtual sensor data such as acceleration and angular velocity for training control strategies or system verification, without using real hardware, which is safe, efficient, and low-cost. The inertial measurement device refers to a sensor used to simulate the inertial navigation system of a real robot, including an accelerometer, a gyroscope, etc., for observing the acceleration, angular velocity, magnetic field data, etc. of the virtual robot in the virtual space. The observation data of the inertial measurement device can be output in a multi-axis form, such as three-axis acceleration and three-axis angular velocity, which can reflect the attitude change, motion trend, and force characteristics of the robot in the virtual environment.

[0060] By collecting the observation data of the virtual robot, the motion state of the virtual robot can be accurately estimated, including information such as its speed, attitude, and displacement, thereby providing accurate and stable input for subsequent deep reinforcement learning, helping to accelerate the model convergence speed and improve the training efficiency. Especially in the high-platform terrain task, accurate state data can more clearly reflect the real interaction between the virtual robot and the environment, which is beneficial to learning the key control features of the jumping action, thereby establishing a more effective initial motion control strategy.

[0061] Moreover, without relying on external positioning systems such as GPS or visual sensing, high-frequency state estimation can be achieved within a short time solely by inertial measurement devices, which has the advantages of strong real-time performance and fast response. Additionally, during simulation training, algorithm verification, and virtual debugging, virtual inertial data can be used to effectively simulate the real response behavior of a robot in a complex dynamic environment, reducing the dependence on actual hardware and improving system development efficiency.

[0062] In step S220, first noise is added to the observation data of the inertial measurement device.

[0063] In the exemplary embodiments of the present disclosure, adding first noise to the observation data of the inertial measurement device means that when processing the sensor data of a virtual robot, certain random perturbations are artificially introduced to simulate the uncertainty generated by a real sensor affected by factors such as environmental interference and hardware errors during operation.

[0064] For example, the observation data of the inertial measurement device includes acceleration observation data and angular velocity observation data, and the first noise can be Gaussian noise, random noise, etc. Correspondingly, the first noise can be added to the acceleration observation data and / or the angular velocity observation data respectively. The processed observation data is closer to the sensor output in a real application scenario, which helps to improve the tolerance of the deep reinforcement learning model to uncertain information, enhance the generalization ability and robustness of the training strategy, and make it more stable and reliable in a real system.

[0065] In some exemplary embodiments, the first noise includes impulse noise, which refers to mutant noise that can simulate the sudden impact on a sensor. Impulse noise can be manifested as a short-time high-amplitude random pulse signal, which is local in time and uncertain in amplitude.

[0066] Optionally, impulse noise can be added to the observation data of the inertial measurement device when it is determined according to the current state data of the virtual robot that the virtual robot lands. For example, at the moment when the virtual robot falls from a high platform and contacts the ground, the observation data of the inertial measurement device such as acceleration and angular velocity will change sharply. At this time, impulse noise is injected into the observation data to simulate the instantaneous perturbation of the sensor caused by factors such as structural impact, ground reaction force, or mechanical vibration when a real robot lands. For example, the impulse noise can be directly superimposed on the original observation data.

[0067] By determining that the virtual robot is in a landed state and adding impact noise to the observation data of the inertial measurement device, not only can the diversity of training data be enhanced, but also the deep reinforcement learning model can learn how to maintain stable control or quickly recover under impact interference, which helps to improve the robustness and reliability of the motion control strategy in actual complex scenarios.

[0068] Optionally, when it is determined according to the current state data of the virtual robot that the virtual robot is in continuous jumping motion, impact noise can be added after each jumping action. That is to say, during the process of identifying that the virtual robot is in a continuous jumping task, at the moment when it completes each jump and lands, by detecting the changes in the current state of the virtual robot, such as the rapid decrease in height, sudden increase in acceleration and other characteristics, it is determined that it has completed a jump and contacted the ground. At this time, impact noise is injected into the observation data of the inertial measurement device to simulate the instantaneous interference or vibration response generated by the sensor due to the landing impact during the frequent jumping process of the real robot.

[0069] By determining that the virtual robot is in continuous jumping motion and has completed each jump, and adding impact noise to the observation data of the inertial measurement device, the authenticity of the virtual training data can be effectively enhanced, enabling the deep reinforcement learning model to have stronger adaptability in the face of frequent and repeated landing impacts, avoiding overfitting to ideal data, and improving the tolerance to attitude fluctuations and sensor disturbances during actual jumping motion, thereby making the robot control strategy more robust and stable.

[0070] In some exemplary embodiments, the first noise further includes vibration noise, which refers to the minute vibration interference that can simulate the sensor under high-frequency repetitive motion states. The vibration noise generally persists and can be manifested as small-amplitude, persistent, high-frequency fluctuations, simulating the background disturbances of the inertial measurement device under the influence of structural resonance, motor jitter or ground micro-vibration during the jumping, takeoff and hovering processes of the real robot.

[0071] Exemplarily, when it is determined according to the current state data of the virtual robot that the virtual robot is in continuous jumping motion, vibration noise can be added to the observation data of the inertial measurement device. Specifically, during the process of identifying that the virtual robot is performing a continuous jumping task, by analyzing the state data of the virtual robot such as periodic displacement changes, attitude adjustment frequency, contact period, etc., it is determined that it is in a continuous jumping state, and vibration noise is superimposed on its acceleration observation data and angular velocity observation data during this period. For example, high-frequency sine waves, randomly filtered noise after band-pass filtering or oscillation models simulating mechanical resonance can be superimposed on the observed values.

[0072] By adding vibration noise, the simulation ability of the training data for vibration interference in the actual scenario can be enhanced, enabling the deep reinforcement learning model to have better robustness when processing real sensor data and improving its control stability and reliability in dynamic complex actions such as continuous jumping.

[0073] In some exemplary embodiments, the first noise includes random Gaussian noise. Random Gaussian noise refers to introducing a random perturbation signal based on the normal distribution into the observed data when simulating the output data of inertial measurement devices to more realistically reflect the measurement errors and background interference existing in their actual working states.

[0074] It should be noted that the mean and variance of the random Gaussian noise are determined according to the data characteristics of the real inertial measurement device. For example, given the zero-bias error, noise density of the accelerometer, and the zero-bias drift and noise density of the gyroscope angular velocity, a corresponding Gaussian distribution model can be constructed based on these characteristic parameters and used as a noise term to be superimposed on the acceleration observed data and angular velocity observed data output by the virtual robot.

[0075] By adopting the Gaussian noise model determined by the characteristics of the real inertial measurement device, the observed data of the virtual robot is closer to the sensor output in actual use in terms of distribution characteristics, thereby improving the adaptability of the training model when deployed to the real robot system.

[0076] It should be noted that in the exemplary embodiments of the present disclosure, the intensity of the first noise can also be dynamically adjusted according to the jumping height of the virtual robot. That is to say, when adding the first noise such as impact noise, vibration noise, and random Gaussian noise to the observed data of the inertial measurement device, a fixed noise intensity parameter is no longer used, but the amplitude of the first noise is adaptively adjusted according to the maximum height or the height change amplitude in the current jumping action. For example, the higher the jumping height, the greater the impact when the robot lands, and the more intense the structural jitter and inertial perturbation, so the observation errors or noises generated by the sensor in these cases are also more obvious. Correspondingly, a stronger noise should be set for higher jumps during simulation to more realistically reflect the dynamic interference process related to height.

[0077] For example, for impact noise, assuming that the amplitude of the impact noise is proportional to the jumping height, there is: (1) Where is the impact noise at time t, is the impact indication function, which takes the value of 1 when the virtual robot lands at time t and 0 at other times, and is used to control the noise to only act at the moment of landing; is a Gaussian distribution with a mean of 0, is the variance, is the basic noise variance, is the linear amplification coefficient of the noise intensity with respect to the jump height, and h is the jump height.

[0078] Equation (1) represents that at the landing moment, a Gaussian noise with a variance increasing with the height is superimposed according to the current jump height h.

[0079] For another example, for vibration noise, assuming that the amplitude of the vibration noise can be adjusted according to the jump height, we have: (2) (3) Where, is the vibration noise at time t, is the amplitude of the vibration noise, is the basic vibration amplitude, is the linear coefficient of the vibration amplitude increasing with the jump height, h is the jump height, is the vibration frequency, is the phase shift.

[0080] Equations (2) and (3) represent that the vibration amplitude increases linearly with the jump height, thereby simulating a more intense high-frequency structural response.

[0081] For another example, for random Gaussian noise, assuming that the variance of the random Gaussian noise can be dynamically scaled according to the jump height, we have: (4) Where, is the current noise standard deviation calculated according to the jump height h, which is used to construct the Gaussian distribution subsequently, is the basic noise standard deviation, representing the minimum noise level without jump height, is the amplification ratio coefficient of the noise varying with the jump height.

[0082] By dynamically adjusting the intensity of the first noise, when the virtual robot performs jump actions with different amplitudes, the observed data of the inertial measurement device will show a more reasonable noise change trend, which not only improves the authenticity of the sensor model, but also enables the deep reinforcement learning to learn a more robust strategy when facing the sensor uncertainty under different motion intensities, thereby improving the generalization ability and real deployment reliability of the control strategy.

[0083] In addition, the intensity range of the first noise can be determined based on the statistical distribution of the measurement errors of the real inertial measurement device at the moment when the robot lands. Specifically, by collecting the observed error data of the inertial measurement device of the real robot during the jump landing process, the error distribution characteristics at the landing moment are statistically analyzed, and the reasonable value range of the first noise intensity parameter in the simulation system is deduced according to the error distribution characteristics.

[0084] For example, first, the difference between the inertial observation value of the robot at the moment of landing and its theoretical acceleration or attitude change is extracted from the measured data to obtain the true error sequence. Subsequently, statistical analysis is performed on this error sequence, including calculating its mean, standard deviation, extreme values, confidence interval, and probability distribution model such as Gaussian distribution, so as to obtain the typical intensity range of the error under the landing impact condition. Further, the standard deviation of the landing error, the maximum disturbance amplitude, etc. can be used as key parameters in the first noise generation function, so that the noise added in the simulation is consistent with the real sensor performance in terms of intensity and distribution characteristics.

[0085] In this example, the determined intensity range of the first noise has a realistic basis, which can effectively improve the physical rationality of the inertial observation data in the simulation, make the output of the virtual sensor closer to the real environment during the training process, and thus enhance the generalization ability and stability of the deep reinforcement learning model during actual deployment.

[0086] In step S230, based on the current state data after adding the first noise, a high-platform terrain is used to train a robot high-platform jumping motion control model based on deep reinforcement learning.

[0087] Among them, a task scenario including complex terrains such as high platforms is constructed in the virtual simulation environment. By obtaining the observation data of the inertial measurement device in the virtual robot after adding the first noise and using it as the training data input, the adaptability of the deep reinforcement learning algorithm to uncertain factors in the real dynamic environment is improved.

[0088] During the training process, the virtual robot continuously tries jumping actions to complete the target task of jumping onto the high platform. Each of its action decisions depends on the current state data, including position, speed, attitude, and the observation information of the inertial sensor. The added first noise effectively simulates the errors and interferences of the real sensor during the jumping and landing processes, thus making the training process closer to reality. Combined with the reward function, the deep reinforcement learning model gradually learns how to maintain attitude stability, take off reasonably, and land precisely on the high platform under noisy state perception, and finally obtains a set of jumping control strategies that are robust to dynamic noises such as impacts and vibrations. This not only improves the training efficiency and strategy generalization ability, but also provides a more stable and reliable control basis for real robots to perform high-dynamic action tasks in complex terrains.

[0089] In some example embodiments, in a simulation environment of a high platform terrain, a policy network for outputting an action policy for controlling the movement of a robot can be trained using a teacher-student model framework to obtain a robot high platform movement control model. This training process can strengthen the learning of the policy network for the behavior of successfully jumping onto the high platform. Additionally, by means of imitation loss or auxiliary supervision, the policy network is kept reasonably close to the teacher's behavior, thereby improving the policy convergence speed and performance stability. Finally, a robot high platform movement control model with both imitation ability and reinforcement adaptability is trained, which is applicable to continuous jumping control tasks in complex terrains.

[0090] Reference Figure 3 As shown, a schematic diagram of a two-stage training-inference framework of a policy network is shown. Among them, the training stage combines a teacher-student encoder structure and a deep reinforcement learning framework. The teacher-student encoder structure includes a teacher encoder 301 and a student encoder 302, and the deep reinforcement learning framework is jointly composed of a policy network 303, a value network 304, and a PPO algorithm. In the training stage, by introducing privileged state information to guide policy learning, the student encoder 302 can stably output high-quality control actions only relying on observable data in the inference stage. The observable data includes the current motion state information of the robot.

[0091] It should be noted that the teacher encoder 301 is only used to extract the high-dimensional semantic representation in the privileged state information in the training stage, which is used as the input of the policy for subsequent learning guidance. The student encoder 302 encodes the historical motion state of the robot itself in the training and inference stages to obtain dynamic features related to action decisions, and then inputs the current motion state information into the policy network 303 together to generate control actions. Additionally, the student encoder 302 finally needs to imitate the representation output by the teacher encoder 301, so in the training process, distillation training is carried out by minimizing the MSE (mean square error).

[0092] The policy network 303 is the core module for action generation. In the training stage, the policy network 303 receives the latent vector from the teacher encoder 301 or the student encoder 302 and the current motion state information, and combines with the value network 304 to update the policy using the PPO algorithm. In the inference stage, the policy network 303 can receive the latent vector from the student encoder 302 and the current motion state information, and output an action policy for controlling the movement of the robot. It should be noted that the current motion state information includes the observation data of the inertial measurement device after adding the first noise.

[0093] Based on Figure 3 the schematic diagram of the framework shown, reference Figure 4 As shown, the process of training a policy network for outputting an action policy for controlling the movement of a robot can include the following steps S401 to step S406: Step S401: Encode the historical motion state information of the robot itself through the student encoder to generate a first latent vector, and encode the privileged state information of the robot through the teacher encoder to generate a second latent vector.

[0094] Among them, the historical motion state information of the robot itself may include joint angles, joint velocities, sole contact states, center-of-gravity trajectories, etc. of several past frames. For example, input the historical motion state sequence of the past 10 frames into the student encoder 302 for encoding to obtain the first latent vector. The first latent vector can capture the continuity and dynamic characteristics of the robot's actions.

[0095] The privileged state information may include the actual contact force between the robot and the ground, the environmental height map, disturbance information, etc. Input the privileged state information into the teacher encoder 301 for encoding to obtain a third latent vector. The third latent vector can highly concentrate key semantics such as terrain structure, obstacle distribution, and the interaction between the robot and the environment, which helps to construct a more complete high-dimensional feature representation of the environmental state.

[0096] It should be noted that the privileged state information can only be obtained during the training phase but is not observable during the testing or deployment phase. Therefore, the teacher encoder 301 can use the complete information to learn the best latent representation, thereby guiding the student encoder 302 to learn. Both the teacher encoder 301 and the student encoder 302 can compress high-dimensional, time-series state information into low-dimensional latent representations, providing behavioral semantic representations from different sources for the policy network 303. For example, the teacher encoder 301 and the student encoder 302 can be multi-layer perceptrons, temporal convolutional networks, or recurrent neural networks, with the ability to extract temporal features and compress them into fixed-length semantic vectors. In addition, the network architectures of the teacher encoder 301 and the student encoder 302 can be the same or different, and the present disclosure does not limit this.

[0097] Step S402: Select the first latent vector or the second latent vector according to a preset policy and input it into the policy network.

[0098] Select to send the first latent vector generated by the student encoder 302 or the second latent vector generated by the teacher encoder 301 into the policy network 303 for decision-making according to a preset policy. Among them, the preset policy can be set according to the real-time environmental state, the training phase, or specific performance metrics, and the specific performance metrics include action execution error thresholds, simulation and real environment difference degrees, etc.

[0099] For example, one preset strategy is to preferentially use the high-quality second latent vectors generated by the teacher encoder 301 at the initial stage of training to guide the policy network 303 to quickly converge to an approximate optimal solution. When the student encoder 302 is optimized through knowledge distillation, it gradually transitions to using only the first latent vectors during the deployment stage to reduce the dependence on privileged information. Another example is that another preset strategy is to select the second latent vectors according to a ratio p, select the first latent vectors according to 1 - p, and gradually decrease the value of p. Of course, the first latent vectors and the second latent vectors can also be concatenated and fed into the policy network 303, and dynamically weighted through an attention mechanism or a gating module, enabling the policy network 303 to flexibly combine historical experience and privileged knowledge in complex scenarios.

[0100] The selective input mechanism can not only utilize the prior knowledge of the ideal state provided by the teacher encoder 301 to accelerate the training process, but also cope with problems such as sensor limitations or environmental disturbances through the generalization ability of the student encoder 302 during actual operation. At the same time, through the conditional processing of the latent vectors by the policy network 303, smooth switching and robust decision-making of motion control are achieved. For example, when the robot encounters unknown terrain, it preferentially adjusts its gait based on the second latent vectors, while relying on the first latent vectors to maintain efficiency during the stable walking stage. Finally, the policy selection rules are optimized through closed-loop feedback, enabling the robot to balance motion performance and adaptability under different stages and environmental conditions.

[0101] In addition, in addition to the selected latent vectors, the current motion state information of the robot can also be input into the policy network 303 for decision-making.

[0102] Step S403, based on the received first latent vector or second latent vector and the current motion state information, output an action policy for controlling the movement of the robot; wherein, the current motion state information includes the observation data of the inertial measurement device.

[0103] Among them, the policy network 303 can be a multi-layer fully connected perceptron, a Transformer structure, etc. The policy network 303 performs multi-modal feature fusion on the first latent vector or second latent vector and the current motion state information, and outputs an action policy, denoted as .

[0104] For example, the policy network can be based on: (5) Output an action policy for controlling the movement of the robot; Among them, represents the probability distribution of taking action under the current motion state , MLP represents a multi-layer perceptron, Represents the first latent vector or the second latent vector selected according to a preset policy.

[0105] For another example, the first latent vector or the second latent vector and the current motion state information are concatenated or weighted interacted in the embedding space, and action policies such as joint angle targets, torque commands, or gait phase parameters are generated through non-linear transformation.

[0106] Step S404, using the privileged state information through the value network to output the value estimate of the current state.

[0107] During the training process of the policy network 303, the long-term return of the current state is also estimated through the value network 304, that is, starting from this state, if the current policy is continuously executed, how much cumulative reward can be obtained in the future, so as to optimize the policy network 303.

[0108] Among them, the value network 304 can be a multi-layer fully connected perceptron. For example, when using the privileged state information through the value network 304 to estimate the value of the current state, the privileged state information is first encoded into a high-dimensional feature vector, such as extracting dynamic features through a convolutional or fully connected layer, and then fused with the current motion state information in the latent space, and then the value estimate representing the current state is output through multi-layer non-linear transformation, denoted as This value estimate is used for policy optimization in reinforcement learning to guide the policy network 303 to learn better behaviors.

[0109] Step S405, based on the difference between the first latent vector and the second latent vector, optimize the parameters of the student encoder.

[0110] This step quantifies the distribution difference between the two in the latent space and backpropagates the difference gradient to update the network weights of the student encoder 302, so as to guide the student encoder 302 to learn to generate latent representations close to the teacher encoder 301, realizing teacher knowledge distillation.

[0111] Exemplarily, during the training stage, the parameters of the teacher encoder 301 are fixed. After using the historical motion state information and the corresponding motion state information as parallel inputs, and generating latent vectors through the student encoder 302 and the teacher encoder 301 respectively, the contrast learning framework is used to minimize the distance between the two, or the output distribution of the student encoder 302 is approximated to the latent space characteristics of the teacher encoder 301 through adversarial training. At the same time, noise injection or data augmentation is introduced to simulate sensor errors in actual deployment, forcing the student encoder 302 to still extract feature expressions compatible with the privileged information encoding under the condition of limited input information.

[0112] For example, to make the output of the student encoder 302 as close as possible to that of the teacher encoder 301, a difference loss function can be constructed, such as the least squares gap or KL divergence. By minimizing this loss, the parameters of the student encoder 302 are optimized so that it can learn to extract feature expressions close to the privileged information from the historical motion states, thereby enhancing the generalization ability of the policy network 303, improving the decision-making quality of the policy network 303 in the real environment, and enabling it to approach the teacher level without relying on privileged information during deployment.

[0113] Step S406: Based on the action policy and value estimation, use the deep reinforcement learning algorithm to update the parameters of the policy network.

[0114] Taking the PPO algorithm as an example. During the training process of the policy network 303, the current policy network 303 is used to interact with the environment. Based on the current state select the action policy , and after executing the action, return the reward and the next state , thus sampling to obtain the interaction trajectory sequence ( ), which is used for subsequent policy optimization. Then, the value network 304 is introduced to evaluate the value of each state .

[0115] Based on this, if an advantage function is constructed : (6) where is the discount factor, which is used to measure the importance of future rewards.

[0116] Then, an optimization objective is constructed according to the clipping objective function of the PPO algorithm, and the policy gradient is calculated through backpropagation to optimize the parameters of the policy network 303, so that it can continuously improve the expected cumulative return under the current policy. This update process can guide the policy network 303 to gradually learn the action policy that maximizes the long-term cumulative reward while maintaining a stable output, thereby improving the robustness and execution efficiency of the control policy in the real environment.

[0117] Refer to Figure 5 shown, which shows a schematic diagram of a two-stage training-inference framework for another policy network. Based on Figure 5 shown in the framework schematic diagram, the training process of the policy network is similar to Figure 4 shown, except that the step of using the privileged state information by the value network 304 to output the value estimation of the current state in this training process is different from the implementation method of step S404. Specifically, in Figure 6In the training phase shown, the value network 304 uses privileged state information and combines the first latent vector or the second latent vector selected by a preset policy to output a value estimate of the current state.

[0118] In this example, the latent features and privileged information are fused and input into the value network 304. This not only retains the high credibility of the privileged information for understanding the environment but also enhances the adaptability of the value estimate to the policy behavior trajectory and motion trend, enabling a better reflection of the long-term reward expectation that the robot can obtain after taking actions in the current state. In addition, the fused latent vector input also improves the representational richness and generalization ability of the value network, making its output more consistent with the value distribution under the real policy behavior, thereby enhancing the stability and performance of the overall policy training.

[0119] Reference Figure 6 As shown, a schematic diagram of the principle of a method for training a robot high-jump motion control model is presented. First, the current state data 601 of the robot is collected, which includes the observation data of the inertial measurement device. Then, noise is added to the observation data of the inertial measurement device to enhance the generalization ability of the model, enabling it to maintain stable decision-making in the case of simulated sensor errors or real-world disturbances. Subsequently, based on the current state data 602 with added noise, the robot high-jump motion control model 603 is trained in the high-platform terrain, enabling the robot high-jump motion control model 603 to learn how to autonomously complete the high-jump task according to the current state. Reference Figure 7 As shown, a schematic diagram of a robot high-jump scenario is presented.

[0120] The exemplary embodiment of the present disclosure also provides a robot motion control method. Reference Figure 8 As shown, the method may include the following steps S801 to step S802: Step S801, obtain the current motion state information of the robot.

[0121] Among them, the robot can be a real robot or a virtual robot. During the process of the robot executing a motion task, the motion state data at its current moment is collected in real time, including but not limited to the position, attitude, speed, etc. of the robot itself. Importantly, the current motion state information includes the observation data from the inertial measurement device, such as including three-axis acceleration and three-axis angular velocity. In addition, the current motion state information may also include sensor readings of the robot body such as joint angles, joint speeds, and end poses, which are not specifically limited in the present disclosure.

[0122] Step S802, based on the current motion state information and the pre-trained robot high-jump motion control model, output an action strategy for controlling the robot's motion.

[0123] During the process of the robot performing the high-platform jump task, the current motion state information obtained in real time is input, and combined with the robot high-platform jump motion control model that has been trained to make action decisions.

[0124] Exemplarily, the robot high-platform jump motion control model consists of a student encoder and a policy network. The student encoder can encode the historical motion state information of the robot itself to generate a first latent vector, which is input to the policy network for further reference by the policy network. Then, based on the received first latent vector and the current motion state information, the policy network outputs an action policy for controlling the robot's motion, such as takeoff time, attitude adjustment, or power distribution, so as to guide the robot to complete the target behavior of jumping onto the high platform.

[0125] In this example, by combining the real-time current motion state information with the historical motion state encoding, the robot high-platform jump motion control model can fuse the historical motion trend on the basis of perceiving the current state and make more accurate action decisions. Among them, the first latent vector extracted by the student encoder provides a dynamic understanding of past motion behaviors, and the policy network uses this latent vector to jointly reason with the current state and outputs a reasonable jumping action policy. Therefore, the robot high-platform jump motion control model can not only identify whether the robot is in the takeoff preparation, airborne, or landing stage, but also adjust the takeoff timing and attitude according to the historical rhythm, thereby effectively improving the coordination, coherence, and success rate of the jumping behavior, and enhancing the control accuracy and robustness in complex task scenarios such as high-platform terrains.

[0126] It can be understood that the robot high-platform jump motion control model is trained according to the robot high-platform jump motion control model training method described in detail in other embodiments of the present disclosure, which will not be elaborated here.

[0127] In the exemplary implementation manner of the present disclosure, a training device for a robot high-platform jump motion control model is also provided. Refer to Figure 9 As shown, the training device 900 for the robot high-platform jump motion control model includes a data acquisition module 901, a noise addition module 902, and a model training module 903, where: The data acquisition module 901 is configured to receive the current state data of the virtual robot, and the current state data includes the observation data of the inertial measurement device; The noise addition module 902 is configured to add first noise to the observation data of the inertial measurement device; The model training module 903 is configured to combine the current state data after adding the first noise and train the robot high-platform jump motion control model based on the high-platform terrain through deep reinforcement learning.

[0128] The specific details of each module in the above robot high-jump motion control model training device have been described in detail in the corresponding robot high-jump motion control model training method, so they will not be elaborated here.

[0129] In an exemplary embodiment of the present disclosure, a robot high-jump motion control device is also provided. Refer to Figure 10 As shown, the robot high-jump motion control device 1000 includes an information acquisition module 1001 and a strategy output module 1002, where: The information acquisition module 1001 is configured to acquire the current motion state information of the robot; The strategy output module 1002 is configured to output an action strategy for controlling the motion of the robot based on the current motion state information and a pre-trained robot high-jump motion control model; Among them, the robot high-jump motion control model is obtained according to the robot high-jump motion control model training method in the embodiments of the present disclosure.

[0130] The specific details of each module in the above robot high-jump motion control device have been described in detail in the corresponding robot high-jump motion control method, so they will not be elaborated here.

[0131] In an exemplary embodiment of the present disclosure, a robot is also provided. The robot includes a processor and a memory, and computer-readable instructions are stored on the memory. When the computer-readable instructions are executed by the processor, the above method is implemented. Among them, the robot includes any one of a legged robot, a quadruped robot, a biped robot, a wheeled robot, a wheel-legged robot, a four-wheel-legged robot, a humanoid robot, a cleaning robot, a transportation robot, a mobile robot, and a robotic arm. Refer to Figures 11 to 14 As shown, schematic diagrams of four different robots are respectively shown.

[0132] Refer to Figure 15 As shown, an electronic device capable of implementing the above method is also provided. Among them, the electronic device 1500 includes a processor 1501 and a memory 1502, and computer-readable instructions are stored on the memory 1502. When the computer-readable instructions are executed by the processor 1501, the method in the embodiments of the present disclosure is implemented.

[0133] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which computer program code instructions are stored. When the computer program code instructions are called by the processor of the robot, the robot is enabled to execute the method as in the embodiment.

[0134] Refer to Figure 16As shown, a program product 1600 for implementing the above method according to an embodiment of the present disclosure is described. It may adopt a portable compact disc read-only memory (CD-ROM), include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0135] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiment of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which may be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiment of the present disclosure.

[0136] Finally, the above preferred embodiments are only used to illustrate the technical solutions of the present application and are not restrictive. Although the present application has been described in detail, those skilled in the art should understand that changes in form and details can be made thereto without departing from the scope defined by the claims of the present application. The dimensions of the drawings have nothing to do with the specific physical objects, and the physical dimensions can be arbitrarily changed.

Claims

1. A robot high jump platform motion control model training method, characterized in that: include: Receiving current state data of the virtual robot, wherein the current state data includes observation data of an inertial measurement device; adding first noise to observation data of the inertial measurement device; Combined with the current state data after adding the first noise, the robot jumping high platform motion control model is trained based on the high platform terrain through deep reinforcement learning.

2. The high jump platform motion control model training method according to claim 1, characterized in that: The first noise includes impact noise; and the adding the first noise to the observation data of the inertial measurement device includes: When it is determined according to the current state data of the virtual robot that the virtual robot has landed, the impact noise is added to the observation data of the inertial measurement device.

3. The high jump platform motion control model training method according to claim 2, characterized in that: The adding a first noise to the observation data of the inertial measurement device further comprises: When it is determined according to the current state data of the virtual robot that the virtual robot is in continuous jumping motion, the impact noise is added after each jumping motion.

4. The high jump platform motion control model training method according to claim 2, characterized in that: The first noise also includes vibration noise; and the adding the first noise to the observation data of the inertial measurement device further includes: When it is determined according to the current state data of the virtual robot that the virtual robot is in continuous jumping motion, the vibration noise is added to the observation data of the inertial measurement device.

5. The robot high jump platform motion control model training method according to claim 1, characterized in that: The first noise includes random Gaussian noise, and a mean value and a variance of the random Gaussian noise are determined according to data characteristics of a real inertial measurement device.

6. The high jump platform motion control model training method according to claim 1, characterized in that: The method further comprises: The intensity of the first noise is dynamically adjusted according to the jumping height of the virtual robot.

7. The high jump platform motion control model training method according to claim 6, characterized in that: The intensity range of the first noise is determined based on the statistical distribution of measurement errors of a real inertial measurement device at the moment the robot lands.

8. The high jump platform motion control model training method according to claim 6, characterized in that: The first noise includes impact noise; and dynamically adjusting the intensity of the first noise according to the jumping height of the virtual robot includes: according to: adjusting the intensity of the first noise; in, is the impact noise at time t, is the shock indicator function, The mean is 0. is a Gaussian distribution with variance, is the basic noise variance, is the linear amplification factor of noise intensity to jump height, and h is the jump height.

9. The high jump platform motion control model training method according to claim 6, characterized in that: The first noise includes vibration noise; and dynamically adjusting the intensity of the first noise according to the jumping height of the virtual robot includes: according to: adjusting the intensity of the first noise; in, is the vibration noise at time t, is the amplitude of vibration noise, is the basic vibration amplitude, is the linear coefficient of the vibration amplitude increasing with the jumping height, h is the jumping height, is the vibration frequency, is the phase shift.

10. The high jump platform motion control model training method according to claim 6, characterized in that: The first noise includes random Gaussian noise; and dynamically adjusting the intensity of the first noise according to the jumping height of the virtual robot includes: according to: adjusting the intensity of the first noise; in, is the current noise standard deviation calculated according to the jump height h, which is used to construct the Gaussian distribution later. is the basic noise standard deviation, It is the amplification coefficient of the noise as the jump height changes.

11. The high jump platform motion control model training method according to claim 1, characterized in that: The adding a first noise to the observation data of the inertial measurement device comprises: The first noise is added to one or more of acceleration observation data and angular velocity observation data of the inertial measurement device.

12. The high jump platform motion control model training method according to any one of claims 1 to 11, characterized in that: The method of training a robot high platform jumping motion control model based on high platform terrain by deep reinforcement learning includes: In a simulation environment of a high platform terrain, a teacher-student model framework is used to train a strategy network for outputting action strategies for controlling robot motion, thereby obtaining the robot high platform jumping motion control model.

13. The high jump platform motion control model training method according to claim 12, characterized in that: The method of using a teacher-student model framework to train a policy network for outputting an action policy for controlling robot motion includes: The robot's historical motion state information is encoded by a student encoder to generate a first latent vector, and the robot's privileged state information is encoded by a teacher encoder to generate a second latent vector; Selecting a first latent vector or a second latent vector according to a preset strategy and inputting the first latent vector into the strategy network; Outputting, through the strategy network, an action strategy for controlling the movement of the robot based on the received first potential vector or second potential vector and current motion state information; wherein the current motion state information includes observation data of an inertial measurement device; Utilizing the privileged state information through a value network, outputting a value estimate of the current state; optimizing parameters of the student encoder based on a difference between the first latent vector and the second latent vector; Based on the action strategy and the value estimate, the parameters of the policy network are updated using a deep reinforcement learning algorithm.

14. The high jump platform motion control model training method according to claim 13, characterized in that: The method of utilizing the privileged state information through the value network to output a value estimate of the current state further includes: The privileged state information is utilized through the value network, and combined with the first latent vector or the second latent vector selected by the preset strategy, a value estimate of the current state is output.

15. A robot high jump platform motion control method, characterized in that: include: Get the current motion state information of the robot; Outputting a motion strategy for controlling the movement of the robot based on the current motion state information and a pre-trained robot high jump platform motion control model; Wherein, the robot high jump platform motion control model is obtained according to the robot high jump platform motion control model training method according to any one of claims 1 to 14.

16. The robot high jump platform motion control method according to claim 15, characterized in that: The pre-trained robot high jump platform motion control model includes a student encoder and a strategy network; The motion strategy for controlling the robot motion based on the current motion state information and the pre-trained robot high jump platform motion control model output includes: Encoding the robot's own historical motion state information through the student encoder to generate a first potential vector, and inputting the first potential vector into the strategy network; The strategy network outputs an action strategy for controlling the movement of the robot based on the received first potential vector and current motion state information.

17. A robot high jump platform motion control model training device, characterized in that: include: A data acquisition module, used to receive current state data of the virtual robot, wherein the current state data includes observation data of an inertial measurement device; A noise adding module, used for adding a first noise to the observation data of the inertial measurement device; The model training module is used to combine the current state data after adding the first noise, and train the robot jumping high platform motion control model based on the high platform terrain through deep reinforcement learning.

18. A robot high jump platform motion control device, characterized in that: include: An information acquisition module is used to obtain the current motion state information of the robot; A strategy output module, used for outputting a motion strategy for controlling the movement of the robot based on the current motion state information and a pre-trained robot high jump platform motion control model; Wherein, the robot high jump platform motion control model is obtained according to the robot high jump platform motion control model training method according to any one of claims 1 to 14.

19. An electronic device, characterized in that: include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the method according to any one of claims 1 to 16.

20. A robot, characterized in that: include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the method according to any one of claims 1 to 16.

21. The robot according to claim 20, characterized in that The robot includes any one of a legged robot, a quadruped robot, a bipedal robot, a wheeled robot, a wheel-legged robot, a quadrupedal robot, a humanoid robot, a cleaning robot, a transport robot, a mobile robot and a robotic arm.

22. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program code instructions, and when the computer program code instructions are called by a processor of the robot, the robot executes the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Body state estimation method of legged robot based on multi-sensor information fusion

    CN108621161A

  • Automatic parking method based on reinforcement learning network training

    CN109492763A

  • Human body behavior recognition method based on deep learning

    CN111860117A

  • Real-time calculation method for angle of anti-position-movement joint based on inertial sensors

    CN111887856A

  • High-speed motion capture and recognition method and system for motion scene

    CN116030533A

Cited By

  • Quadruped robot low-noise gait control method and system based on soft landing reward function

    CN121209388A