Biped robot control method and system based on head-mounted display control equipment
The head-mounted display control device obtains the user's head movements and converts them into a bipedal robot action indication information. Combined with the deep reinforcement learning model output action strategy, the problem of unnatural and inaccurate robot control in the existing technology is solved, and more efficient and intuitive robot standing height control is achieved.
Patent Information
- Application Number
- CN202510527011.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
AI Technical Summary
Existing robot control technology is difficult to achieve natural and accurate human-computer interaction, especially in scenarios where flexibly adjusting postures or finely controlling motion parameters, and the height control method of bipedal robot standing is not intuitive and inconvenient.
The user's head movement is obtained through the head-mounted display control device, and the action indication information is converted into a bipedal robot, including standing height indication information, and this information is input to the robot motion control model based on deep reinforcement learning to output the action strategy.
It realizes a more natural, intuitive and conveniently deployed robot standing height control, improves the accuracy and user experience of robot control, and overcomes the limitations of traditional handle input.
Smart Images

Figure CN120044865A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of sensors and robots, and more particularly to a biped robot control method and system based on a head mounted display control device. Background Art
[0002] Existing robot control technology mainly relies on handles or similar input devices, and inputs commands through joysticks, buttons, touchpads, etc. to achieve robot movement control and standing height adjustment. Users usually push joysticks or press buttons to move the robot in a specific direction, and can adjust the standing height with the help of additional control inputs. However, this control method has certain limitations and is difficult to meet the needs of more natural and precise human-computer interaction, especially in scenarios that require flexible posture adjustment or fine control of motion parameters. Its performance is limited by many factors.
[0003] In response to the above problems, some existing technologies, such as the paper "ROS Reality: A Virtual Reality Framework Using Consumer Grade Hardware for ROS Enabled Robots" and the Chinese patent application with publication number CN115185368A, disclose methods for robot motion control combined with virtual reality technology. This type of technical solution obtains the user's motion information through a virtual reality device to improve the robot control method and enhance the immersion and intuitiveness of the interaction. However, in these technical solutions, the standing height control of the bipedal robot still relies on the handle input, and the user needs to adjust the standing height through a specific button, rocker or touchpad. This method makes it difficult for the user to intuitively and naturally control the height of the robot during operation, especially when the height needs to be adjusted quickly or finely controlled. The discreteness of the handle input may cause the command execution to be unsmooth, affecting the coherence of the control. In addition, due to differences in users' adaptability to the handle control method, some users may be more accustomed to a certain control logic, while other users may find it difficult to adapt to the layout of a specific handle, increasing the learning cost and reducing the user experience.
[0004] In the Chinese patent application with publication number CN114227679A, a robot motion control method combined with virtual reality technology is also disclosed. This method uses a user knee sensor to collect human knee joint movement information and maps it to the standing height adjustment of the bipedal robot. Compared with handle control, this solution can make the robot's standing height adjustment method more in line with the natural movements of the human body and improve the intuitiveness of the interaction. However, this solution requires the installation of additional sensors on the user's knees, which increases the hardware cost, and in actual applications, the user needs to wear or fix the sensor additionally, making the deployment inconvenient and not conducive to large-scale application and promotion.
[0005] Therefore, how to provide a more natural, intuitive and easy-to-deploy method for controlling the robot's standing height and improve the accuracy of robot control and user experience is a problem that needs to be urgently solved in current technology. Summary of the invention
[0006] The present disclosure provides a bipedal robot control method and system based on a head-mounted display control device to achieve more natural, intuitive and easy-to-deploy robot standing height control, thereby improving the accuracy of robot control and user experience.
[0007] Additional aspects and advantages of the disclosure will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the disclosure.
[0008] According to a first aspect of the present disclosure, a biped robot control method based on a head mounted display control device is provided, comprising: Acquire the user's head movements through a head-mounted display control device; Convert the user's head movement into action instruction information of the biped robot, wherein the action instruction information includes standing height instruction information; The action instruction information is input into a robot motion control model based on deep reinforcement learning, so that the robot motion control model outputs an action strategy to the bipedal robot in combination with the action instruction information.
[0009] In an exemplary embodiment of the present disclosure, the user's head motion includes a head pitch motion; Convert the user's head movements into action instructions for the bipedal robot, including: Convert head pitch movements into standing height instructions.
[0010] In an exemplary embodiment of the present disclosure, converting the head pitching action into standing height indication information includes: Based on the head pitch movement, the pitch angle of the user's head relative to the horizontal plane is obtained through the sensor; Maps pitch angle to standing height indication.
[0011] In an exemplary embodiment of the present disclosure, Maps pitch angle to stand-alone height indication, including: according to: Get standing height indication information; in, It is the standing height indication information; is the pitch angle; , , are preset parameters.
[0012] In an exemplary embodiment of the present disclosure, the user's head motion includes a head rotation motion, and the action indication information includes a travel direction indication information; Convert the user's head movements into action instructions for the bipedal robot, including: Convert head rotations into direction-of-travel instructions for bipedal robots.
[0013] In an exemplary embodiment of the present disclosure, converting the head rotation action into the moving direction indication information of the biped robot includes: Based on the head rotation action, the horizontal rotation angle of the user's head is obtained through the sensor; Map the horizontal rotation angle to the direction of travel indication information.
[0014] In an exemplary embodiment of the present disclosure, the method further includes: Obtain image information of the environment where the biped robot is located; The image information is integrated with the display screen of the head-mounted display control device to generate a digital twin scene, and then displayed in real time on the head-mounted display control device.
[0015] In an exemplary embodiment of the present disclosure, before inputting the action instruction information into the robot motion control model based on deep reinforcement learning, the method further includes: Using the teacher-student model framework, the policy network used to output the action strategy for controlling the motion of the biped robot is trained to obtain the robot motion control model. Among them, the input data of the teacher-student model framework includes action instruction information.
[0016] In an exemplary embodiment of the present disclosure, a policy network for outputting an action policy for controlling the motion of a biped robot is trained using a teacher-student model framework, including: The student encoder encodes the historical motion state information of the biped robot itself and generates a first latent vector, and the teacher encoder encodes the privileged state information of the biped robot and generates a second latent vector; Selecting a first latent vector or a second latent vector according to a preset strategy and inputting the vector into a strategy network; Outputting, through the strategy network, a motion strategy for controlling the motion of the biped robot based on the received first latent vector or the second latent vector, the action instruction information, and the current motion state information; Utilize privileged state information through the value network and output a value estimate of the current state; optimizing parameters of the student encoder based on a difference between the first latent vector and the second latent vector; Based on the action strategy and value estimation, the parameters of the policy network are updated using a deep reinforcement learning algorithm.
[0017] In an exemplary embodiment of the present disclosure, the robot motion control model outputs an action strategy to the biped robot in combination with the action instruction information, including: Encoding the historical motion state information of the biped robot itself through a student encoder to generate a first latent vector, and inputting the first latent vector into a policy network; The action strategy for controlling the movement of the biped robot is outputted through the strategy network based on the received first potential vector, the action instruction information and the current movement state information.
[0018] In an exemplary embodiment of the present disclosure, the head-mounted display control device includes a virtual reality head display, an augmented reality head display, or a mixed reality head display.
[0019] According to a second aspect of the present disclosure, there is provided a bipedal robot control system based on a head mounted display control device, comprising: A head-mounted display control device for acquiring the user's head movements; An instruction information generating module, used for converting the user's head movement into action instruction information of the biped robot, the action instruction information including standing height instruction information; The action strategy generation module is used to input action instruction information into a robot motion control model based on deep reinforcement learning, so that the robot motion control model outputs an action strategy to the bipedal robot in combination with the action instruction information.
[0020] According to a third aspect of the present disclosure, there is provided an electronic device, including: Processor; and A memory stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the method of the above embodiment is implemented.
[0021] According to a fourth aspect of the present disclosure, there is provided a bipedal robot, comprising: Processor; and A memory stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the method of the above embodiment is implemented.
[0022] In an exemplary embodiment of the present disclosure, the bipedal robot includes any one of a footed robot, a wheeled robot, a wheel-footed robot, a humanoid robot, a cleaning robot, a transport robot, and a mobile robot.
[0023] According to a fifth aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program code instructions are stored. When the computer program code instructions are called by a processor of a robot, the robot executes the method as described in the above embodiment.
[0024] It can be seen from the above technical solution that the present disclosure has at least one of the following advantages and positive effects: The present disclosure provides a biped robot control method based on a head mounted display control device, wherein the head movement of the user is obtained through the head mounted display control device, and the head movement of the user is converted into action indication information of the biped robot, wherein the action indication information includes standing height indication information. In addition, the action indication information is also input into a robot motion control model based on deep reinforcement learning, so that the robot motion control model outputs an action strategy to the biped robot in combination with the action indication information, thereby realizing more intelligent robot motion control. On the one hand, the present disclosure obtains the head movement of the user through the head mounted display control device and determines the standing height indication information based on the action information, so that the robot can respond to the user's interaction needs more naturally, abandons the dependence on the handle and additional sensors, makes the control of the robot's standing height more in line with the user's natural movement mode, improves the interaction intuition and control fluency, and overcomes the limitations of the existing methods. In addition, the existing head mounted display control device usually directly maps the user's head pitch angle to the pitch of the robot's head. Although this control method conforms to the natural intuition of human body movements, due to the fact that the robot has very little control demand for the head pitch in actual applications, the control action mapping of the standing height is missing, which affects the overall motion control of the robot. The present disclosure overcomes the limitations of the prior art, optimizes the action mapping method, and enables the user's operation in a virtual reality environment to more accurately control the robot's standing height, making the robot's motion control more in line with the human body interaction logic, and improving the rationality and intuitiveness of the control. On the other hand, the present disclosure is based on a robot motion control model based on deep reinforcement learning, which enables the bipedal robot to autonomously optimize motion decisions in a complex environment. Compared with the rule-based control method, the deep reinforcement learning model can improve the recognition ability of different user actions through adaptive learning, and automatically adjust the action strategy in different motion scenes, reducing the control instability problem caused by user misoperation. In addition, the deep reinforcement learning method can optimize the robot's analysis of action instruction information during the training process, so that it can more accurately understand and execute the user's standing height adjustment instructions, improve the stability and accuracy of the robot's movement, overcome the problems of discrete input signals and control lag in traditional methods, and make the robot's movement more natural, stable and efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0026] Figure 1A system architecture diagram is shown to which the biped robot control method based on a head mounted display control device in an embodiment of the present disclosure can be applied.
[0027] Figure 2 A flow chart of a biped robot control method based on a head mounted display control device in an embodiment of the present disclosure is shown.
[0028] Figure 3 A schematic diagram of a process for generating standing height indication information in an embodiment of the present disclosure is shown.
[0029] Figure 4 A schematic diagram of a process for generating travel direction indication information in an embodiment of the present disclosure is shown.
[0030] Figure 5 A schematic diagram of a two-stage training-inference framework of a robot motion control model in an embodiment of the present disclosure is shown.
[0031] Figure 6 A schematic diagram of a process for training a robot motion control model in an embodiment of the present disclosure is shown.
[0032] Figure 7 A schematic diagram of a flow chart of an output action strategy in an embodiment of the present disclosure is shown.
[0033] Figure 8 A schematic diagram of a scenario in which a biped robot control method based on a head mounted display control device in an embodiment of the present disclosure is applied is shown.
[0034] Fig. 9 A block diagram of a bipedal robot control system based on a head mounted display control device in an embodiment of the present disclosure is shown.
[0035] Fig.10 A schematic diagram of a robot in an embodiment of the present disclosure is shown.
[0036] Fig.11 Another schematic diagram of a robot in an embodiment of the present disclosure is shown.
[0037] Fig.12 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present disclosure is shown.
[0038] Fig.13 A schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0039] In the present disclosure, the terms "first" and "second" are used for description only and do not indicate relative importance or imply the number of technical features. Therefore, the "first" and "second" features may explicitly or implicitly include at least one of the features. "Multiple" means at least two, unless otherwise explicitly defined.
[0040] Figure 1 The system architecture diagram of the biped robot control method based on the head mounted display control device in the embodiment of the present disclosure is shown. Figure 1 As shown, the system architecture 100 may include a head mounted display control device 101, a robot 102, a terminal device 103, a server 104 and a network 105.
[0041] Among them, the head mounted display control device 101 is used as the user's human-computer interaction terminal, and has a built-in IMU (Inertial Measurement Unit) sensor, etc., for real-time acquisition of the user's head motion data. The collected head motion data can be sent to the terminal device 103 through the network 105. For example, the terminal device 103 is used as an intermediate processing unit to receive, pre-process and convert the user's head motion data, and map it into action indication information recognizable by the robot 102, such as the direction of travel, standing height, gait control and other parameters. Of course, the collected head motion data can also be sent to the server 104 through the network 105 for centralized processing to complete the intention recognition and control instruction generation, and then sent to the robot 102 through the network 105. Among them, the server 104 undertakes the functions of centrally trained deep learning model deployment, large-scale historical data analysis, etc., and regularly sends the optimized deep learning model to the robot 102 to improve its motion control performance. At the same time, the server 104 supports model lightweight processing, so that the optimized model can adapt to the computing power of the robot 102, thereby ensuring the effective deployment and operation of the model on the robot 102.
[0042] As the execution end, the robot 102 receives control instructions from the terminal device 103 or the server 104, completes the corresponding physical actions, and realizes the user's remote or real-time control intention. In the example implementation of the present disclosure, the robot 102 is a bipedal robot. It is understandable that the head-mounted display control device 101 can also directly transmit the collected head action data to the robot 102. The robot 102 has local computing and intelligent processing capabilities, and can independently complete action feature extraction, user intention recognition (for example, judging whether the user wants the robot to turn, move or adjust the standing height) and generate control instructions, thereby realizing closed-loop control and real-time response, which is particularly suitable for scenarios with unstable networks or high low-latency requirements.
[0043] The terminal device 103 includes but is not limited to desktop computers, portable computers, smart phones, tablet computers, etc. The terminal device 103 can be used as a buffer node to temporarily perform local data preprocessing and command forwarding when the server 104 is unreachable or the network 105 has a high latency. It can also provide users with auxiliary functions such as local operation interface, status feedback, and visual information feedback. It can also be used as a debugging platform or human-machine interface device to assist in monitoring the status of the robot 102, configuring control parameters, or switching interaction modes.
[0044] The network 105 is used as a medium for providing a communication link between the head mounted display control device 101, the robot 102, the terminal device 103 and the server 104. The network 105 may include various connection types, such as wired or wireless communication links or optical fiber cables, etc. By connecting various devices, the network 105 ensures that the robot 102 can obtain the latest model update from the server 104 when necessary, and ensures that the robot 102 can perform independent reasoning locally, reducing dependence on the server 104 and reducing data transmission requirements, thereby improving the real-time and deployability of the system.
[0045] It should be understood that Figure 1 The number and type of head mounted display control devices, terminal devices, bipedal robots, networks and servers in the embodiment are only for illustration. According to the implementation requirements, any number and type of head mounted display control devices, terminal devices, bipedal robots, networks and servers may be provided.
[0046] The present disclosure provides a method for controlling a bipedal robot based on a head mounted display control device. Figure 2 As shown, the method may include the following steps S201 to S203: Step S201, obtaining a user's head movement through a head mounted display control device; Step S202, converting the user's head movement into action instruction information of the biped robot, where the action instruction information includes standing height instruction information; Step S203: input the action instruction information into a robot motion control model based on deep reinforcement learning, so that the robot motion control model outputs an action strategy to the bipedal robot in combination with the action instruction information.
[0047] The bipedal robot control method based on the head-mounted display control device provided by the exemplary embodiment of the present disclosure is implemented. On the one hand, the present disclosure obtains the user's head movement through the head-mounted display control device and determines the standing height indication information based on the movement information, so that the robot can respond to the user's interaction needs more naturally, abandons the dependence on the handle and additional sensors, makes the control of the robot's standing height more in line with the user's natural movement mode, improves the interaction intuitiveness and control fluency, and overcomes the limitations of the existing methods. In addition, the present disclosure overcomes the limitations of the prior art, optimizes the action mapping method, enables the user's operation in the virtual reality environment to more accurately control the robot's standing height, makes the robot motion control more in line with the human body interaction logic, and improves the rationality and intuitiveness of the control. On the other hand, the robot motion control model based on deep reinforcement learning in the present disclosure enables the bipedal robot to autonomously optimize motion decisions in complex environments. Compared with the rule-based control method, the deep reinforcement learning model can improve the recognition ability of different user actions through adaptive learning, and automatically adjust the action strategy in different motion scenes to reduce the control instability caused by user misoperation. In addition, the deep reinforcement learning method can optimize the robot's analysis of action instruction information during the training process, enabling it to more accurately understand and execute the user's standing height adjustment instructions, improve the smoothness and accuracy of the robot's movement, overcome the problems of discrete input signals and control lag in traditional methods, and make the robot's movement more natural, stable and efficient.
[0048] Next, the biped robot control method based on the head mounted display control device in this example embodiment will be described in detail.
[0049] In step S201, the user's head movement is obtained through a head mounted display control device.
[0050] In the example implementation of the present disclosure, a head-mounted display control device refers to a device that has a display function and can sense user actions. The device can manipulate objects or interfaces in a virtual environment by sensing user intentions. Among them, the head-mounted display control device can be a virtual reality head display, an augmented reality head display, or a mixed reality head display. The present disclosure does not limit the type of head-mounted display control device. The device integrates a head posture perception module, can capture the user's head movement in real time and interact with virtual content. Moreover, according to different application requirements, a suitable type of head-mounted display control device can be selected to give full play to its advantages of head movement perception and immersive interaction.
[0051] User head movements refer to head space change signals generated by the user's natural movements. For example, they may include head rotations, such as yaw, pitch, and roll, and may also include changes in head position, such as front-to-back, up-down, left-to-right translation, and the like.
[0052] Taking the head-mounted display control device as a virtual reality headset as an example, the virtual reality headset is equipped with an IMU sensor, which can be used to track the head posture. Among them, the IMU sensor includes an accelerometer, a gyroscope, and a magnetometer, which are used to collect the acceleration, angular velocity, and direction information of the user's head in three-dimensional space in real time, respectively, to accurately capture and analyze the user's head movements. Specifically, the accelerometer captures the translational movement of the head by detecting the motion acceleration in three-dimensional space, the gyroscope senses the rotation change through the angular velocity, and the magnetometer assists in calibrating the direction through the earth's magnetic field data.
[0053] After collecting the raw data from the IMU sensor, data fusion algorithms such as Kalman filtering or complementary filtering can be used for processing to eliminate noise and compensate for the errors of a single sensor, such as the drift of the gyroscope, thereby improving the accuracy and stability of the user's head posture estimation.
[0054] The head movement information obtained in this example can drive the change of perspective in the virtual scene or the operation of the control object in real time, thereby achieving immersive interaction. Moreover, it can also improve the naturalness and intuitiveness of human-computer interaction, allowing users to perform hands-free control through head movements, further enhancing the immersion in the head-mounted display control device, and making the visual feedback highly synchronized with the head movement.
[0055] In step S202, the user's head movement is converted into action instruction information of the biped robot, where the action instruction information includes standing height instruction information.
[0056] After obtaining the user's head movement, the user's head movement information can be parsed into instructions for controlling the movement of the bipedal robot. For example, when the user's head tilts forward at a certain angle and moves forward horizontally, the head movement can be mapped to a "step forward" instruction; for another example, when the user raises his head and moves upward, indicating that he wants the robot to stand higher, the head movement can be mapped to a control instruction to adjust the "standing height".
[0057] For example, a mapping model between the user's head movement and the bipedal robot's action can be pre-established. The mapping model can be based on predefined rules, machine learning methods, or motion intention recognition algorithms, and can be trained with historical data and sensor input data to obtain the optimal conversion strategy, and then the mapping model can be used to convert the user's head movement into the action instruction information of the bipedal robot. Among them, the action instruction information is an intermediate control layer instruction for the bipedal robot, including but not limited to parameters such as standing height, moving direction, gait mode, step frequency, step length, and steering angle. These parameters will be used as inputs to the robot's motion control module to drive the lower control system to perform specific actions.
[0058] In an example implementation, the user's head movement includes a head pitch movement, and the action instruction information includes a standing height instruction information. The head pitch movement refers to the angle change of the user's head in the pitch direction. The standing height instruction information refers to an instruction for controlling the height change of the upper and lower body postures of the bipedal robot, and the continuous control of the height of the center of gravity of the body is finally achieved by jointly adjusting the angles of the hip joint, knee joint and ankle joint of the bipedal robot.
[0059] Accordingly, converting the user's head movement into action instruction information of the bipedal robot may refer to converting the head pitch movement into the standing height instruction information of the bipedal robot. Figure 3 As shown, the process of converting the head pitching action into standing height indication information may include the following steps S301 and S302: Step S301, based on the head pitch movement, obtaining the pitch angle of the user's head relative to the horizontal plane through a sensor.
[0060] The pitch angle refers to the angle at which the user's head is raised or lowered relative to the horizontal plane. For example, when the user raises his head, the pitch angle is positive, and when the user lowers his head, the pitch angle is negative.
[0061] For example, when the user's head moves, the IMU sensor collects the angular changes of the user's head in the pitch direction in real time. Among them, the accelerometer in the IMU sensor can sense the components of the head's gravity in each axis direction, thereby estimating the tilt angle of the head in a static state, and the gyroscope measures the angular velocity to track the angular changes of the head in a dynamic process. The data collected by the gyroscope and accelerometer are fused, such as through complementary filtering or Kalman filtering algorithms, to obtain the current pitch angle of the user's head, and update it in real time as a key input parameter for identifying the user's action intention and generating control instructions.
[0062] Step S302: Map the pitch angle to standing height indication information.
[0063] After obtaining the pitch angle of the user's head relative to the horizontal plane, the user's control intention can be inferred based on the pitch angle, and the angle value can be converted into a control parameter for controlling the standing height of the bipedal robot, that is, the standing height indication information is obtained. For example, when the user raises his head, the pitch angle is positive, indicating that the user may want the bipedal robot to stand higher or enter a fully standing state; and when the user lowers his head, the pitch angle is negative, indicating that the user may want the bipedal robot to lower its posture or enter a half-squat state.
[0064] For example, the pitch angle value can be mapped to the standing height indication information according to a preset mapping relationship, including linear function mapping, nonlinear function mapping, lookup table mapping, neural network mapping, etc. Among them, when linear function mapping and nonlinear function mapping are adopted, the mapping function can be personalized according to the user's habits, angle variation range and the movement ability of the bipedal robot, and smoothing filtering can be used to avoid frequent changes in standing height caused by small head shaking. This mapping process enables users to quickly and intuitively control the height of the robot through simple and natural head pitch movements, thereby improving the efficiency and immersion of human-computer interaction.
[0065] For example, when a nonlinear function is used to map the pitch angle value to the standing height indication information, there is: (1) in, It is the standing height indication information; is the pitch angle; , , For preset parameters, parameters a Controls the amplitude of the output height change, that is, the standing height relative to the reference value c The maximum offset of b Controls the sensitivity of the angle input to output changes, b The larger the value, the steeper the function changes near the zero point and the more sensitive the response. b The smaller the value, the smoother the response; parameter c Indicates the base value of standing height, such as θ =0 (i.e. head level), parameter c This is the normal standing height that a bipedal robot should maintain.
[0066] From formula (1), we can see that when the user looks up, θ >0, , so that H > c , indicating that the bipedal robot needs to stand higher; when the user lowers his head, θ <0, , so that H < c , indicating that the bipedal robot needs to lower its posture.
[0067] Different users can choose different a , b , c Therefore, user personalized adaptation can be achieved through parameter adjustment to obtain ideal control sensitivity and posture response.
[0068] It should be noted that in order to make the mapping method shown in formula (1) more adaptable, the neural network can automatically learn and output the parameters a , b , c , making the nonlinear function an adjustable and dynamically adaptive control model.
[0069] For example, build a neural network model that takes the user's current state or features as input and outputs parameters a , b , c Among them, the input features may include the current pitch angle, the historical sequence or change rate of the pitch angle, the user ID, the current robot state such as the center of gravity height, the environmental context such as the terrain flatness, the task mode, etc. The network structure of the model can adopt a multi-layer perceptron, including several fully connected layers and nonlinear activation functions, and the output layer is three neurons, corresponding to the parameters a , b , c During the model training process, a data set containing multiple sets of data samples can be used for supervised learning of the model. Each set of data includes input features and the corresponding ideal standing height. The network weights are optimized through the back propagation algorithm so that the three output prediction parameters can approach the ideal mapping.
[0070] In addition, to ensure numerical stability, the output value can be restricted. Specifically, using activation functions such as sigmoid Function to restrict prediction parameters a Positive value to control the range; limit the prediction parameters b The range of the function should avoid being too steep or too flat to control the sensitivity; c Add a constant offset to satisfy the standing height constraint.
[0071] In this example, a natural nonlinear mapping is achieved, which makes the control of small-range head movements delicate and the response of large-scale movements gradual, improving the continuity and comfort of interaction. In addition, the output range of the function is controllable, and the parameters are set to ensure that the standing height varies within a safe range, thereby enhancing the robustness of the system. In formula (1), the nonlinear characteristics of the hyperbolic tangent function tanh(⋅) are introduced to make the mapping relationship have a smooth, finite and asymptotic response effect, which is suitable for processing the natural limitations and small jitters in the user's pitch movement. Importantly, the mathematical form is simple and the computational overhead is low, which is suitable for deployment in the real-time control module of the robot body, supports high-frequency updates and low-latency responses, and can achieve more effective and stable motion control.
[0072] For another example, the user's pitch angle and related action information can be input into a trained neural network model, and the model automatically learns the mapping relationship between the pitch angle and the standing height, thereby outputting the corresponding standing height indication information. Among them, the neural network model can be a multi-layer perceptron, a recurrent neural network such as a long short-term memory network, a one-dimensional convolutional neural network, etc., and the present disclosure does not limit the specific type of the neural network model. The pitch angle used to input the neural network model includes the pitch angle of the current frame and the pitch angle sequence of the past several frames. In addition, auxiliary features such as the pitch angle change rate and the user's height can also be included. In addition, a variety of samples including different users, different postures, and different speeds can be collected to supervise the learning of the neural network model. The present disclosure does not elaborate on the training process of the model.
[0073] The resulting standing height indication information will be sent to the bipedal robot as control input to adjust the angles of the hip, knee, ankle and other joints, thereby changing its overall standing height to meet the user's wishes. Figure 1 The action feedback.
[0074] In an example implementation, the user's head movement also includes a head rotation movement, and the action indication information also includes travel direction indication information. The head rotation movement refers to the rotational movement of the user's head around a vertical axis, such as when the user turns his head to the left or right. The travel direction indication information refers to an instruction for controlling the movement direction of the bipedal robot, which defines the direction in which the bipedal robot should move next. The travel direction indication information can be an absolute direction angle, such as a rotation angle relative to an initial direction, or a direction offset value relative to a current motion state, which is not limited in the present disclosure. The travel direction indication information serves as an important input for the bipedal robot's path planning and gait generation, and determines the robot's movement trajectory in a plane space.
[0075] Accordingly, converting the user's head movement into action instruction information of the biped robot may refer to converting the head rotation movement into the moving direction instruction information of the biped robot. Figure 4 As shown, the process of converting the head pitching action into standing height indication information may include the following steps S401 and S402: Step S401, based on the head rotation action, obtain the horizontal rotation angle of the user's head through the sensor.
[0076] Among them, the horizontal rotation angle refers to the offset angle of the user's facing direction relative to the initial direction, which reflects the user's head turning movement in the horizontal direction, such as looking left or right. For example, when the user's head moves, the IMU sensor detects the rotational movement of the user's head around the vertical axis (usually the Z axis) in real time to obtain the horizontal rotation angle of the user's head. For example, the gyroscope in the IMU sensor can detect the angle change of the user's head in the horizontal direction per second, and the current horizontal rotation angle can be obtained by integrating over time. The accelerometer mainly assists in obtaining the direction of gravity, which is used to correct the posture when the device is stationary.
[0077] The horizontal rotation angle of the user's head can be used as one of the core inputs for identifying the user's intention and used to control the bipedal robot's direction of travel, turning action or perspective synchronization, so that the bipedal robot can automatically adjust its movement direction according to the user's head orientation, achieving natural and intuitive interactive control.
[0078] Step S402: Map the horizontal rotation angle into traveling direction indication information.
[0079] After obtaining the horizontal rotation angle of the user's head, the user's control intention can be inferred based on the horizontal rotation angle, and the angle value can be converted into a control parameter for controlling the direction of travel of the bipedal robot, that is, the direction of travel indication information is obtained. For example, when the user's head rotates to the right, the horizontal rotation angle is positive, indicating that the user may want the bipedal robot to move to the right front or turn right; and when the user's head rotates to the left, the horizontal rotation angle is negative, indicating that the user may want the robot to move to the left front or turn left.
[0080] Exemplarily, a variety of mapping methods can also be used to map the horizontal rotation angle to the direction of travel indication information, including linear function mapping, nonlinear function mapping, lookup table mapping or neural network mapping. Among them, linear function mapping can directly convert the horizontal rotation angle into the direction of travel angle. Nonlinear function mapping can introduce a smoothly changing response curve, so that the bipedal robot remains stable for small angle rotation and responds quickly to large angle rotation. Lookup table mapping can correspond the horizontal rotation angle to discrete direction control instructions in segments to simplify the control logic. Neural network mapping can learn the complex relationship between the horizontal rotation angle and the desired direction of travel based on historical actions, user habits and real-time status, and realize adaptive adjustment. This mapping process enables users to quickly and intuitively control the direction of travel or turning behavior of the bipedal robot through simple and natural head rotation movements, which improves the freedom, intuitiveness and immersion of human-computer interaction.
[0081] In order to avoid frequent changes in the robot's direction due to slight head shaking or unintentional movements, angle thresholds, dead zone processing, and smoothing filtering mechanisms can usually be introduced during the mapping process to trigger the adjustment of the travel direction only when the user's movements exceed a certain amplitude or duration.
[0082] In this example, linking the head rotation action with the travel direction indication information is a natural and efficient way of human-computer interaction. The user only needs to turn his head to express the direction he wants the robot to go, without using a handle or voice commands. It is fast in response, highly intuitive, and has a higher degree of control freedom.
[0083] In the example implementation of the present disclosure, the action instruction information also includes terminal position instruction information, wherein the terminal includes the wrist or robot arm of the biped robot. At this time, it is necessary to input the hand position relative to the head through the head-mounted display control device or the handle, and the position is mapped to the terminal position relative to the robot body. Furthermore, the terminal position instruction information can be used as the input target of the reinforcement learning policy network to learn how to control the angles of each joint of the biped robot to drive the terminal to approach the target position, thereby enhancing the policy network's ability to understand and execute the terminal's motion intention during task execution.
[0084] In step S203, the action instruction information is input into a robot motion control model based on deep reinforcement learning, so that the robot motion control model outputs an action strategy to the bipedal robot in combination with the action instruction information.
[0085] In an example implementation of the present disclosure, action indication information such as standing height indication information, travel direction indication information, etc. is passed as an input signal to a robot motion control model obtained through deep reinforcement learning training, so that the model can generate an action strategy suitable for the bipedal robot based on the action indication information based on the perception of the current environmental state. The output action strategy refers to the output of specific actions that the bipedal robot should take at each moment, such as control instructions for each joint.
[0086] Among them, the robot motion control model based on deep reinforcement learning is a policy function learned through a large amount of interactive experience. Deep reinforcement learning includes but is not limited to algorithms such as PPO (Proximal Policy Optimization), DDPG (Deep Deterministic Policy Gradient), and SAC (Soft Actor-Critic). The present disclosure does not limit the specific algorithm type of deep reinforcement learning.
[0087] Exemplarily, the state space, action space and reward function in deep reinforcement learning can be defined. The state space may include the current posture information, environmental information and action instruction information of the bipedal robot, wherein the current posture information includes joint angles, speeds, center of gravity positions, etc.; the environmental information includes terrain height maps, obstacle distribution, etc. The action space is the motion control instructions that the bipedal robot can execute, such as the target angular velocity or torque of each joint. The reward function is set according to the goal of the bipedal robot's behavior, such as maintaining balance, following the user's intention to move, etc., which is used as the basis for the model learning strategy.
[0088] By training the policy network, the network can output continuous and stable control actions when facing different input instructions and states. During the training process, a simulator is used to build a simulation environment, repeatedly sample actions and evaluate rewards, and optimize the policy network through gradient updates. When actually deployed, the trained policy network is used as part of the robot control system, receiving action instruction information in real time and sending it to the policy network, and finally outputting the corresponding action strategy.
[0089] It is understandable that before inputting the action instruction information into the robot motion control model, the robot motion control model needs to be pre-trained so that the model has the ability to understand and respond to different action instruction information.
[0090] In an example implementation, a teacher-student model framework can be used to train a policy network for outputting an action strategy for controlling the motion of a bipedal robot, thereby obtaining a robot motion control model. The input data of the teacher-student model framework includes action instruction information, and the action instruction information includes at least one of standing height instruction information, travel direction instruction information, and end position instruction information. The teacher model is a reference policy source with superior performance, the student model is a target policy network, and the policy network is a neural network structure that maps states and instructions to low-level control actions.
[0091] Specifically, a teacher model is first constructed. The model can be a high-performance policy network trained by deep reinforcement learning, or a traditional controller based on expert experience, which has the ability to generate optimal or approximately optimal action strategies under given states and instructions. In the model training stage, input samples are constructed, which include action instruction information from the user and the current robot state information, such as joint angles, speeds, center of gravity position, ground contact state, etc. The teacher model can output the corresponding high-quality action strategy and provide it to the student model as a supervisory signal. The student model is the policy network to be trained. The network structure can be a multi-layer perceptron, a convolutional network, or a module with an attention mechanism. It is trained by imitating the behavior of the teacher model, and a supervised learning method is used to minimize the output difference between the student strategy and the teacher strategy. During the training process, the output of the teacher model is used as label data, and the student model learns the mapping relationship through back propagation, so as to obtain efficient training without relying on environmental interaction. After the training is completed, the student model becomes a robot motion control model, which can be used independently in actual deployment, and can quickly output action strategies only by relying on action instruction information and robot state.
[0092] For example, reference Figure 5 As shown, a schematic diagram of a two-stage training-inference framework of a robot motion control model is shown. Among them, the training stage combines the teacher-student encoder structure and the deep reinforcement learning framework. The teacher-student encoder structure includes a teacher encoder 501 and a student encoder 502. The deep reinforcement learning framework is composed of a policy network 503, a value network 504 and a PPO algorithm. In the training stage, by introducing privileged state information to guide policy learning, the student encoder 502 can stably output high-quality control actions only by relying on observable data in the inference stage. The observable data includes current motion state information and action instruction information.
[0093] It should be noted that the teacher encoder 501 is only used to extract high-dimensional semantic representations from the privileged state information during the training phase, which is used as the input of the strategy to guide learning. That is to say, the privileged state information and the teacher encoder 501 will no longer be used during the reasoning phase. The student encoder 502 encodes the historical motion state of the biped robot itself during the training and reasoning phases to obtain dynamic features related to action decisions, and then connects the current motion information and action instruction information to the strategy network 503 to generate control actions.
[0094] In addition, the student encoder 502 ultimately needs to imitate the representation output by the teacher encoder 501, so distillation training is performed by minimizing MSE (mean square error) during the training process.
[0095] The policy network 503 is the core module of action generation. In the training phase, the policy network 503 receives three inputs, namely, the potential vector, current motion information, and action instruction information from the teacher encoder 501 or the student encoder 502, and combines with the value network 504 to use the PPO algorithm to update the policy. In the inference phase, the policy network 503 still receives three inputs, namely, the potential vector, current motion information, and action instruction information from the student encoder 502, and outputs the action strategy for controlling the movement of the biped robot.
[0096] based on Figure 5 The model structure diagram shown in Figure 6 As shown, the process of training a policy network for outputting an action policy for controlling the motion of a biped robot using a teacher-student model framework may include the following steps S601 to S606: Step S601, encode the historical motion state information of the biped robot itself through the student encoder and generate a first latent vector, and encode the privileged state information of the biped robot through the teacher encoder and generate a second latent vector.
[0097] The historical motion state information of the biped robot itself may include the joint angles, joint speeds, sole contact states, center of gravity trajectories, etc. of the past several frames. For example, the historical motion state sequence of the past 10 frames is input into the student encoder 502 for encoding to obtain the first latent vector, which is recorded as The first latent vector can capture the continuity and dynamic characteristics of the biped robot's motion.
[0098] The privileged state information may include the actual contact force between the biped robot and the ground, the environmental height map, disturbance information, etc. The privileged state information obtained by the biped robot in the simulation training is input into the teacher encoder 501 for encoding to obtain a second latent vector, which is recorded as The second latent vector contains more precise motion semantic information.
[0099] It should be noted that the privileged state information can only be obtained during the training phase but cannot be observed during the testing or deployment phase. Therefore, the teacher encoder 501 can use the complete information to learn the best potential expression, thereby guiding the student encoder 502 to learn. Both the teacher encoder 501 and the student encoder 502 can compress the high-dimensional, time-series state information into a low-dimensional potential representation, providing the policy network 503 with behavioral semantic representations from different sources. For example, the teacher encoder 501 and the student encoder 502 can be multi-layer perceptrons, time-series convolutional networks, or recurrent neural networks, and have the ability to extract time-series features and compress them into fixed-length semantic vectors. In addition, the network architecture of the teacher encoder 501 and the network architecture of the student encoder 502 can be the same or different, and the present disclosure does not limit this.
[0100] Step S602: select a first latent vector or a second latent vector according to a preset strategy and input it into a strategy network.
[0101] According to the preset strategy, the first latent vector generated by the student encoder 502 or the second latent vector generated by the teacher encoder 501 is sent to the strategy network 503 for decision making. The preset strategy can be set according to the real-time environment state, the training stage or a specific performance indicator, and the specific performance indicator includes an action execution error threshold, a difference between the simulation and the real environment, etc.
[0102] For example, one preset strategy is to preferentially use the high-quality second latent vector generated by the teacher encoder 501 to guide the strategy network 503 to quickly converge to a near-optimal solution at the beginning of training. After the student encoder 502 is optimized through knowledge distillation, it gradually transitions to using only the first latent vector in the deployment phase to reduce the reliance on privileged information. For another example, another preset strategy is to use the first latent vector in proportion to the number of latent vectors generated by the teacher encoder 501. p Select the second latent vector according to 1- p Select the first latent vector and gradually reduce p Of course, the first latent vector and the second latent vector may be concatenated and sent to the policy network 503, and dynamically weighted through an attention mechanism or a gating module, so that the policy network 503 can flexibly combine historical experience and privileged knowledge in complex scenarios.
[0103] The selective input mechanism can not only accelerate the training process by utilizing the ideal state prior knowledge provided by the teacher encoder 501, but also cope with sensor limitations or environmental disturbances through the generalization ability of the student encoder 502 in actual operation. At the same time, the policy network 503 conditions the latent vector to achieve smooth switching and robust decision-making of motion control. For example, when the robot encounters unknown terrain, it preferentially adjusts the gait based on the second latent vector, while relying on the first latent vector to maintain efficiency in the stable walking stage. Finally, the closed-loop feedback optimizes the strategy selection rules, so that the bipedal robot can balance motion performance and adaptability in different stages and environmental conditions.
[0104] In addition, in addition to the selected latent vector, action instruction information and current motion state information of the biped robot, such as joint position and speed, etc., may also be input into the strategy network 503 for decision making.
[0105] Step S603: outputting a motion strategy for controlling the motion of the biped robot based on the received first latent vector or the second latent vector, the action instruction information, and the current motion state information through the strategy network.
[0106] The strategy network 503 may be a multi-layer fully connected perceptron, a Transformer structure, etc. The strategy network 503 performs multimodal feature fusion on the first latent vector or the second latent vector, the action instruction information, and the current motion state information, and outputs an action strategy, which is recorded as .
[0107] For example, the first latent vector or the second latent vector is concatenated or weightedly interacted with the current motion state information in the embedding space, and combined with the action instruction information as conditional input, and action strategies such as joint angle targets, torque instructions or gait phase parameters are generated through nonlinear transformation.
[0108] Step S604, using the privileged state information through the value network, outputs the value estimate of the current state.
[0109] During the training of the policy network 503 , the long-term return of the current state is estimated through the value network 504 , that is, starting from this state, if the current strategy is continuously executed, how much cumulative reward can be obtained in the future, so as to optimize the policy network 503 .
[0110] The value network 504 may be a multi-layer fully connected perceptron. For example, when the value network 504 uses the privileged state information to estimate the value of the current state, the privileged state information is first encoded into a high-dimensional feature vector, such as extracting dynamic features through a convolution or fully connected layer, and then fused with the current state information in the latent space, and then outputted through a multi-layer nonlinear transformation to represent the value estimate of the current state, which is recorded as This value estimate is used for policy optimization in reinforcement learning to guide the policy network 503 to learn better behaviors.
[0111] Step S605 , optimizing parameters of the student encoder based on the difference between the first latent vector and the second latent vector.
[0112] This step quantifies the distribution difference between the two in the latent space and back-propagates the difference gradient to update the network weights of the student encoder 502, thereby guiding the student encoder 502 to learn to generate a latent representation close to the teacher encoder 501, thereby achieving teacher knowledge distillation.
[0113] Exemplarily, during the training phase, the parameters of the teacher encoder 501 are fixed, and the historical motion state information and the corresponding motion state information are taken as parallel inputs. After the latent vectors are generated by the student encoder 502 and the teacher encoder 501 respectively, the distance between the two is minimized using a contrastive learning framework, or adversarial training is used to make the output distribution of the student encoder 502 approach the latent space characteristics of the teacher encoder 501. At the same time, noise injection or data enhancement is introduced to simulate sensor errors in actual deployment, forcing the student encoder 502 to extract feature expressions compatible with privileged information encoding under the condition of limited input information.
[0114] For example, a difference loss function, such as least squares difference or KL divergence, can be constructed to make the output of the student encoder 502 as close as possible to the teacher encoder 501. By minimizing the loss, the parameters of the student encoder 502 are optimized so that it can learn to extract feature expressions close to the privileged information from the historical motion state, thereby enhancing the generalization ability of the policy network 503 and improving the decision quality of the policy network 503 in the real environment, so that it can be close to the teacher level when deployed without relying on privileged information.
[0115] Step S606: Based on the action strategy and value estimation, the parameters of the policy network are updated using a deep reinforcement learning algorithm.
[0116] Take the PPO algorithm as an example. During the training of the policy network 503, the current policy network 503 is used to interact with the environment. Select action strategy , returns the reward after executing the action and the next state , and the interaction trajectory sequence is obtained by sampling ( ), which is used for subsequent strategy optimization. Next, the value network 504 is introduced to evaluate the value of each state In order to optimize the strategy more stably and efficiently, the policy gradient calculation can be performed through the advantage function to obtain the parameter update direction of the policy network 503.
[0117] For example, the advantage function is defined as: (2) in, is the advantage function, is the odds approximation obtained using the generalized odds estimate.
[0118] Then, the optimization target is constructed according to the clipping objective function of the PPO algorithm, and the policy gradient is calculated through back propagation to optimize the parameters of the policy network 503 so that it can continuously improve the expected cumulative return under the current policy. At the same time, the state-reward pair ( ) Perform supervised learning on the value network 504 to minimize its output With real returns The mean square error between them is thus improved, thereby improving the estimation accuracy of the state value by the value network 504.
[0119] The student encoder 502 performs latent space alignment with the output of the teacher encoder 501 by contrastive learning or knowledge distillation, forcing the student network to infer a latent expression close to the encoding result of the privileged information through historical states in actual deployment scenarios lacking privileged information, thereby improving the robustness and adaptability of the motion strategy. In the example implementation of the present disclosure, the bipedal robot can simulate motion decisions optimized based on privileged information in a real environment by relying only on its own sensor data, effectively solving the performance degradation problem caused by unobservable states or noise interference in actual deployment, and at the same time enhancing the generalization ability of the motion strategy through implicit knowledge transfer of the latent vector.
[0120] Furthermore, the action instruction information is input into the robot motion control model based on deep reinforcement learning, and the robot motion control model can output the action strategy to the bipedal robot in combination with the action instruction information. Figure 5 The model structure diagram shown is explained in detail. Figure 7 As shown, the process of outputting the action strategy in the reasoning phase may include the following steps S701 and S702: Step S701, encoding the historical motion state information of the biped robot itself through the student encoder to generate a first latent vector, and inputting the first latent vector into the policy network.
[0121] In this step, the first latent vector output by the student encoder 502 carries rich historical behavior semantics, such as "whether to prepare to take a step", "whether there is a center of gravity shift", "whether there is an unstable posture", etc., thereby helping the policy network 503 to perform more context-aware control. The first latent vector will be input into the policy network 503 as a high-level abstract input together with other information for subsequent action strategy generation.
[0122] Step S702: outputting a motion strategy for controlling the motion of the biped robot based on the received first potential vector, the action instruction information, and the current motion state information through the strategy network.
[0123] After the policy network 503 completes training, the action instruction information obtained in real time can be input into the trained policy network 503, and the control action that meets the instruction can be output immediately to ensure the real-time and stability of the system response.
[0124] Exemplarily, the policy network 503 receives the first latent vector, the current motion state, and the action instruction information, and jointly encodes the three through the pre-trained multimodal fusion layer to generate a unified feature representation of comprehensive historical experience, real-time environmental perception, and task objectives, and then outputs the action strategy through the multi-layer nonlinear transformation of the policy network 503, such as the joint target angle and expected torque at the next moment. Furthermore, after obtaining the action strategy, the strategy directly drives the underlying controller to generate motor commands to achieve real-time motion control of the bipedal robot.
[0125] The robot motion control model based on deep reinforcement learning in this example enables the bipedal robot to autonomously optimize motion decisions in complex environments. Moreover, the deep reinforcement learning model can improve the ability to recognize different user actions through adaptive learning, and automatically adjust the motion strategy in different motion scenarios to reduce the control instability caused by user misoperation. In particular, by introducing detailed instructions such as standing height, the motion stability and flexibility of the bipedal robot are optimized, which helps its performance in complex terrain or dynamic tasks.
[0126] It should be noted that in the exemplary implementation of the present disclosure, information transmission between functional modules can be realized through a publish-subscribe model. For example, user input information is obtained from a head-mounted display control device by subscribing to an mros topic, and interactive information such as device location and user head movement is published as a message to a preset mros topic by connecting to the local area network where the bipedal robot is located. The robot control end, as a topic subscriber, receives these input data in real time and parses them to obtain action instruction information, such as the direction of travel, standing height adjustment or a specific trigger action. Subsequently, the action instruction information is input into the strategy network together with the current robot state information, and the action strategy at the current moment is output. These control commands are executed to each joint of the bipedal robot body through a real-time control channel such as a CAN (controller area network) bus inside the bipedal robot to realize actual action execution. At the same time, the RGB image stream obtained by the top camera of the bipedal robot is transmitted back to the head-mounted display control device in real time through the local area network, and is performed in the virtual screen in the form of a video window or a virtual screen, so that the operator can perceive the real environment state of the robot from a first-person perspective, thereby realizing efficient and immersive closed-loop interactive control.
[0127] In addition, it is also possible to obtain image information of the environment in which the bipedal robot is located, and merge the image information with the display screen of the head-mounted display control device to generate a digital twin scene, and display it in real time on the head-mounted display control device.
[0128] The image information of the environment where the bipedal robot is located refers to the visual sensor data obtained by the bipedal robot, which can reflect the real physical environment. The display screen of the head-mounted display control device is the visual output interface generated by the system for the user. The fusion of image information and the display screen of the head-mounted display control device refers to the process of visually matching the real image and the virtual image in space and displaying them uniformly. The digital twin scene refers to a virtual-real synchronized three-dimensional visualization space constructed based on the bipedal robot and the environmental state.
[0129] Exemplarily, the visual perception device deployed on the biped robot can obtain image information of its environment, including but not limited to RGB (three-channel image) images, depth images, stereoscopic vision data or three-dimensional point cloud data. Among them, the visual perception device can be a front camera, a depth camera or a panoramic camera, etc., which can obtain the actual environment structure, terrain status, dynamic obstacles, etc. around the biped robot. After obtaining the original image information, the original image is processed, such as image distortion correction, depth registration, semantic segmentation, posture estimation, etc., to build a digital model of the real world.
[0130] Then, the image information is transmitted to the head-mounted display control device in real time so as to be integrated with the original virtual display screen of the device, thereby generating a digital twin scene and displaying it in real time in the head-mounted display control device. For example, through coordinate transformation, projection mapping and lighting consistency processing, the collected real environment image and the virtual scene model are aligned in space and vision, thereby constructing a dynamic space scene with a real physical background and enhanced virtual information. Users can observe the state of the bipedal robot, environmental details and interactive targets from a first-person perspective or a free perspective in the head-mounted display control device, realizing a space perception and control interface shared by humans and machines.
[0131] In this example, on the one hand, the user's perception of the bipedal robot's motion environment is significantly enhanced, allowing it to perform precise remote operations in complex or invisible scenes; on the other hand, the superposition of virtual information and the real world is achieved, allowing users to more intuitively understand the bipedal robot's task progress and environmental feedback, improving interaction efficiency and safety. In addition, this process can also work in conjunction with voice control, motion capture, and other methods to build a more natural and efficient human-computer interaction system.
[0132] refer to Figure 8 As shown, a schematic diagram of a scenario in which a biped robot control method based on a head mounted display control device in an embodiment of the present disclosure is applied is given. Figure 8In the embodiment, the user wears a head mounted display control device, and the built-in sensor of the head mounted display control device collects the head posture data of the user in real time, such as obtaining the pitch motion 801 of the user's head. Then, the pitch motion 801 is converted into standing height indication information 802 understandable to the bipedal robot 804. Further, the standing height indication information 802 is input into the robot motion control model 803 based on deep reinforcement learning, and the corresponding action strategy is output to the bipedal robot 804, and the bipedal robot 804 performs the corresponding action, thereby adjusting the standing height 805 of the bipedal robot 804, such as controlling the bipedal robot to enter a fully standing state or a half-squatting state.
[0133] By obtaining the user's head movements through the head-mounted display control device and determining the standing height indication information based on the movement information, the robot can respond to the user's interaction needs more naturally, eliminating the dependence on handles and additional sensors, making the control of the robot's standing height more in line with the user's natural movement pattern, improving the interactive intuitiveness and control fluency.
[0134] Furthermore, in the exemplary implementation of the present disclosure, a bipedal robot control system based on a head mounted display control device is also provided. Fig. 9 As shown, a biped robot control system 900 based on a head mounted display control device may include a head mounted display control device 101, an indication information generating module 901 and an action strategy generating module 902, wherein: A head mounted display control device 101, used to obtain the user's head movement; An instruction information generating module 901 is used to convert the user's head movement into action instruction information of the biped robot, where the action instruction information includes standing height instruction information; The action strategy generation module 902 is used to input the action instruction information into the robot motion control model based on deep reinforcement learning, so that the robot motion control model outputs the action strategy to the bipedal robot in combination with the action instruction information.
[0135] The specific details of each module in the above-mentioned bipedal robot control system based on head-mounted display control device have been described in detail in the corresponding bipedal robot control method based on head-mounted display control device, so they will not be repeated here.
[0136] In the exemplary implementation of the present disclosure, a bipedal robot is also provided, the bipedal robot includes a processor and a memory, the memory stores computer-readable instructions, and the computer-readable instructions implement the above method when executed by the processor. The bipedal robot includes any one of a legged robot, a wheeled robot, a wheel-footed robot, a humanoid robot, a cleaning robot, a transport robot, and a mobile robot. Fig.10 and Fig.11As shown, schematic diagrams of two different bipedal robots are shown respectively.
[0137] refer to Fig.12 As shown, an electronic device capable of implementing the above method is also provided. The electronic device 1200 includes a processor 1201 and a memory 1202, and the memory 1202 stores computer-readable instructions, which implement the method in the embodiment of the present disclosure when executed by the processor 1201.
[0138] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is further provided, on which computer program code instructions are stored. When the computer program code instructions are called by a processor of a robot, the robot executes the method in the embodiment.
[0139] refer to Fig.13 As shown, a program product 1300 for implementing the above method according to an embodiment of the present disclosure is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, apparatus, or device.
[0140] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the embodiment of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CDROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiment of the present disclosure.
[0141] Finally, the above preferred embodiments are only used to illustrate the technical solutions of the present application and are not restrictive. Although the present application has been described in detail, those skilled in the art should understand that changes in form and details can be made without departing from the scope defined by the claims of the present application. The dimensions of the drawings have nothing to do with the specific objects, and the dimensions of the objects can be changed arbitrarily.
Claims
1. A biped robot control method based on a head mounted display control device, characterized in that: include: Acquire the user's head movements through a head-mounted display control device; Converting the user's head movement into action instruction information of the biped robot, wherein the action instruction information includes standing height instruction information; The action instruction information is input into a robot motion control model based on deep reinforcement learning, so that the robot motion control model outputs an action strategy to the bipedal robot in combination with the action instruction information.
2. The biped robot control method based on a head mounted display control device according to claim 1, characterized in that: The user's head movement includes a head pitch movement; The step of converting the user's head movement into action instruction information of the biped robot includes: The head pitching action is converted into the standing height indication information.
3. The biped robot control method based on a head mounted display control device according to claim 2, characterized in that: The converting the head pitching action into the standing height indication information comprises: Based on the head pitch action, obtaining the pitch angle of the user's head relative to the horizontal plane through a sensor; The pitch angle is mapped to the standing height indication information.
4. The biped robot control method based on a head mounted display control device according to claim 3, characterized in that: The mapping the pitch angle to the standing height indication information includes: according to: Obtaining the standing height indication information; in, The standing height indication information; is the pitch angle; , , are preset parameters.
5. The biped robot control method based on a head mounted display control device according to claim 1, characterized in that: The user head movement includes a head rotation movement, and the action indication information includes a travel direction indication information; The step of converting the user's head movement into action instruction information of the biped robot includes: The head rotation action is converted into traveling direction indication information of the biped robot.
6. The biped robot control method based on a head mounted display control device according to claim 5, characterized in that: The step of converting the head rotation action into the moving direction indication information of the biped robot includes: Based on the head rotation action, obtaining the horizontal rotation angle of the user's head through a sensor; The horizontal rotation angle is mapped into the traveling direction indication information.
7. The biped robot control method based on a head mounted display control device according to claim 1, characterized in that: The method further comprises: Obtain image information of the environment where the biped robot is located; The image information is integrated with the display screen of the head-mounted display control device to generate a digital twin scene, and the scene is displayed in real time on the head-mounted display control device.
8. The biped robot control method based on a head mounted display control device according to any one of claims 1 to 7, characterized in that: Before inputting the action instruction information into the robot motion control model based on deep reinforcement learning, the method further includes: Using a teacher-student model framework, a policy network for outputting an action strategy for controlling the motion of a biped robot is trained to obtain the robot motion control model; Among them, the input data of the teacher-student model framework includes the action instruction information.
9. The biped robot control method based on a head mounted display control device according to claim 8, characterized in that: The method of using a teacher-student model framework to train a policy network for outputting an action policy for controlling the movement of a biped robot includes: The student encoder encodes the historical motion state information of the biped robot itself and generates a first latent vector, and the teacher encoder encodes the privileged state information of the biped robot and generates a second latent vector; Selecting a first latent vector or a second latent vector according to a preset strategy and inputting the first latent vector into the strategy network; Outputting, through the strategy network, a motion strategy for controlling the movement of the biped robot based on the received first latent vector or second latent vector, the action instruction information, and the current motion state information; Utilizing the privileged state information through a value network, outputting a value estimate of the current state; optimizing parameters of the student encoder based on a difference between the first latent vector and the second latent vector; Based on the action strategy and the value estimate, the parameters of the policy network are updated using a deep reinforcement learning algorithm.
10. The biped robot control method based on a head mounted display control device according to claim 8, characterized in that: The robot motion control model outputs a motion strategy to the biped robot in combination with the action instruction information, including: Encoding the historical motion state information of the biped robot itself through a student encoder to generate a first latent vector, and inputting the first latent vector into the policy network; The action strategy for controlling the movement of the biped robot is outputted through the strategy network based on the received first potential vector, the action instruction information and the current motion state information.
11. The biped robot control method based on a head mounted display control device according to claim 1, characterized in that: The head mounted display control device includes a virtual reality head mounted display, an augmented reality head mounted display or a mixed reality head mounted display.
12. A bipedal robot control system based on a head mounted display control device, characterized in that: include: A head-mounted display control device for acquiring the user's head movements; an indication information generating module, used for converting the user's head movement into action indication information of the biped robot, wherein the action indication information includes standing height indication information; An action strategy generation module is used to input the action instruction information into a robot motion control model based on deep reinforcement learning, so that the robot motion control model outputs an action strategy to the bipedal robot in combination with the action instruction information.
13. An electronic device, characterized in that: include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the bipedal robot control method based on a head-mounted display control device as described in any one of claims 1 to 11.
14. A bipedal robot, characterized in that: include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the bipedal robot control method based on a head-mounted display control device as described in any one of claims 1 to 11.
15. The biped robot according to claim 14, characterized in that: The bipedal robot includes any one of a footed robot, a wheeled robot, a wheel-footed robot, a humanoid robot, a cleaning robot, a transport robot, and a mobile robot.
16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program code instructions, and when the computer program code instructions are called by the processor of the robot, the robot executes the bipedal robot control method based on the head-mounted display control device as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Remote robot control method and system based on digital virtual human driving
CN114227679A
Mobile robot interaction operation system based on Hololens
CN115185368A
Wireless electroencephalogram-based control system for controlling crawler type mobile robot
CN106406297A
Immersive remote interaction system and method based on natural walking
CN111716365A
Robot reinforcement learning control method and system based on gating circulation unit
CN119536333A
Cited By
Four-wheel-foot robot shock resistance control method and system based on time sequence contrast learning
CN122221903A