Multi-style vehicle automatic driving method for humanoid robot

By constructing a multi-style vehicle autonomous driving model, combining deep reinforcement learning and physical control strategies, the problem of humanoid robots achieving driving performance in complex environments is solved, and the steering wheel is achieved compliant control and adaptive driving are achieved.

CN120039279AActive Publication Date: 2025-05-27NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510532931.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The prior art is difficult to enable humanoid robots to achieve driving performance comparable to human drivers in the physical field, especially in complex, dynamic multimodal environments.

Method used

The multi-style vehicle autonomous driving method is adopted, and by constructing a multi-style vehicle autonomous driving model, the decision-making problem is modeled as a Markov decision-making process, and the feature extraction module, the autonomous driving preference alignment module and the humanoid robot admission control module are used to combine deep reinforcement learning and physical control strategies to optimize the action strategy and achieve the compliance control of the steering wheel.

Benefits of technology

It realizes the compliant and adaptive control of the steering wheel by the humanoid robot, and can generate reasonable driving instructions in complex environments, adapt to different environments and driving needs, and improves the complexity and flexibility of the robot driving the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120039279A_ABST
    Figure CN120039279A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-style vehicle automatic driving method for a humanoid robot. The method comprises the following steps: constructing a multi-style vehicle automatic driving model by an automatic driving preference alignment module, a humanoid robot admittance control module and a feature extraction module; the feature extraction module processes environment and vehicle state data by using a VAE encoder to obtain state representation. And the automatic driving preference alignment module models the decision problem into a Markov decision process, an action containing an expected front wheel steering angle and an accelerator is generated through a high-level planner, and a strategy is optimized according to a PPO algorithm in combination with a reward value output by a reward model. A humanoid robot admittance control module generates a steering wheel actual turning angle according to the expected front wheel steering angle, a bottom layer controller fuses impedance, fuzzy control and adjusting parameters to obtain speed control input, and finally multi-style vehicle automatic driving is achieved by integrating the steering wheel actual turning angle, an accelerator and the speed control input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of robots, and particularly to a multi-style vehicle autonomous driving method for humanoid robots. Background Art

[0002] With the development of robot technology, the humanoid design of humanoid robots, advanced perception capabilities, and inherent ability to interact with a human-centered environment are being increasingly widely applied in home service, healthcare, and industrial production. Their bipedal structure and dexterous multi-joint arms enable them to operate in spaces optimized for ergonomics. However, despite these advancements, the current applications of humanoid robots are mostly limited to well-structured, static tasks in controlled environments, restricting them to simple, repetitive, or low-complexity operations. To fully realize the potential of humanoid robots, it is crucial to expand their ability to handle complex, dynamic, real-world tasks that challenge the boundaries of perception, decision-making, and physical interaction. Vehicle driving is such an ambitious challenge as it requires dynamic multi-modal environmental perception and highly flexible motion control. The ability of humanoid robots to drive vehicles highlights their adaptability to complex human-centered workplaces. In addition, by integrating with vehicles, robots can utilize the speed and agility of the vehicles to expand their application scope to areas currently dominated by human drivers, such as disaster response, remote exploration, or extreme driving conditions. However, existing solutions, especially software-based autonomous driving frameworks, lack the physical interaction capabilities required to replicate human control. This highlights the urgent need to develop control strategies based on humanoid robots that combine physical operation with intelligent decision-making. Therefore, the core challenge lies in enabling humanoid robots to achieve driving performance comparable to that of human drivers in the physical domain.

[0003] However, compared with autonomous vehicles that have been developed for over 50 years, the development of humanoid robots capable of driving vehicles has not been deeply explored. Controlling a vehicle requires a tightly integrated system: perceiving and interpreting the environment, planning driving operations, and using physical actuators (such as steering wheels, pedals) to execute precise and compliant control actions. In humanoid robots, this complexity is amplified by the need to coordinate a high degree-of-freedom (DoF) system (such as a robotic arm for steering), while ensuring smooth, robust, and human-like motion behavior. In addition, connecting simulation to real-world platforms introduces further challenges, such as accurately modeling system dynamics, transferring deep reinforcement learning (RL) policies to physical hardware, and safety issues when verifying performance under real-world uncertainties. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a multi-style vehicle autonomous driving method for humanoid robots that enables humanoid robots to achieve compliant and adaptive manipulation of the steering wheel.

[0005] A multi-style vehicle autonomous driving method for humanoid robots, the method comprising: Obtain environmental data and vehicle state data during vehicle driving; construct a multi-style vehicle autonomous driving model; the multi-style vehicle autonomous driving model includes an autonomous driving preference alignment module, a humanoid robot admittance control module, and a feature extraction module; the autonomous driving preference alignment module includes a high-level planner and a reward model; the high-level planner includes an evaluation network and an action network; the humanoid robot admittance control module includes a low-level controller and a steering skill generation module; In the feature extraction module, the environmental data and vehicle state data are feature-extracted according to the VAE encoder to obtain a state representation; In the autonomous driving preference alignment module, the decision-making problem of autonomous driving is modeled as a Markov decision process. The state representation is input into the high-level planner. The evaluation network is used to evaluate the state value, and the action network generates an action policy according to the state; the action policy and the corresponding state representation are input into a pre-trained reward model, and the corresponding reward value is output; according to the PPO algorithm, the advantage function is calculated by combining generalized advantage estimation with temporal error and coordinated with the reward value to update and optimize the action policy, and the updated action policy is output; the updated action policy includes the expected front wheel steering angle and throttle at the current time. In the humanoid robot admittance control module, the expected front wheel steering angle is input into the steering skill generation module to obtain the actual steering wheel angle; in the low-level controller, an impedance control strategy is combined with fuzzy control, and the controller parameters are dynamically adjusted through fuzzy inference to obtain a speed control input; multi-style vehicle autonomous driving is achieved according to the actual steering wheel angle, throttle, and speed control input.

[0006] For the above multi-style vehicle autonomous driving method for humanoid robots, in this application, the decision-making problem is modeled as a Markov decision process through the multi-style vehicle autonomous driving model. After the feature extraction module obtains the state representation, actions are generated by the high-level planner, and then optimized through the reward model and PPO algorithm, and an action including the expected front wheel steering angle is output, providing a reasonable instruction for steering wheel manipulation and being able to adapt to different environments and driving requirements. The expected front wheel steering angle is input into the steering skill generation module to obtain the actual steering wheel angle, and this process converts and processes the angle, making the manipulation actions of the robot more in line with the actual driving needs.

[0007] With the help of the low-level controller, the humanoid robotic arm manipulates the steering wheel according to the output steering wheel angle, and the output throttle signal is directly applied to the vehicle. The trajectory tracking of the end effector of the robotic arm is achieved through a hybrid control strategy that combines impedance control and fuzzy control, thus achieving precise motion control. Impedance control can adjust the posture based on force feedback, and fuzzy control can handle complex situations, ensuring that the robot is both precise and flexible in adjusting according to the actual situation when manipulating the steering wheel, achieving compliant manipulation. Brief Description of the Drawings

[0008] Figure 1 It is a schematic flowchart of a multi-style vehicle autonomous driving method for a humanoid robot in one embodiment; Figure 2 It is a schematic diagram of a multi-style vehicle autonomous driving model in one embodiment; Figure 3 It is a schematic diagram of the upper limb of a 7-degree-of-freedom humanoid robot in one embodiment; Figure 4 It is a schematic diagram of the reward model training process in another embodiment. Detailed Embodiment

[0009] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0010] In one embodiment, as Figure 1 shown, a multi-style vehicle autonomous driving method for a humanoid robot is provided, including the following steps: Step 102, obtain the environmental data and vehicle state data during vehicle driving; construct a multi-style vehicle autonomous driving model; the multi-style vehicle autonomous driving model includes an autonomous driving preference alignment module, a humanoid robot admittance control module, and a feature extraction module; the autonomous driving preference alignment module includes a high-level planner and a reward model; the high-level planner includes an evaluation network and an action network; the humanoid robot admittance control module includes a low-level controller and a steering skill generation module.

[0011] The environmental data during vehicle driving includes road data and path status, and the vehicle state data includes the vehicle driving state. The humanoid robot driving task can be formalized as follows: Given the desired steering wheel angle , the robotic arm must apply a rotational force to drive the steering wheel to rotate at an angular velocity within a specified time. Ensure that the actual steering wheel angle satisfies the error constraint , where is the error threshold. The rotation of the steering wheel can be modeled as circular motion with an angular velocity of , and it is required that the end effector of the robotic arm moves synchronously with it during rotation, that is . In addition, the desired steering wheel angle and speed of the vehicle must satisfy an implicit relationship. Specifically, when , the lateral error between the vehicle and the target path should satisfy , where is the maximum allowable lateral error.

[0012] To enhance the adaptability of the humanoid robot driver in various complex scenarios, the structure of the multi-style vehicle autonomous driving model constructed in this application is as shown in Figure 2 . The high-level DRL model generates the desired steering wheel angle based on vehicle environment observations and then provides it to the low-level controller. This structure effectively integrates the adaptability of DRL and the stability of the situation-specific control method. To more realistically simulate the robot's steering wheel manipulation, the autonomous driving system of the CARLA simulation platform is used, and DRL is adopted for decision-making. Subsequently, the humanoid robotic arm manipulates the steering wheel according to the steering wheel angle output by DRL with the help of the low-level controller, and the throttle signal output by DRL is directly applied to the vehicle. Specifically, the trajectory tracking of the end effector of the robotic arm is achieved through a hybrid control strategy that combines impedance control and fuzzy control, thus achieving precise motion control.

[0013] Step 104, in the feature extraction module, the environmental data and vehicle state data are feature-extracted according to the VAE encoder to obtain the state representation.

[0014] The decision-making problem of autonomous driving is modeled as a Markov decision process (MDP). The proximal policy optimization (PPO) algorithm based on deep reinforcement learning is adopted to solve this problem. Based on the state information in the CARLA simulator, the reinforcement learning agent outputs actions , including the desired front wheel steering angle at time t and the throttle .

[0015] The state consists of multimodal features . The road feature : Obtained by reducing the dimension of the RGB image at time t using a pre-trained variational autoencoder (VAE). The vehicle state feature : Consists of the throttle input at the moment of , the smoothed desired steering angle and the actual steering angle at time tVehicle speed Composed, expressed as Local path feature : A sequence representing the coordinates (x, y) of n predefined waypoints in the vehicle coordinate system, dynamically updated to construct a trajectory tracking target.

[0016] Step 106, in the autonomous driving preference alignment module, the decision-making problem of autonomous driving is modeled as a Markov decision process. The state representation is input into the high-level planner. The evaluation network is used to evaluate the state value, and the action network generates an action policy based on the state. The action policy and the corresponding state representation are input into a pre-trained reward model to output the corresponding reward value. According to the PPO algorithm, the advantage function is calculated by combining generalized advantage estimation with temporal error and coordinated with the reward value to update and optimize the action policy, and the updated action policy is output. The updated action policy includes the expected front-wheel steering angle and throttle at the current time.

[0017] The evaluation network and the action network in the high-level planner cooperate with each other. The evaluation network evaluates the input state representation and provides valuable reference information for the action network. The action network then optimizes the action policy based on these evaluation results. This multi-layer network structure can fully explore the potential features in the state information, thereby generating better actions. The evaluation network can evaluate the values corresponding to different steering angles and throttle sizes in the current state. The action network selects the optimal combination based on this value information and outputs the accurate expected front-wheel steering angle and throttle, providing accurate guidance for the humanoid robot to operate the steering wheel.

[0018] The pre-trained reward model combines human preference feedback to provide an effective evaluation criterion for the action policy. The reward value reflects the pros and cons of the current action policy in achieving goals (such as safe driving, efficient driving, etc.). Through the PPO algorithm, the advantage function is calculated by combining generalized advantage estimation with temporal error and coordinated with the reward value to update and optimize the action policy. This reward-based optimization mechanism can continuously guide the model to learn action policies that better meet the actual needs. If the steering strategy adopted in a certain state leads to a risk of near collision, the reward model will give a lower reward value, and the PPO algorithm will adjust the subsequent strategy according to this feedback, making the humanoid robot more cautious and safe when operating the steering wheel in the subsequent process, achieving compliant operation of the steering wheel.

[0019] Among them, PPO uses the probability ratio to adjust the new policy and the old policy for policy update between them. To ensure training stability, PPO restricts policy update through the objective function, restricting the probability ratio within the range:[[]] ; Among them, is a predefined constant, is the advantage function estimated by Generalized Advantage Estimation (GAE). GAE combines Temporal Difference (TD) , where is the discount factor.

[0020] State value function is optimized using mean squared error loss: ; where N is the number of samples.

[0021] As Figure 2 shown, the policy network consists of an encoder of a Variational Autoencoder (VAE) and a Multi-Layer Perceptron (MLP). The RGB image is compressed to 64 dimensions by the VAE and then stacked with the vehicle state feature and the local path feature as the input of the MLP. The network outputs the mean and variance of the actions, and these actions follow a Gaussian distribution. The evaluation network has the same architecture as the action network but does not share parameters. The VAE is offline trained using 10,000 pairs of RGB images and semantic images collected by controlling the vehicle to drive in the simulation through the keyboard. As Figure 2 shown, the encoder of the VAE adopts 4 convolutional layers followed by two parallel fully connected layers to generate latent variables that follow a Gaussian distribution. The reparameterization sampling connects the encoder and the decoder composed of 4 transposed convolutional layers for semantic image reconstruction. The training combines the reconstruction loss and Kullback-Leibler (KL) divergence regularization. After training, the encoder is frozen and used for road feature extraction.

[0022] Step 108, in the admittance control module of the humanoid robot, the desired front wheel steering angle is input to the steering skill generation module to obtain the actual steering wheel angle; in the underlying controller, an impedance control strategy is combined with fuzzy control, and the controller parameters are dynamically adjusted through fuzzy inference to obtain the speed control input; multi-style vehicle autonomous driving is achieved according to the actual steering wheel angle, throttle, and speed control input.

[0023] The steering skill generation module converts the desired front-wheel steering angle into the actual steering wheel angle. This process involves accurate modeling of vehicle dynamics and kinematics. In actual driving, there is a complex relationship between the rotation of the steering wheel and the steering of the front wheels, which is affected by various factors such as vehicle type, tire characteristics, and steering system design. The steering skill generation module can consider these factors and accurately convert the desired front-wheel steering angle into the actual steering wheel angle through precise calculations and mappings. For vehicles of different models, their steering ratios may be different. The steering skill generation module can convert the desired front-wheel steering angle into an appropriate steering wheel rotation angle according to the steering ratio of the specific model, ensuring that the humanoid robot can accurately manipulate the steering wheel and achieve adaptability to different vehicles. The design of this module aims to more realistically simulate the process of the robot manipulating the steering wheel. It not only considers the conversion of the steering angle but also may consider factors such as the rotation speed and force of the steering wheel. In actual driving, a human driver adjusts the rotation speed and force according to different situations to achieve a smooth and comfortable driving experience. The steering skill generation module can simulate this human driving behavior and generate the actual steering wheel angle that meets the actual driving requirements, making the manipulation of the humanoid robot more natural and compliant.

[0024] The underlying controller adopts a new generation of 7-degree-of-freedom humanoid robot upper limb based on the designed upper limb exoskeleton, as Figure 3 shown. A dexterous hand is integrated on the end effector, enabling the humanoid robot to manipulate the steering wheel, which provides an ideal experimental platform for simulating human driving behavior. The forward kinematics model uses the Denavit-Hartenberg (D-H) method to map the joint angles to the pose of the end effector. The inverse kinematics uses the Jacobian matrix to relate the joint velocity to the end effector velocity through . To solve the kinematic redundancy problem in the 7-degree-of-freedom manipulator, the pseudo-inverse is calculated through singular value decomposition (SVD) to achieve the inverse solution .

[0025] The dynamic modeling uses the Newton-Euler equations: ; where , and represent the mass matrix, the Coriolis and centrifugal force vectors, and the gravity term respectively. is the joint torque. maps the end effector force F to the joint space.

[0026] As Figure 2As shown in the upper right corner, an impedance control strategy is combined with fuzzy control. The controller parameters are dynamically adjusted through fuzzy inference, and the force feedback of the end effector is used to adapt to the posture changes under different operating conditions, so as to achieve compliant manipulation of the steering wheel. Define the desired posture of the end effector and the actual posture The deviation between them is where is the control period. The second-order dynamic equation of the impedance model is expressed as: ; where , and represent the inertia, damping and stiffness coefficient matrices respectively. represents the equivalent environmental force output by the impedance model. Based on the external force actually acting on the end effector, calculate the velocity control input : ; where is the desired velocity trajectory of the end effector. In order to satisfy the circular motion constraint when manipulating the steering wheel, define , satisfying: ; where R is the radius of the steering wheel. and represent the linear velocities along the x-axis and y-axis in the base coordinate system B of the humanoid robot respectively. represents the angular velocity around the x-axis in the local coordinate system E of the end effector, which is equal to the angular velocity of the steering wheel .

[0027] A fuzzy regulator based on the joint angle is designed to solve the controller instability problem when the humanoid manipulator is close to the singular configuration. Specifically, use as the input, where is the minimum joint angle, and improve the flexibility and stability of the system by dynamically adjusting the impedance parameters: ; ; .

[0028] The impedance control strategy adopted by the underlying controller obtains external force information in real time through a force sensor. When the humanoid robot manipulates the steering wheel, impedance control can dynamically adjust the posture and movement of the robotic arm according to the reaction force of the steering wheel on the end effector of the robotic arm. When the resistance to turning the steering wheel is large, the force sensor will detect an increase in the reaction force, and the impedance control strategy will correspondingly increase the force exerted by the robotic arm to ensure that the steering wheel can be turned smoothly. This force feedback mechanism enables the humanoid robot to adapt to different steering wheel characteristics and external disturbances and achieve compliant manipulation of the steering wheel. The fuzzy control strategy dynamically adjusts the controller parameters through fuzzy inference. In actual driving, there are many uncertain factors, such as road bumps and wind direction changes, which may affect the manipulation of the steering wheel. Fuzzy control can handle these uncertainties and flexibly adjust the controller parameters according to different situations. When it is detected that the vehicle is driving on a bumpy road, fuzzy control can increase the damping coefficient of the controller to make the movement of the robotic arm more stable and reduce the impact of road bumps on the steering wheel manipulation. The combination of fuzzy control and impedance control further improves the adaptability of the humanoid robot to complex driving environments.

[0029] Dynamically adjusting the controller parameters through fuzzy inference enables the underlying controller to optimize the control strategy in real time according to the actual situation. In different driving scenarios, such as high-speed driving and low-speed turning, the required control parameters may be different. Fuzzy inference can dynamically adjust the coefficient matrices such as inertia, damping, and stiffness in the impedance model according to the current driving state and environmental information, thereby achieving precise control and adaptive manipulation of the steering wheel.

[0030] In the above multi-style vehicle autonomous driving method for humanoid robots, in this application, the decision-making problem is modeled as a Markov decision process through a multi-style vehicle autonomous driving model. After the feature extraction module obtains the state representation, an action is generated by the high-level planner, and then optimized through the reward model and the PPO algorithm, and an action containing the expected front-wheel steering angle is output, providing a reasonable instruction for steering wheel manipulation and being able to adapt to different environments and driving requirements. The expected front-wheel steering angle is input into the steering skill generation module to obtain the actual steering angle of the steering wheel. This process converts and processes the angle to make the manipulation actions of the robot more in line with the actual driving needs.

[0031] With the help of the underlying controller, the humanoid robotic arm manipulates the steering wheel according to the output steering wheel angle, and the output throttle signal is directly applied to the vehicle. The trajectory tracking of the end effector of the robotic arm is achieved through a hybrid control strategy that combines impedance control and fuzzy control, thereby achieving precise motion control. Impedance control can adjust the posture based on force feedback, and fuzzy control can handle complex situations, ensuring that the robot is both precise and flexible in adjusting according to the actual situation when manipulating the steering wheel, achieving compliant manipulation.

[0032] In one embodiment, the feature extraction module extracts features from the environmental data and vehicle state data according to the VAE encoder to obtain a state representation, including: When the feature extraction module extracts features from the environmental data and vehicle state data according to the VAE encoder, the state representation obtained is: ; Among them, the road feature : Obtained by reducing the dimensionality of the RGB image at time t using a pre-trained variational autoencoder. The vehicle state feature : Consists of the throttle input at time , the smoothed desired steering angle and the actual steering angle at time t the vehicle speed The local path feature : Represents a sequence of the coordinates (x, y) of n predefined waypoints in the vehicle coordinate system.

[0033] In one embodiment, the pre-training process of the reward model includes: For each target , construct K reward models with the same structure , where , train through the ranking loss to align the estimated reward with the human preference label: ; Among them, and represent state frames, is automatically generated by the machine according to the agent's target deviation; calculate the mean and standard deviation from the estimated outputs of the K·M reward models to construct the reward model as: ; Among them, is the weight coefficient, is the decay rate, is a constant.

[0034] In a specific embodiment, such as Figure 4As shown, the research on RL in the field of autonomous driving faces the challenge of balancing the design of the reward function among safety, efficiency, and adaptability. This application addresses this issue by using the PbRL method to dynamically optimize the reward function and leveraging human feedback. However, the existing PbRL methods have two key limitations in complex real-world driving scenarios, which affect their effectiveness. First, the goal of RL is to maximize the cumulative reward by optimizing the policy. However, when the reward model takes trajectory segments as input, the single-frame reward output is affected by historical frames, resulting in biased reward estimation. Second, human annotators face difficulties in multi-objective optimization annotation. For example, if the lateral error of trajectory is small, while the vehicle speed of trajectory is closer to the target speed, the annotator may have difficulty making annotations due to hesitation, leading to low efficiency of preference feedback. To address these problems, this application combines ensemble learning with PbRL to achieve self-training of the reward model through multi-objective decomposition and dynamic exploration mechanisms. As Figure 4 shown, for each objective , K reward models with the same structure are constructed, where , and they are trained through ranking loss to align the estimated reward with the human preference label: ; where, and represent state frames, and is the preference label automatically generated by the machine according to the agent's target deviation: ; where, is the calculation function of the m-th metric, is the expected target value, and is the sign function. Subsequently, the mean and standard deviation are calculated from the estimated outputs of the K·M reward models to construct the composite reward function: ; where, is the weight coefficient, is the decay rate, and is a constant. This method establishes an implicit mapping between states and rewards and achieves accurate estimation through an ensemble of reward models. The reward model is consistent with the value network in structure, and the final sigmoid activation layer restricts the output to .

[0035] In the section on driving style learning, the vehicle is required to drive steadily, that is, the lateral error and speed tracking error (the difference between the actual speed of the vehicle and the desired speed) should be as small as possible. To drive the vehicle to achieve aggressive driving, the preference label generation function is modified as follows: ; In this way, if the vehicle speed is higher, its label value is 1 and the reward is higher; otherwise, it is -1 and the reward is lower.

[0036] In one of the embodiments, the generation function of the preference label is: ; where is the calculation function of the m-th index, is the sign function.

[0037] In one of the embodiments, the desired front wheel steering angle is input into the steering skill generation module to obtain the actual steering wheel angle, including: To achieve smooth manipulation of the steering wheel by the humanoid robotic arm and improve the trajectory tracking accuracy, the desired front wheel steering angle is input into the steering skill generation module. The desired front wheel steering angle is smoothed to: ; where is the smoothing factor; Assume that the left and right extreme steering angles of the steering wheel are symmetric, , establish and the mapping relationship with the desired steering wheel angle ; ; Calculate the speed trajectory according to the circular motion constraint. The controller drives the robot to manipulate the steering wheel until the RL decision cycle terminates, and the actual steering wheel angle is obtained.

[0038] In a specific embodiment, through the inverse mapping model the is converted into the front wheel angle of the vehicle .

[0039] In one of the embodiments, the impedance control strategy is combined with fuzzy control in the underlying controller, and the controller parameters are dynamically adjusted through fuzzy inference to obtain the speed control input, including: Define the deviation between the desired attitude of the end effector and the actual attitude as , where is the control period; Calculate the equivalent environmental force of the impedance model according to the deviation and calculate the desired speed trajectory of the end effector; Calculate the velocity control input using the equivalent environmental force of the impedance model and the desired velocity trajectory of the end effector.

[0040] In one embodiment, calculating the equivalent environmental force of the impedance model according to the deviation includes: The equivalent environmental force of the impedance model calculated according to the deviation is: ; Wherein, , and respectively represent the inertia coefficient matrix, the damping coefficient matrix, and the stiffness coefficient matrix, represents the deviation between the desired attitude and the actual attitude of the end effector, represents the first derivative of the deviation between the desired attitude and the actual attitude of the end effector, represents the second derivative of the deviation between the desired attitude and the actual attitude of the end effector.

[0041] In one embodiment, calculating the desired velocity trajectory of the end effector includes: The desired velocity trajectory of the end effector calculated is: , satisfying: ; Wherein, R is the radius of the steering wheel, and respectively represent the linear velocities along the x-axis and y-axis in the base coordinate system B of the humanoid robot, represents the angular velocity about the x-axis in the local coordinate system E of the end effector, equal to the angular velocity of the steering wheel .

[0042] In one embodiment, calculating the velocity control input using the equivalent environmental force of the impedance model and the desired velocity trajectory of the end effector includes: The velocity control input calculated using the equivalent environmental force of the impedance model and the desired velocity trajectory of the end effector is: ; Wherein, is the desired velocity trajectory of the end effector, is the equivalent environmental force of the impedance model, is the external force actually acting on the end effector, is the inertia coefficient matrix.

[0043] In one embodiment, design a fuzzy regulator based on the joint angle , using as the input, wherein is the minimum joint angle, and the parameters of the impedance model are dynamically adjusted: ; ; .

[0044] It should be understood that although Figure 1 each step in the flowchart of Figure 1 is shown sequentially according to the indication of the arrow, these steps are not necessarily executed sequentially in the order indicated by the arrow. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0045] the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0046] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A multi-style vehicle autonomous driving method for a humanoid robot, characterized in that: The method comprises: Acquire environmental data and vehicle status data during vehicle driving; construct a multi-style vehicle autonomous driving model; the multi-style vehicle autonomous driving model includes an autonomous driving preference alignment module, a humanoid robot admittance control module and a feature extraction module; the autonomous driving preference alignment module includes a high-level planner and a reward model; the high-level planner includes an evaluation network and an action network; the humanoid robot admittance control module includes a low-level controller and a steering skill generation module; In the feature extraction module, feature extraction is performed on the environment data and the vehicle state data according to the VAE encoder to obtain a state representation; In the autonomous driving preference alignment module, the decision problem of autonomous driving is modeled as a Markov decision process, the state representation is input into the high-level planner, the evaluation network is used to evaluate the state value, and the action network generates an action strategy based on the state; the action strategy and the corresponding state representation are input into a pre-trained reward model, and the corresponding reward value is output; according to the PPO algorithm, the generalized advantage estimation is combined with the time error to calculate the advantage function and the action strategy is updated and optimized in conjunction with the reward value, and the updated action strategy is output; the updated action strategy includes the expected front wheel steering angle and throttle at the current time; In the humanoid robot admittance control module, the desired front wheel steering angle is input into the steering skill generation module to obtain the actual steering wheel angle; in the underlying controller, an impedance control strategy is combined with fuzzy control, and controller parameters are dynamically adjusted through fuzzy reasoning to obtain speed control input; multi-style vehicle automatic driving is realized according to the actual steering wheel angle, throttle and speed control input.

2. The method according to claim 1, characterized in that In the feature extraction module, feature extraction is performed on the environment data and the vehicle state data according to the VAE encoder to obtain a state representation, including: In the feature extraction module, the environment data and vehicle state data are subjected to feature extraction according to the VAE encoder, and the state representation is obtained as follows: ; Among them, road characteristics :The vehicle state features are obtained by reducing the dimension of the RGB image at time t using a pre-trained variational autoencoder :Depend on Throttle input at any time , the desired steering angle after smoothing and the actual steering angle at time t Vehicle speed Composition, local path features : Represents a sequence of n predefined waypoint coordinates (x, y) in the vehicle coordinate system.

3. The method according to claim 1, characterized in that The pre-training process of the reward model includes: For each target , construct K reward models with the same structure ,in , trained with a ranking loss to align the estimated rewards with human preference labels: ; in, and Indicates the status frame. It is automatically generated by the machine based on the agent's goal deviation; the mean is calculated from the estimated outputs of K M reward models and standard deviation The reward model is constructed as: ; in, is the weight coefficient, is the decay rate, is a constant.

4. The method according to claim 3, characterized in that The generation function of the preference label is: ; in, is the calculation function of the mth indicator, is a sign function.

5. The method according to claim 1, characterized in that The desired front wheel steering angle is input into the steering skill generation module to obtain the actual steering wheel angle, including: The desired front wheel steering angle is input into the steering skill generation module. is smoothed to: ; in, is the smoothing factor; Assuming that the left and right extreme turning angles of the steering wheel are symmetrical, ,Establish Expected steering wheel angle The mapping relationship ; The velocity trajectory is calculated based on the circular motion constraints, and the controller drives the robot to manipulate the steering wheel until the RL decision cycle ends, and the actual steering wheel angle is obtained. .

6. The method according to claim 1, characterized in that The underlying controller uses a combination of impedance control strategy and fuzzy control, dynamically adjusts controller parameters through fuzzy reasoning, and obtains speed control input, including: Define the desired pose of the end effector With actual posture The deviation between ,in, is the control cycle; Calculate the equivalent environmental force of the impedance model according to the deviation and calculate the desired velocity trajectory of the end effector; The velocity control input is calculated using the equivalent environmental force of the impedance model and the desired velocity trajectory of the end effector.

7. The method according to claim 6, characterized in that The equivalent environmental force of the impedance model is calculated according to the deviation, including: The equivalent environmental force of the impedance model calculated according to the deviation is: ; in, , and denote the inertia coefficient matrix, damping coefficient matrix and stiffness coefficient matrix respectively, represents the deviation between the desired posture and the actual posture of the end effector, The first derivative representing the deviation between the desired and actual posture of the end effector, The second derivative represents the deviation between the desired pose and the actual pose of the end effector.

8. The method according to claim 1, characterized in that Calculate the desired velocity trajectory of the end effector, including: Calculate the desired velocity trajectory of the end effector as: ,satisfy: ; Where R is the steering wheel radius, and They represent the linear velocities along the x-axis and y-axis in the humanoid robot base coordinate system B, represents the angular velocity of the end effector around the x-axis in the local coordinate system E, which is equal to the angular velocity of the steering wheel .

9. The method according to claim 6, characterized in that The velocity control input is calculated using the equivalent environmental force of the impedance model and the desired velocity trajectory of the end effector, including: The speed control input is calculated using the equivalent environmental force of the impedance model and the desired speed trajectory of the end effector as: ; in, is the desired velocity trajectory of the end effector, is the equivalent environmental force of the impedance model, is the external force actually acting on the end effector, is the inertia coefficient matrix.

10. The method according to claim 7, characterized in that The method further comprises: Design based on joint angles The fuzzy regulator uses As input, is the minimum joint angle, dynamically adjusting the parameters of the impedance model: ; ; 。

Citation Information

Patent Citations

  • Trajectory tracking control method used for automatic driving robot of vehicle

    CN102358287A

  • Integrated coordination control method for robot driving vehicle

    CN109991856A

  • Automatic driving robot control system and method

    CN113635918A

  • Safe driving decision-making method for highly human-like automatic driving commercial vehicle

    CN114407931A

  • Unmanned driving safety control method

    CN118778640A

Cited By

  • Humanoid robot unmanned driving control system based on humanoid actions

    CN121341216A