Multi-style Vehicle Autopilot Method for Humanoid Robots
By constructing a multi-style vehicle autonomous driving model, combining feature extraction, Markov decision-making process and hybrid control strategies, the problem of insufficient driving performance of humanoid robots in complex environments is solved, and precise steering wheel handling and environmental adaptability are achieved.
Patent Information
- Application Number
- CN202510532931.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The prior art lacks control strategies that enable humanoid robots to achieve driving performance comparable to human drivers in complex and dynamic environments, especially in vehicle driving, lack of a combination of physical interaction capabilities and complex environmental perception and motion control.
Build a multi-style vehicle autonomous driving model, including the autonomous driving preference alignment module and the humanoid robot admission control module, and generate accurate steering wheel control instructions through feature extraction, Markov decision-making process, deep reinforcement learning and hybrid control strategies, and combine impedance and fuzzy control to achieve the robot's compliant control of the steering wheel.
It realizes adaptive manipulation of humanoid robots under different environments and driving needs, can accurately simulate human driving behavior, adapt to complex driving environments, and ensure safe and flexible motion control.
Smart Images

Figure CN120039279B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of robots, and particularly to a multi-style vehicle autonomous driving method for humanoid robots. Background Art
[0002] With the development of robot technology, the humanoid design of humanoid robots, advanced perception capabilities, and the inherent ability to interact with a human-centered environment are being increasingly widely applied in home service, healthcare, and industrial production. Their bipedal structure and dexterous multi-joint arms enable them to operate in spaces optimized for ergonomics. However, despite these advancements, the current applications of humanoid robots are mostly limited to well-structured, static tasks in controlled environments, restricting them to simple, repetitive, or low-complexity operations. To fully realize the potential of humanoid robots, it is crucial to expand their ability to handle complex, dynamic, real-world tasks that challenge the boundaries of perception, decision-making, and physical interaction. Vehicle driving is such an ambitious challenge as it requires dynamic multi-modal environmental perception and highly flexible motion control. The ability of humanoid robots to drive vehicles highlights their adaptability to complex human-centered workplaces. Additionally, by integrating with vehicles, robots can leverage the speed and agility of vehicles to extend their application scope to areas currently dominated by human drivers, such as disaster response, remote exploration, or extreme driving conditions. However, existing solutions, especially software-based autonomous driving frameworks, lack the physical interaction capabilities required to replicate human control. This highlights the urgent need to develop control strategies based on humanoid robots that combine physical operation with intelligent decision-making. Therefore, the core challenge lies in enabling humanoid robots to achieve driving performance comparable to that of human drivers in the physical domain.
[0003] However, compared with autonomous vehicles that have been developed for over 50 years, the development of humanoid robots capable of driving vehicles has not been explored in depth. Controlling a vehicle requires a tightly integrated system: perceiving and interpreting the environment, planning driving operations, and using physical actuators (such as steering wheels, pedals) to execute precise and compliant control actions. In humanoid robots, this complexity is amplified by the need to coordinate a high degree-of-freedom (DoF) system (such as a robotic arm for steering), while ensuring smooth, robust, and human-like motion behavior. Additionally, connecting simulation to real-world platforms introduces further challenges, such as accurately modeling system dynamics, transferring deep reinforcement learning (RL) policies to physical hardware, and safety issues when validating performance under real-world uncertainties. Summary of the Invention
[0004] Based on this, it is necessary to provide a multi-style vehicle autonomous driving method for humanoid robots to enable the humanoid robot to achieve compliant and adaptive manipulation of the steering wheel in response to the above technical problems.
[0005] A multi-style vehicle autonomous driving method for humanoid robots, the method comprising:
[0006] Obtain environmental data and vehicle state data during vehicle driving; construct a multi-style vehicle autonomous driving model; the multi-style vehicle autonomous driving model includes an autonomous driving preference alignment module, a humanoid robot admittance control module, and a feature extraction module; the autonomous driving preference alignment module includes a high-level planner and a reward model; the high-level planner includes an evaluation network and an action network; the humanoid robot admittance control module includes a low-level controller and a steering skill generation module;
[0007] In the feature extraction module, the environmental data and vehicle state data are subjected to feature extraction according to the VAE encoder to obtain a state representation;
[0008] In the autonomous driving preference alignment module, the decision-making problem of autonomous driving is modeled as a Markov decision process, the state representation is input into the high-level planner, the evaluation network is used to evaluate the state value, and the action network generates an action policy according to the state; the action policy and the corresponding state representation are input into a pre-trained reward model, and the corresponding reward value is output; according to the PPO algorithm, the advantage function is calculated by combining generalized advantage estimation with temporal error and cooperating with the reward value to update and optimize the action policy, and the updated action policy is output; the updated action policy includes the expected front-wheel steering angle and throttle at the current time;
[0009] In the humanoid robot admittance control module, the expected front-wheel steering angle is input into the steering skill generation module to obtain the actual steering wheel angle; in the low-level controller, an impedance control strategy is combined with fuzzy control, and the controller parameters are dynamically adjusted through fuzzy inference to obtain a speed control input; multi-style vehicle autonomous driving is achieved according to the actual steering wheel angle, throttle, and speed control input.
[0010] For the above multi-style vehicle autonomous driving method for humanoid robots, in this application, the decision-making problem is modeled as a Markov decision process through the multi-style vehicle autonomous driving model. After the feature extraction module obtains the state representation, actions are generated by the high-level planner, and then optimized through the reward model and the PPO algorithm, and an action including the expected front-wheel steering angle is output, providing a reasonable instruction for steering wheel manipulation and being able to adapt to different environments and driving requirements. The expected front-wheel steering angle is input into the steering skill generation module to obtain the actual steering wheel angle, and this process converts and processes the angle, making the manipulation actions of the robot more in line with the actual driving needs.
[0011] With the help of the underlying controller, the humanoid robotic arm manipulates the steering wheel according to the output steering wheel angle, while the output throttle signal is directly applied to the vehicle. The trajectory tracking of the end effector of the robotic arm is achieved through a hybrid control strategy that combines impedance control and fuzzy control, thus realizing precise motion control. Impedance control can adjust the posture based on force feedback, and fuzzy control can handle complex situations, ensuring that the robot is both precise and flexible in adjusting according to the actual situation when manipulating the steering wheel, achieving compliant manipulation. Description of the Drawings
[0012] Figure 1 It is a schematic flowchart of a multi-style vehicle autonomous driving method for a humanoid robot in one embodiment;
[0013] Figure 2 It is a schematic diagram of a multi-style vehicle autonomous driving model in one embodiment;
[0014] Figure 3 It is a schematic diagram of the upper limb of a 7-degree-of-freedom humanoid robot in one embodiment;
[0015] Figure 4 It is a schematic diagram of the reward model training process in another embodiment. Detailed Embodiment
[0016] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further elaborates on this application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0017] In one embodiment, as Figure 1 shown, a multi-style vehicle autonomous driving method for a humanoid robot is provided, including the following steps:
[0018] Step 102, obtain the environmental data and vehicle state data during vehicle driving; construct a multi-style vehicle autonomous driving model; the multi-style vehicle autonomous driving model includes an autonomous driving preference alignment module, a humanoid robot admittance control module, and a feature extraction module; the autonomous driving preference alignment module includes a high-level planner and a reward model; the high-level planner includes an evaluation network and an action network; the humanoid robot admittance control module includes an underlying controller and a steering skill generation module.
[0019] The environmental data during vehicle driving includes road data and path status, and the vehicle state data includes the vehicle driving state. The humanoid robot driving task can be formalized as follows: Given the desired steering wheel angle , the robotic arm must apply a rotational force to drive the steering wheel to rotate at an angular velocity within a specified time. Ensure the actual steering wheel angle Meet the error constraint , where is the error threshold. The rotation of the steering wheel can be modeled as circular motion with an angular velocity of , and it is required that the end effector of the robotic arm moves synchronously with it during rotation, that is . In addition, the desired steering wheel angle and speed of the vehicle must satisfy an implicit relationship. Specifically, when , the lateral error between the vehicle and the target path should satisfy , where is the maximum allowable lateral error.
[0020] To enhance the adaptability of the humanoid robot driver in various complex scenarios, the structure of the multi-style vehicle autonomous driving model constructed in this application is as Figure 2 shown. The advanced DRL model generates the desired steering wheel angle based on vehicle environment observations and then provides it to the underlying controller. This structure effectively integrates the adaptability of DRL and the stability of the context-specific control method. To more realistically simulate the robot's steering wheel manipulation, the autonomous driving system of the CARLA simulation platform is used, and DRL is adopted for decision-making. Subsequently, the humanoid robotic arm manipulates the steering wheel according to the steering wheel angle output by DRL with the help of the underlying controller, and the throttle signal output by DRL is directly applied to the vehicle. Specifically, the trajectory tracking of the end effector of the robotic arm is achieved through a hybrid control strategy that combines impedance control and fuzzy control, thus achieving precise motion control.
[0021] Step 104, in the feature extraction module, the environmental data and vehicle state data are feature-extracted according to the VAE encoder to obtain the state representation.
[0022] The decision-making problem of autonomous driving is modeled as a Markov decision process (MDP). The proximal policy optimization (PPO) algorithm based on deep reinforcement learning is adopted to solve this problem. Based on the state information in the CARLA simulator, the reinforcement learning agent outputs actions , including the desired front wheel steering angle at time t and the throttle .
[0023] State consists of multi-modal features . Road feature : Obtained by reducing the dimension of the RGB image at time t using a pre-trained variational autoencoder (VAE). Vehicle state feature : Consists of the throttle input at time , the smoothed desired steering angle and the actual steering angle at time t Vehicle speed are combined and expressed as . Local path feature : It represents a sequence of coordinates (x, y) of n predefined waypoints in the vehicle coordinate system, which is dynamically updated to construct a trajectory tracking target.
[0024] Step 106, in the autonomous driving preference alignment module, the decision-making problem of autonomous driving is modeled as a Markov decision process. The state representation is input into the high-level planner. The evaluation network is used to evaluate the state value, and the action network generates an action policy based on the state. The action policy and the corresponding state representation are input into a pre-trained reward model to output the corresponding reward value. According to the PPO algorithm, the advantage function is calculated by combining generalized advantage estimation with temporal error and is coordinated with the reward value to update and optimize the action policy, and the updated action policy is output. The updated action policy includes the desired front-wheel steering angle and throttle at the current time.
[0025] The evaluation network and the action network in the high-level planner cooperate with each other. The evaluation network evaluates the input state representation and provides valuable reference information for the action network. The action network then optimizes the action policy based on these evaluation results. This multi-layer network structure can fully explore the potential features in the state information, thereby generating better actions. The evaluation network can evaluate the values corresponding to different steering angles and throttle sizes in the current state. The action network selects the optimal combination based on this value information and outputs the accurate desired front-wheel steering angle and throttle, providing accurate guidance for the humanoid robot to operate the steering wheel.
[0026] The pre-trained reward model combines human preference feedback and provides an effective evaluation criterion for the action policy. The reward value reflects the pros and cons of the current action policy in achieving goals (such as safe driving, efficient driving, etc.). Through the PPO algorithm, the advantage function is calculated by combining generalized advantage estimation with temporal error and is coordinated with the reward value to update and optimize the action policy. This reward-based optimization mechanism can continuously guide the model to learn action policies that are more in line with actual needs. If the steering strategy adopted in a certain state leads to a risk of near collision, the reward model will give a lower reward value, and the PPO algorithm will adjust the subsequent strategy according to this feedback, making the humanoid robot more cautious and safe when operating the steering wheel in the future, and achieving compliant operation of the steering wheel.
[0027] Among them, PPO uses the probability ratio to adjust the new policy and the old policy for policy update between them. To ensure training stability, PPO limits the policy update through the objective function and limits the probability ratio to Range:
[0028] ;
[0029] where is a predefined constant, is the advantage function estimated by Generalized Advantage Estimation (GAE). GAE combines Temporal Difference (TD) where is the discount factor.
[0030] State value function is optimized using mean squared error loss:
[0031] ;
[0032] where N is the number of samples.
[0033] As Figure 2 shown, the policy network consists of an encoder of a Variational Autoencoder (VAE) and a Multi-Layer Perceptron (MLP). The RGB image is compressed to 64 dimensions by the VAE and then stacked with the vehicle state feature and the local path feature as the input of the MLP. The network outputs the mean and variance of the actions, and these actions follow a Gaussian distribution. The evaluation network has the same architecture as the action network but does not share parameters. The VAE is offline trained using 10,000 pairs of RGB images and semantic images collected by controlling the vehicle to drive in simulation through the keyboard. As Figure 2 shown, the encoder of the VAE adopts 4 convolutional layers followed by two parallel fully connected layers to generate latent variables following a Gaussian distribution. The reparameterization sampling connects the encoder and the decoder composed of 4 transposed convolutional layers for semantic image reconstruction. The training combines the reconstruction loss and Kullback-Leibler (KL) divergence regularization. After training, the encoder is frozen and used for road feature extraction.
[0034] Step 108, in the admittance control module of the humanoid robot, the desired front wheel steering angle is input to the steering skill generation module to obtain the actual steering wheel angle; in the underlying controller, an impedance control strategy is combined with fuzzy control, and the controller parameters are dynamically adjusted through fuzzy inference to obtain the speed control input; multi-style vehicle autonomous driving is achieved according to the actual steering wheel angle, throttle, and speed control input.
[0035] The steering skill generation module converts the desired front-wheel steering angle into the actual steering wheel angle, and this process involves the precise modeling of vehicle dynamics and kinematics. In actual driving, there is a complex relationship between the rotation of the steering wheel and the steering of the front wheels, which is affected by various factors such as vehicle type, tire characteristics, and steering system design. The steering skill generation module can consider these factors and accurately convert the desired front-wheel steering angle into the actual steering wheel angle through precise calculations and mappings. For vehicles of different models, their steering ratios may be different. The steering skill generation module can convert the desired front-wheel steering angle into an appropriate steering wheel rotation angle according to the steering ratio of the specific model, ensuring that the humanoid robot can accurately manipulate the steering wheel and achieve adaptability to different vehicles. The design of this module aims to more realistically simulate the process of the robot manipulating the steering wheel. It not only considers the conversion of the steering angle but also may consider factors such as the rotation speed and force of the steering wheel. In actual driving, human drivers adjust the rotation speed and force according to different situations when manipulating the steering wheel to achieve a smooth and comfortable driving experience. The steering skill generation module can simulate this human driving behavior and generate the actual steering wheel angle that meets the actual driving requirements, making the manipulation of the humanoid robot more natural and compliant.
[0036] The underlying controller adopts a new generation of 7-degree-of-freedom humanoid robot upper limb based on the designed upper limb exoskeleton, as Figure 3 shown. A dexterous hand is integrated on the end effector, enabling the humanoid robot to manipulate the steering wheel, which provides an ideal experimental platform for simulating human driving behavior. The forward kinematics model uses the Denavit-Hartenberg (D-H) method to map the joint angles to the pose of the end effector. The inverse kinematics uses the Jacobian matrix , to relate the joint velocity to the end effector velocity through . To solve the kinematic redundancy problem in the 7-degree-of-freedom manipulator, the pseudo-inverse is calculated through singular value decomposition (SVD) to achieve the inverse solution .
[0037] The dynamic modeling uses the Newton-Euler equations:
[0038] ;
[0039] where, , and represent the mass matrix, the Coriolis and centrifugal force vectors, and the gravity term respectively. is the joint torque. Maps the end effector force F to the joint space.
[0040] As shown in the upper right corner of Figure 2 , an impedance control strategy is combined with fuzzy control. The controller parameters are dynamically adjusted through fuzzy inference, and the force feedback of the end effector is used to adapt to the posture changes under different operating conditions, so as to achieve compliant manipulation of the steering wheel. Define the desired posture of the end effector and the actual posture . The deviation between them is , where is the control period. The second-order dynamic equation of the impedance model is expressed as:
[0041] ;
[0042] Among them, , and represent the inertia, damping, and stiffness coefficient matrices respectively. represents the equivalent environmental force output by the impedance model. Based on the external force actually acting on the end effector, calculate the velocity control input :
[0043] ;
[0044] Among them, is the desired velocity trajectory of the end effector. To satisfy the circular motion constraint when manipulating the steering wheel, define , satisfying:
[0045] ;
[0046] Among them, R is the radius of the steering wheel. and represent the linear velocities along the x-axis and y-axis in the base coordinate system B of the humanoid robot respectively. represents the angular velocity around the x-axis in the local coordinate system E of the end effector, which is equal to the angular velocity of the steering wheel.
[0047] A fuzzy regulator based on the joint angle is designed to solve the controller instability problem when the humanoid manipulator approaches the singular configuration. Specifically, use as the input, where is the minimum joint angle, and improve the flexibility and stability of the system by dynamically adjusting the impedance parameters:
[0048] ;
[0049] ;
[0050] 。
[0051] The impedance control strategy adopted by the underlying controller obtains the external force information in real time through the force sensor. When the humanoid robot manipulates the steering wheel, the impedance control can dynamically adjust the posture and movement of the robotic arm according to the reaction force of the steering wheel on the end effector of the robotic arm. When the resistance to the rotation of the steering wheel is large, the force sensor will detect an increase in the reaction force, and the impedance control strategy will correspondingly increase the force exerted by the robotic arm to ensure that the steering wheel can be rotated smoothly. This force feedback mechanism enables the humanoid robot to adapt to different steering wheel characteristics and external disturbances and achieve compliant manipulation of the steering wheel. The fuzzy control strategy dynamically adjusts the controller parameters through fuzzy inference. In actual driving, there are many uncertain factors, such as road bumps and wind direction changes, which may affect the manipulation of the steering wheel. Fuzzy control can handle these uncertainties and flexibly adjust the controller parameters according to different situations. When it is detected that the vehicle is driving on a bumpy road, the fuzzy control can increase the damping coefficient of the controller to make the movement of the robotic arm more stable and reduce the impact of road bumps on the steering wheel manipulation. The combination of fuzzy control and impedance control further improves the adaptability of the humanoid robot to complex driving environments.
[0052] Dynamically adjusting the controller parameters through fuzzy inference enables the underlying controller to optimize the control strategy in real time according to the actual situation. In different driving scenarios, such as high-speed driving and low-speed turning, the required control parameters may be different. Fuzzy inference can dynamically adjust the coefficient matrices such as inertia, damping, and stiffness in the impedance model according to the current driving state and environmental information, so as to achieve precise control and adaptive manipulation of the steering wheel.
[0053] In the above multi-style vehicle autonomous driving method for humanoid robots, this application models the decision-making problem as a Markov decision process through a multi-style vehicle autonomous driving model. After the feature extraction module obtains the state representation, it generates an action through the high-level planner, and then optimizes it through the reward model and the PPO algorithm, and outputs an action containing the expected front wheel steering angle, providing a reasonable instruction for the steering wheel manipulation and being able to adapt to different environments and driving requirements. Inputting the expected front wheel steering angle into the steering skill generation module to obtain the actual steering angle of the steering wheel. This process converts and processes the angle to make the manipulation action of the robot more in line with the actual driving needs.
[0054] With the help of the underlying controller, the humanoid robotic arm manipulates the steering wheel according to the output steering wheel angle, while the output throttle signal is directly applied to the vehicle. The trajectory tracking of the end effector of the robotic arm is achieved through a hybrid control strategy that combines impedance control and fuzzy control, thus enabling precise motion control. Impedance control can adjust the posture based on force feedback, and fuzzy control can handle complex situations, ensuring that the robot is both precise and flexible in adjusting according to the actual situation when manipulating the steering wheel, achieving compliant manipulation.
[0055] In one embodiment, the feature extraction module extracts features from the environmental data and vehicle state data according to the VAE encoder to obtain a state representation, including:
[0056] When the feature extraction module extracts features from the environmental data and vehicle state data according to the VAE encoder, the state representation is:
[0057] ;
[0058] Among them, the road feature : Obtained by reducing the dimension of the RGB image at time t using a pre-trained variational autoencoder, the vehicle state feature : Consists of the throttle input at time , the smoothed desired steering angle and the actual steering angle at time t the vehicle speed The local path feature : Represents a sequence of n predefined waypoint coordinates (x, y) in the vehicle coordinate system.
[0059] In one embodiment, the pre-training process of the reward model includes:
[0060] For each target , construct K reward models with the same structure , where , and train them through ranking loss to align the estimated rewards with the human preference labels:
[0061] ;
[0062] Among them, and represent state frames, is automatically generated by the machine according to the agent target deviation; calculate the mean and standard deviation from the estimated outputs of the K·M reward models, and construct the reward model as:
[0063] ;
[0064] Among them, is the weight coefficient, is the decay rate, is a constant.
[0065] In a specific embodiment, as Figure 4 shown, the research of RL in the field of autonomous driving faces the challenge of balancing the design of the reward function among safety, efficiency, and adaptability. This application solves this problem by adopting the PbRL method to dynamically optimize the reward function and using human feedback. However, there are two key limitations in the existing PbRL methods in complex real-world driving scenarios, which affect their effectiveness. First, the goal of RL is to maximize the cumulative reward by optimizing the policy. However, when the reward model takes a trajectory segment as input, the single-frame reward output is affected by historical frames, resulting in biased reward estimation. Second, human annotators face difficulties in multi-objective optimization annotation. For example, if the lateral error of the trajectory is small, while the vehicle speed of the trajectory is closer to the target speed, the annotator may have difficulty making annotations due to hesitation, resulting in low preference feedback efficiency. To solve these problems, this application combines ensemble learning with PbRL to achieve self-training of the reward model through multi-objective decomposition and dynamic exploration mechanisms. As Figure 4 shown, for each target , construct K reward models with the same structure, where , and train them through the ranking loss to align the estimated reward with the human preference label:
[0066] ;
[0067] Among them, and represent the state frames, is the preference label automatically generated by the machine according to the agent's target deviation:
[0068] ;
[0069] Among them, is the calculation function of the m-th metric, is the expected target value, is the sign function. Subsequently, calculate the mean and the standard deviation from the estimated outputs of the K·M reward models to construct the composite reward function:
[0070] ;
[0071] Among them, is the weight coefficient, is the decay rate, is a constant. This method establishes an implicit mapping between the state and the reward, and achieves accurate estimation by integrating the reward model. The reward model is consistent with the value network in structure, and the final sigmoid activation layer restricts the output to .
[0072] In the section of driving style learning, it is required that the vehicle drives steadily, that is, the smaller the lateral error and the speed tracking error (the difference between the actual speed of the vehicle and the desired speed) are, the better. To drive the vehicle to achieve aggressive driving, modify the preference label generation function:
[0073] ;
[0074] In this way, if the vehicle speed is higher, its label value is 1 and the reward is higher, otherwise it is -1 and the reward is lower.
[0075] In one of the embodiments, the generation function of the preference label is:
[0076] ;
[0077] Among them, is the calculation function of the m-th index, is the sign function.
[0078] In one of the embodiments, input the desired front wheel steering angle into the steering skill generation module to obtain the actual steering wheel angle, including:
[0079] To achieve the smooth manipulation of the steering wheel by the humanoid robotic arm and improve the trajectory tracking accuracy, input the desired front wheel steering angle into the steering skill generation module, and the desired front wheel steering angle is smoothed to:
[0080] ;
[0081] Among them, is the smoothing factor;
[0082] Assume that the left and right limit steering angles of the steering wheel are symmetric, , establish the mapping relationship with the desired steering wheel angle ;
[0083] Calculate the speed trajectory according to the circular motion constraint, and the controller drives the robot to manipulate the steering wheel until the RL decision cycle terminates, and obtain the actual steering wheel angle .
[0084] In a specific embodiment, through the inverse mapping model to be converted into the steering angle of the vehicle's front wheels .
[0085] In one of the embodiments, the impedance control strategy is combined with fuzzy control in the underlying controller, and the controller parameters are dynamically adjusted through fuzzy inference to obtain the speed control input, including:
[0086] Define the desired posture of the end effector and the actual posture The deviation between them is , where is the control period;
[0087] Calculate the equivalent environmental force of the impedance model according to the deviation and calculate the desired speed trajectory of the end effector;
[0088] Calculate the speed control input by using the equivalent environmental force of the impedance model and the desired speed trajectory of the end effector.
[0089] In one of the embodiments, calculating the equivalent environmental force of the impedance model according to the deviation includes:
[0090] Calculate the equivalent environmental force of the impedance model according to the deviation as:
[0091] ;
[0092] where , and respectively represent the inertia coefficient matrix, the damping coefficient matrix and the stiffness coefficient matrix, represents the deviation between the desired posture and the actual posture of the end effector, represents the first derivative of the deviation between the desired posture and the actual posture of the end effector, represents the second derivative of the deviation between the desired posture and the actual posture of the end effector.
[0093] In one of the embodiments, calculating the desired speed trajectory of the end effector includes:
[0094] Calculate the desired speed trajectory of the end effector as:
[0095] , satisfying:
[0096] ;
[0097] where R is the radius of the steering wheel, and respectively represent the linear velocities along the x-axis and y-axis in the base coordinate system B of the humanoid robot, represents the angular velocity of the end effector around the x-axis in the local coordinate system E, which is equal to the angular velocity of the steering wheel .
[0098] In one embodiment, calculating the velocity control input using the equivalent environmental force of the impedance model and the desired velocity trajectory of the end effector includes:
[0099] The speed control input is calculated using the equivalent environmental force of the impedance model and the desired speed trajectory of the end effector:
[0100] ;
[0101] in, is the desired velocity trajectory of the end effector, is the equivalent environmental force of the impedance model, is the actual external force acting on the end effector, is the inertia coefficient matrix.
[0102] In one embodiment, the design is based on the joint angle The fuzzy regulator uses As input, is the minimum joint angle, dynamically adjusting the parameters of the impedance model:
[0103] ;
[0104] ;
[0105] .
[0106] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0107] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0108] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A multi-style vehicle autonomous driving method for humanoid robots, characterized in that The method includes: Obtaining environmental data and vehicle state data during vehicle driving; constructing a multi-style vehicle autonomous driving model; the multi-style vehicle autonomous driving model includes an autonomous driving preference alignment module, a humanoid robot admittance control module, and a feature extraction module; the autonomous driving preference alignment module includes a high-level planner and a reward model; the high-level planner includes an evaluation network and an action network; the humanoid robot admittance control module includes a low-level controller and a steering skill generation module; In the feature extraction module, the environmental data and vehicle state data are feature-extracted according to the VAE encoder to obtain a state representation; In the autonomous driving preference alignment module, the decision-making problem of autonomous driving is modeled as a Markov decision process. The state representation is input into the high-level planner. The evaluation network is used to evaluate the state value, and the action network generates an action policy according to the state; the action policy and the corresponding state representation are input into a pre-trained reward model to output a corresponding reward value; according to the PPO algorithm, the advantage function is calculated by combining generalized advantage estimation with temporal error and cooperating with the reward value to update and optimize the action policy, and the updated action policy is output; the updated action policy includes the expected front-wheel steering angle and throttle at the current time; In the humanoid robot admittance control module, the expected front-wheel steering angle is input into the steering skill generation module to obtain the actual steering wheel angle; in the low-level controller, an impedance control strategy is combined with fuzzy control, and the controller parameters are dynamically adjusted through fuzzy inference to obtain a speed control input; multi-style vehicle autonomous driving is realized according to the actual steering wheel angle, throttle, and speed control input; In the low-level controller, an impedance control strategy is combined with fuzzy control, and the controller parameters are dynamically adjusted through fuzzy inference to obtain a speed control input, including: Define the desired attitude of the end effector and the actual attitude The deviation between them is , where is the control period, " " represents the desired attitude of the end effector and the actual attitude The deviation between them, " " represents the first derivative of the desired attitude of the end effector; Calculating the equivalent environmental force of the impedance model according to the deviation and calculating the expected speed trajectory of the end effector; Calculating the speed control input by using the equivalent environmental force of the impedance model and the expected speed trajectory of the end effector.
2. The method according to claim 1, wherein In the feature extraction module, the environmental data and vehicle state data are feature-extracted according to the VAE encoder to obtain a state representation, including: In the feature extraction module, the environmental data and vehicle state data are feature-extracted according to the VAE encoder, and the obtained state representation is: ; Among them, road features : Obtained by reducing the dimension of the RGB image at time t using a pre-trained variational autoencoder, vehicle state features : Consist of the throttle input at time t - 1 , the smoothed desired steering angle and the actual steering angle at time t vehicle speed Local path features : Represents a sequence of the coordinates (x, y) of n predefined waypoints in the vehicle coordinate system, " " represents the state representation.
3. The method according to claim 1, characterized in that, The pre-training process of the reward model includes: For each target , construct K reward models with the same structure , where , train through ranking loss to align the estimated rewards with the human preference labels: ; Among them, and represent status frames, which are automatically generated by the machine according to the agent target deviation; calculate the mean and standard deviation to construct the reward model as: ; Among them, is the weight coefficient, is the attenuation rate, is a constant, " " represents the estimated reward, " " represents the reward model.
4. The method according to claim 3, wherein The generation function of the preference label is: ; Among them, is the calculation function of the m-th index, is the sign function.
5. The method according to claim 1, characterized in that, Inputting the expected front-wheel steering angle into the steering skill generation module to obtain the actual steering wheel angle, including: Input the desired front wheel steering angle into the steering skill generation module, the desired front wheel steering angle is smoothed to be: ; Among them, is the smoothing factor; Assume that the left and right extreme steering wheel angles are symmetric. , establish the mapping relationship with the desired steering wheel angle . " " represents the extreme steering wheel angle, " " represents the extreme left steering wheel angle, " " represents the extreme right steering wheel angle; represents the extreme right steering wheel angle; Calculate the velocity trajectory according to the circular motion constraint. The controller drives the robot to manipulate the steering wheel until the RL decision-making cycle terminates, and the actual steering angle of the steering wheel is obtained .
6. The method according to claim 1, characterized in that Calculating the equivalent environmental force of the impedance model according to the deviation, including: Calculating the equivalent environmental force of the impedance model according to the deviation as: ; Among them, , and represent the inertia coefficient matrix, the damping coefficient matrix, and the stiffness coefficient matrix respectively, represents the deviation between the desired attitude and the actual attitude of the end effector, represents the first derivative of the deviation between the desired attitude and the actual attitude of the end effector, represents the second derivative of the deviation between the desired attitude and the actual attitude of the end effector, " " represents the equivalent environmental force of the impedance model.
7. The method according to claim 1, wherein Calculating the expected speed trajectory of the end effector, including: Calculating the expected speed trajectory of the end effector as , satisfying: ; where R is the radius of the steering wheel, and respectively represent the linear velocities along the x-axis and y-axis in the base coordinate system B of the humanoid robot, represents the angular velocity about the x-axis in the local coordinate system E of the end effector, which is equal to the angular velocity of the steering wheel ," " represents the desired velocity vector of the end effector, represents the actual rotation angle of the steering wheel at t time, represents the actual rotation angle of the steering wheel at t time - 1.
8. The method according to claim 1, characterized in that Calculating the speed control input by using the equivalent environmental force of the impedance model and the expected speed trajectory of the end effector, including: Calculating the speed control input by using the equivalent environmental force of the impedance model and the expected speed trajectory of the end effector as: ; Among them, is the desired velocity trajectory of the end effector, is the equivalent environmental force of the impedance model, is the external force actually acting on the end effector, is the inertia coefficient matrix, " " represents the velocity control input.
9. The method according to claim 6, wherein The method further includes: The design is based on a fuzzy regulator for joint angles and uses as the input, where is the minimum joint angle, and dynamically adjusts the parameters of the impedance model: 。
Citation Information
Patent Citations
Integrated coordination control method for robot driving vehicle
CN109991856A
Unmanned driving safety control method
CN118778640A