A method for obstacle avoidance of humanoid robots based on model predictive control
Through the humanoid robot obstacle avoidance method based on model prediction control, combined with laser scanner and depth camera to obtain environmental information, judge and perform inverted flip span or detour operations, the existing obstacle avoidance method is solved, and the efficient and stable obstacle avoidance effect is achieved.
Patent Information
- Application Number
- CN202411097297.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-08-12
AI Technical Summary
Existing humanoid robot obstacle avoidance methods are inefficient in extremely narrow or complex environments, and there is instability and risk of injury when jump obstacle avoidance.
A humanoid robot obstacle avoidance method based on model prediction control is adopted. The road topographic map and obstacle height are obtained through a laser scanner and depth camera to determine whether it can be crossed in an upside down. If possible, the inverted flip operation is performed, and if not, the way around will be taken.
It realizes efficient obstacle avoidance in narrow or complex environments, reduces movement distance and time, improves movement efficiency, and improves stability and flexibility through model prediction algorithms and adaptive nonlinear controllers.
Smart Images

Figure CN118915773B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of humanoid robot control, and particularly relates to a method for obstacle avoidance of a humanoid robot based on model predictive control. Background Art
[0002] Currently, the obstacle avoidance actions adopted by humanoid robots are as follows: 1. Side-step obstacle avoidance: that is, the robot avoids obstacles in front by moving its feet to one side; however, when it is necessary to quickly pass through a narrow space or avoid an emergency obstacle, side-step obstacle avoidance is not as direct and efficient as the method of forward handstand to cross the obstacle. In an extremely narrow environment, side-step obstacle avoidance may be restricted by the space and cannot effectively avoid obstacles; 2. Turning obstacle avoidance: that is, when the robot detects an obstacle in front, it can change its traveling direction by turning, so as to avoid the obstacle; however, turning obstacle avoidance means that the robot needs to change its original traveling direction, which may increase the path length and time to reach the destination. And in some complex environments, turning obstacle avoidance will be restricted by the flexibility of the robot or the surrounding obstacles, while the method of forward handstand to cross the obstacle does not require changing the direction, saves time, and improves the movement efficiency; 3. Jumping obstacle avoidance: that is, for some relatively high obstacles, the robot crosses them by jumping; however, this method requires the robot to have a powerful power system and precise jump control, and has great instability. There may be a risk of losing control or falling during the jumping process, causing damage to the robot and the surrounding environment. Summary of the Invention
[0003] The present invention discloses a method for obstacle avoidance of a humanoid robot based on model predictive control. The specific method is as follows:
[0004] Scan the road in front of the humanoid robot through a laser scanner set on the head of the humanoid robot, form coordinate point clouds according to the scanning results, and draw a topographic map of the road in front of the humanoid robot.
[0005] Calculate the height of the obstacle through a depth camera set on the knee of the humanoid robot.
[0006] Judge whether the height of the obstacle is greater than the leg length of the humanoid robot.
[0007] If so, control the humanoid robot to walk in front of the obstacle and perform the operation of bypassing the obstacle.
[0008] If not, control the humanoid robot to walk in front of the obstacle and perform the operation of handstand flipping to cross the obstacle.
[0009] Further, to form the coordinate point clouds, the specific formula is as follows:
[0010] X = Lcosαcosβ
[0011] Y = Lcosαcosβ
[0012] Z = Lcosαcosβ
[0013] Wherein, L is the distance from the target scanning point to the laser scanner after the emitted laser is reflected back, α and β are the rotation angles of the laser mirror of the laser scanner in the horizontal and vertical directions respectively, and X, Y, and Z are the x, y, and z axis coordinate values of the target scanning point relative to the laser scanner.
[0014] Further, calculate the height of the obstacle, and the calculation formula is as follows:
[0015]
[0016] Wherein, Z is the distance between the imaging point at the bottom of the obstacle and the depth camera mounted on the hip of the humanoid robot, f is the focal length of the depth camera, T is the distance between the imaging point at the top of the obstacle and the depth camera mounted on the hip of the humanoid robot, x 0l is the horizontal axis coordinate of the scanning point of the obstacle in front imaged on the left camera head of the depth camera, x 0r is the horizontal axis coordinate of the scanning point of the obstacle in front imaged on the right camera head of the depth camera.
[0017] Further, perform the operation of inverting and flipping over the obstacle, and the specific method is as follows:
[0018] Construct a dynamic model of the humanoid robot;
[0019] Use a humanoid robot controller based on the model predictive control algorithm to complete the attitude adjustment of the humanoid robot during the inverting and flipping action.
[0020] Further, construct a dynamic model of the humanoid robot, and the specific method is as follows:
[0021] Simplify the humanoid robot into a single-mass inverted pendulum with a spring and a damper, and the motion equation is:
[0022]
[0023] Wherein, m is the mass of the robot, s is the length from the support point of the inverted pendulum to the mass point, that is, the height from the ground to the mass point when the robot is in the ready posture before walking, c is the damping constant, k is the spring constant, g is the acceleration due to gravity, τ is the torque applied to the robot, δ is the angle of the inverted pendulum, corresponding to the inclination angle of the robot's leg in the direction perpendicular to the ground, u δ is the output angle of the inverted pendulum, input by the robot controller;
[0024] u is the displacement of the controller with respect to the center of mass, approximately:
[0025]
[0026] where m is the mass of the robot, s is the length from the support point of the inverted pendulum to the mass point, i.e., the height from the ground to the mass point when the robot is in the ready position before walking, c is the damping constant, k is the spring constant, g is the acceleration due to gravity, τ is the torque applied to the robot, δ is the angle of the inverted pendulum, corresponding to the inclination angle of the robot's leg with respect to the vertical direction of the ground, and u δ is the output angle of the inverted pendulum, which is input by the robot controller;
[0027] The zero moment point ZMP of the robot in the x - direction is calculated as follows:
[0028]
[0029] where F y is the ground reaction force in the y - direction, and F z is the ground reaction force in the z - direction;
[0030] Since F y is also the reaction torque of the spring and shock absorber, it can be written as:
[0031]
[0032] where s is the length from the support point of the inverted pendulum to the mass point, i.e., the height from the ground to the mass point when the robot is in the ready position before walking, c is the damping constant, k is the spring constant, and δ is the angle of the inverted pendulum, corresponding to the inclination angle of the robot's leg with respect to the vertical direction of the ground;
[0033] When there is no interference, the state - space equations of and can be derived:
[0034]
[0035] where m is the mass of the robot, s is the length from the support point of the inverted pendulum to the mass point, i.e., the height from the ground to the mass point when the robot is in the ready position before walking, c is the damping constant, k is the spring constant, g is the acceleration due to gravity, δ is the angle of the inverted pendulum, corresponding to the inclination angle of the robot's leg with respect to the vertical direction of the ground, and u is the displacement of the center of mass by the controller; the states of the two equations are the angle δ and angular velocity of the robot
[0036] Furthermore, the humanoid robot controller encodes the distance input signal measured by the depth camera into a neuron. After the input distance is presented, the neuron that best matches the input distance is selected as the optimal neuron;
[0037] After that, the weights of the optimal neurons are updated. Under the guidance of the reward function, the correct control actions are autonomously found using the input pattern information and the performance metrics to be optimized.
[0038] Furthermore, the humanoid robot controller is specifically set as follows:
[0039] Define the network structure and hyperparameters, and randomly initialize the input and output weights;
[0040] At time I(t), present the input pattern and measure the distance using a depth camera, and select the optimal neuron N S , which minimizes the distance between I(t) and the input weight vector connecting N S to the input nodes; where I(t) is the input pattern;
[0041] For the optimal neuron N S , the corresponding output weights are used to generate control actions; that is, the perturbation version is adopted Add a Gaussian zero-mean random variable λ to the weights of the optimal neuron , multiply by the gain α S related to the optimal neuron N S ; α S starts from an initial value, fixes a final value, and decreases linearly between these two values as the learning phase progresses; where O(t) is the output control signal; S The actual reward function R(t) is evaluated together with the reward increment ΔR = R(t) - R(t - 1); if ΔR > b
[0042] , where b s is the average increment related to the neuron N s ; the update rule of b S is as follows: s
[0043]
[0044] where σ > 0, is the new average increment, is the old average increment, and ΔR is the average increment;
[0045] The input and output weights connecting to the optimal neuron N S and its adjacent neurons are updated according to the following rules:
[0046] w i,in (t + 1) = w i,in (t) + η in ζ(I(t) - w i,in (t))
[0047] w i,out (t + 1) = w i,out (t) + η out ζ(O(t) - w i,out (t))
[0048] where η in is the input learning rate, η out is the output learning rate, w i,in is the input weight, w i,out is the output weight, ζ is the neighborhood function of the neuron. If N R represents a neighborhood with radius R, then around the current optimal neuron N S there is:
[0049]
[0050] ζ establishes a radius neighborhood around N S and the weights are updated within this range;
[0051] In each learning iteration, an output pattern is given, an optimal neuron N S is selected and the output signal is provided. Then the reward is evaluated and the learning phase is executed, that is, w i,in , w i,out , b s and α S are updated. After several iterations, the update of other learning hyperparameters is performed.
[0052] Furthermore, after the inverted flip over the obstacle operation, the humanoid controller calculates the forces required to maintain the body standing position, and then uses these forces as the input signals of the controller. The controller provides stable joint torques;
[0053] If the change in the robot structure causes the center of gravity to shift at this time, an adaptive nonlinear controller MMC is added as a feedforward error compensator to adjust the robot's attitude. This adaptive nonlinear controller MMC provides an additional nonlinear feedforward control action. After transforming the joint torques through the Jacobian matrix, it is added to the underlying leg joint controller in the form of forces; A reward function is introduced:
[0054]
[0055] where Θ ref is the reference pitch velocity of the robot, Θ is the actual pitch velocity of the robot, is the reference pitch acceleration of the robot, is the actual pitch acceleration of the robot, and β is the introduced gain, which establishes the priority of the pitch velocity over the pitch position error: as the pitch error decreases, the correlation of the pitch velocity error becomes greater;
[0056] The algorithm for controlling the robot updates the output weights by improving the reward function to ensure the stability of completing the striding action. The reward function takes into account the error between the reference and the actual pitch of the robot.
[0057] Due to the adoption of the above technical solution, the present invention has the following beneficial effects:
[0058] 1. The way of forward handstand striding over obstacles can directly stride over obstacles without achieving the goal by detouring or avoiding like other methods, avoiding the situation that traditional obstacle avoidance methods will increase the moving distance and duration, and can significantly improve the moving efficiency in specific scenarios.
[0059] 2. The way of forward handstand striding over obstacles can adapt to different types of obstacles lower than the leg height by adjusting the height and posture of the handstand, with a certain degree of flexibility; in rough or narrow spaces where traditional obstacle avoidance methods are difficult to effectively handle, this way can overcome these limitations and achieve stable movement.
[0060] 3. The model predictive algorithm introduced in the present invention can predict the motion trajectory in the next period of time according to the current state of the system and the prediction model, with strong robustness and flexibility; and an adaptive non-linear controller is introduced, which does not need to know the physical characteristics of the robot a priori, the controller does not require internal iteration, and can quickly reach the maximum value of the reward function after appropriately mapping the input signal, providing an endless learning opportunity;
[0061] 4. By combining the way of forward handstand striding over obstacles with the model predictive algorithm, the present invention can improve the work efficiency and stability while completing specific environmental tasks, and has a broad application prospect.
[0062] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The brief description of the drawings of the present invention is as follows.
[0064] Figure 1 is a schematic diagram of the control flow of the present invention.
[0065] Figure 2 is a schematic diagram of the control loop for the MPC algorithm learning iteration. DETAILED DESCRIPTION OF THE INVENTION
[0066] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0067] A method for obstacle avoidance of a humanoid robot based on model predictive control is as Figure 1 shown, and the specific steps are as follows:
[0068] S1. Use the laser scanner installed on the head of the humanoid robot to scan the road in front of the humanoid robot, form a coordinate point cloud according to the scanning results, and draw a topographic map of the road in front of the humanoid robot.
[0069] In step S1, the laser scanner installed on the head of the humanoid robot is used to scan the road surface in front of the humanoid robot and receive the reflected optical signal. The generator can calculate the distance between the laser scanner on the humanoid robot and the front scanning point according to the reflection data of the collected optical signal. A coordinate point cloud map is formed according to the coordinates of multiple front scanning points. The calculation formula of the coordinate point cloud is as follows:
[0070] X = Lcosαcosβ
[0071] Y = Lcosαcosβ
[0072] Z = Lcosαcosβ
[0073] In the formula, L is the distance from the target scanning point to the laser scanner after the emitted laser is reflected, α and β are the rotation angles of the laser mirror of the laser scanner in the horizontal and vertical directions respectively, and X, Y, and Z are the x, y, and z axis coordinate values of the target scanning point relative to the laser scanner.
[0074] S2. Use the depth camera installed on the knee of the humanoid robot to calculate the height of the obstacle.
[0075] In step S2, when the humanoid robot walks to a stop in front of the obstacle, calculating the height of the obstacle measured by the depth camera installed on its hip specifically means: the depth camera forms images through its left and right two cameras, and sends the data measured by the depth camera to the main controller. The main controller uses the principle of triangulation for calculation, and obtains the distance between the imaging point of the bottom of the obstacle in the depth camera and the imaging point of the top of the obstacle as the height of the obstacle. The calculation formula is as follows:
[0076]
[0077] In the formula, Z is the distance between the imaging point of the bottom of the obstacle and the depth camera installed on the hip of the humanoid robot, f is the focal length of the depth camera, in pixel units, T is the distance between the imaging point of the top of the obstacle and the depth camera installed on the hip of the humanoid robot, x 0l is the horizontal axis coordinate of the scanning point of the front obstacle imaged on the left camera of the depth camera, x 0r$x$ is the horizontal axis coordinate of the scanning point of the obstacle in front imaging on the right camera of the depth camera. After measuring the height of the obstacle, it is compared with the leg length of the humanoid robot for subsequent steps.
[0078] According to the height of the obstacle measured by the depth camera and the coordinate point cloud map formed by the laser scanner, the terrain map of the road ahead can be obtained, which is convenient for the robot to select the next operation of bypassing or forward handstand to cross the obstacle.
[0079] S3. Determine whether the height of the obstacle is greater than the leg length of the humanoid robot.
[0080] If so, control the humanoid robot to walk in front of the obstacle and perform the operation of bypassing the obstacle.
[0081] If not, control the humanoid robot to walk in front of the obstacle and perform the operation of handstand flipping to cross the obstacle.
[0082] In step S3, compare the height of the obstacle measured by using the triangulation principle with the depth camera and the leg length of the humanoid robot; if the height of the obstacle is greater than the leg length of the humanoid robot and the humanoid robot cannot cross the obstacle by forward handstand, then take the bypass operation, that is, the main controller plans the footstep sequence of the humanoid robot to bypass the obstacle and return to the original path according to the terrain map of the road ahead drawn by the laser scanner; if the height of the obstacle is less than the leg length of the humanoid robot, the humanoid robot can cross the obstacle by forward handstand to complete obstacle avoidance; the humanoid robot is controlled by dynamic model derivation: model the humanoid robot, simplify the humanoid robot into a single-mass inverted pendulum with springs and dampers, and the motion equation is:
[0083]
[0084] In the formula, $m$ is the mass of the robot, $s$ is the length from the support point of the inverted pendulum to the mass point, that is, the height from the ground to the mass point when the robot is in the ready position before walking, $c$ is the damping constant, $k$ is the spring constant, $g$ is the acceleration due to gravity, $\tau$ is the torque applied to the robot, $\delta$ is the angle of the inverted pendulum, corresponding to the inclination angle of the robot's leg in the vertical direction to the ground, and $u$ δ is the output angle of the inverted pendulum, input by the robot controller;
[0085] Therefore, $u$ is the displacement of the controller for the center of mass, approximately:
[0086]
[0087] Where, m is the mass of the robot, s is the length from the support point of the inverted pendulum to the mass point, i.e., the height from the ground to the mass point when the robot is in the preparatory posture before walking, c is the damping constant, k is the spring constant, g is the acceleration due to gravity, τ is the torque applied to the robot, δ is the angle of the inverted pendulum, corresponding to the inclination angle of the robot's leg with respect to the vertical direction to the ground, u δ is the output angle of the inverted pendulum, which is input by the robot controller;
[0088] The zero moment point (ZMP) of the robot in the x direction is calculated as follows:
[0089]
[0090] Where, F y is the ground reaction force in the y direction, F z is the ground reaction force in the z direction;
[0091] Since F y is also the reaction torque of the spring and shock absorber, it can be written due to the action of the spring and shock absorber as:
[0092]
[0093] Where s is the length from the support point of the inverted pendulum to the mass point, i.e., the height from the ground to the mass point when the robot is in the preparatory posture before walking, c is the damping constant, k is the spring constant, δ is the angle of the inverted pendulum, corresponding to the inclination angle of the robot's leg with respect to the vertical direction to the ground;
[0094] When the system model is not disturbed, it can be derived that
[0095] and The state space equations of the two equations:
[0096]
[0097] Where m is the mass of the robot, s is the length from the support point of the inverted pendulum to the mass point, i.e., the height from the ground to the mass point when the robot is in the preparatory posture before walking, c is the damping constant, k is the spring constant, g is the acceleration due to gravity, δ is the angle of the inverted pendulum, corresponding to the inclination angle of the robot's leg with respect to the vertical direction to the ground, u is the displacement of the controller with respect to the center of mass; the states of the two equations are the angle δ and angular velocity of the robot
[0098] After completing the dynamic modeling of the robot and determining that the robot can cross obstacles by performing a forward handstand, the Model Predictive Control (MPC) algorithm is introduced to adjust the posture of the robot during the crossing action. That is, based on the rapid calculation of the robot's dynamic model, it is used to predict the system output within a limited time range and optimize according to the control action to minimize the specified performance index. The MPC algorithm encodes the distance input signal measured by the depth camera into a neuron, which is achieved by considering a learning rule. After the input distance is presented, the neuron that best matches the input distance is selected as the optimal neuron. Then, the weight of the optimal neuron is updated, and under the guidance of the reward function, the correct control action is autonomously found using the input pattern information and the performance index to be optimized.
[0099] The following are the design of the humanoid robot controller and the application steps of the algorithm:
[0100] S31. Define the network structure and hyperparameters, and randomly initialize the input and output weights;
[0101] S32. At time I(t), present the input pattern and measure the distance using the depth camera, and select the optimal neuron N S , which minimizes the distance between I(t) and the input weight vector connecting N S to the input nodes; where I(t) is the input pattern;
[0102] S33. For the optimal neuron N S , the corresponding output weight is used to generate the control action; that is, the perturbation version adds a Gaussian zero-mean random variable λ to the optimal neuron weight and multiplies it by the gain α S related to the optimal neuron N S ; α S starts from an initial value, fixes a final value, and as the learning stage progresses, α S linearly decreases between these two values; where O(t) is the output control signal;
[0103] S34. Evaluate the actual reward function R(t) together with the reward increment ΔR = R(t) - R(t - 1); if ΔR > b s , where b s is the average increment related to the neuron N S ; the update rule of b s is as follows:
[0104]
[0105] where σ > 0, is the new average increment, R is the old average increment, and ΔR is the average increment;
[0106] Connect to the optimal neuron N S The input and output weights connected to it and its adjacent neurons are updated according to the following rules:
[0107] w i,in (t + 1)= w i,in (t)+ η in ζ(I(t)- w i,in (t))
[0108] w i,out (t + 1)= w i,out (t)+ η out ζ(O(t)- w i,out (t))
[0109] where η in is the input learning rate, η out is the output learning rate, w i,in is the input weight, w i,out is the output weight, ζ is the neighborhood function of the neuron. If N R represents the neighborhood with a radius of R, then around the current optimal neuron N S there is:
[0110]
[0111] Therefore, ζ establishes a neighborhood with a radius around N S and the weights are updated within this range;
[0112] In each learning iteration, an output pattern is given, an optimal neuron N S is selected and an output signal is provided. Then the reward is evaluated and the learning phase is executed, that is, w i,in , w i,out , b s and α S are updated. After several iterations, the update of other learning hyperparameters is executed; Figure 2 is the flow chart of the control loop for the i-th learning iteration;
[0113] S4. After completing the striding motion, the MPC algorithm calculates the forces required to maintain the body's standing position, and then uses these forces as the input signals of the controller. The controller provides stable joint torques. If the robot's structure changes at this time, resulting in a center-of-gravity shift, an adaptive non-linear controller MMC is added as a feed-forward error compensator to adjust the robot's posture. This adaptive non-linear controller MMC provides an additional non-linear feed-forward control action. After transforming the joint torques through the Jacobian matrix, it is added to the underlying leg joint controller in the form of forces. A reward function is introduced:
[0114]
[0115] where Θ ref is the reference pitch velocity of the robot, Θ is the actual pitch velocity of the robot, is the reference pitch acceleration of the robot, is the actual pitch acceleration of the robot, and β is the introduced gain. The priority of the pitch velocity with respect to the pitch position error is established: as the pitch error decreases, the correlation of the pitch velocity error becomes greater;
[0116] The algorithm for controlling the robot realizes the update of the output weights by improving the reward function, thereby ensuring the stability of the completion of the striding motion. For simplicity, only the pitch angle in the sagittal plane is restricted. The control objective is to keep the pitch angle of the robot as close as possible to the reference angle. Therefore, the reward function takes into account the error between the reference and the actual pitch of the robot.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for humanoid robot obstacle avoidance based on model predictive control, characterized in that: The specific method is as follows: The laser scanner arranged on the head of the humanoid robot is used to scan the road in front of the humanoid robot, and a coordinate point cloud is formed according to the scanning result to draw a topographic map of the road in front of the humanoid robot; The height of the obstacle is calculated by using a depth camera installed on the patella of the humanoid robot; Determine whether the obstacle height is greater than the leg length of the humanoid robot; If yes, the humanoid robot is controlled to walk to the obstacle and perform an obstacle bypassing operation; If not, the humanoid robot is controlled to walk to the obstacle and perform an inverted flip over the obstacle; To perform an inverted flip over an obstacle, the specific steps are as follows: Construct a dynamic model of a humanoid robot; A humanoid robot controller based on a model predictive control algorithm is used to complete the posture adjustment of the humanoid robot during the inverted flip stride action. Construct a humanoid robot dynamics model. The specific method is as follows: The humanoid robot is simplified as a single-mass inverted pendulum with a spring and a damper, and the equation of motion is: In the formula, is the mass of the robot, is the length from the support point of the inverted pendulum to the mass point, that is, the height from the ground to the mass point when the robot is in the preparation posture before walking, is the damping constant, is the spring constant, is the acceleration due to gravity, is the torque applied to the robot, is the angle of the inverted pendulum, corresponding to the inclination angle of the robot's legs in the vertical direction to the ground, is the output angle of the inverted pendulum, which is input by the robot controller; is the displacement of the controller relative to the center of mass, which is approximately: Zero moment point of the robot in the x-direction The calculation is as follows: In the formula, is the ground reaction force in the y direction, is the ground reaction force in the z direction; because It is also the reaction torque of the spring and shock absorber, which can be written as: When undisturbed, it can be deduced that and The state space equations of the two formulas are: The humanoid robot controller encodes the distance input signal measured by the depth camera into a neuron, and after the input distance is presented, the neuron that best matches the input distance is selected as the optimal neuron; After that, the weights of the optimal neurons are updated, and the correct control actions are found autonomously under the guidance of the reward function using the input pattern information and the performance indicators to be optimized. Humanoid robot controller, the specific setting method is as follows: Define the network structure and hyperparameters, and randomly initialize the input and output weights; In time The input pattern is presented at the same time, and the depth camera is used to measure the distance and select the optimal neuron , it will Connect with The distance between the input weight vectors to the input nodes is minimized; where is the input mode; For the optimal neuron , the corresponding output weight Used to generate control actions; that is, the orbiting version , at the optimal neuron weight Add a Gaussian zero-mean random variable to , multiplied by the optimal neuron Related Gains ; Starting from an initial value, a final value is fixed, and as the learning phase progresses, It decreases linearly between these two values; is the output control signal; Actual Reward Function With reward increment Evaluate together; if ,in For neurons The average increment of the correlation; The update rules are as follows: in , is the new average increment, is the old average increment, is the average increment; Connect to the best neuron The input and output weights of and its neighboring neurons are updated according to the following rules: in is the input learning rate, is the output learning rate, Input weights, is the output weight, is the neighborhood function of the neuron, if Represents a neighborhood with a radius of R, then in the current optimal neuron Around, there are: in Represents the optimal neuron Any surrounding neuron, Built around The radius neighborhood, the weight is updated within this range; In each learning iteration, given an output pattern, an optimal neuron is selected and provide an output signal, after which the reward is evaluated and the learning phase is performed, i.e., updating , , and , after several iterations, perform updates of other learning hyperparameters; After the inverted flip over obstacle maneuver, the humanoid controller calculates the forces required to maintain the body's standing position, and then uses these forces as input signals to the controller, which provides stable joint torques; If the robot structure changes and causes the center of gravity to shift, an adaptive nonlinear controller MMC is added as a feedforward error compensator to adjust the robot's posture. This adaptive nonlinear controller MMC provides an additional nonlinear feedforward control action, which is added to the underlying leg joint controller in the form of force after transforming the joint torque through the Jacobian matrix; a reward function is introduced: In the formula is the reference pitch velocity of the robot, is the actual pitch velocity of the robot, is the reference pitch acceleration of the robot, is the actual pitch acceleration of the robot, For the introduced gain, the priority of pitch velocity to pitch position error is established: as the pitch error decreases, the relevance of the pitch velocity error becomes greater; The algorithm that controls the robot updates the output weights to ensure the stability of the stride by improving the reward function, which takes into account the error between the reference and the actual pitch of the robot.
2. The method for avoiding obstacles of a humanoid robot based on model predictive control as claimed in claim 1, characterized in that: The coordinate point cloud is formed, and the specific formula is as follows: Where L is the distance from the target scanning point to the laser scanner after the emitted laser is reflected back. and are the rotation angles of the laser reflector of the laser scanner in the horizontal and vertical directions, respectively; X, Y and Z are the x, y and z axis coordinates of the target scanning point relative to the laser scanner, respectively.
3. The method for avoiding obstacles of a humanoid robot based on model predictive control as claimed in claim 1, characterized in that: To calculate the height of the obstacle, you need to calculate the distance Z between the obstacle and the depth camera. The specific calculation formula is as follows: Where Z is the distance between the imaging point at the bottom of the obstacle and the depth camera mounted on the hip of the humanoid robot, f is the focal length of the depth camera, T is the distance between the left and right cameras of the depth camera, is the horizontal coordinate of the scanning point of the obstacle in front of us imaged on the left camera of the depth camera, It is the horizontal coordinate of the scanning point of the obstacle in front of the image on the right camera of the depth camera.
Citation Information
Patent Citations
Autonomous navigation method and device for quadruped robot, computer equipment and storage medium
CN111752285A