Robot safe near-optimal motion planning method based on hybrid reinforcement learning

By combining a hybrid reinforcement learning method with a multi-layer neural network and HOCBF, the problems of vanishing Lie derivatives and insufficient local optimization in high-order nonlinear systems are solved, global near-optimal planning and real-time safety constraints for robot motion control are achieved, ensuring the stability and safety of the system.

CN120779747APending Publication Date: 2025-10-14SHANGHAI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510979930.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing robot motion control methods suffer from the vanishing Lie derivative problem, insufficient local optimization, lack of neural network convergence guarantee, and insufficient safety verification framework in high-order nonlinear systems, resulting in the inability to effectively achieve global optimality and stability.

Method used

A hybrid reinforcement learning method is adopted, combined with multi-layer neural networks and high-order control barrier functions (HOCBF). Through the deep integration of the target navigator and the safety controller, a hierarchical assurance architecture is constructed to achieve global near-optimal path planning and real-time safety constraints.

Benefits of technology

It achieves the synchronization of global near-optimal path planning and safety constraints in high-order systems, overcomes the vanishing Lie derivative problem, breaks through the local optimization bottleneck, and ensures the stability and safety of the control system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120779747A_ABST
    Figure CN120779747A_ABST
Patent Text Reader

Abstract

The invention discloses a robot safe near-optimal motion planning method based on hybrid reinforcement learning, and belongs to the technical field of robot motion. The method comprises the following steps: initializing an initial state of a robot, a multilayer neural network weight and a high-order control barrier function HOCBF parameter; generating a target navigator through forward propagation of the multi-layer neural network; screening a safety controller from a safety control set defined by the HOCBF; fusing the target navigator and the safety controller to generate a final control strategy; calculating a Bellman error; and updating the initial state of the robot and the weight of the multi-layer neural network based on the final control strategy and the Bellman error. According to the invention, a neural network approximation target navigator and an HOCBF safety controller are deeply fused, so that a hybrid control system with a layered guarantee framework is constructed. The problem of safety guarantee failure caused by Lie derivative disappearance of a traditional CBF in a high-order system is solved, and meanwhile, the technical bottleneck that an existing HOCBF scheme is limited to local optimization is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of robot motion, and particularly relates to a robot safe near-optimal motion planning method based on hybrid reinforcement learning. BACKGROUND

[0002] With the rapid development of artificial intelligence, the use scenarios of robots are more and more, and the corresponding motion control systems are also emerging in an endless stream. Robots perform path planning and obstacle avoidance automatic driving through autonomous motion control systems.

[0003] However, the existing similar safety control schemes have the following defects: 1. Application limitations of traditional control barrier function (CBF) in complex systems: the traditional CBF method requires that the control affine system dynamics must satisfy the Lie derivative existence condition. However, in high-order systems, as the system order increases, the Lie derivative disappears, and the control quantity cannot effectively act on the state dimension represented by the CBF, which makes the classic CBF scheme not directly applicable to typical high-order nonlinear systems such as unmanned aerial vehicles and robotic arms. 2. Optimization limitations of high-order control barrier function (HOCBF): the existing HOCBF scheme proposes a QP framework that can handle time-varying systems, but its optimization process only focuses on local optimal solutions, lacks global optimality guarantee. At the same time, this method does not consider the initial state constraint condition, which has stability risks in actual engineering applications. 3. Convergence defects of neural network approximation scheme: the HJB-HOCBF hybrid method proposed by the existing HOCBF scheme can construct a constrained optimal controller, but the neural network approximation algorithm used lacks strict mathematical convergence proof, which makes the reliability of the algorithm in complex dynamic environments cannot be theoretically guaranteed. 4. Empirical limitations of safety verification: the HOCBF scheme based on reinforcement learning proposed by the existing HOCBF scheme verifies safety in experiments, but lacks a formal verification framework, and cannot establish a universal safety theoretical system. SUMMARY

[0004] The application aims to provide a robot safe near-optimal motion planning method based on hybrid reinforcement learning to improve the disconnection between theory and practice of high-order system safety control.

[0005] The technical solution adopted by the application is as follows: a robot safe near-optimal motion planning method based on hybrid reinforcement learning, the method comprising the following steps: Initializing the initial state of the robot, the weights of the multi-layer neural network, and the parameters for constructing the high-order control barrier function HOCBF ; Generating a target navigator through forward propagation of the multi-layer neural network; Real-time screening of a safety controller from a safety control set defined by the high-order control barrier function HOCBF; Fuse the target navigator and safety controller to generate the final control strategy; Calculate Bellman error based on synchronous sampling of multi-layer neural network; Based on the final control strategy and Bellman error, the robot's initial state and multi-layer neural network weights are updated.

[0006] Furthermore, the multi-layer neural network includes a hidden layer and an output layer. The weight of the multi-layer neural network is , For the The network weight matrix of the layer, is the number of layers of the multi-layer neural network; Generating a target navigator includes the following: Hidden layer activation: ; in, is a nonlinear activation function, For the The bias vector of the layer, together with the linear transformation, constitutes the input of the current hidden layer. For the The feature vector of the layer is used to pass it to the next layer; Output layer control amount: ; in, is the control weight matrix in the cost function, is the transpose of the Jacobian matrix of the control input term in the system dynamics, that is, the direction of the system control influence, is the characteristic function The gradient transpose describes the rate of change of the value function, is the output of the last layer L of the neural network, that is, the feature map of the entire network to the input state x. is the depth of the neural network, that is, the layer number of the output layer; Get the target navigator .

[0007] Furthermore, the security control set defined by the high-order control barrier function HOCBF is , in the security control set Real-time screening security controller Among them, the control input Satisfaction form is of The order control barrier function inequality, is the state vector of the system, is the control input, that is, the control strategy under the current state.

[0008] Furthermore, the fusion target navigator and safety controllers , generate the final control strategy .

[0009] Furthermore, based on the final control strategy Update the robot's initial state , where the initial state Follow the kinetic equation , where the initial state Follow the kinetic equation ,in is the state vector of the system, is the control input, is the drift term of the system, is the control input matrix, is the time derivative of the system state, that is, the rate of change of the state over time.

[0010] Further, calculate the Bellman error : ; in, is the state weight matrix in the cost function, is the transpose of the state vector, is the output of the last layer L of the neural network, that is, the feature map of the entire network to the input state x. for , represents the characteristic function gradient in the current state By controlling the direction The strength structure acting on the system, is the characteristic function The gradient transpose describes the rate of change of the value function, is the drift term of the system, is the control weight matrix in the cost function, For safety controller, For security control items, is the initial state of the robot.

[0011] Furthermore, using the update law Update the weights of a multi-layer neural network Layer weights for : , ; in, is the weight of the multi-layer neural network before updating, For update The weight update law of the multi-layer neural network is is the time step, is the normalized gradient, is the learning rate, is the stratified normalization factor, is the Bellman error, is a positive definite matrix, is the damping term; ; ; ; in, is the depth of the neural network, that is, the layer number of the output layer, is the feature vector of the i-th layer, is the feature vector of the kth layer, which is used to pass to the next layer. It is the output of the last L layer of the neural network, that is, the feature map of the entire network to the input state x.

[0012] Furthermore, the high-order control barrier function HOCBF is recursively defined by , elevate the safety constraints to the control input space and directly limit The feasible domain of for: ; in, is a security function, For the final control strategy, is the initial state of the robot, For function Drift term along the system The kth order Lie derivative of , To control the direction of the system The Lie derivative of is the residual term, which serves as the constant bias term in the HOCBF constraint. For class Function, used to introduce relaxation terms and enhance robustness, needs to be continuous and strictly increasing, and ; By constructing a security control set is a convex set and minimizes the safety controller Target Navigator Minimize interference.

[0013] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: The application fuses a neural network NN approximation target navigator and a HOCBF safety controller in depth, and constructs a hybrid control system with a layered guarantee architecture. BRIEF DESCRIPTION OF DRAWINGS Figure 1 is a method flowchart of the application. DETAILED DESCRIPTION

[0014] The application will be described in detail below with reference to the accompanying drawings.

[0015] In order to make the purpose, technical scheme and advantages of the application clearer, the application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are only used to explain the application and do not limit the application.

[0016] The existing robot autonomous motion safety control method has the following problems: (1) in a high-order system, as the order of the system increases, the traditional CBF method will cause the problem of vanishing of the Lie derivative, the control quantity cannot effectively act on the state dimension represented by the CBF, so that the classical CBF scheme cannot be directly applied to typical high-order nonlinear systems such as unmanned aerial vehicles and mechanical arms; (2) the QP framework proposed by the existing HOCBF scheme can handle time-varying systems, but its optimization process only focuses on local optimal solution, lacks global optimality guarantee, and does not consider the initial state constraint condition, which has stability risk in actual application; (3) the HJB-HOCBF hybrid method proposed by the existing HOCBF scheme can construct a constraint optimal controller, but the neural network approximation algorithm used by it lacks strict mathematical convergence proof, which leads to the reliability of the algorithm in complex dynamic environment cannot be theoretically guaranteed; (4) the HOCBF scheme based on reinforcement learning proposed by the existing HOCBF scheme verifies the safety in the experiment, but lacks a formal verification framework, and cannot establish a universal safety theoretical system.

[0017] Therefore, the application proposes a robot safety near-optimal motion planning method based on hybrid reinforcement learning, which expects to plan the approximate optimal motion of the robot based on the hybrid reinforcement learning method. The application fuses a neural network approximation navigator and a HOCBF guarantee mechanism to construct a hybrid reinforcement learning architecture with double guarantee, and realizes the following technical breakthroughs: for the first time, the optimal navigation and strict satisfaction of safety constraints are simultaneously achieved in motion planning, and based on Lyapunov stability theory, the triple mathematical guarantees of approximate optimality proof, convergence verification and collision-free safety proof of the controller are established.

[0018] like Figure 1 As shown, a robot safe near-optimal motion planning method based on hybrid reinforcement learning includes the following steps: Step S100: Initialize the robot's initial state, the multi-layer neural network weights, and the parameters used to construct the high-order control barrier function HOCBF ; Step S200: Generate a target navigator through multi-layer neural network forward propagation; Step S300: Filtering safety controllers in real time from the safety control set defined by the high-order control barrier function HOCBF; Step S400: Fusing the target navigator and the safety controller to generate a final control strategy; Step S500: Calculating Bellman error based on multi-layer neural network synchronous sampling; Step S600: Based on the final control strategy and Bellman error, update the robot's initial state and the multi-layer neural network weights.

[0019] This invention deeply integrates a target navigator approximated by a neural network (NN) with a HOCBF safety controller to construct a hybrid control system with a layered assurance architecture. The NN navigator implements global near-optimal path planning, while the HOCBF controller enforces real-time safety constraints. This overcomes the safety failure issue of traditional CBF solutions in high-order systems due to the vanishing Lie derivatives, while also breaking through the technical bottleneck of existing HOCBF solutions, which are limited to local optimization.

[0020] In step S200, the multi-layer neural network includes a hidden layer and an output layer. The weights of the multi-layer neural network are , represents the neural network weight matrix of the kth layer, is the number of layers of the multi-layer neural network; Generating a target navigator includes the following: Hidden layer activation: ; in, is a nonlinear activation function, For the The bias vector of the layer, together with the linear transformation, constitutes the input of the current hidden layer. For the The feature vector of the layer is used to pass it to the next layer; Output layer control amount: ; in, is the control weight matrix in the cost function, is the transpose of the Jacobian matrix of the control input term in the system dynamics, that is, the direction of the system control influence, is the characteristic function the rate of change of the value function, is the output of the last layer L of the neural network, i.e., the feature map of the entire network to the input state x, is the depth of the neural network, i.e., the layer number of the output layer; get the target navigator .

[0021] The hidden layer nonlinear activation improves the expression ability of the strategy table; enhances the approximation ability of high-dimensional / nonlinear motion planning problems, and reduces the suboptimal trajectory generation error caused by single-layer networks.

[0022] In step S300, the safety control set defined by the high-order control barrier function HOCBF is , and the safety controller is filtered in real time in the safety control set . Wherein, the control input satisfies the high-order control barrier function inequality of the form , where is the state vector of the system, is the control input, i.e., the control strategy under the current state. The safety control set defined by the high-order control barrier function HOCBF converts the safety constraint into a linear programming problem of the safety controller , ensuring that the control input clock satisfies the constraint of the HOCBF , solving the problem of conflict between safety and optimality in traditional methods.

[0023] In step S400, the target navigator and the safety controller are fused to generate the final control strategy . The target navigator and the safety controller are linearly superimposed instead of being connected in series or in switching mode, avoiding strategy oscillation and preserving optimality.

[0024] In steps S500-S600, based on the Bellman error calculation and weight update based on synchronous sampling, the multi-layer neural network training is synchronized with the system state evolution, solving the weight update lag problem caused by traditional hierarchical training, and making the gradient information of the error feedback to each hidden layer in real time, improving the training efficiency.

[0025] The hidden layer extracts the abstract features of the state layer by layer through the nonlinear activation function , so that the navigation strategy Able to adapt to complex environments, the deep network module and the HOCBF module are coupled in parallel architecture, and the target navigator output by NN With safety controller Avoid signal delay by linear superposition rather than series filtering; network depth It can be adjusted dynamically to realize on-demand allocation of computing resources. The reverse propagation path is cross-layer chain propagation, making the error Gradient information of each layer is used to correct the weights of each layer at the same time , improving the efficiency of parameter updating.

[0026] The multi-layer network approximates complex optimal strategies through composite nonlinear mapping, supporting scenarios such as dynamic obstacles and non-convex constraints; HOCBF constraints screen safety controllers through real-time linear programming. , ensuring safety while using Minimize and reduce the interference of safety compensation on optimality; the synchronous sampling mechanism synchronizes state updates with weight learning, eliminating the experience replay buffer of traditional asynchronous RL; Based on the final control strategy Update the robot's initial state , where the initial state Follow the kinetic equation .

[0027] Calculating Bellman Error : ; in, is the state weight matrix in the cost function, is the transpose of the state vector, is the output of the last layer L of the neural network, that is, the feature map of the entire network to the input state x. for , represents the characteristic function gradient in the current state By controlling the direction The strength structure acting on the system, is the characteristic function The gradient transpose describes the rate of change of the value function, is the drift term of the system, is the control weight matrix in the cost function, For safety controller, For security control items, is the initial state of the robot.

[0028] Using the update law Update the weights of a multi-layer neural network Layer weights for : ; ; ; ; ; wherein, is the multi-layer neural network weight before update, is the multi-layer neural network weight update law for updating , is the time step, is the normalized gradient, is the learning rate, is the hierarchical normalization factor, is the Bellman error, is a positive definite matrix, is a damping term; is the depth of the neural network, i.e., the layer number of the output layer, is the feature vector of the kth layer, is the feature vector of the kth layer for passing to the next layer, is the output of the last layer L of the neural network, i.e., the feature mapping of the entire network to the input state x. The principle of the HOCBF safety controller is as follows:

[0029] The high-order control barrier function HOCBF is defined recursively to lift the safety constraints to the control input space, directly limit the feasible region of , guarantee the forward invariance of the system trajectory, and thus guarantee the safety of the system, is: ; wherein, is a safety function, is the final control strategy, is the initial state of the robot, is the kth-order Lie derivative of the function along the system drift term , is the Lie derivative along the system control direction , is a residual term as a constant bias term in the HOCBF constraint, is a class function for introducing a relaxation term to enhance robustness, which needs to be continuous and strictly increasing, and ; By constructing the safety control set is a convex set and minimizes the safety controller Target Navigator The interference is minimized and the Pareto equilibrium of security and optimality is achieved.

[0030] The principle of the policy superposition mechanism is as follows: NN output target navigator , and smoothed by the activation function, while the safety controller Projected to The complementary space of the two still satisfies Lipschitz continuity after superposition, ensuring the safety set The convexity of remains unchanged, and the superposition operation only requires linear algebra to run, which reduces the computational complexity compared to existing optimization-based policy coordination methods.

[0031] use The differential characteristics of the order HOCBF transform the safety constraints into the affine form of the control input to ensure The convexity and closure of , thus ensuring the feasibility of real-time solution. According to the HOCBF theorem, there is a constructed safety control set for the disturbed system , so that for any bounded perturbation , can be controlled by safety compensation Pull the system back to the safe zone Inside. exist Under the action of Reciprocal order constraint Always true to ensure that the state trajectory does not cross the safety boundary , that is, to ensure safety.

[0032] The target navigator works as follows: Based on the universal approximation theorem, multi-layer neural networks are constructed by composite nonlinear activation functions. Continuous optimal strategies can be approximated with arbitrary precision , and calculate the cross-layer gradient through the complete chain rule, and use the current state to instantaneously calculate the error , eliminating the timing error caused by experience replay in traditional asynchronous RL, and Explicitly include security controls , so that the weight update process can simultaneously optimize the target navigation cost and safety control energy consumption, avoiding suboptimal solutions. Each layer of gradient is normalized separately to prevent the weight mutation caused by the difference in gradient magnitude in the deep network; the weight update law Introducing the damping term , suppressing gradient explosion and ensuring weight convergence.

[0033] In order to prove the convergence and weight boundedness of the present invention, a composite Lyapunov function is constructed: ; wherein, is the optimality term, is the safety term, is the weighted error term; the state vector is constructed, by Lyapunov direct method, it can be proved that Lyapunov function is less than 0, guaranteeing state convergence and weight boundedness.

[0034] Embodiment Set a scene, there are two fixed circular obstacles, the radius is 2, the center is respectively located and Coordinate; the robot needs to plan the path from the starting point to the target point, while avoiding collision with the obstacle. In the scene, it is necessary to verify the convergence of the neural network controller weight And the stability of the control strategy, the system needs to maintain the control performance in a long time running.

[0035] The robot can smoothly bypass the obstacle through the robot safety near-optimal motion planning method based on hybrid reinforcement learning provided by the application, the distance between the nearest obstacle in the path and the safety boundary (the radius of the obstacle is 2) is always greater than the safety boundary, the angular velocity control input and the acceleration input are adjusted cooperatively, ensuring that the robot avoids the obstacle with the minimum turning radius, while maintaining the stability of the linear velocity. The target navigator generated by the neural network controller is linearly superimposed with the HOCBF safety controller , guaranteeing that the control frequency and the system dynamics are synchronized. The parameters of the neural network controller are stable during the training process, and the robot finally reaches the target point accurately, and the terminal position error approaches zero.

[0036] The application can significantly improve the strategy optimization capability of high-dimensional complex systems while maintaining the safety of the original algorithm through the modular architecture and the hierarchical optimization mechanism.

[0037] The above only describes the preferred embodiments of the application and is not intended to limit the application, any modification, equivalent replacement and improvement made within the spirit and principle of the application shall be included in the protection scope of the application.

Claims

1. A safe near-optimal motion planning method for robots based on hybrid reinforcement learning, characterized by: Methods include the following: Initialize the robot's initial state, multi-layer neural network weights, and parameters for constructing the high-order control barrier function HOCBF ; Generate target navigator through multi-layer neural network forward propagation; Real-time screening of safety controllers from the safety control set defined by the high-order control barrier function HOCBF; Fuse the target navigator and safety controller to generate the final control strategy; Calculate Bellman error based on synchronous sampling of multi-layer neural network; Based on the final control strategy and Bellman error, the robot's initial state and multi-layer neural network weights are updated.

2. The robot safe near-optimal motion planning method based on hybrid reinforcement learning according to claim 1 is characterized in that: The multi-layer neural network includes hidden layers and output layers. The weights of the multi-layer neural network are , represents the neural network weight matrix of the kth layer, is the number of layers of the multi-layer neural network; Generating a target navigator includes the following: Hidden layer activation: ; in, is a nonlinear activation function, For the The bias vector of the layer, together with the linear transformation, constitutes the input of the current hidden layer. For the The feature vector of the layer is used to pass it to the next layer; Output layer control amount: ; in, is the control weight matrix in the cost function, is the transpose of the Jacobian matrix of the control input term in the system dynamics, that is, the direction of the system control influence, is the characteristic function The gradient transpose describes the rate of change of the value function, is the output of the last layer L of the neural network, that is, the feature map of the entire network to the input state x. is the depth of the neural network, that is, the layer number of the output layer; Get the target navigator .

3. The robot safe near-optimal motion planning method based on hybrid reinforcement learning according to claim 2 is characterized in that: The security control set defined by the high-order control barrier function HOCBF is , in the security control set Real-time screening security controller ; Among them, the control input Satisfaction form is of The order control barrier function inequality, is the state vector of the system, is the control input, that is, the control strategy under the current state.

4. The robot safe near-optimal motion planning method based on hybrid reinforcement learning according to claim 3 is characterized in that: Fusion Target Navigator and safety controllers , generate the final control strategy .

5. The robot safe near-optimal motion planning method based on hybrid reinforcement learning according to claim 4 is characterized in that: Based on the final control strategy Update the robot's initial state , where the initial state Follow the kinetic equation ,in is the state vector of the system, is the control input, is the drift term of the system, is the control input matrix, is the time derivative of the system state, that is, the rate of change of the state over time.

6. The robot safe near-optimal motion planning method based on hybrid reinforcement learning according to claim 5 is characterized in that: Calculating Bellman Error : ; in, is the state weight matrix in the cost function, is the transpose of the state vector, is the output of the last layer L of the neural network, that is, the feature map of the entire network to the input state x. for , represents the characteristic function gradient in the current state By controlling the direction The strength structure acting on the system, is the characteristic function The gradient transpose describes the rate of change of the value function, is the drift term of the system, is the control weight matrix in the cost function, For safety controller, For security control items, is the initial state of the robot.

7. The robot safe near-optimal motion planning method based on hybrid reinforcement learning according to claim 6 is characterized in that: Using the update law Update the weights of a multi-layer neural network Layer weights for : ; ; in, is the weight of the multi-layer neural network before updating, For update The weight update law of the multi-layer neural network is is the time step, is the normalized gradient, is the learning rate, is the stratified normalization factor, is the Bellman error, is a positive definite matrix, is the damping term; ; ; ; in, is the depth of the neural network, that is, the layer number of the output layer, For the The feature vector of the layer, is the feature vector of the kth layer, which is used to pass to the next layer. It is the output of the last L layer of the neural network, that is, the feature map of the entire network to the input state x.

8. The robot safe near-optimal motion planning method based on hybrid reinforcement learning according to claim 3 is characterized in that: The high-order control barrier function HOCBF is defined recursively , elevate the safety constraints to the control input space and directly limit The feasible domain of for: ; in, is a security function, For the final control strategy, is the initial state of the robot, For function Drift term along the system The kth order Lie derivative of , To control the direction of the system The Lie derivative of is the residual term, which serves as the constant bias term in the HOCBF constraint. For class Function, used to introduce relaxation terms and enhance robustness, needs to be continuous and strictly increasing, and ; By constructing a security control set is a convex set and minimizes the safety controller Target Navigator Minimize interference.

Citation Information

Cited By

  • Series-parallel robot self-adaptive motion control method and system

    CN121187140A

  • Hybrid robot adaptive motion control method and system

    CN121187140B