Method and system for controlling reinforcement learning formation of incomplete constraint robots

By constructing error functions and fault-tolerant control models and optimizing the control strategy of non-holonomic constrained robots, the problems of insufficient control performance and poor stability of formations in complex environments are solved, and efficient and stable formation control is achieved.

CN120704331APending Publication Date: 2025-09-26SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510859982.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing non-holonomic constrained robot formations suffer from insufficient control performance, poor environmental adaptability, and poor formation stability in complex environments, making it difficult to achieve efficient and stable task execution.

Method used

Construct an error function and fault-tolerant control model, including a virtual controller model, a nonlinear filter model, and an adaptive compensation mechanism model. Optimize the control strategy through the Hamilton-Jacobi-Bellman equation and the actor-critic algorithm to achieve precise control of the non-holonomic constrained robot.

Benefits of technology

The control accuracy and adaptability of non-holonomic constrained robot formations are improved, the robustness and fault tolerance of the system are enhanced, and efficient and stable control performance is ensured in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704331A_ABST
    Figure CN120704331A_ABST
Patent Text Reader

Abstract

The invention discloses a control method and system for reinforcement learning formation of incomplete constraint robots, and relates to the technical field of robot formation control, and the method comprises the steps: constructing an error function; constructing a fault-tolerant control model which comprises a virtual controller model, a nonlinear filter model and a self-adaptive compensation mechanism model; constructing a cost function based on the error function and the compensation mechanism; a virtual optimal control strategy is obtained by solving a Hamiltonian Jacobibellman equation; and performing approximation solution on an unknown function in the virtual optimal strategy by using an actor-commentator algorithm to obtain an optimal control strategy, and controlling the formation according to the optimal control strategy. Through a fault-tolerant mechanism and a reinforcement learning algorithm, the control performance and stability of the formation in a complex environment are improved, the incomplete constraint robot formation better adapts to the complex environment and task requirements, actuator faults and external interference are effectively handled, and the method is suitable for various application scenes such as environment monitoring, resource exploration and logistics transportation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of robot formation control, and in particular to a control method and system for a reinforcement learning formation of non-holonomically constrained robots. Background Art

[0002] With the continuous advancement of technology, unmanned systems have experienced rapid development. As a key branch of the unmanned systems field, nonholonomic robots have become a hot research focus. Currently, nonholonomic robot technology is undergoing rapid development, with technological integration and innovation as its main trends. In the future, nonholonomic robots will become more intelligent and autonomous, capable of performing tasks in more complex and harsh environments. At the same time, the application areas of nonholonomic robots will continue to expand, from traditional fields such as environmental monitoring and resource exploration to logistics and transportation, infrastructure inspection, disaster relief, and other fields. The rapid development of science and technology has made the use of a single nonholonomic robot insufficient for complex tasks. Therefore, the use of multiple nonholonomic robots to perform tasks simultaneously is inevitable.

[0003] When performing tasks such as environmental monitoring, resource exploration, and rescue, nonholonomic robots are often subject to uncertainties such as external interference, placing increasingly stringent demands on their control performance. The convergence speed and transient performance of formation errors are crucial during these tasks, and thus, performance-based control has garnered widespread attention in engineering applications. However, factors such as harsh environments, limited resources, and actuator failures in nonholonomic robots can affect system control performance and hinder the completion of control tasks, resulting in suboptimal optimal control performance for existing nonholonomic robot formations.

[0004] Nowadays, the application prospects of non-holonomic constrained robots in many fields have become apparent. There is an urgent need for a reinforcement learning formation control method for non-holonomic constrained robots that can solve the problems of insufficient control performance, poor environmental adaptability, and poor formation stability existing in related technologies. Summary of the Invention

[0005] The purpose of this application is to provide a control method and system for a reinforcement learning formation of non-holonomic constrained robots, which can realize efficient and stable operation of non-holonomic constrained robot formations in complex environments and significantly improve the fault tolerance and adaptability of the system.

[0006] To achieve the above objectives, this application provides the following solutions:

[0007] In a first aspect, the present application provides a control method for a reinforcement learning formation of non-holonomic constrained robots, comprising:

[0008] Constructing an error function for a nonholonomic constrained robot formation; the error function is used to calculate the error between the actual state and the desired state of the nonholonomic constrained robot formation;

[0009] Constructing a fault-tolerant control model; the fault-tolerant control model includes a virtual controller model, a nonlinear filter model, and an adaptive compensation mechanism model; the virtual controller model is used to generate a virtual control law for each nonholonomic constrained robot; the nonlinear filter model is used to calculate the derivative of the virtual control law for each nonholonomic constrained robot; the adaptive compensation mechanism model is used to calculate the compensation error of each nonholonomic constrained robot based on the derivative of the virtual control law;

[0010] Constructing a cost function for a nonholonomic constrained robot formation based on the error function and the adaptive compensation mechanism model;

[0011] The Hamilton-Jacobi-Bellman equation is obtained based on the cost function, and the Hamilton-Jacobi-Bellman equation is solved to obtain the virtual optimal control strategy;

[0012] Based on the fault-tolerant control model, the actor-critic algorithm is used to approximate and solve the unknown continuous function in the virtual optimal control strategy to obtain the optimal control strategy.

[0013] Control the nonholonomic constrained robot formation according to the optimal control strategy.

[0014] In a second aspect, the present application provides a control system for a reinforcement learning formation of non-holonomic constrained robots, comprising:

[0015] An error function construction module is used to construct an error function of a nonholonomic constrained robot formation; the error function is used to calculate the error between the actual state and the expected state of the nonholonomic constrained robot formation;

[0016] a fault-tolerant control model construction module, configured to construct a fault-tolerant control model; the fault-tolerant control model comprising a virtual controller model, a nonlinear filter model, and an adaptive compensation mechanism model; the virtual controller model being configured to generate a virtual control law for each nonholonomic constrained robot; the nonlinear filter model being configured to calculate a derivative of the virtual control law for each nonholonomic constrained robot; and the adaptive compensation mechanism model being configured to calculate a compensation error for each nonholonomic constrained robot based on the derivative of the virtual control law;

[0017] A cost function construction module, configured to construct a cost function for a nonholonomic constrained robot formation based on the error function and the adaptive compensation mechanism model;

[0018] A virtual optimal control strategy solving module is used to obtain the Hamilton-Jacobi-Bellman equation based on the cost function and solve the Hamilton-Jacobi-Bellman equation to obtain the virtual optimal control strategy;

[0019] The optimal control strategy solving module is used to approximate and solve the unknown continuous function in the virtual optimal control strategy based on the fault-tolerant control model and adopt the actor-critic algorithm to obtain the optimal control strategy;

[0020] The control execution module is used to control the non-holonomic constrained robot formation according to the optimal control strategy.

[0021] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the control method for a reinforcement learning formation of non-holonomic constrained robots as described in any one of the above.

[0022] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the control method for the reinforcement learning formation of non-holonomic constrained robots described in any one of the above.

[0023] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the control method of the reinforcement learning formation of non-holonomic constrained robots described in any one of the above.

[0024] According to the specific embodiments provided in this application, this application has the following technical effects:

[0025] This application provides a control method and system for a reinforcement learning formation of non-holonomically constrained robots. By constructing an error function for the non-holonomically constrained robot formation, the deviation between the actual and desired states of the formation is accurately measured, solving the problem of control direction deviation caused by inaccurate error calculation, and achieving accurate monitoring and feedback of the formation state. By constructing a fault-tolerant control model that includes a virtual controller model, a nonlinear filter model, and an adaptive compensation mechanism model, the virtual controller model generates a preliminary control law, the nonlinear filter model calculates its derivative to make the control smoother, and the adaptive compensation mechanism model calculates compensation errors based on the derivatives to cope with interference. This series of operations solves the problem of system control performance degradation due to uncertain factors in complex environments, achieves preliminary optimization of the control strategy and dynamic error compensation, and enhances the robustness of the system. By constructing a cost function based on the error function and the adaptive compensation mechanism model, and solving the Hamilton-Jacobi-Bellman equation accordingly to obtain a virtual optimal control strategy, the actor-critic algorithm is then used to approximate the unknown continuous function in the virtual optimal control strategy to obtain the optimal control strategy. This solves the problem that existing control methods are difficult to achieve optimal control effects in complex and changing environments, and achieves efficient, stable and precise control of non-completely constrained robot formations, thereby significantly improving the overall control performance and task execution efficiency of the formation. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0027] Figure 1 A flowchart of a control method for a reinforcement learning formation of non-holonomic constrained robots provided in one embodiment of the present application.

[0028] Figure 2 This is a diagram of a nonholonomic constrained robot formation of a leader-follower provided by an embodiment of the present application; wherein, Figure 2 (a) Spatial layout diagram of a leader-follower nonholonomic constrained robot formation. Figure 2 (b) Communication and collision avoidance area diagram for a non-holonomic constrained robot formation.

[0029] Figure 3 This is a communication topology diagram of a non-holonomic constrained robot formation provided in one embodiment of the present application.

[0030] Figure 4 This is a snapshot of the trajectory of a non-holonomic constrained robot formation provided by an embodiment of the present application.

[0031] Figure 5 This is a distance tracking error diagram of a non-holonomic constrained robot formation provided by an embodiment of the present application.

[0032] Figure 6 This is a diagram of the angle tracking errors of 1 and 0 and 1 and 3 in a nonholonomic constrained robot formation provided by an embodiment of the present application.

[0033] Figure 7 This is a diagram of the angle tracking errors of the robots No. 2 and 0, and No. 2 and 4 in a nonholonomic constrained robot formation provided by an embodiment of the present application.

[0034] Figure 8 This is a control force control curve diagram of a non-holonomic constrained robot formation provided by an embodiment of the present application.

[0035] Figure 9 This is a control torque control curve diagram of a non-holonomic constrained robot formation provided by an embodiment of the present application.

[0036] Figure 10 A schematic diagram of the functional modules of a control system for a reinforcement learning formation of non-holonomic constrained robots provided in another embodiment of the present application.

[0037] Figure 11 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0039] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0040] In an exemplary embodiment, Figure 1 As shown, a control method for a reinforcement learning formation of non-holonomic constrained robots is provided, comprising the following steps 201 to 206. In which:

[0041] Step 201 : constructing an error function of a nonholonomic constrained robot formation; the error function is used to calculate the error between an actual state and a desired state of the nonholonomic constrained robot formation.

[0042] Step 202, construct a fault-tolerant control model; the fault-tolerant control model includes a virtual controller model, a nonlinear filter model and an adaptive compensation mechanism model; the virtual controller model is used to generate a virtual control law for each non-complete constrained robot; the nonlinear filter model is used to calculate the derivative of the virtual control law for each non-complete constrained robot; the adaptive compensation mechanism model is used to calculate the compensation error of each non-complete constrained robot based on the derivative of the virtual control law.

[0043] Step 203: constructing a cost function for the non-holonomic constrained robot formation based on the error function and the adaptive compensation mechanism model.

[0044] Step 204 : obtaining the Hamilton-Jacobi-Bellman equation based on the cost function, and solving the Hamilton-Jacobi-Bellman equation to obtain a virtual optimal control strategy.

[0045] Step 205 : Based on the fault-tolerant control model, the actor-critic algorithm is used to approximate and solve the unknown continuous function in the virtual optimal control strategy to obtain the optimal control strategy.

[0046] Step 206 : Control the nonholonomic constrained robot formation according to the optimal control strategy.

[0047] By implementing the above steps 201 to 206, the present application can improve the accuracy and adaptability of the control strategy, solve the problems of insufficient control performance, actuator failure, and poor formation stability in the existing technology, and enable the system to maintain efficient and stable control performance in complex and changing environments, meeting the application requirements in various mission scenarios such as environmental monitoring, resource exploration, and rescue, and providing strong support for the widespread application of non-complete constrained robots in multiple fields.

[0048] In another exemplary embodiment of the present application, step 201 specifically includes:

[0049] A dynamic model of a nonholonomic constrained robot is constructed, and a neural network is used to calculate unknown nonlinear terms in the dynamic model.

[0050] Based on the dynamic model of nonholonomic constrained robots, a nonholonomic constrained robot formation model is constructed using a leader-follower structure method. The nonholonomic constrained robot formation model includes at least one nonholonomic leader robot and multiple nonholonomic follower robots. In this application, any nonholonomic robot can serve as both a leader and a follower, so the same symbols are used to represent them. The leader and follower here are only related to the communication topology.

[0051] An error function is constructed based on a nonholonomic constraint robot formation model and an obstacle function; the obstacle function is constructed based on a relative distance constraint and a relative angle constraint between any two nonholonomic constraint robots.

[0052] As an optional implementation, the dynamic model of the nonholonomic constrained robot is:

[0053]

[0054] υ i =[v i ,ω i ] T .

[0055] in, represents the acceleration of the i-th nonholonomic constrained robot; M i represents the inertia matrix of the i-th nonholonomic constrained robot; J i represents the total Coriolis and centripetal acceleration matrices of the i-th nonholonomic constrained robot; υ i represents the generalized velocity of the i-th nonholonomic constrained robot; D i represents the damping matrix of the i-th nonholonomic constrained robot; represents the control input of the i-th nonholonomic constrained robot; v i represents the linear velocity of the i-th nonholonomic constrained robot; ω i represents the angular velocity of the i-th nonholonomic constrained robot.

[0056] As an optional implementation, the nonholonomic constrained robot formation model is:

[0057]

[0058] Δx ij =|x j -x i |.

[0059] Δy ij =|y j -y i |.

[0060] Among them, l ij and θ ij denote the relative distance and angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; Δx ij represents the absolute value of the difference between the horizontal coordinates of the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot in the earth coordinate system; Δy ijrepresents the absolute value of the difference between the ordinates of the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot in the earth coordinate system; atan2 represents the inverse tangent function; and Represents Δx ij and Δy ij The horizontal and vertical coordinates in the local coordinate system with the heading angle of the i-th nonholonomic constrained robot relative to the j-th nonholonomic constrained robot as the reference direction; x i and y i They represent the horizontal and vertical coordinates of the i-th nonholonomic constrained robot in the earth coordinate system; x j and y j denote the horizontal and vertical coordinates of the jth nonholonomic constrained robot in the earth coordinate system; ψ ij represents the heading angle of the i-th nonholonomic constrained robot relative to the j-th nonholonomic constrained robot.

[0061] As an optional implementation, the barrier function is:

[0062]

[0063]

[0064] Among them, ij,m represents the obstacle function of type m between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; m = {l, θ}; l and θ represent the relative distance and relative angle, respectively; represents the normalized error of m constraint types between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; represents the upper bound of the conversion error of the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot with m constraints; represents the lower bound of the conversion error of the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot with respect to the m-constraint type; z ij,m represents the actual error of the m-constraint type between the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; Q ij,m (t) represents the preset performance function of the m-constraint type of the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot at time t.

[0065] In another exemplary embodiment of the present application, step 204 specifically includes:

[0066] The Hamilton Jacobi-Bellman equation is obtained by the following formula:

[0067]

[0068] in, The Hamilton-Jacobi-Bellman equation for the velocity order 2f of the i-th nonholonomic constrained robot; f = {1, 2}; s i,2f represents the error of the 2f velocity order of the i-th nonholonomic constrained robot; represents the virtual optimal control strategy of the i-th nonholonomic constrained robot with velocity type q; q = {v, ω}; v and ω represent the linear velocity and angular velocity respectively; represents the optimal cost function of the 2f velocity order of the i-th nonholonomic constrained robot; represents the error change rate of the 2f velocity order of the i-th nonholonomic constrained robot; Ω represents the control input set of the nonholonomic constrained robot; Represents the second virtual control input of velocity type q for the i-th nonholonomic constrained robot.

[0069] Let the derivative be equal to 0, and take the derivative of the Hamilton-Jacobi-Bellman equation to obtain the virtual optimal control strategy:

[0070]

[0071] Among them, k i,2f represents the adjustment coefficient of the control input of the i-th nonholonomic constraint robot; H ij,m represents the transformation error function of the m-constraint type between the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; Q ij,m (t) represents the preset performance function of the m-constraint type between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot at time t; m = {l, θ}; l and θ represent the relative distance and relative angle, respectively; θ ij represents the relative angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; W iq represents the actual weight vector of the velocity type q of the i-th nonholonomic constrained robot; W iq The transpose of The radial basis function representing the velocity type q of the i-th nonholonomic constrained robot. represents the filtering error of the velocity type q of the i-th nonholonomic constraint robot; represents the filter coefficient of the velocity type q of the i-th nonholonomic constrained robot; represents the parameter estimate of the filter compensator of the velocity type q of the i-th nonholonomic constrained robot; tanh() represents the hyperbolic tangent function; Δ iq represents the control input coefficient of the velocity type q of the ith nonholonomic constrained robot; Δ represents the parameter estimate of the actuator fault of the i-th nonholonomic constrained robot with velocity type q; i,2frepresents the control input parameter for the 2f-th velocity order of the i-th nonholonomic constrained robot; represents an unknown continuous function.

[0072] In another exemplary embodiment of the present application, a control method for a reinforcement learning formation of non-holonomic constrained robots is provided, comprising the following steps:

[0073] Step S1: Build a dynamic model for the multi-nonholonomically constrained robot. This model accurately describes the robot's motion characteristics in complex environments, encompassing its dynamic behavior and kinematic constraints. This precise dynamic model lays a solid foundation for subsequent control strategy design.

[0074] Consider N nonholonomically constrained robots. The kinematic model of the i-th nonholonomically constrained robot is shown below. In this application, any robot can serve as a leader and a follower, so the same symbols are used to represent them.

[0075]

[0076] Where i represents the i-th nonholonomic constrained robot and i = 1, 2, ..., N; x i and y i denote the horizontal and vertical coordinates of the i-th nonholonomic constrained robot in the earth coordinate system; ψ i represents the heading angle of the i-th nonholonomic constrained robot, v i and ω i Represent the linear velocity and angular velocity of the i-th nonholonomic constrained robot, and They represent the rate of change of the position of the i-th nonholonomic constrained robot in the x-axis direction, the rate of change of the position of the i-th nonholonomic constrained robot in the y-axis direction, and the rate of change of the heading angle.

[0077] Order i =[v i ,ω i ] T , then the dynamic model of the i-th nonholonomic constrained robot is:

[0078]

[0079] in, represents the acceleration of the i-th nonholonomic constrained robot; M i represents the inertia matrix of the i-th nonholonomic constrained robot; J i represents the total Coriolis and centripetal acceleration matrices of the i-th nonholonomic constrained robot; υ i =[v i ,ω i ]T ,υ i represents the generalized velocity of the i-th nonholonomic constrained robot; D i represents the damping matrix of the i-th nonholonomic constrained robot; represents the control input of the i-th nonholonomic constrained robot, and They represent the linear velocity control input component and angular velocity control input component of the i-th nonholonomic constrained robot respectively; v i represents the linear velocity of the i-th nonholonomic constrained robot; ω i represents the angular velocity of the i-th nonholonomic constrained robot.

[0080] In order to facilitate the design and analysis of subsequent control strategies, the dynamic equations of the nonholonomic constrained robot are transformed. Through a series of mathematical derivations and variable substitutions, the dynamic equations of the robot can be transformed into the following form, which is more concise and easier to apply control theory:

[0081]

[0082] in, and They represent the linear velocity change rate and angular velocity change rate of the i-th nonholonomic constrained robot respectively. iv and f iω They represent the unknown nonlinear terms of linear velocity and angular velocity of the i-th nonholonomic constrained robot. iv and m iω They represent the system parameters related to the linear velocity characteristics and the angular velocity characteristics of the i-th nonholonomic constrained robot respectively; m iv =r i / 2(m 1i +m 2i ), m iω =r i / 2c i (m 1i -m 2i ), r i represents the wheel radius of the i-th nonholonomic constrained robot; c i represents the centripetal force and Coriolis force coefficients of the i-th nonholonomic constrained robot; m 1i and m 2i denote the first constant and the second constant of the i-th nonholonomic constrained robot, respectively.

[0083] Step S2: Use a neural network to approximate the unknown nonlinear terms in the dynamic model. With their superior nonlinear approximation capabilities, neural networks can effectively address the uncertainty and complexity of the model. The introduction of neural networks significantly improves the system's adaptability and robustness, enabling it to maintain excellent control performance even in complex environments.

[0084] As an optional implementation, a neural network is used to compensate and approximate the unknown nonlinear terms, specifically:

[0085]

[0086] Among them, W iv and W iω They represent the linear velocity weight and angular velocity weight of the i-th nonholonomic constraint robot respectively; f iv and f iω They represent the unknown nonlinear terms of linear velocity and angular velocity of the i-th nonholonomic constrained robot respectively; and They represent the radial basis functions of the linear velocity and angular velocity of the i-th nonholonomic constrained robot respectively; ε iv (X) and ε iω (X) represents the linear velocity approximation error and angular velocity approximation error of the neural network approximation function of the i-th nonholonomic constrained robot S, respectively; X represents the input of the radial basis function.

[0087] Step S3: Based on the leader-follower architecture, a formation model is established that maintains connectivity while avoiding collisions. This formation model carefully designs the relative positions and motion relationships between the leader and followers, ensuring that the robots maintain their formation while avoiding collisions. This formation structure not only enhances formation stability but also improves the overall performance of the system.

[0088] As an optional implementation, a formation model is established to maintain connectivity and avoid collisions. The relative distances and angles are as follows:

[0089]

[0090] Δx ij =|x j -x i |.

[0091] Δy ij =|y j -y i |.

[0092] Among them, l ij and θ ijdenote the relative distance and angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; Δx ij represents the absolute value of the difference between the horizontal coordinates of the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot in the earth coordinate system; Δy ij represents the absolute value of the difference between the ordinates of the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot in the earth coordinate system; atan2 represents the inverse tangent function; and Represents Δx ij and Δy ij The horizontal and vertical coordinates in the local coordinate system with the heading angle of the i-th nonholonomic constrained robot relative to the j-th nonholonomic constrained robot as the reference direction; x i and y i They represent the horizontal and vertical coordinates of the i-th nonholonomic constrained robot in the earth coordinate system; x j and y j denote the horizontal and vertical coordinates of the jth nonholonomic constrained robot in the earth coordinate system; ψ ij represents the heading angle of the i-th nonholonomic constrained robot relative to the j-th nonholonomic constrained robot.

[0093] Wherein, in implementing this embodiment, the constraints of relative distance and relative angle are designed as follows:

[0094] 0<l ij,min <l ij <l ij,max .

[0095] -π / 2<-θ ij,max <θ ij <θ ij,max <π / 2.

[0096] Among them, l ij,min and l ij,max They represent the minimum distance to avoid collision and the maximum distance to maintain communication between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot, respectively; θ ij,max It represents the maximum range of the search field of view between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot, that is, the maximum azimuth angle.

[0097] In this embodiment, the distance formation error and the angle formation error are calculated as follows:

[0098] z ij,l =l ij -l ij,des .

[0099] z ij,θ =θij -θ ij,des .

[0100] Among them, l ij,des and θ ij,des represent the target relative distance and target relative angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot, respectively.

[0101] As an optional implementation, based on a control method of a preset performance function, the above-mentioned distance formation error and angle formation error are transformed into an asymmetric constraint with preset performance.

[0102] -Y ij,m Q ij,m (t)<z ij,m <Q ij,m (t),m={l,θ}.

[0103]

[0104] Among them, the performance function Q ij,m (t) is expressed as:

[0105]

[0106] Q ij,m0 =Q ij,l0 ,Q ij,θ0 .

[0107] Q ij,l0 =l ij,max -l ij,des .

[0108] Q ij,θ0 =θ ij,max -θ ij,des .

[0109] Among them, Y ij,m Y represents the function related to the m constraint type of the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; ij,l Y represents the function related to the relative distance between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; ij,θ The function representing the relative angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; z ij,m represents the actual error of the m-constraint type between the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; Q ij,m (t) represents the m-constraint type preset performance function of the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot at time t; Q ij,m0is the initial value of the preset performance function of the m-constraint type of the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; Q ij,m∞ is the steady-state value to be designed of the preset performance function of the m-constraint type of the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; T m,i represents the preset convergence time of constraint type m of the i-th nonholonomic constrained robot, which is a human-designed convergence time parameter; σ m,j The coefficient representing the convergence rate of the preset performance function of the m constraint type of the i-th nonholonomic constrained robot; l ij and θ ij are the relative distance and angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; Q ij,l0 and Q ij,θ0 They represent the initial values ​​of the preset performance function of the relative distance constraint type and the initial values ​​of the preset performance function of the relative angle constraint type between the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot respectively; l ij,max and θ ij,max are the maximum communication distance and maximum azimuth angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; l ij,des and θ ij,des denote the target relative distance and target relative angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot, respectively; t is the time constant.

[0110] In the control method based on the preset performance function, in order to more effectively analyze and process the error, the error is normalized. Here we introduce the concept of normalized error. It represents the normalized error of the m constraint type between the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot, and its calculation formula is:

[0111] Through this normalization process, errors of different magnitudes and ranges can be unified into a relatively standardized range, which facilitates subsequent control design and stability analysis.

[0112] To ensure that certain constraints are met during the control process and to prevent the system from violating the constraints, a barrier function is constructed. When the system approaches the constraint boundary, a larger value is generated to prevent the system from moving any closer, thus ensuring that the system always operates within the safe region. Based on the normalized error, the following barrier function is constructed for the i-th nonholonomically constrained robot and the j-th nonholonomically constrained robot:

[0113] Among them, ij,mrepresents the obstacle function of type m constraint between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; represents the upper bound of the conversion error of the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot with m constraints; It represents the lower bound of the conversion error of the degree of m constraints between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot.

[0114] By rationally designing the form and parameters of the obstacle function, the system constraints can be effectively integrated into the control strategy to achieve safe control of non-holonomically constrained robot formations.

[0115] Step S4: Design a new nonlinear filter and construct a barrier function. Simultaneously, use this filter to calculate the derivative of the virtual control law. By introducing this new nonlinear function, the accuracy and response speed of the virtual control law are effectively improved, thereby optimizing both the transient and steady-state performance of the system. This design not only ensures system performance but also effectively mitigates collision risks and further enhances the stability of the formation.

[0116] In order to achieve accurate formation tracking and stable motion control in nonholonomic robot formation control, it is necessary to construct a suitable error function to measure the deviation between the actual state and the desired state of the robot, and to design a nonlinear filter and adaptive compensation mechanism to deal with the uncertainty and interference in the system. The error function constructed is:

[0117] s i,11 =ζ ij,l ,s i,12 =ζ ij,θ s.

[0118]

[0119] Among them, s i,11 and s i,12 Respectively represent the relative distance error and relative angle error after conversion; ζ ij,l and ζ ij,θ are the obstacle functions of the relative distance constraint type and the relative angle constraint type between the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; s represents the error; s i,21 and s i,22 They represent the linear velocity order error and angular velocity order error of the i-th nonholonomic constrained robot respectively; and They represent the linear velocity output value and angular velocity output value of the first-order filter of the i-th nonholonomic constrained robot respectively; and They represent the linear velocity filtering error and angular velocity filtering error of the i-th nonholonomic constrained robot respectively; αiv and α iω They represent the first virtual control input of linear velocity and angular velocity of the i-th nonholonomic constrained robot respectively;

[0120] Constructing nonlinear filters and adaptive compensation mechanisms:

[0121]

[0122] in, Represents the filter coefficient of the velocity type q of the i-th nonholonomic constrained robot; q = {v, ω}; represents the rate of change of the filter output of the velocity type q of the i-th nonholonomic constrained robot; f = {1, 2}; represents the filtering error of the velocity type q of the i-th nonholonomic constraint robot; represents the parameter estimate of the velocity type q of the i-th nonholonomic constrained robot; Δ iq Parameters representing the control input of velocity type q for the i-th nonholonomic constrained robot; Represents the output value of the filter of velocity type q for the i-th nonholonomic constrained robot.

[0123] Step S5: Build an adaptive filtering compensation mechanism to compensate for the errors introduced by the filtering process. By adaptively adjusting the compensation mechanism, the system's tracking accuracy can be significantly improved, ensuring a more precise trajectory for the robot in complex environments. Based on this, a virtual controller is designed to achieve preliminary motion control.

[0124] By taking the derivative of the error function in step S4, we can get the following expression:

[0125]

[0126] in, and They represent the derivative of the transformed relative distance error and the derivative of the relative angle error respectively; and They respectively represent the rate of change of the preset performance function of relative distance and relative angle of the i-th non-holonomic constrained robot.

[0127] According to the Lyapunov function and the above expression after derivation of the error function, the virtual controller and adaptive law are designed.

[0128] The virtual controllers are:

[0129]

[0130] Among them, α i,v and α i,ωThey represent the linear velocity virtual control input and angular velocity virtual control input of the i-th nonholonomic constrained robot respectively; l ij and θ ij are the relative distance and angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; k i,11 and k i,12 They represent the relative distance control coefficient and relative angle control coefficient to be designed respectively; and Respectively represent the linear velocity related function and angular velocity related function of the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; tanh() represents the hyperbolic tangent function; represents the estimated velocity of the j-th nonholonomic constrained robot; Q ij,l and Q ij,θ denote the preset performance function of the relative distance constraint and the preset performance function of the relative angle constraint between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot, respectively; δ denotes the first parameter to be designed; and They represent the change rates of the preset performance function of the relative distance constraint and the preset performance function of the relative angle constraint respectively; ij,l and z ij,ω They represent the actual relative distance error and the actual relative angle error between the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot respectively; H ij,l and H ij,θ represent the conversion error function of the relative distance and the conversion error function of the relative angle of the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot, respectively.

[0131] The adaptive law is:

[0132] in, represents the rate of change of the velocity estimate of the j-th nonholonomic constrained robot; Γ i1 Separately i1 represents the second and third parameters to be designed; H ij,m ={H ij,l ,H ij,θ} represents the transformation error function of the m-constraint type between the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; represents the estimated velocity of the j-th nonholonomic constrained robot.

[0133] Step S6: Use a reinforcement learning algorithm to solve the time-limited fault-tolerant formation controller. This algorithm dynamically adjusts the control strategy based on environmental feedback, achieving optimal formation control. With reinforcement learning, the system can maintain good control performance even in the face of abnormal conditions such as actuator failures, achieving optimal formation control.

[0134] First, according to the error function in step S4 and the derivative of the dynamic model of the nonholonomic constrained robot in step S2, we can obtain:

[0135]

[0136] in, and They represent the rate of change of linear velocity order error and angular velocity order error of the i-th nonholonomic constraint robot respectively; m iv and m iω They represent the system parameters related to the linear velocity characteristics and the angular velocity characteristics of the i-th nonholonomic constrained robot respectively; and They represent the failure rate of the linear velocity actuator and the angular velocity actuator of the i-th nonholonomic constraint robot respectively; τ ivd and τ iωd They represent the unknown function parts in the linear velocity actuator fault and the unknown function parts in the angular velocity actuator fault of the i-th nonholonomic constraint robot, respectively.

[0137] Then construct the cost function:

[0138] Among them, J i,2f represents the cost function of the 2f velocity order of the i-th nonholonomic constrained robot; represents the rate of change of the 2f-order velocity error of the i-th nonholonomic constrained robot; represents the virtual optimal control input of the i-th nonholonomic constrained robot with velocity type q.

[0139] Select the intermediate controller as the optimal control The optimal control cost function can be obtained:

[0140]

[0141] The following HJB equation (Hamilton-Jacobi-Bellman equation) is derived:

[0142]

[0143] in, The Hamilton-Jacobi-Bellman equation for the velocity order 2f of the i-th nonholonomic constrained robot; f = {1, 2}; s i,2f represents the error of the 2f velocity order of the i-th nonholonomic constrained robot; represents the virtual optimal control strategy of the i-th nonholonomic constrained robot with velocity type q; q = {v, ω}; v and ω represent the linear velocity and angular velocity respectively; represents the optimal cost function of the 2f velocity order of the i-th nonholonomic constrained robot; represents the error change rate of the 2f velocity order of the i-th nonholonomic constrained robot; Ω represents the control input set of the nonholonomic constrained robot; Represents the second virtual control input of velocity type q for the i-th nonholonomic constrained robot.

[0144] Let the derivative be equal to 0, and take the derivative of the Hamilton-Jacobi-Bellman equation to obtain the virtual optimal control strategy:

[0145]

[0146] right Decomposition can be obtained:

[0147]

[0148] Among them, k i,2f represents the adjustment coefficient of the control input of the i-th nonholonomic constraint robot; H ij,m represents the transformation error function of the m-constraint type between the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; Q ij,m (t) represents the preset performance function of the m-constraint type between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot at time t; m = {l, θ}; l and θ represent the relative distance and relative angle, respectively; θ ij represents the relative angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; W iq represents the actual weight vector of the velocity type q of the i-th nonholonomic constrained robot; W iq The transpose of The radial basis function representing the velocity type q of the i-th nonholonomic constrained robot. represents the filtering error of the velocity type q of the i-th nonholonomic constraint robot; represents the filter coefficient of the velocity type q of the i-th nonholonomic constrained robot; represents the parameter estimate of the filter compensator of the velocity type q of the i-th nonholonomic constrained robot; tanh() represents the hyperbolic tangent function; Δ iqrepresents the control input coefficient of the velocity type q of the ith nonholonomic constrained robot; Δ represents the parameter estimate of the actuator fault of the i-th nonholonomic constrained robot with velocity type q; i,2f represents the adjustment control input parameter of the 2f velocity order of the i-th nonholonomic constrained robot; represents an unknown continuous function.

[0149]

[0150] Therefore, the optimal controller is obtained:

[0151]

[0152] Because J2 0 It is an unknown continuous function, so a neural network is needed to approximate it.

[0153] Finally, the optimal controller is obtained:

[0154]

[0155] in, Network for speakers.

[0156] Among them, implementing this implementation method, the update rate of the speaker network is designed to be:

[0157]

[0158] in, represents the update rate of the orator network of the i-th nonholonomic constrained robot; φ i represents the radial basis function of the i-th nonholonomic constrained robot critic network; Γ iq,5 represents the design constant; represents the estimated value of the orator network weight of the i-th nonholonomic constrained robot with velocity type q; represents the estimated value of the critic network weight of the i-th nonholonomic constrained robot with velocity type q; Γ iq,4 Represents the second design constant of the velocity type q of the i-th nonholonomic constrained robot.

[0159] The reviewer's network update rate is:

[0160]

[0161] in, The network update rate of the critic for the i-th nonholonomically constrained robot.

[0162] Finally, the actual fault-tolerant control and adaptive law are obtained as follows:

[0163]

[0164] Among them, τ iv and τ iω They represent the actual control force of linear velocity and angular velocity of the i-th nonholonomic constraint robot respectively; sign() represents the sign function; Γ iq,1 , Γ iq,2 and Γ iq,3 The first, second and third adjustment parameters respectively represent the actual control force; iq,1 、ι iq,2 and ι iq,3 Respectively represent the first adaptive update rate adjustment parameter, the second adaptive update rate adjustment parameter and the third adaptive update rate adjustment parameter; represents the rate of change of the weighted estimate of the velocity type q of the i-th nonholonomic constrained robot; represents the rate of change of the parameter estimate for the actuator fault of velocity type q of the i-th nonholonomic constrained robot; represents the update rate of the function related to the optimal control input of velocity type q for the i-th nonholonomic constrained robot.

[0165] So far, the design of the time-limited reinforcement learning fault-tolerant formation controller has been completed.

[0166] In another exemplary embodiment of the present application, first, the fault module of the actuator is set to:

[0167]

[0168] Five nonholonomic constrained robots are used, where robot 0 represents the leader and its trajectory is as follows:

[0169] x0=60sin(0.02t), y0=-60cos(0.02t)+60.

[0170] The initial position of the robot is designed to be:

[0171] η0=[0,0,0] T .

[0172] η1=[-4.1,2,-π / 4] T .

[0173] η2=[-2,-3.1,-π / 3] T .

[0174] η3=[-8.2,1,π / 9] T .

[0175] η4=[-6,-5,π / 6]T .

[0176] The initial speed is designed to be 0.

[0177] Adaptive law:

[0178] The safety distance and collision avoidance distance are designed as follows: ij,max =5m,l ij,min =3m,θ ij,max =π / 2.

[0179] The desired distance and desired angle are set as:

[0180] l 10,des =l 20,des =l 13,des =l 24,des =4,θ 10,des =θ 13,des =-π / 6,θ 20,des =θ 24,des =π / 6.

[0181] The control parameters are selected as follows:

[0182] k i,11 =k i,12 =k i,21 =k i,22 =0.2Γ i1 =0.3,ι i1 =0.4.

[0183]

[0184] Γ iv,2 =0.2,Γ iω,2 =20,Γ iv,3 =0.1,Γ iω,3 =10,ι 1v,2 =ι 3v,2 =4,ι 2v,2 =ι 4v,2 =3,ι iω,3 =0.4.

[0185] Using MATLAB software, the control method of this application is used to control Figure 2 The leader-follower nonholonomic constrained robot formation architecture shown in the figure is modeled and simulated. Figure 2 Medium x b, y b Indicates the robot's own coordinate system, X e ,Y e Represents the geodetic coordinate system. Figure 2 (a) Spatial layout diagram of a leader-follower nonholonomic constrained robot formation. Figure 2 (b) is the communication and collision avoidance area diagram of the nonholonomic constrained robot formation. The communication topology diagram of the nonholonomic constrained robot formation is shown in Figure 2. Figure 3 The simulation results are shown as follows. Figure 4-Figure 9 shown. Figure 4 A snapshot of five nonholonomically constrained robots forming a triangle formation performing circular motion is given. Figure 4 As can be clearly seen in the figure, the five nonholonomically constrained robots successfully formed a triangular formation and smoothly completed circular motion. This demonstrates that the proposed control method can effectively coordinate the motion of multiple nonholonomically constrained robots, enabling them to move according to a predetermined formation, validating the feasibility and effectiveness of the control method in achieving robot formation motion tasks. The robots can maintain a relatively stable spatial relationship, demonstrating that the control strategy can precisely adjust the motion trajectory of each robot to meet the requirements of the formation. Figure 5 The distance evolution trajectory of the formation error under the constraints of the specified performance function is shown, namely the distance tracking error between robots 1 and 0, 2 and 0, 1 and 3, and 2 and 4. It can be seen that the distance tracking error is always confined within the specified performance function envelope and gradually stabilizes over time. This fully demonstrates the excellent performance of the control method in distance tracking, ensuring that the relative distance between robots meets the formation requirements. Even during motion, it maintains good distance control accuracy and avoids formation failure caused by excessive or insufficient distance between robots. Figure 6 and Figure 7 The evolution diagram of the angular formation error within the envelope under the action of the specified performance function is given, where Figure 6 This is the angle tracking error diagram between numbers 1 and 0, and between numbers 1 and 3; Figure 7 This is the angle tracking error diagram of number 2 and 0 and number 2 and 4. Figure 6 and Figure 7 As can be seen, the angular tracking error is also effectively confined within the envelope and converges to a smaller range. This demonstrates that the control method can not only accurately control the distance between robots, but also accurately adjust the relative angle between them, ensuring the consistency and stability of the formation's posture, allowing the robot formation to move in the expected geometric shape. Figure 8 and Figure 9 The control force diagram under fault conditions is given, where Figure 8 It is the control force curve of the non-holonomic constrained robot formation. Figure 9 This is the control torque curve of the nonholonomic constrained robot formation. Figure 8 and Figure 9As can be seen from the figure, even when the system is disturbed by a fault, the control force and torque can be quickly adjusted to maintain the stability of the formation. This shows that the proposed optimal control scheme has a certain degree of fault tolerance. In the face of sudden faults, it can adaptively adjust the control input to ensure the stability and reliability of the robot formation system, and avoid the loss of control or disbanding of the formation due to faults.

[0186] Based on the above simulation results, all signals in the system are uniformly and ultimately bounded. This means that under the proposed control method, the state of the robot formation system (including position, velocity, angle, etc.) does not diverge, but instead reaches a stable state within a finite time and fluctuates within a reasonable range. This further demonstrates the effectiveness of the proposed optimal control scheme, which can provide stable and reliable control for nonholonomically constrained robot formation systems.

[0187] In summary, MATLAB modeling and simulation fully validate the effectiveness of the proposed optimal control scheme for nonholonomically constrained robot formation control, including aspects such as formation motion implementation, error constraint and tracking, fault tolerance, and system stability. This control scheme provides a feasible solution for nonholonomically constrained robot formation control in practical applications.

[0188] The present application also provides an application scenario, which applies the above-mentioned control method of the reinforcement learning formation of non-holonomic constrained robots. Specifically: the control method of the reinforcement learning formation of non-holonomic constrained robots provided in this embodiment can be applied to the scenario of cargo handling and sorting in large warehouses. This scenario includes a task planning link, a robot formation formation link, a path tracking and obstacle avoidance link, and a task completion evaluation link. The task planning link formulates a detailed handling task plan, and the robots enter the robot formation formation link from the standby area, exchange information with each other through the communication module, and form a team according to the preset formation strategy. After determining their respective positions and roles in the formation, they enter the path tracking and obstacle avoidance link to plan the optimal driving path. When the robot formation completes the cargo handling task, it enters the task completion evaluation link to evaluate the quality of the task completion and feed back the evaluation results to the warehouse management system for subsequent task adjustment and optimization. The control method of the reinforcement learning formation of non-holonomic constrained robots provided in this embodiment belongs to the path tracking and obstacle avoidance link in the scenario of cargo handling and sorting in large warehouses. By precisely controlling the robot's linear and angular velocity and processing the relative distance and angular relationship between robots, the robot formation can travel stably and efficiently in complex warehouse environments, effectively improving the efficiency and accuracy of cargo handling and reducing the need for human intervention.

[0189] This application aims to provide a fault-tolerant control method for nonholonomically constrained robot formations using time-limited reinforcement learning, addressing existing issues such as insufficient control performance, poor environmental adaptability, and poor formation stability. First, by establishing a dynamic model for the multiple nonholonomically constrained robots to be controlled, the motion characteristics of each robot in a complex environment, including its dynamic behavior and kinematic constraints, can be accurately described. This precise dynamic model provides a foundation for subsequent control strategy design. A neural network is used to approximate the unknown nonlinear terms in the dynamic model. Neural networks have powerful nonlinear approximation capabilities and can effectively handle the uncertainty and complexity in the model. The introduction of neural networks improves the adaptability and robustness of the system, enabling it to maintain good control performance in complex environments. Based on a leader-follower architecture, a formation model is established that maintains connectivity and avoids collisions. By rationally designing the relative positions and motion relationships between the leader and followers, this formation model ensures that the robots maintain their formation and avoid collisions during formation. This formation structure not only improves formation stability but also enhances the overall system performance.

[0190] Secondly, an obstacle function is constructed and a new nonlinear filter is designed to calculate the derivatives of the virtual control. The introduction of this new nonlinear function effectively improves the accuracy and response speed of the virtual control law, thereby enhancing the system's transient and steady-state performance. This design not only ensures system performance, but also effectively avoids collisions and enhances formation stability. Adaptive filtering and compensation mechanisms are constructed to compensate for errors introduced by the filtering process. By adaptively adjusting the compensation mechanism, the system's tracking accuracy is effectively improved, ensuring a more precise robot trajectory in complex environments. The introduction of this compensation mechanism further enhances the system's robustness and adaptability.

[0191] Then, a reinforcement learning algorithm is used to solve the time-limited fault-tolerant formation controller. This algorithm dynamically adjusts the control strategy based on environmental feedback, achieving optimal formation control. Through reinforcement learning, the system can maintain good control performance even in the face of abnormal conditions such as actuator failures, achieving optimal formation control.

[0192] In summary, through the above technical solutions, the present application can realize the efficient and stable operation of non-holonomic constrained robot formations in complex environments, and significantly improve the fault tolerance and adaptability of the system.

[0193] Compared with the existing technology, the beneficial effects of this application are:

[0194] The technical solution of this application designs a performance function with a specified time. Compared with a standard boundary function in the form of an exponential function, this function can artificially pre-define the convergence time and reduce chattering, so that the formation error converges to a neighborhood of the origin within a predetermined time. The constructed nonlinear filter integrates an adaptive compensation mechanism to achieve adaptive compensation of the filtering error, reducing tracking error and improving tracking accuracy. Within the framework of backstepping optimization, a new formation control strategy is obtained that minimizes energy consumption while ensuring that the formation error evolves within a given region. Because the X-axis and Y-axis in the coordinate systems of two consecutive nonholonomic constrained robots are coupled, the Hamilton-Jacobi-Bellman equations are jointly designed, and two optimal virtual controllers are then obtained simultaneously. The reinforcement learning algorithm utilizes the negative gradient of a simple positive definite function of the partial derivatives of the Hamilton-Jacobi-Bellman equations to greatly simplify the optimal control and achieve optimal formation control.

[0195] Based on the same inventive concept, an embodiment of the present application also provides a control system for a reinforcement learning formation of nonholonomically constrained robots for implementing the control method for the reinforcement learning formation of nonholonomically constrained robots involved above. The implementation solution provided by this system is similar to the implementation solution described in the above method. Therefore, the specific limitations of the control system embodiment of one or more reinforcement learning formations of nonholonomically constrained robots provided below can be found in the limitations of the control method for the reinforcement learning formation of nonholonomically constrained robots above, and will not be repeated here.

[0196] In an exemplary embodiment, Figure 10 As shown, a control system for a reinforcement learning formation of non-holonomic constrained robots is provided, including:

[0197] The error function construction module 301 is used to construct an error function of the nonholonomic constrained robot formation; the error function is used to calculate the error between the actual state and the expected state of the nonholonomic constrained robot formation.

[0198] The fault-tolerant control model construction module 302 is used to construct a fault-tolerant control model; the fault-tolerant control model includes a virtual controller model, a nonlinear filter model and an adaptive compensation mechanism model; the virtual controller model is used to generate a virtual control law for each non-complete constrained robot; the nonlinear filter model is used to calculate the derivative of the virtual control law for each non-complete constrained robot; the adaptive compensation mechanism model is used to calculate the compensation error of each non-complete constrained robot based on the derivative of the virtual control law.

[0199] The cost function construction module 303 is used to construct a cost function of the non-holonomic constrained robot formation based on the error function and the adaptive compensation mechanism model.

[0200] The virtual optimal control strategy solving module 304 is used to obtain the Hamilton-Jacobi-Bellman equation based on the cost function, and solve the Hamilton-Jacobi-Bellman equation to obtain the virtual optimal control strategy.

[0201] The optimal control strategy solving module 305 is used to approximate and solve the unknown continuous function in the virtual optimal control strategy based on the fault-tolerant control model using the actor-critic algorithm to obtain the optimal control strategy.

[0202] The control execution module 306 is used to control the non-holonomic constrained robot formation according to the optimal control strategy.

[0203] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 11 As shown. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. Those skilled in the art will understand that Figure 11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0204] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0205] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0206] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A control method for reinforcement learning formation of nonholonomic constrained robots, characterized in that: The control method of the reinforcement learning formation of the non-holonomic constrained robots includes: Constructing an error function for a nonholonomic constrained robot formation; the error function is used to calculate the error between the actual state and the desired state of the nonholonomic constrained robot formation; Constructing a fault-tolerant control model; the fault-tolerant control model includes a virtual controller model, a nonlinear filter model, and an adaptive compensation mechanism model; the virtual controller model is used to generate a virtual control law for each nonholonomic constrained robot; the nonlinear filter model is used to calculate the derivative of the virtual control law for each nonholonomic constrained robot; the adaptive compensation mechanism model is used to calculate the compensation error of each nonholonomic constrained robot based on the derivative of the virtual control law; Constructing a cost function for a nonholonomic constrained robot formation based on the error function and the adaptive compensation mechanism model; The Hamilton-Jacobi-Bellman equation is obtained based on the cost function, and the Hamilton-Jacobi-Bellman equation is solved to obtain the virtual optimal control strategy; Based on the fault-tolerant control model, the actor-critic algorithm is used to approximate and solve the unknown continuous function in the virtual optimal control strategy to obtain the optimal control strategy. Control the nonholonomic constrained robot formation according to the optimal control strategy.

2. The control method of reinforcement learning formation of nonholonomic constrained robots according to claim 1, characterized in that: Construct the error function of the nonholonomic constrained robot formation, including: Constructing a dynamic model of a nonholonomic constrained robot and calculating unknown nonlinear terms in the dynamic model using a neural network; Based on the dynamic model of nonholonomic constrained robots, a nonholonomic constrained robot formation model is constructed using the leader-follower structure method. An error function is constructed based on a nonholonomic constraint robot formation model and an obstacle function; the obstacle function is constructed based on a relative distance constraint and a relative angle constraint between any two nonholonomic constraint robots.

3. The control method of reinforcement learning formation of non-holonomic constrained robots according to claim 2, characterized in that: The dynamic model of the nonholonomic constrained robot is: u i =[v i ,oh i ] T ; in, represents the acceleration of the i-th nonholonomic constrained robot; M i represents the inertia matrix of the i-th nonholonomic constrained robot; J i represents the total Coriolis and centripetal acceleration matrices of the i-th nonholonomic constrained robot; υ i represents the generalized velocity of the i-th nonholonomic constrained robot; D i represents the damping matrix of the i-th nonholonomic constrained robot; represents the control input of the i-th nonholonomic constrained robot; v i represents the linear velocity of the i-th nonholonomic constrained robot; ω i represents the angular velocity of the i-th nonholonomic constrained robot.

4. The control method of reinforcement learning formation of non-holonomic constrained robots according to claim 2, characterized in that: The nonholonomic constrained robot formation model is: Δx ij =|x j -x i |; Δy ij =|y j -y i |; Among them, e ij and θ ij denote the relative distance and angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; Δx ij represents the absolute value of the difference between the horizontal coordinates of the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot in the earth coordinate system; Δy ij represents the absolute value of the difference between the ordinates of the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot in the earth coordinate system; atan2 represents the inverse tangent function; and Represents Δx ij and Δy ij The horizontal and vertical coordinates in the local coordinate system with the heading angle of the i-th nonholonomic constrained robot relative to the j-th nonholonomic constrained robot as the reference direction; x i and y i They represent the horizontal and vertical coordinates of the i-th nonholonomic constrained robot in the earth coordinate system; x j and y j denote the horizontal and vertical coordinates of the jth nonholonomic constrained robot in the earth coordinate system; ψ ij represents the heading angle of the i-th nonholonomic constrained robot relative to the j-th nonholonomic constrained robot.

5. The control method of reinforcement learning formation of non-holonomic constrained robots according to claim 2, characterized in that: The barrier function is: Among them, ij,m represents the obstacle function of type m between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; m = {l, θ}; l and θ represent the relative distance and relative angle, respectively; represents the normalized error of m constraint types between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; represents the upper bound of the conversion error of the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot with m constraints; represents the lower bound of the conversion error of the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot with respect to the m-constraint type; z ij,m represents the actual error of the m-constraint type between the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; Q ij,m (t) represents the preset performance function of the m-constraint type of the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot at time t.

6. The control method of reinforcement learning formation of nonholonomic constrained robots according to claim 1, characterized in that: Based on the cost function, the Hamilton-Jacobi-Bellman equation is obtained, and the Hamilton-Jacobi-Bellman equation is solved to obtain the virtual optimal control strategy, which specifically includes: The Hamilton Jacobi-Bellman equation is obtained by the following formula: in, The Hamilton-Jacobi-Bellman equation for the velocity order 2f of the i-th nonholonomic constrained robot; f = {1, 2}; s i,2f represents the error of the 2f velocity order of the i-th nonholonomic constrained robot; represents the virtual optimal control strategy of the i-th nonholonomic constrained robot with velocity type q; q = {v, ω}; v and ω represent the linear velocity and angular velocity respectively; represents the optimal cost function of the 2f velocity order of the i-th nonholonomic constrained robot; represents the error change rate of the 2f velocity order of the i-th nonholonomic constrained robot; Ω represents the control input set of the nonholonomic constrained robot; represents the virtual control input of velocity type q for the i-th nonholonomic constrained robot; Let the derivative be equal to 0, and take the derivative of the Hamilton-Jacobi-Bellman equation to obtain the virtual optimal control strategy: Among them, k i,2f represents the adjustment coefficient of the control input of the i-th nonholonomic constraint robot; H ij,m represents the transformation error function of the m-constraint type between the i-th nonholonomic constraint robot and the j-th nonholonomic constraint robot; Q ij,m (t) represents the preset performance function of the m-constraint type of the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot at time t; m = e, θ; e and θ represent the relative distance and relative angle, respectively; θ ij represents the relative angle between the i-th nonholonomic constrained robot and the j-th nonholonomic constrained robot; W iq represents the actual weight vector of the velocity type q of the i-th nonholonomic constrained robot; W iq The transpose of The radial basis function representing the velocity type q of the i-th nonholonomic constrained robot. represents the filtering error of the velocity type q of the i-th nonholonomic constraint robot; represents the filter coefficient of the velocity type q of the i-th nonholonomic constrained robot; represents the parameter estimate of the filter compensator of the velocity type q of the i-th nonholonomic constrained robot; tanh() represents the hyperbolic tangent function; Δ iq represents the control input coefficient of the velocity type q of the ith nonholonomic constrained robot; Δ represents the parameter estimate of the actuator fault of the i-th nonholonomic constrained robot with velocity type q; i,2f represents the adjustment control input parameter of the 2f velocity order of the i-th nonholonomic constrained robot; represents an unknown continuous function.

7. A control system for a reinforcement learning formation of nonholonomically constrained robots, characterized in that: The control system of the reinforcement learning formation of the non-holonomic constrained robots includes: An error function construction module is used to construct an error function of a nonholonomic constrained robot formation; the error function is used to calculate the error between the actual state and the expected state of the nonholonomic constrained robot formation; a fault-tolerant control model construction module, configured to construct a fault-tolerant control model; the fault-tolerant control model comprising a virtual controller model, a nonlinear filter model, and an adaptive compensation mechanism model; the virtual controller model being configured to generate a virtual control law for each nonholonomic constrained robot; the nonlinear filter model being configured to calculate a derivative of the virtual control law for each nonholonomic constrained robot; and the adaptive compensation mechanism model being configured to calculate a compensation error for each nonholonomic constrained robot based on the derivative of the virtual control law; A cost function construction module, configured to construct a cost function for a nonholonomic constrained robot formation based on the error function and the adaptive compensation mechanism model; A virtual optimal control strategy solving module is used to obtain the Hamilton-Jacobi-Bellman equation based on the cost function and solve the Hamilton-Jacobi-Bellman equation to obtain the virtual optimal control strategy; The optimal control strategy solving module is used to approximate and solve the unknown continuous function in the virtual optimal control strategy based on the fault-tolerant control model and adopt the actor-critic algorithm to obtain the optimal control strategy; The control execution module is used to control the non-holonomic constrained robot formation according to the optimal control strategy.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the control method for a reinforcement learning formation of a non-holonomic constrained robot according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the control method of the reinforcement learning formation of the non-holonomic constrained robots according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the control method of the reinforcement learning formation of the non-holonomic constrained robots according to any one of claims 1 to 6 is implemented.