Distributed Active Disturbance Rejection Elastic Control Method for Straddle-Type Monorail Trains Oriented to Virtual Formation
Through multi-agent modeling and adaptive ephemeral optimization methods, a distributed self-immune disturbance controller is designed, which solves the formation recovery and stability problems of air-rail trains in complex operating scenarios, and realizes the fast adaptive control of trains under disturbances.
Patent Information
- Application Number
- CN202411152526.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-08-21
AI Technical Summary
The prior art cannot effectively solve the problem that air-rail trains can quickly recover and maintain the desired formation in complex and time-varying operating scenarios, especially under large-scale uncertain disturbances, and the stability and comfort of train speed control are difficult to ensure.
The multi-agent modeling method is used to construct a distributed dynamic model of an air-rail train, and combined with the adaptive mayfly optimization method to optimize the hyperparameters of the depth deterministic strategy gradient algorithm. The coordination and collision avoidance control protocol based on the two-way-pilot communication topology is designed. The distributed self-immunication controller is trained using the optimized depth deterministic strategy gradient algorithm to realize adaptive control of virtual marshalling trains.
The rapid recovery of the virtual formation of trains and the maintenance of the expected formation under different degrees of external disturbances is achieved, ensuring the stability and comfort of the train, and improving the adaptive ability of train control.
Smart Images

Figure CN119037511B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sky train control, and particularly to a distributed active disturbance rejection elastic control method for sky trains oriented to virtual formation. Background Art
[0002] In order to alleviate problems related to ground transportation, improve the capacity and speed of urban transportation, and effectively reduce the frequency of car use, many countries have focused on the development and application of suspended monorail transportation systems (i.e., sky trains).
[0003] Sky trains are grouped in a virtual coupling form, which is an effective way to improve the flexibility and transportation efficiency of rail transit systems. Due to the complexity and time-variability of the train operation scenario, the resistance experienced by the train cannot be accurately expressed, which will inevitably cause uncertain disturbances to the entire control system. Therefore, the tracking controllers of the unit trains within the virtual formation need to be able to cancel or restrain uncertain disturbances.
[0004] The following documents introduce some research on disturbance suppression in train tracking control.
[0005] In "X. Wang, L. Zhu, H. Wang, et al., 'Robust Distributed Cruise Control of Multiple High-Speed Trains Based on Disturbance Observer', IEEE Trans. Intell. Transp. Syst., vol. 22, no. 1, pp. 267 - 279, 2021.", to solve the problem of robust distributed cruise control of multiple high-speed trains under external disturbances, a disturbance observer was designed to approximate the disturbances. In "Z. Q. Long, Y. Li, W. Xu, 'Research on Automatic Driving Algorithm for Maglev Trains Based on Active Disturbance Rejection Control', Proceedings of the 27th China Control Conference two thousand and eight, 2008.", based on the constraints of maglev train track operation and passenger comfort index, an active disturbance rejection control algorithm was proposed and its feasibility and superiority were verified by simulation. Considering time-varying external disturbances. In "M. G. Suo, X. R. Li. 'The application of disturbance observer in train speed prediction and tracking control', Mechanical Science and Technology, vol. 293, no. 7, pp. 123 - 130, 2019.", an optimal predictive control algorithm combined with a disturbance observer was designed to constrain uncertain disturbances and finally achieved the prediction and tracking of train running speed. In "Y. Liu, Y. Zhou, S. Su, J. Xun, and T. Tang. 'Tube-based control strategy for stable formation of high-speed virtually coupled trains with disturbances and delays', Computer-Aided Civil and Infrastructure Engineering, vol. 38 no. 13 pp. 1093 - 9687, 2023.", a Tube model predictive control strategy combining optimization and feedback control methods was designed to solve the complexity of the virtual train formation system under uncertain disturbances.To achieve the virtual formation control of permanent magnet maglev trains, an effective cooperative tracking control method based on distributed active disturbance rejection control was proposed in "Z.Y. Guo, Z.Q. Li, "Virtual coupling of permanent magnetic maglev trains: An improved cooperative tracking and collision avoidance control protocol", IET Intell. Transp. Syst., pp. 1–14, 2023." Compared with traditional control methods, this method has higher tracking accuracy and stronger anti-interference ability. However, these disturbance rejection methods have certain limitations and are only effective for uncertain disturbances that vary within a small range. If the disturbance is large, the controller parameters cannot be adjusted online, and the stability of the train speed control will deteriorate, resulting in reduced comfort. Therefore, designing an effective and flexible distributed cooperative train controller has become an urgent issue in the transportation field.
[0006] There is little research work on the online parameter adjustment method of elastic controllers in the literature. In "A. Savran, 'A multivariable predictive fuzzy PID control system', 'Applied Soft Computing', vol. 13, no. 5, pp. 2658 - 2667, 2013.", a fuzzy adaptive proportional integral derivative (PID) controller was designed, which adjusts its parameters by querying the fuzzy matrix table. In "C. C. Li, C. F. Zhang, 'Adaptive PID control based on minimum resource allocation network', Application Research of Computers, vol. 32, no. 1, pp. 167 - 169, 2015.", an adaptive PID control method based on neural network was proposed, which uses the good approximation ability of neural network to achieve effective control without the need to identify complex nonlinear controlled objects. In "J. L. Chang, Y. P. Li, X. P. Ma, 'PID optimization design based on improved differential evolution algorithm', Control Engineering, vol. 17, no. 6, pp. 807 - 810, 2010.", an adaptive PID controller based on evolutionary algorithm was designed. Although this method requires some prior knowledge, it cannot achieve real-time control in practical engineering applications. In "J. J. Liu, M. R. Hao, M. W. Sun, X. Guo, Z. Q. Chen, 'PID parameter adjustment method for attitude control of cruise missile based on reinforcement learning', Tactical Missile Technology, vol. 5, no. 1, pp. 58 - 63, 2019.", a PID parameter adjustment method for attitude control of cruise missile based on reinforcement learning was proposed. The evaluation network is used to recursively approximate the control objective function, and the action network is used to obtain the PID control parameters.In “JHZhang, “Research on adaptive PID control strategy based on Actor-Critic learning”, YanShan University, 2018.”, an adaptive PID control strategy based on an actor-critic (AC) structure and a radial basis function neural network is proposed. The controller uses the AC structure to adjust its parameters online and uses a neural network to approximate the value and decision function in the Markov decision process. In “QFSun, H.Ren, YXDuan, “Adaptive PID control design based on asynchronous dominant actuator evaluator learning”, Information and Control, vol.48, no.3, pp.70-74, 2019.”, a new adaptive PID controller based on the asynchronous dominant actuator evaluator (A3C) algorithm is proposed. The controller uses the multi-threaded asynchronous learning capability of the A3C algorithm to train agents based on multiple actuator evaluator structures in parallel. Among the above methods, evolutionary algorithms perform offline parameter adjustment with poor real-time performance, while other methods involve online parameter adjustment but require optimization of multiple hyperparameters.
[0007] Moreover, the distributed cooperative control method proposed in the document "ZY Guo, ZQ Li, "Distributed auto disturbances rejection resilient control of permanent magnetic maglev trains based on the optimized deep deterministic policy gradient algorithm", IET Control Theory and Applications, 2024." only considers the running status information of the adjacent front and rear vehicles, which can accelerate the restoration of a stable formation to a certain extent, but the input of the status information of the front and rear vehicles needs to be weighed. If the controller is not designed properly, it is easy to produce speed and spacing oscillations, and the stability of the queue cannot be guaranteed while taking into account the stability of the single vehicle. In addition, if the formation state deviation of the adjacent rear vehicles is too large, the stability of the queue will deteriorate.
[0008] In summary, the existing technology has the pain points of too many hyperparameters and difficult tuning. When encountering external disturbances of different degrees, it is unable to achieve adaptive adjustment control, and the existing train tracking control methods cannot be effectively applied. These shortcomings lead to the inability of the existing technology to ensure the rapid recovery of the virtual formation of trains under the action of interference and to maintain the operation in the desired formation. Summary of the Invention
[0009] The purpose of the present invention is to provide a distributed active disturbance rejection elastic control method for suspended monorail trains for virtual formation, which can ensure the rapid recovery of the virtual formation of trains under the action of interference and maintain the operation in the desired formation.
[0010] To achieve the above object, the present invention provides the following solutions.
[0011] A distributed active disturbance rejection elastic control method for suspended monorail trains for virtual formation, comprising: constructing a distributed dynamic model of the suspended monorail trains in the virtual formation by using a multi-agent modeling method; determining a formation state error model of the virtual formation according to the expected relative running speed and the expected running interval between adjacent trains and between the following trains and the leader train in the virtual formation; smoothing the formation state error of adjacent trains by using the tanh function according to the distributed dynamic model and the formation state error model, and establishing a cooperation and collision avoidance control protocol based on a two-way-leader communication topology; constructing a distributed active disturbance rejection controller based on the cooperation and collision avoidance control protocol; optimizing the hyperparameters of the deep deterministic policy gradient algorithm by using an adaptive mayfly optimization method; the adaptive mayfly optimization method adds a mayfly adaptive mutation mechanism based on Steffensen; training the distributed active disturbance rejection controller by using the optimized deep deterministic policy gradient algorithm to obtain a trained distributed active disturbance rejection controller; adaptively adjusting the parameters of the trained distributed active disturbance rejection controller online according to time-varying disturbances, and controlling the suspended monorail trains in the virtual formation, so that the suspended monorail trains in the virtual formation maintain the operation in the desired formation under the action of interference.
[0012] Optionally, a distributed dynamics model of the straddle-type monorail train in the virtual formation is constructed using a multi-agent modeling method, specifically including: considering a multi-agent virtual formation system composed of N trains, where each train is regarded as a multi-agent rigid mass point, numbered 0 to N-1 in sequence along the running direction. Let the train numbered 0 be the leader train, and the trains numbered 1 to N-1 be the follower trains, and the follower trains only receive the state information of the trains before and after in their respective neighborhoods; among them, let E represent the set of all trains, E = {0, 1, …, N-1}; F represent the set of follower trains, F = {1, 2, 3, …, N-1}; G represent the set of neighboring trains of the train numbered i, G = {i-1, i+1}; the train is a straddle-type monorail train; the distributed dynamics model of the straddle-type monorail train in the multi-agent virtual formation system is constructed using a multi-agent modeling method as follows: In the formula, x i (t) represents the position of the train numbered i at time t, represents the differential of x i (t), v i (t) represents the speed of the train numbered i at time t, represents the differential of v i (t), f(v i , t) represents the resistance suffered by the train numbered i at time t, b represents the controller coefficient, u i (t) represents the traction / braking force of the train numbered i at time t, represents the external environmental disturbance suffered by the train numbered i at time t, α v v i (t) represents the known system linear term, α v represents the system linear term coefficient, represents the non-linear internal disturbance with uncertain system parameter perturbation, σ i (v i (t), t) represents the total disturbance suffered by the system at time t.
[0013] Optionally, according to the distributed dynamics model and the formation state error model, the hyperbolic tangent function is used to smooth the formation state error between adjacent trains, and a cooperative and collision avoidance control protocol based on a two-way-leader communication topology is established, specifically including: designing a tracking differentiator for smooth acceleration or smooth deceleration of the train as follows: In the formula, d i,0 (t) represents the actual relative state of the train numbered i and the leader train at time t, d i,0 (t) = [x i,0 (t), v i,0 (t)] T , x i,0(t) represents the actual interval between the train numbered i and the leading train at time t, v i,0 (t) represents the actual relative speed between the train numbered i and the leading train at time t; represents d i,0 (t) jumps to the step input signal of represents x i,0 (t) jumps to the step input signal of represents the desired interval between the train numbered i and the leading train at time t, represents v i,0 (t) jumps to the step input signal of represents the desired relative speed between the train numbered i and the leading train at time t; represents the desired relative state between the train numbered i and the leading train at time t, represents the transition signal of represents the running interval transition signal at time t, represents the relative speed transition signal at time t; and both represent the differential signal of represents the differential signal of; r0, r1 both represent the transition speed of the signal, h0, h1 both represent the terms that filter the noise of the signal; fhan() represents the fhan function; in order to observe all unknown terms including internal interference and external interference, a nonlinear extended state observer based on a high-order sliding mode differentiator is designed as: In the formula, ψ1(χ1) represents the estimated value of the train speed χ1, χ1 = v i , v i represents the current speed of the train numbered i; ψ2(χ1) represents the estimated value of the total system interference χ2, χ2 = σ i (v i ), σ i (v i ) represents the total interference suffered by the system; e1 represents the error of the train speed, e2 represents the error of the total system interference; represents the differential of ψ1(χ1), represents the differential of ψ2(χ1); η1(χ1) represents the internal state of the nonlinear extended observer; f(v i ) represents the resistance suffered by the train numbered i; b represents the controller coefficient, u iDenote the traction / braking force of train numbered \(i\); \(a_1, b_1, q_1\) and \(p_1\) are all adjustable observer parameters, \(a_1\gt0, b_1\gt0, 0\lt q_1\lt p_1\); \(sgn()\) represents the sign function; introduce the improved repulsive force field function as: In the formula, represents the improved repulsive force field of the neighboring train numbered \(j\) relative to the train numbered \(i\), \(\rho\) j represents the radius of the repulsive force field of the neighboring train numbered \(j\), \(r\) j represents the actual size radius of the neighboring train numbered \(j\), \(\mu\) represents the repulsive force scale factor, \(\rho\) i,j represents the actual interval between the train numbered \(i\) and the neighboring train numbered \(j\), \(\Delta\rho\) i,j represents \(\rho\) i,j and \(r\) j The deviation of, \(\Delta\rho\) represents \(\rho\) j and \(r\) j The deviation of, represents the expected interval between the train numbered \(i\) and the neighboring train numbered \(j\), represents \(\rho\) i,j and The absolute value of the deviation, \(F\) represents the set of follower trains, \(G\) represents the set of neighboring trains of the train numbered \(i\); \(\rho\) j represents the radius of the repulsive force field of the neighboring train numbered \(j\); According to the improved repulsive force field function, the repulsive force exerted on each train by a neighboring train is: In the formula, represents the repulsive force generated by the neighboring train numbered \(j\) on the train numbered \(i\), \(i\in F, j\in G\), \(n\) j→i represents the unit repulsive force vector of the neighboring train numbered \(j\) pointing in the direction of the train numbered \(i\), \(n\) i→B represents the unit vector that makes the train numbered \(i\) move in the direction of the expected formation \(B\), represents the train numbered \(i\) \(n\) j→i The repulsive force received in the direction of, represents the train numbered \(i\) \(n\) i→B The repulsive force received in the direction of; According to the repulsive force exerted on each train by a neighboring train, determine the total repulsive force of each train from the front and rear neighboring trains as: In the formula, \(F\) i rep represents the total repulsive force of the train numbered \(i\) from the front and rear neighboring trains, \(i\in F\); represents the repulsive force of the train numbered \(i\) from the front neighboring train, represents the repulsive force of the train numbered \(i\) from the rear neighboring train; The repulsive force exerted on each train by a neighboring train and the total repulsive force of each train from the front and rear neighboring trains constitute the collision avoidance control, and the collision avoidance control, the tracking differentiator and the nonlinear extended state observer together constitute the cooperation and collision avoidance control protocol.
[0014] Optionally, the Steffensen-based mayfly adaptive mutation mechanism specifically includes: taking the ratio of the number of weight nodes adjacent to the current optimal solution of mayflies to the total number of weight nodes as the evaluation index of the diversity of the mayfly population; if the diversity evaluation index is less than a preset threshold, the Steffensen method is used, and according to the formula to mutate the mayfly population; where X 1,k represents the k-th dimension of the mayfly X1 to be mutated, and X 1,k ′ is the value after mutation of X 1,k ; θ(X 1,k ) represents the deviation between the distance between X1 and Z1 and the expected distance between X 1,k ′ and Z1, θ(X 1,k ) = L(X1, Z1) - l * , L(X1, Z1) represents the distance between the mayfly individual X1 to be mutated and the weight node Z1, and l * represents the expected distance between X 1,k ′ and Z1, l * = α·L(X1, Z1), α ∈ [0, 1], and α represents a random number in the interval [0, 1].
[0015] Optionally, according to the expected relative running speed and the expected running interval between adjacent trains in the virtual formation and the leading train, the formation state error model of the virtual formation is determined. After that, the following assumptions are included: Assumption 1: Assume that each train obtains the actual speed and position information of the front and rear trains and the leading train in the neighborhood in real time through vehicle-to-vehicle wireless communication; Assumption 2: Assume that the initial time is t0, and the initial state deviation between the trains in the neighborhood satisfies ||ξ i (t0) - ξ j (t0)|| ≤ δ1, j ∈ G, and the external disturbance received by the system satisfies ||σ i (t)|| ≤ δ2; where δ1 and δ2 are both bounded positive constants, ξ i (t0) represents the state vector of train numbered i at the initial time t0, ξ j (t0) represents the state vector of the adjacent train numbered j at the initial time t0, and σ i (t) represents the external disturbance received by the system at time t; Assumption 3: Assume that the actual speed of the leading train is consistent with the target speed trajectory; the ultimate control objective of the multi-train virtual formation collaborative tracking is expressed as: Among them, and are both extremely small positive values, and v0 represents the speed of train numbered i at the initial time.
[0016] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention.
[0017] In the embodiment of the present invention, an adaptive mayfly optimization method is used to optimize the hyperparameters of the deep deterministic policy gradient algorithm. The optimized deep deterministic policy gradient algorithm is used to continuously train a distributed controller based on a cooperation and collision avoidance control protocol to obtain a trained distributed active disturbance rejection resilient controller, which is used to control each unit train in the virtual formation. The trained distributed active disturbance rejection resilient controller has good active disturbance rejection performance. When encountering external disturbances of different degrees, the parameters of the trained distributed active disturbance rejection resilient controller can be adaptively adjusted online to ensure that the virtual formation of the trains can quickly recover and maintain the expected formation under the action of disturbances. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a schematic flowchart of the method for distributed active disturbance rejection resilient control of an empty-rail train for virtual formation provided by the embodiment of the present invention.
[0020] Figure 2 It is a schematic diagram of the virtual formation provided by the embodiment of the present invention.
[0021] Figure 3 It is a schematic diagram of the weight node provided by the embodiment of the present invention.
[0022] Figure 4 It is a schematic diagram of the mayfly mutation based on Steffensen provided by the embodiment of the present invention.
[0023] Figure 5 It is a schematic diagram of the distance L(X1, Z1) provided by the embodiment of the present invention.
[0024] Figure 6 It is a schematic diagram of the virtual formation in the simulation experiment provided by the embodiment of the present invention.
[0025] Figure 7 It is a schematic diagram of the iteration result provided by the embodiment of the present invention.
[0026] Figure 8 It is a schematic diagram of the noise standard deviation provided by the embodiment of the present invention.
[0027] Figure 9 It is a schematic diagram of the noise standard deviation attenuation rate provided by the embodiment of the present invention.
[0028] Figure 10 Schematic diagram of the network learning rate provided by the embodiments of the present invention.
[0029] Figure 11 Schematic diagram of the return function curve provided by the embodiments of the present invention.
[0030] Figure 12 Schematic diagram of the actor network structure provided by the embodiments of the present invention.
[0031] Figure 13 Schematic diagram of the critic network structure provided by the embodiments of the present invention.
[0032] Figure 14 Schematic diagram of the episode return result provided by the embodiments of the present invention.
[0033] Figure 15 Schematic diagram of the average return result provided by the embodiments of the present invention.
[0034] Figure 16 Schematic diagram of the control ability verification result provided by the embodiments of the present invention.
[0035] Figure 17 Schematic diagram of the generalization ability verification result provided by the embodiments of the present invention.
[0036] Figure 18 Schematic diagram of the train speed change provided by the embodiments of the present invention.
[0037] Figure 19 Schematic diagram of the train speed change provided by the embodiments of the present invention.
[0038] Figure 20 Schematic diagram of the train control parameter change provided by the embodiments of the present invention.
[0039] Figure 21 Schematic diagram of the train operation interval change provided by the embodiments of the present invention. Detailed implementation manners
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] Due to the uncertainty and time-variability of the external disturbances suffered by the train during operation, conventional distributed controllers are difficult to meet the requirements of the virtual formation stable control of the suspended monorail train. The present invention relates to a distributed active disturbance rejection elastic control method and system for a suspended monorail train for virtual formation. First, a multi-agent modeling method is used to construct a distributed dynamic model of the suspended monorail train. At the same time, the formation state error is calculated according to the expected running speed and running interval of the unit train. Based on the multi-agent train model and the formation state error, a cooperation and collision avoidance control protocol based on a bidirectional-leader communication topology is designed for the virtual formation unit train. The hyperparameters of the Deep deterministic policy gradient algorithm (DDPG) are optimized by using the Adaptive Mayfly algorithm (AMA). The optimized DDPG algorithm is used to continuously train the distributed controller based on the cooperation and collision avoidance control protocol to obtain a trained distributed active disturbance rejection elastic controller to control each unit train of the virtual formation. This method has good active disturbance rejection performance. When encountering external disturbances of different degrees, the system controller parameters can be adaptively adjusted online to ensure that the virtual formation of the train quickly recovers and maintains the expected formation during the action of the disturbance.
[0042] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0043] As Figure 1 shown, the distributed active disturbance rejection elastic control method for a suspended monorail train for virtual formation in this embodiment includes the following steps 1 to 7.
[0044] Step 1: Use a multi-agent modeling method to construct a distributed dynamic model of the suspended monorail train for virtual formation.
[0045] Considering the internal disturbance of system parameter perturbation and the uncertain external environmental disturbance, a multi-agent modeling method is used to construct a dynamic model of the suspended monorail train based on virtual formation.
[0046] Consider a multi-agent virtual formation system consisting of N sky train cars. Each car is regarded as a multi-agent rigid mass point, numbered from 0 to N-1 in sequence. Let the car numbered 0 be the leader car; the cars numbered 1 to N-1 be the follower cars, and the follower cars only receive the state information of the cars within their respective neighborhoods. The interaction between cars can be described by a directed graph. Let E = {0, 1, …, N-1} represent the set of all cars, F = {1, 2, 3, …, N-1} represent the set of follower cars, and G = {i-1, i+1} represent the set of neighboring cars of the car numbered i. The dynamic model of the sky train car for multi-agent is expressed as Equation (1).
[0047]
[0048] In the formula, i ∈ E, x i , v i , u i , f i and M i represent the current position, speed, traction / braking force, resistance suffered, and mass of the car numbered i, respectively. The resistance f i suffered by the car numbered i is composed of the basic resistance f bi , the additional resistance f wi on the curve, the additional resistance f si on the slope, and the friction force f ci between the running wheels and the track. Among them, the additional resistance f wi on the curve and the additional resistance f si on the slope are related to the curve radius and slope angle, respectively; while the basic resistance f bi has a non-linear relationship with the train speed and can be expressed as Equation (2) according to the empirical formula.
[0049]
[0050] In the formula, α0, α1, and α2 are all resistance coefficients, and all three change with the change of the train operation environment. g is the acceleration due to gravity.
[0051] Since the sky train is inevitably affected by the external environment during operation, the system model has the characteristics of time-varying, uncertainty, and non-linearity. Therefore, the multi-agent train model can also be expressed in the following form.
[0052]
[0053] In the formula, x i (t) represents the position of the car numbered i at time t, represents the differential of x i (t), v i(t) represents the speed of train numbered i at time t. represents v i the differential of (t), f(v i ,t) represents the resistance force on train numbered i at time t, b represents the controller coefficient, u i (t) represents the traction / braking force of train numbered i at time t. represents the external environmental disturbance on train numbered i at time t, α v v i (t) represents the known system linear term, α v represents the system linear term coefficient. represents the uncertain non - linear internal disturbance of system parameter perturbation, σ i (v i (t),t) represents the total disturbance on the system at time t.
[0054] Step 2: Determine the formation state error model of the virtual formation according to the expected relative running speed and expected running interval between adjacent follower trains and the leader train in the virtual formation.
[0055] This invention studies the elastic control problem of the virtual formation of the suspended monorail train based on the leader - follower structure considering uncertainty disturbances. By using the deep reinforcement learning method to continuously train the train controller, the controller parameters can be continuously adjusted online with the change of disturbances to ensure the stable control of the virtual formation of the train. The schematic diagram of the virtual formation composed of multiple suspended monorail trains is as Figure 2 shown. Figure 2 where LT: leader train, FT: follower train, h is the train length.
[0056] In the virtual formation, the leader train runs according to the target speed trajectory given by the ground signal control center, while the follower train regards the actual speed of the preceding train as the expected running speed, thus forming a stable virtual formation.
[0057] The state vector of train numbered i is expressed as Equation (4).
[0058] ξ i =[x i ,v i T , i ∈ E (4).
[0059] The expected relative state between train numbered i and its neighboring train numbered j i,j and the actual relative state d
[0060]
[0061] In the formula, when \(i\in F\), the train numbered \(i\) is a follower train. Denotes the expected relative state between the train numbered \(i\) and the neighboring train numbered \(j\), \(d\) i,j Denotes the actual relative state between the train numbered \(i\) and the neighboring train numbered \(j\), \(\vert\vert\) represents the absolute value symbol, \(F\) represents the set of follower trains, \(F = \{1, 2, 3, \ldots, N - 1\}\); \(G\) represents the set of neighboring trains of the train numbered \(i\), \(G=\{i - 1, i + 1\}\), and \(j = i - 1\) or \(j = i + 1\). Denotes the expected headway between the train numbered \(i\) and the neighboring train numbered \(j\), \(x\) i,j Denotes the actual headway between the train numbered \(i\) and the neighboring train numbered \(j\). Denotes the expected relative speed between the train numbered \(i\) and the neighboring train numbered \(j\), generally \(v\) i,j Denotes the actual relative speed between the train numbered \(i\) and the neighboring train numbered \(j\), \(x\) i Denotes the current position of the train numbered \(i\), \(x\) j Denotes the current position of the neighboring train numbered \(j\), \(v\) i Denotes the current speed of the train numbered \(i\), \(v\) j Denotes the current speed of the neighboring train numbered \(j\), and the superscript \(T\) represents the transpose. The expected headway Can be expressed as Equation (6).
[0062]
[0063] Among them, \(S\) t Denotes the safety margin at time \(t\), \(S\) brake (\(v\) i ) Denotes the braking distance when the speed of the train numbered \(i\) is \(v\) i At this time, and \(S\) brake (\(v\) j ) Denotes the braking distance when the speed of the neighboring train numbered \(j\) is \(v\) j At this time. When the speeds of the leader train and the follower train are the same, the expected headway Is only affected by the train length \(h\) and the safety margin \(S\) t At time \(t\). The expected relative state Between the train numbered \(i\) and the leader train and the actual relative state \(d\) i,0 Are as shown in Equation (7).
[0064]
[0065] In the formula, Denotes the expected headway between the train numbered \(i\) and the leader train, \(x\) i,0 Denotes the actual headway between the train numbered \(i\) and the leader train. Denote the expected relative speed between train numbered \(i\) and the leader train as \(v\). i,0 Denote the actual relative speed between train numbered \(i\) and the leader train, \(x_0\) represents the current position of the leader train, and \(v_0\) represents the current speed of the leader train.
[0066] The formation state error model can be expressed as Equations (8) and (9).
[0067]
[0068] In the formula, \(e\) i,j Denote the formation state error between train numbered \(i\) and the neighboring train numbered \(j\), and \(e\) i,0 Denote the formation state error between train numbered \(i\) and the leader train.
[0069] For the convenience of research, the following assumptions are made without loss of generality.
[0070] Assumption 1: Assume that train numbered \(i\) (\(i\in F\)) obtains the positions and speeds of the two adjacent trains and the leader train in the neighborhood in real time through vehicle-to-vehicle wireless communication with low latency and low packet loss rate.
[0071] Assumption 2: Assume that the initial time is \(t_0\), and the initial state deviation \(\|\xi\) i (t_0)-\xi j (t_0)\|\leq\delta_1\), \(j\in G\), and the external disturbance \(\|\sigma\) i (t)\|\leq\delta_2\), where \(\delta_1\) and \(\delta_2\) are bounded positive constants. \(\xi\) i (t_0)\) represents the state vector of train numbered \(i\) at the initial time \(t_0\), and \(\xi\) j (t_0)\) represents the state vector of the neighboring train numbered \(j\) at the initial time \(t_0\), and \(\sigma\) i (t)\) represents the external disturbance suffered by the system at time \(t\).
[0072] Represents the norm.
[0073] Assumption 3: Assume that the speed of the leader train is consistent with the target speed trajectory.
[0074] The final control objective of multi-train virtual formation cooperative tracking can be expressed as Equation (10). In Equation (10), and are both extremely small positive values. \(v_0\) represents the speed of train numbered \(i\) at the initial time.
[0075]
[0076] Step 3: According to the distributed dynamic model and the formation state error model, use the tanh function to smooth the formation state error of adjacent trains, and establish a collaborative and collision avoidance control protocol based on a two-way leader-follower communication topology.
[0077] Due to internal disturbances such as uncertain external disturbances and system parameter perturbations, conventional control methods are difficult to meet the requirements of the collaborative operation stability and comfort of virtual formation trains. Therefore, a distributed active disturbance rejection control strategy is proposed.
[0078] 3.1 Design of the tracking differentiator.
[0079] Since there is a large deviation between the initial formation of multiple trains and the desired formation, the jerk of the follower trains is too large during the dynamic formation stage, resulting in poor ride comfort of the trains. A tracking differentiator (TD) is designed to provide a transition process for the train state deviation. The state deviation gradually increases smoothly from zero, and the train accelerates or decelerates smoothly, improving the ride comfort of the train. The formation state deviation consists of two parts: the running interval deviation and the speed deviation.
[0080] The tracking differentiator of train numbered i (i ∈ F) is as shown in Equation (11).
[0081]
[0082] In the formula, d i,0 (t) = ξ i (t) - ξ0(t) represents the actual relative state of the follower train numbered i and the leader train at time t; d i,0 (t) = [x i,0 (t), v i,0 (t)] T , x i,0 (t) represents the actual interval between train numbered i and the leader train at time t, and v i,0 (t) represents the actual relative speed between train numbered i and the leader train at time t; represents the differential signal of. represents d i,0 (t) jumps to the step input signal, represents x i,0 (t) jumps to the step input signal, represents v i,0 (t) jumps to the step input signal, represents the desired relative speed between train numbered i and the leader train at time t.
[0083] denotes the expected relative state between the train numbered \(i\) and the leader train at time \(t\). denotes the transition signal of denotes the running interval transition signal at time \(t\). denotes the relative speed transition signal at time \(t\); and both denote the differential signal of denotes the differential signal of ; \(r_0\) and \(r_1\) both denote the transition speed of the signal, and \(h_0\) and \(h_1\) both denote the terms that filter the noise of the signal; \(fhan()\) denotes the fhan function. The specific form of is as shown in Equation (12).
[0084]
[0085] In the formula, \(s\), \(y\), \(w_0\), \(w_1\), \(w_2\), \(s\) y , \(s\) a are the intermediate variables of the fhan function.
[0086] 3.2 Design of Extended State Observer.
[0087] The purpose of designing an extended state observer (ESO) is to observe all unknown terms including internal disturbances and external disturbances. However, the disturbance observation accuracy of the ESO will determine the effect of distributed cooperative operation of multiple trains. Therefore, it is particularly important to design an ESO with better observation performance. The parameters of the conventional ESO are severely coupled, resulting in repeated oscillations and the disadvantage that the convergence time gradually increases with the increase of the initial error. Therefore, a nonlinear ESO based on a high-order sliding mode differentiator is designed as shown in Equation (13).
[0088]
[0089] In the formula, \(\psi_1\) and \(\psi_2\) respectively denote the estimated values of the train speed \(\chi_1 = v\) i and the total system disturbance \(\chi_2=\sigma\) i (v i ). For any initial errors \(e_1(t_0)\) and \(e_2(t_0)\), there is always \(e_1(t)\) and \(e_2(t)\) converging to an arbitrarily small error interval within a finite time. \(\psi_1(\chi_1)\) denotes the estimated value of the train speed \(\chi_1\), \(\chi_1 = v\) i , \(v\) i denotes the current speed of the train numbered \(i\); \(\psi_2(\chi_1)\) denotes the estimated value of the total system disturbance \(\chi_2\), \(\chi_2=\sigma\) i (v i ), \(\sigma\) i (vi ) represents the total interference suffered by the system; e1 represents the error of the train speed, and e2 represents the error of the total system interference; represents the differential of ψ1(χ1), represents the differential of ψ2(χ1); η1(χ1) represents the internal state of the nonlinear extended observer; f(v i ) represents the resistance suffered by train No. i; b represents the controller coefficient, and u i represents the traction / braking force of train No. i; a1, b1, q1, and p1 are all adjustable observer parameters, a1 > 0, b1 > 0, 0 < q1 < p1; sgn() represents the sign function.
[0090] 3.3 Collision avoidance control.
[0091] When multiple empty rail trains operate in a virtual formation mode, the biggest safety hazard is the collision between trains. Therefore, it is very necessary to design an effective collision avoidance control strategy. Here, the idea of the artificial potential field method is used to achieve effective collision avoidance in the coordinated operation of multiple trains.
[0092] Considering the actual size of the train, use ρ j to represent the repulsive field radius of the neighboring train No. j, and r j to represent the actual radius of the neighboring train No. j. When the distance between train No. i and the neighboring train No. j is less than ρ j , train No. i will be subject to a repulsive force, and when the distance is equal to or less than r j , it is considered that the two trains have collided. The traditional repulsive field function of the artificial potential field method has certain limitations. When there are obstacles near the target point, there will be a problem that the target is unreachable. Therefore, an improved repulsive field function is introduced as shown in Equation (14).
[0093]
[0094] In the formula, represents the improved repulsive field of the neighboring train No. j relative to train No. i, r j represents the actual size radius of the neighboring train No. j, μ represents the repulsive force scale factor, ρ i,j represents the actual interval between train No. i and the neighboring train No. j, Δρ i,j represents ρ i,j and r j deviation, Δρ represents ρ j and r j deviation, represents the expected interval between train No. i and the neighboring train No. j, represents ρ i,j and The absolute value of the deviation, F represents the set of follower trains, and G represents the set of neighboring trains of train numbered i.
[0095] According to the improved repulsive force field function, the repulsive force exerted on each train by a neighboring train is as shown in Eqs. (15) and (16).
[0096]
[0097] In the formula, represents the repulsive force generated by the neighboring train numbered j on the train numbered i, i ∈ F, j ∈ G, n j→i represents the unit repulsive force vector from the neighboring train numbered j to the train numbered i, n i→B represents the unit vector that makes the train numbered i move in the direction B of the expected formation, represents the train numbered i, n j→i the repulsive force received in the direction, represents the train numbered i, n i→B the repulsive force received in the direction.
[0098] According to the repulsive force exerted on each train by a neighboring train, determine the total repulsive force of each train from the front and rear neighboring trains as follows.
[0099]
[0100] In the formula, F i rep represents the total repulsive force of the train numbered i from the front and rear neighboring trains, i ∈ F; represents the repulsive force of the train numbered i from the front neighboring train, represents the repulsive force of the train numbered i from the rear neighboring train.
[0101] The repulsive force exerted on each train by a neighboring train and the total repulsive force of each train from the front and rear neighboring trains constitute the collision avoidance control. The collision avoidance control, the tracking differentiator, and the nonlinear extended state observer together constitute the cooperation and collision avoidance control protocol.
[0102] Step 4: Construct a distributed active disturbance rejection controller based on the cooperation and collision avoidance control protocol.
[0103] Convert the multi-train cooperative control problem in the virtual formation mode into a multi-agent consensus problem. To better eliminate the adverse effects of interference on the system, the distributed active disturbance rejection cooperative controller for train numbered i is as follows.
[0104]
[0105] In the formula, e i,0Denote the formation state error between train numbered \(i\) and the leader train as \(e\). i,j Denote the formation state error between train numbered \(i\) and the neighboring trains. Denote as the transition signal. Denote the desired relative state between train numbered \(i\) and the neighboring train ahead, \(i\in F\); \(\xi\). i Denote the state vector of train numbered \(i\) as \(\xi\). i \(=\left[x\right.\) i , \(v\) i \(\left]\right.\) T , where \(x\) i denotes the current position of train numbered \(i\), and \(v\) i denotes the current speed of train numbered \(i\); \(\xi_0\) denotes the state vector of the leader train, \(\xi\). j Denote the state vector of the neighboring trains. Denote the collaborative part of train numbered \(i\), \(k\), and both denote the adjustable coefficient matrices of the collaborative and collision avoidance control protocols, \(\psi_2\) denotes the estimated value of the total system disturbance; \(k = [k_1,k_2]\), where \(k_1\) denotes the component in the position direction of \(k\), and \(k_2\) denotes the component in the speed direction of \(k\). Denote the component in the position direction of Denote the component in the speed direction of Denote the component in the position direction of Denote the component in the speed direction of; \(b\) denotes the controller coefficient; \(u\). i Denote the traction / braking force of train numbered \(i\) as \(F\). i rep Denote the total repulsive force exerted on train numbered \(i\) by the neighboring trains before and after, \(t\) denotes the moment \(t\), \(\tanh()\) denotes the hyperbolic tangent function, \(t\) denotes the moment \(t\). \(\psi_2\) is used as the total disturbance compensation term. \(F\). i rep is obtained from Equation (17). Equation (18) is related to the transition signal in Equation (11). Equation (18) is related to the total disturbance compensation term \(\psi_2\) in Equation (13).
[0106] Step 5: Use the adaptive mayfly optimization method to optimize the hyperparameters of the deep deterministic policy gradient algorithm; the adaptive mayfly optimization method adds a mayfly adaptive mutation mechanism based on Steffensen.
[0107] Due to the time-varying and uncertain characteristics of the disturbances suffered by the system, it is difficult for fixed distributed controller parameters to adaptively meet the actual control requirements. Therefore, the present invention proposes a distributed active disturbance rejection control algorithm based on AMA-DDPG to improve the adaptive anti-interference ability of the system.
[0108] 5.2 Hyperparameter optimization based on the adaptive mayfly optimization method.
[0109] Due to the problems of excessive hyperparameters and difficult tuning in the DDPG algorithm, an adaptive mayfly optimization method is designed to optimize multiple hyperparameters of the DDPG algorithm in the search space. Since the noise standard deviation, the noise standard deviation decay rate, and the network learning rate have a greater impact on the training effect of the DDPG algorithm, they are taken as the objects of hyperparameter optimization.
[0110] If in a D-dimensional target search space, a mayfly population Q is constructed by Y mayflies, where the number of male mayflies and female mayflies is both Y / 2. Xm i′ =(xm i′,1 ,xm i′,2 ,…,xm i′,D ) represents the position of the i'-th male mayfly; Vm i′ =(vm i′,1 ,vm i′,2 ,…vm i′,D ) represents the search speed of the i'-th male mayfly; Xf i′ =(xf i′,1 ,xf i′,2 ,…,xf i′,D ) represents the position of the i'-th female mayfly; Vf i′ =(vf i′,1 ,vf i′,2 ,…,vf i′,D ) represents the search speed of the i'-th female mayfly, i' = 1, 2, …, Y.
[0111] Since the conventional mayfly algorithm (MA) has the defect of being easily trapped in local optimal points, the present invention proposes an adaptive MA algorithm. The adaptive MA algorithm is mainly reflected in the mayfly adaptive mutation mechanism based on Steffensen.
[0112] 5.2.1 State update and mating and reproduction behavior of male and female mayflies.
[0113] 5.2.1.1 State update of male mayflies.
[0114] The position update of the i'-th male mayfly in the n-th generation population is as follows:
[0115]
[0116] In the formula, respectively represent the position and velocity of the j'-th dimension of the male mayfly. i' = 1, 2, …, Y, j' = 1, 2, …, D.
[0117] If the male mayfly is not the optimal individual in the mayfly population, its velocity is updated as follows:
[0118]
[0119] In the formula, pbest i′,j′ represents the individual optimal point searched by the i'-th male mayfly in the n-th generation population, gbest represents the global optimal point searched by the i'-th male mayfly in the n-th generation population, λ1, λ2, β are constants, r p is the Euclidean distance between the current position and pbest i′,j′ and r g is the Euclidean distance between the current position and gbest.
[0120] If the male mayfly is the optimal individual in the mayfly population, its velocity is updated as follows:
[0121]
[0122] In the formula, λ is a constant, and r1 is a random number within [-1, 1].
[0123] 5.2.1.2 Female mayfly state update.
[0124] The position of the i'-th female mayfly in the n-th generation population is updated as follows:
[0125]
[0126] In the formula, respectively represent the position and velocity of the j'-th dimension of the female mayfly.
[0127] If the female mayfly is inferior to the paired male mayfly, its velocity is updated as follows:
[0128]
[0129] In the formula, xm i ″ ,j′ is the position of the male mayfly paired with the female mayfly, is the position of the female mayfly, and r mf is the position between the two male and female mayflies.
[0130] If the female mayfly is superior to the paired male mayfly, its velocity is updated as follows:
[0131]
[0132] In the formula, fl is a constant, and r2 is a random number within [-1, 1].
[0133] 5.2.1.3 Mating and reproduction behavior.
[0134] In the mating behavior of mayflies, a pair of female and male mayflies will produce two offspring, and the two offspring produced can be expressed as:
[0135]
[0136] In the formula, L′ is a random number in the range of (0, 1), and represent two offspring produced by a pair of male and female mayflies in the nth generation of mayfly population. The mating behavior of mayflies provides the global search ability for the mayfly algorithm, and selecting the better individuals among the offspring can accelerate the convergence speed of the algorithm.
[0137] 5.2.2 Mayfly adaptive mutation mechanism based on Steffensen.
[0138] In the process of the conventional mayfly optimization algorithm searching for the global optimal point, the population diversity of mayflies will show a deteriorating state, and the phenomenon of premature convergence may occur. Therefore, the present invention proposes a mayfly adaptive mutation mechanism based on Steffensen.
[0139] 5.2.2.1 Diversity evaluation index based on weight nodes.
[0140] The present invention proposes a diversity evaluation index based on weight nodes to evaluate the diversity of mayfly population distribution. The core idea of this evaluation index: select a group of weight nodes uniformly distributed in the one-dimensional objective space. The distribution of the current mayfly population is evaluated by calculating the ratio κ of the number of weight nodes adjacent to the current mayfly optimal solution to the total number of weight nodes.
[0141] The schematic diagram of the weight nodes is as Figure 3 shown. Z1 to Z3 represent a group of weight nodes uniformly distributed in the one-dimensional objective space R (adaptive fitness space), and X1 to X3 are the current mayfly individuals. It can be seen that X1 to X3 are respectively closest to the Z2 node, and the only weight node adjacent to the entire mayfly population is Z2. Therefore, the current κ value is 1 / 3. κ ∈ [0, 1]. The smaller the κ value, the more concentrated the distribution of the current population of mayflies.
[0142] 5.2.2.2 Adaptive mayfly mutation strategy based on Steffensen method.
[0143] The Steffensen method is an iterative root-finding method for solving the non-linear equation f(x) = 0. The Steffensen formula is as follows.
[0144]
[0145] k = 0, 1, 2…
[0146] where x0 is the initial root value, and x k and x k+1 are the root values before and after the (k + 1)-th iteration respectively.
[0147] The present invention adopts a mutation strategy based on the Steffensen method to enhance the diversity of the mayfly population. If the κ index of the mayfly population is less than the given threshold ξ, then each mayfly in the population is mutated.
[0148] As Figure 4 shown, considering the mayfly individual X1, first randomly select a weight node (assumed to be Z1) from several weight nodes closest to it, and then mutate the mayfly individual X1 with the goal of shortening the distance L(X1, Z1) between the mayfly individual X1 and Z1, obtaining a mutated mayfly individual X1' that is closer to Z1. By performing the above operations on each mayfly individual in the population one by one, a group of mutated mayfly individuals that are more evenly distributed in the target space is obtained.
[0149] As Figure 5 shown, take the mayfly individual X1 to be mutated and the corresponding weight node Z1 as an example to illustrate the operation process of the Steffensen method.
[0150] First, according to Figure 4 calculate the distance L(X1, Z1) = l between the mayfly individual X1 and Z1. The smaller L(X1, Z1) is, the closer the mayfly individual X1 is to Z1. Then, calculate l * = α·L(X1, Z1), α ∈ [0, 1]. Obviously, l * < L(X1, Z1). Randomly perform a mutation operation on a certain dimension of the mayfly individual X1, so that the mutated mayfly individual X1' satisfies L(X1, Z1) - l * = 0. If the k-th dimension of the mayfly individual X1 is selected for mutation, it is equivalent to solving the problem of the non-linear equation L(X 1,k , Z1) - l * = 0. Use the Steffensen method to solve the root of L(X 1,k , Z1) - l * = 0, and take the first iteration approximate solution as the value after mutation of the mayfly individual X 1,k . According to the iteration formula of the Steffensen method, obtain the mayfly individual X1,k The mutated values are as follows.
[0151]
[0152] In the formula, X 1,k represents the k-th dimension of the mayfly individual X1 to be mutated, and X 1,k ' is the value of X 1,k after mutation; θ(X 1,k ) represents the deviation between the distance between the mayfly individual X1 to be mutated and the weight node Z1 and the expected distance between X 1,k ' and Z1, θ(X 1,k ) = L(X1, Z1) - l * , L(X1, Z1) represents the distance between the mayfly individual X1 to be mutated and the weight node Z1, and l * represents the expected distance between X 1,k ' and Z1, l * = α·L(X1, Z1), α ∈ [0, 1], and α represents a random number in the interval [0, 1].
[0153] 5.3 AMA optimization process.
[0154] The AMA optimization method is used to optimize the four hyperparameters of the DDPG algorithm in the specified search space, namely the noise standard deviation σ, the noise standard deviation decay rate ψ, the learning rate α u of the Actor network and the learning rate α Q of the Critic network.
[0155] Select the fitness function based on the ITAE (Integral of Time multiplied by the Absolute Error) index as follows.
[0156]
[0157] In the formula, e(k) represents the deviation between the current cumulative return value R(k) and the cumulative return target value R o in the k-th iteration (episode) of the process of training the agent, where the value of R o is related to the reward function and the maximum number of training steps in a single episode. The AMA optimization method is used to optimize the DDPG hyperparameters to finally stabilize the fitness function J corresponding to the training results of the DDPG algorithm at the minimum value.
[0158] The specific process of optimizing the hyperparameters of the Deep Deterministic Policy Gradient algorithm using the Adaptive Mayfly Optimization method is as follows:
[0159] 5.3.1: Initialize the mayfly population and generate a set of weight nodes evenly distributed in the target space.
[0160] 5.3.2: Assign values to the hyperparameters to be optimized in the deep deterministic policy gradient algorithm using the initialized mayfly population; the hyperparameters to be optimized include the noise standard deviation, the noise standard deviation decay rate, the Actor network learning rate, and the Critic network learning rate.
[0161] 5.3.3: Calculate the diversity evaluation index of the current mayfly population.
[0162] 5.3.4: If the diversity evaluation index of the current mayfly population is less than the preset threshold, then based on the weight nodes, use the Steffensen method to mutate each mayfly in the current population, and select the best from the mutated mayfly population and the original mayfly population, and recombine them into a new mutated mayfly population with the same number as the original population.
[0163] 5.3.5: Divide the new mutated mayfly population into half females and half males, calculate the fitness values of the mayfly individuals included in the new mutated population according to the fitness function based on the ITAE index, and select the individual best mayfly and the global best mayfly from all male mayflies.
[0164] 5.3.6: According to the individual best mayfly and the global best mayfly selected so far, use the male mayfly state update formula to update the position and speed of the male mayflies, and update the individual best mayfly and the global best mayfly; according to the female mayfly state update formula to update the position and speed of the female mayflies, and calculate the fitness values of all mayflies in the mayfly population.
[0165] 5.3.7: Each pair of females and males in the mayfly population mates, and two offspring are produced according to the mating and reproduction behavior formula, and excellent individuals with the same number as the original population are selected from the mated mayflies and reserved for the next generation.
[0166] 5.3.8: Detect whether the end condition is reached; if so, output the optimal hyperparameters; if not, replace the initialized mayfly population with the latest mutated population, and return to the step "Assign values to the hyperparameters to be optimized in the deep deterministic policy gradient algorithm using the initialized mayfly population".
[0167] The brief steps of the AMA to optimize the DDPG hyperparameters are as follows.
[0168] step1: Initialize the position and speed states of the mayfly population with the quantity Y, and generate a set of weight nodes Z = {Z1, Z2, …, Z k} that are uniformly distributed in the target space according to the method proposed by Das and Dennis.
[0169] step2: The mayfly population assigns values to the hyperparameters to be optimized in the DDPG algorithm. The DDPG algorithm trains the maglev train agent to realize the multi-train adaptive cooperative tracking control process.
[0170] Step 3: Calculate the diversity κ index of the current mayfly population. If κ is less than the given threshold ζ, mutate each mayfly in the mayfly population, and select the better ones from the mutated mayfly population and the original mayfly population to form a new mutated mayfly population of quantity Y.
[0171] Step 4: Divide the new mutated mayfly population into half females and half males, and at the same time calculate the fitness value J of each mayfly individual in the population. Select the individual best mayfly PBEST and the global best mayfly GBEST from all male mayflies.
[0172] Step 5: Update the positions and velocities of male mayflies according to Formulas (19)-(21), and update the individual best mayfly and the global best mayfly; update the positions and velocities of female mayflies according to Formulas (22)-(24); calculate the fitness values of all mayflies in the mayfly population.
[0173] Step 6: A pair of female and male mayflies in the mayfly population will produce two offspring according to Formula (25); select Y better individuals from the 2Y mayflies after mating and retain them for the next generation.
[0174] Step 7: Finally, detect whether the end condition is reached. If it is reached, give the result of parameter optimization. If it is not reached, go back to Step 2 to continue the next iteration.
[0175] Aiming at the pain points of too many hyperparameters and difficult tuning in the DDPG algorithm, the present invention proposes an adaptive MA algorithm (AMA) to optimize hyperparameters such as the noise standard deviation, the noise standard deviation attenuation rate, and the network learning rate in the search space, improving the learning effect of the DDPG algorithm. The subsequent simulation results show that: compared with the MA-AC algorithm, the MA-DQN, and the MA-DDPG algorithms, the AMA-DDPG algorithm shows better effects in both the training and validation stages. This method effectively realizes the adaptive online adjustment of the controller parameters for different degrees of disturbances, ensuring the stability control of the virtual formation of trains under the action of interference.
[0176] Step 6: Train the distributed active disturbance rejection controller using the optimized deep deterministic policy gradient algorithm to obtain the trained distributed active disturbance rejection controller.
[0177] Continuously train each maglev train agent using the DDPG algorithm. The trained train agent can adaptively adjust the parameters of the distributed controller according to time-varying disturbances. First, define the learning environment of the train agent, including the state space, the action space, and the reward function.
[0178] 6.1.1 Learning environment.
[0179] (1) State space.
[0180] In the virtual formation mode, the leader train state vector is expressed as ξ0 = [x0, v0] T , when i ∈ F, the follower train state vector is expressed as ξ i = [x i , v i T . The comfort of train riding is an important indicator of high-quality train operation. Here, the jerk J of the follower train needs to be considered i . Therefore, the state space S of the follower train can be expressed as: In the formula, represents the second derivative of the follower train speed with respect to time, that is, the train jerk.
[0181] (2) Action space.
[0182] The core idea of the distributed active disturbance rejection controller is to observe the disturbance suffered by the system through NESO (nonlinear extended state observer) and perform control compensation. However, when the amplitude and frequency of the disturbance suffered by the system change, selecting appropriate adjustable parameters a of NESO 1,i can effectively suppress the disturbance. Therefore, the action space A can be expressed as follows.
[0183] A = {[a 1,i}, i ∈ F (30).
[0184] (3) State transition equation.
[0185] The state s of the agent train at time t t takes the action a t and transfers to the state s at time t + 1 t+1 with a state transition probability P s a = 1. According to equations (1), (2) and (3), the motion state transition equation of the straddle-type monorail train can be expressed as follows.
[0186]
[0187] In the formula, g i represents the average acceleration of the follower train within ΔT time, represents the change rate of the average jerk of the follower train within ΔT time, Since the g i term contains the traction / braking force u of train numbered i i , u iAs the output item of the controller, that is, the control input of train numbered i, and u i is related to the NESO parameters. ΔT represents a small time interval. It can be seen from formula (13) that by changing the NESO parameters, the transfer direction of the system state can be directly affected.
[0188] (4) Reward function.
[0189] The reward function is designed based on the principle of heavy reward and light punishment. Here, the normal distribution reward function r is selected.
[0190]
[0191] In the formula, σ1 and μ1 respectively represent the standard deviation and mathematical expectation of the normal distribution reward function. η′1 and η2 both represent adjustable parameters of the reward function, which are used to adjust the degree of heavy reward and light punishment of the reward function. x = J1 represents the train jerk as the random variable of the normal distribution reward function. The closer J1 is to μ1, the greater the positive reward value. On the contrary, the smaller the negative reward value.
[0192] 6.1.2 DDPG algorithm.
[0193] The DDPG algorithm is a deep deterministic policy gradient algorithm. It adds an experience replay mechanism and a dual-network structure on the basis of the Actor-Critic (AC) algorithm framework. The DDPG framework mainly includes a learning environment, an experience pool for storing samples, an actor network, and a critic network. The agent continuously interacts with the environment to generate sample data and stores these samples in an experience pool. A certain number of samples (minibatch samples) are randomly selected from the experience pool to train the network parameters. To improve the learning efficiency of the algorithm, the DDPG algorithm adopts a dual-network structure. Whether it is the actor network or the critic network, it includes an eval network and a target network. The target network and the eval network are a pair of neural networks with exactly the same structure. The eval network updates the network parameters (θ u , θ Q ) during each step of training, while the target network parameters (θ u′ , θ Q′ ) are periodically copied from the eval network parameters by soft update.
[0194]
[0195] In the formula, τ is the soft update coefficient. To enhance the agent's exploration ability of the environment, it is necessary to add the OU noise N′ that follows the normal distribution to the output action μ(s t ) of the actor-eval network as follows.
[0196]
[0197] In the formula, μ2 represents the expected mean value, μ2 = 0, σ0 represents the initial noise standard deviation, and n represents the current iteration step of a certain round. The noise standard deviation σ2 gradually decreases as the iteration step n increases, reducing the exploration of the environment. ψ represents the standard deviation decay rate. The final action a t is superimposed with OU noise.
[0198] a t = μ(s t ) + N′(35).
[0199] The actor network and the critic network are trained using different loss functions. When using batch data, the loss function of the critic network is as follows.
[0200]
[0201] In the formula, K is the number of samples in the batch data; Q(s t , a t |θ Q ) is the value evaluated by the eval network for the current moment state and the action generated by the eval network in the actor.
[0202] y t = r(s t , a t ) + γQ′(s t+1 , μ′(s t+1 |θ u′ )|θ Q′ ) (37).
[0203] In the formula, r(s t , a t ) is the immediate reward after executing the action a t at the current moment state; γ is the reward decay coefficient; Q′(s t+1 , μ′(s t+1 |θ u′ )|θ Q′ ) is the value evaluated by the target network for the next moment state s t+1 and the action μ(s t+1 ) selected by the target network in the actor for the next moment state.
[0204] The loss function of the actor network is as follows.
[0205]
[0206] The parameter update formula of the critic network is as follows.
[0207]
[0208] where α Q is the learning rate of the critic network. represents the gradient of the current network loss function of the critic.
[0209] The parameter update formula of the actor network is as shown in Equation (40).
[0210]
[0211] where α u is the learning rate of the actor network. represents the gradient of the current network loss function of the Actor.
[0212] 6.1.3 DDPG Training Process.
[0213] Each straddle train is regarded as an agent, and the train controller is continuously trained using the DDPG algorithm. The trained train controller can adaptively and online adjust the controller parameters for different disturbances, effectively suppressing the adverse effects brought by external disturbances.
[0214] The steps for adaptive training of the parameters of the distributed cooperative controller based on the DDPG algorithm are as follows.
[0215] 1) Initialize the experience pool D with a storage capacity of M. 2) Initialize the policy networks eval of the actor and the critic: μ(s; θ u ) and Q(s t , a t |θ Q ); copy the parameters of the eval network to the corresponding target network: θ u′ ← θ u , θ Q′ ← θ Q ; the initialization of the target networks target of the actor and the critic is completed: μ(s; θ u′ ) and Q(s t , a t |θ Q′ ). 3) For episode = 1 to MaxEpisode do. 4) Initialize the OU noise process N t . 5) Randomly initialize the initial state of the straddle train and the external disturbance, and obtain the initial state of the simulation environment. 6) For t = 1 to MaxStep do. 7) Select the controller parameter a t as the action through the train state s t = π θ (s t ) + Nt . Among them, π θ (s t ) represents the selected strategy in state s t . 8) Establish a simulation environment according to the controlled object model (3), and obtain the control quantity through formula (16). Input the control quantity into the simulation environment to obtain the state of the straddle-type monorail train at the next moment, that is, the state s of the environment at the next moment t+1 . 9) Obtain the immediate reward value r of the environment according to the current state t . 10) Store the experience sample [s t , a t , r t , s t+1 in the experience replay pool D. 11) Randomly sample from the experience replay pool D to obtain a sample set {[s t , a t , r t , s t+1} of size BatchSize. 12) Update the policy network eval parameter θ of the critic action network Q . 13) Update the policy network eval parameter θ of the actor action network u . 14) Update the target network parameters θ u′ , θ Q′ in the two networks. 15) EndFor. If the episode end mechanism is satisfied, Break
[0216] The nonlinear extended state observer in the distributed active disturbance rejection controller is used to observe and suppress external disturbances. Here, different degrees of external disturbances that the train often encounters are selected as the training environment, and the AMA-DDPG algorithm is used to continuously train the train controller. The trained train controller can suppress different degrees of disturbances by online adjusting the NESO parameters
[0217] Step 7: Adaptively adjust the parameters of the trained distributed active disturbance rejection controller online according to the time-varying disturbance, and control the straddle-type monorail trains in the virtual formation, so that the straddle-type monorail trains in the virtual formation maintain the expected formation under the action of the disturbance
[0218] The specific process of Step 7 is as follows
[0219] 7.1: Obtain the actual speed and position information of one of the straddle-type monorail trains at the current moment
[0220] 7.2: Obtain the actual speeds and position information of the adjacent train in front, the adjacent train behind, and the leader train of a given suspended monorail train at the current moment, and calculate the actual intervals between a given suspended monorail train and the adjacent train in front, between a given suspended monorail train and the adjacent train behind, as well as the actual interval and actual relative speed between a given suspended monorail train and the leader train through deviation calculation;
[0221] 7.3: Use the actual interval and actual relative speed between a given suspended monorail train and the leader train at the current moment as the input signals of the tracking differentiator. The tracking differentiator provides a transition signal for the train formation state error and outputs the running interval transition signal and relative speed transition signal at the current moment;
[0222] 7.4: Take the control quantity and actual speed of a given suspended monorail train at the current moment as the input of the nonlinear extended state observer, and estimate the total system disturbance through state observation output.
[0223] 7.5: Use the actual intervals between a given suspended monorail train and the adjacent train in front, and between a given suspended monorail train and the adjacent train behind at the current moment as the input of the collision avoidance control. If the inputs of the collision avoidance control are both less than the set train repulsion field radius, a given suspended monorail train will be subject to the repulsive forces exerted by the adjacent train in front and the adjacent train behind respectively. A given suspended monorail train prevents collisions between trains through the total repulsive force exerted by the adjacent train in front and the adjacent train behind.
[0224] 7.6: Based on the state information of a given suspended monorail train, the adjacent train in front, the adjacent train behind, and the leader train at the current moment, use the running interval transition signal and relative speed transition signal output by the tracking differentiator, the estimated value of the total system disturbance output by the nonlinear extended state observer, and the total repulsive force signal output by the collision avoidance control as the current inputs of the distributed active disturbance rejection controller, and calculate the output control quantity of the distributed active disturbance rejection controller for a given suspended monorail train at the next moment; the state information includes actual speed and position information.
[0225] 7.7: Use the output control quantity at the next moment to drive a given suspended monorail train to track the adjacent train in front and run in coordination, and obtain the actual speed and position information of a given suspended monorail train at the next moment.
[0226] 7.8: Update the current moment to the next moment and return to step 7.2.
[0227] The distributed active disturbance rejection controller ensures the stability control of the virtual formation of trains under the action of disturbances by controlling the traction and braking forces of the trains.
[0228] Taking train numbered i as an example, the specific control process of the distributed active disturbance rejection control is as follows.
[0229] Step1: Taking t0 as the initial moment, initialize the output control quantity u i (t0) of the controller of train numbered i.
[0230] Step2: Train numbered i obtains its own actual speed v i (t) and position information x i (t) at the current moment t through the speed measurement and positioning device, and obtains the actual speeds v i-1 (t), v i+1 (t), v0(t) and position information x i-1 (t), x i+1 (t), x0(t) of the leading train at the current moment t of the preceding train numbered i - 1, the following train numbered i + 1 and the leading train by using the vehicle - to - vehicle wireless communication method. Calculate the actual intervals x i,i-1 (t) and relative speeds v i,i-1 (t) between train numbered i and the preceding train numbered i - 1, the actual intervals x i,i+1 (t) and relative speeds v i,i+1 (t) between train numbered i and the following train numbered i + 1, and the actual intervals x i,0 (t) and relative speeds v i,0 (t) between train numbered i and the leading train.
[0231] Step3: Take the actual interval x i,0 (t) and relative speed v i,0 (t) between the above - mentioned train numbered i and the leading train as the input signals of the tracking differentiator, and the tracking differentiator provides a transition process for the train formation state error and outputs the running interval transition signal and relative speed transition signal at the current moment t.
[0232] Step4: Take the control quantity u i (t) of train numbered i and the obtained actual speed v i (t) at the current moment t as the input of the nonlinear extended state observer, and estimate the signal ψ2 of the external disturbance σ i (v i ) through the state observation.
[0233] Step5: Take the actual intervals x i,i-1 (t), x i,i+1(t) serves as the input for collision avoidance control. If at this time x i,i-1 (t), x i,i+1 (t) are both less than the set repulsive force field radius ρ j of the train, the train numbered i will be subject to repulsive forces exerted by the train numbered i - 1 and the train numbered i + 1 respectively The total repulsive force F obtained by the train numbered i above i rep is used to prevent collisions between trains.
[0234] Step6: Based on the state information of the train numbered i, the leading train numbered i - 1, the following train numbered i + 1, and the leader train at the current time t, the running interval transition signal output by the tracking differentiator and the relative speed transition signal the disturbance estimation signal ψ2 output by the nonlinear extended state observer and the total repulsive force signal F output by the collision avoidance control i rep are used as the current inputs of the distributed active disturbance rejection controller. After calculation, the output control quantity u i (t + 1) of the controller for the train numbered i at the next time t + 1 is obtained.
[0235] Step7: The output control quantity u i (t + 1) obtained above drives the train numbered i to track the leading train and run in coordination, and the actual speed v i (t + 1) and position signal x i (t + 1) of the train numbered i at the next time t + 1 are obtained.
[0236] Step8: The output control quantity and state information at the next time t + 1 obtained above are used as the output control quantity and state information of the train numbered i at the current time t, that is, u i (t) = u i (t + 1), v i (t) = v i (t + 1), x i (t) = x i (t + 1).
[0237] Step9: Return to Step2 and loop through the above steps to finally achieve the stable formation running of each unit train in the virtual formation.
[0238] The method proposed in the existing literature "Z.Y. Guo, Z.Q. Li, "Distributed auto disturbances rejection resilient control of permanent magnetic maglev trains based on the optimized deep deterministic policy gradient algorithm", IET Control Theory and Applications, 2024." only considers the operation state information of adjacent front and rear trains. To a certain extent, it can accelerate the restoration of a stable formation. However, the input of the state information of the front and rear trains needs to be weighed and processed. If the controller is not designed properly, oscillations in speed and spacing are likely to occur, and the stability of a single vehicle cannot be guaranteed while taking into account the stability of the queue. Moreover, too large a formation state deviation of the adjacent rear train will cause the deterioration of the queue stability. The method proposed in the present invention introduces the operation state information of the leading train on the basis of the existing literature to improve the oscillation situation. By taking the operation state information of the leading train and the front and rear trains as control inputs, the stability of both the queue and the single vehicle can be taken into account simultaneously.
[0239] The main beneficial effects of the present invention are as follows.
[0240] (1) Compared with the existing literature, the distributed control method proposed in the present invention is designed based on a two-way - leader communication topology, introduces the operation state information of the leading train, can take into account the stability of both the queue and the single vehicle simultaneously, and achieves a satisfactory control effect. Moreover, the tanh function is used to smooth the formation state error with the adjacent train to prevent improper control caused by too large an error.
[0241] (2) The present invention proposes an adaptive mayfly optimization method (AMA), which effectively solves the problem that the conventional mayfly method is prone to fall into local optimal points.
[0242] (3) The present invention uses the AMA method to optimize the hyperparameters of the DDPG algorithm, improving the learning efficiency of the DDPG algorithm.
[0243] (4) The present invention proposes a distributed auto disturbances rejection resilient control method based on AMA-DDPG for sky trains to achieve adaptive cooperative control, which is the main innovation in terms of theoretical development.
[0244] To solve the adaptive control problem for external disturbances of different degrees and achieve effective virtual formation of sky trains, the present invention proposes a distributed active disturbance rejection elastic control method based on the deep deterministic policy gradient algorithm and adaptive mayfly optimization (AMA-DDPG). A nonlinear extended state observer (NESO) based on a high-order sliding mode differentiator is designed to observe disturbances, and a deep deterministic policy gradient algorithm (DDPG) is designed to adaptively adjust the parameters of the NESO to adapt to changes in external disturbances. However, the DDPG algorithm has a large number of hyperparameters, which makes it difficult to calculate the optimal values. Therefore, to optimize hyperparameters such as the noise standard deviation (SD), noise SD decay rate, and learning rate of the network, an adaptive mayfly optimization method (AMA) is proposed. Each train in the virtual formation is regarded as an agent. The unique feature of this method is that the AMA-DDPG algorithm is designed to continuously train all agent trains. After successful training, the controller parameters of the sky train can be adaptively adjusted online according to different degrees of external interference. Compared with other methods, the AMA-DDPG algorithm exhibits stronger anti-interference performance.
[0245] The method of the present invention is verified by simulation below.
[0246] (1) Simulation environment.
[0247] Here, 3 sky trains form a virtual formation for simulation tests. Train No. 0# is used as the leader train, and Train No. 1# and Train No. 2# are used as follower trains. Train No. 1# follower train runs in the middle position of the virtual formation. The present invention mainly focuses on the research of Train No. 1# follower train. The virtual formation in this simulation test is as Figure 6 shown.
[0248] Assume that the initial states of the leader train and each follower train are as shown in Table 1.
[0249] Table 1 Initial states of agent trains
[0250] Agent train Position (m) Speed (m / s) Pilot train 135 0 No. 1 follower train 95 0 No. 2 follower train 55 0
[0251] Internal disturbances of follower sky trains can be expressed as follows.
[0252]
[0253] The sampling period h = 0.01, controller coefficient b = 0.00059, and α v = -0.0078 are set for system simulation. During the simulation, the expected running interval between adjacent trains changes at 950 s, from 8 m to 30 m.
[0254] Design parameter selection of the differential tracker (TD) for the follower train: r0 = 0.1, h0 = 0.01. Selection of other parameters of the NESO for the follower train: b1 = 1.0, q1 = 1.37, p1 = 2.35.
[0255] The parameters of the distributed active disturbance rejection controller (18) are selected as follows.
[0256]
[0257] Parameter selection of the artificial potential field method: the influence radius ρ of the train repulsive force = 10 m, the actual radius r of the train = 3 m, and the repulsive force scale factor μ = 5.9.
[0258] (2) Optimization by the AMA method.
[0259] The AMA method is used to optimize multiple hyperparameters of the DDPG algorithm. The number of mayfly populations Y = 100, the maximum number of iterations N d = 50, the adjustable parameters λ1 = λ2 = β = 2, λ = fl = 0.1, the mayfly dimension D = 3, and here α is set u = α Q . [v min , v max = [-0.1, 0.1], and the search space for 4 hyperparameters is set: the standard deviation σ ∈ [1, 20], the standard deviation decay rate ψ ∈ [0.01, 1], and the network learning rate α u , α Q ∈ [0.0001, 0.01].
[0260] Table 2 Iteration results
[0261] Optimization algorithm Convergence iteration number Convergence fitness MA 32 280 AMA 12 240
[0262] According to Figure 7 and Table 2, it can be seen that compared with the conventional MA algorithm, the AMA algorithm has a faster convergence speed and a smaller fitness after convergence, and can effectively find the global optimal point, making up for the defect that the conventional MA algorithm is prone to fall into the local optimal point.
[0263] From Figures 8 to 10 it can be known that the optimal hyperparameters of the DDPG algorithm are: the noise standard deviation σ = 15.66, the noise standard deviation decay rate ψ = 0.49, and the Actor and Critic network learning rates α u = α Q = 0.00205.
[0264] (3) Training and verification of the DDPG algorithm.
[0265] The DDPG algorithm is used to train the maglev train intelligent agent to realize the adaptive online adjustment of the parameters of the distributed active disturbance rejection controller under different external disturbance conditions.
[0266] Other hyperparameter settings of the DDPG algorithm: The simulation step size is 0.01, the soft update coefficient τ = 0.001, the size of the experience replay pool M = 100000, the discount factor γ = 0.99, the number of samples drawn BatchSize = 300, the maximum number of steps per episode MaxStep = 50, and the maximum number of training episodes MaxEpisode = 5000. The initial value of the parameter action a = 5.1. AdamOptimizer is used as the optimizer for network parameters during training. The reward function curve is as Figure 11 shown.
[0267] The target network and the eval network of the actor adopt a feedforward neural network with multiple hidden layers. The relu function is used as the activation function for the hidden layers, and the tanh function is used as the activation function for the output layer, so that the output parameter action of the agent is limited within a certain range: [a min , a max = [-5, 5]. The structure of the actor network is as Figure 12 shown, Figure 12 where the numbers represent the number of neurons in a certain layer.
[0268] The target network and the eval network of the critic also adopt a feedforward neural network with multiple hidden layers. The input layer of the network consists of two parts: the state of the train agent and the parameter action generated by the actor network. The output layer of the network outputs a state-action value, and its network structure is as Figure 13 shown, Figure 13 where the numbers represent the number of neurons in a certain layer.
[0269] 3.1 Training process.
[0270] To better train the adaptive ability of the suspended monorail train to external disturbances and increase the credibility of training, the initial external disturbance training environment setting of the present invention is represented by Equation (43). For each training episode, the system randomly gives an initial external disturbance to train the agent train.
[0271]
[0272] where, is the external disturbance.
[0273] The MA-PG, MA-AC, MA-DDPG, and AMA-DDPG algorithms are used to train the train agent respectively to achieve the function of adaptive adjustment of control parameters. The results of their training processes are as Figures 14 to 15 shown.
[0274] From Figure 14 and Figure 15It can be seen that compared with the MA-PG and MA-AC algorithms, the AMA-DDPG algorithm shows faster and more stable convergence during the training process. The MA-DDPG algorithm gets stuck in a local optimum during the training process. However, the AMA-DDPG algorithm well avoids this problem, finds the global optimum, and can obtain a greater episodic reward.
[0275] 3.2 Verification process.
[0276] To verify the performance of the trained distributed elastic controller, the present invention randomly gives initial external disturbances to form 500 verification episodes, and verifies the results of four algorithms, namely MA-PG, MA-AC, MA-DDPG, and AMA-DDPG. The initial external disturbances during the verification process are as follows:
[0277]
[0278] Figure 16 and Figure 17 represent the verification results of the control ability and generalization ability of the four algorithms. Assume that an episode with an episodic reward exceeding 12 is called a successful episode. Then the control success rate and generalization success rate of the four learning algorithms are shown in Table 3.
[0279] Table 3 Verification results
[0280] Learning algorithm Control success rate Generalization success rate MA-PG 73.6% 60.6% MA-AC 78.0% 63.4% MA-DDPG 91.6% 84.0% AMA-DDPG 100% 95.8%
[0281] From Figure 16 , Figure 17 and Table 3, it can be seen that in terms of control ability, the verification effect of MA-PG is the worst, the control success rate is the lowest, MA-AC is the second, MA-DDPG performs well, and the verification effect of AMA-DDPG is the best, with a control success rate as high as 100%. In terms of generalization ability, the generalization ability of MA-PG is the worst, MA-AC is the second, MA-DDPG has good generalization ability, and the generalization ability of AMA-DDPG is the best, with a generalization success rate as high as 95.8%.
[0282] To further demonstrate the effectiveness of the AMA-DDPG algorithm, three different degrees of external disturbance environments are set during the collaborative tracking process of multiple maglev trains to verify the effect of the train agent's adaptive adjustment of control parameters with the change of external disturbances. The time-varying external disturbance environment can be expressed as follows.
[0283]
[0284] The intelligent agent train is trained using the AMA-DDPG algorithm, Figure 18 and Figure 19 respectively represent the speed change trends of the train before and after training. Figure 20It represents the changing trend of the running interval between trains after training. And Figure 21 It represents the comparison of the control parameter changes of the follower train before and after training.
[0285] Combined with Figures 18 to 20 It can be seen that compared with the untrained agent train, the trained agent train has better adaptability. Despite encountering external disturbances of three different levels, the system control parameters can still be adaptively adjusted according to the changing disturbances, and the speed control curve fluctuates less. However, the control parameters of the untrained agent train remain unchanged, and the speed control curve fluctuates greatly, making it difficult to meet the high-precision speed control requirements of the suspended monorail train. From Figure 21 It can be seen that by using the distributed adaptive cooperative control algorithm based on AMA-DDPG, it is possible to achieve high-precision cooperative operation of 3 suspended monorail trains with a certain safety interval in the virtual formation mode.
[0286] This invention studies the problem of virtual formation cooperative control of suspended monorail trains under uncertain disturbances, and proposes a distributed active disturbance rejection elastic control algorithm based on AMA-DDPG. During the training process, the adaptive ability of the train agent to different levels of disturbances is gradually improved. The main conclusions of this study are as follows.
[0287] 1) Compared with the conventional MA (Mayfly algorithm), AMA shows outstanding advantages in optimizing multiple hyperparameters of DDPG, with a faster convergence speed and better optimization effect, converging only after 12 iterations.
[0288] 2) Compared with the three algorithms of MA-PG, MA-AC, and MA-DDPG, the AMA-DDPG algorithm shows more stable and faster convergence during the training process; during the verification process, it shows a higher control success rate and stronger generalization ability.
[0289] 3) Compared with the conventional method, the distributed disturbance rejection elastic controller based on AMA-DDPG can adaptively adjust the controller parameters online for different levels of disturbances, thus showing stronger active disturbance rejection performance.
[0290] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0291] In the present invention, specific examples are used to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A distributed active disturbance rejection resilient control method for an empty rail train facing virtual formation, characterized in that Including: Constructing a distributed dynamic model of the suspended monorail train in a virtual formation using a multi-agent modeling method; Determining the formation state error model of the virtual formation according to the expected running relative speed and the expected running interval between adjacent follower trains and the leader train in the virtual formation; According to the distributed dynamic model and the formation state error model, using the tanh function to smooth the formation state error of adjacent trains, and establishing a cooperative and collision avoidance control protocol based on a bidirectional-leader communication topology; Constructing a distributed active disturbance rejection controller based on the cooperative and collision avoidance control protocol; Using an adaptive mayfly optimization method to optimize the hyperparameters of the deep deterministic policy gradient algorithm; the adaptive mayfly optimization method adds a mayfly adaptive mutation mechanism based on Steffensen; Training the distributed active disturbance rejection controller using the optimized deep deterministic policy gradient algorithm to obtain a trained distributed active disturbance rejection controller; Adapting and online adjusting the parameters of the trained distributed active disturbance rejection controller according to time-varying disturbances, and controlling the suspended monorail trains in the virtual formation, so that the suspended monorail trains in the virtual formation maintain the expected formation under the action of disturbances; Adapting and online adjusting the parameters of the trained distributed active disturbance rejection controller according to time-varying disturbances, specifically including: Obtaining the actual speed and position information of a suspended monorail train at the current moment; Obtaining the actual speed and position information of the leading train, the rear neighborhood train, and the leader train of a suspended monorail train at the current moment, and calculating the actual interval between a suspended monorail train and the leading train in the front neighborhood, the actual interval between a suspended monorail train and the rear neighborhood train, and the actual interval and actual relative speed between a suspended monorail train and the leader train at the current moment through deviation calculation; Taking the actual interval and actual relative speed between a suspended monorail train and the leader train at the current moment as the input signal of the tracking differentiator, and the tracking differentiator provides a transition signal for the train formation state deviation, and outputs the running interval transition signal and relative speed transition signal at the current moment; Taking the control quantity and actual speed of a suspended monorail train at the current moment as the input of the nonlinear extended state observer, and estimating the total disturbance of the system through state observation; Taking the actual interval between a suspended monorail train and the leading train in the front neighborhood, and the actual interval between a suspended monorail train and the rear neighborhood train at the current moment as the input of the collision avoidance control. If the inputs of the collision avoidance control are both less than the set train repulsive force field radius, a suspended monorail train will be respectively subjected to repulsive forces exerted by the leading train in the front neighborhood and the rear neighborhood train, and a suspended monorail train prevents collisions between trains through the total repulsive force exerted by the leading train in the front neighborhood and the rear neighborhood train.
2. The distributed active disturbance rejection resilient control method for the suspended monorail train facing virtual formation according to claim 1, characterized in that Constructing a distributed dynamic model of the suspended monorail train in a virtual formation using a multi-agent modeling method, specifically including: Consider a multi-agent virtual formation system composed of N trains. Each train is regarded as a multi-agent rigid point mass, numbered from 0 to N-1 in sequence along the running direction. Let the train numbered 0 be the leader train, and the trains numbered 1 to N-1 be the follower trains. The follower trains only receive the state information of the trains before and after in their respective neighborhoods. Here, let E denote the set of all trains, E = {0, 1, …, N-1}; F denote the set of follower trains, F = {1, 2, 3, …, N-1}; G denote the set of neighboring trains of train numbered i, G = {i-1, i+1}; and the trains are maglev trains. Adopt a multi-agent modeling method to construct the distributed dynamic model of the maglev trains in the multi-agent virtual formation system as follows: where x i (t) represents the position of train No. i at time t, represents the differential of x i (t), v i (t) represents the speed of train No. i at time t, represents the differential of v i (t), f(v i , t) represents the resistance force suffered by train No. i at time t, b represents the controller coefficient, u i (t) represents the traction / braking force of train No. i at time t, represents the external environmental interference suffered by train No. i at time t, α v v i (t) represents the known system linear term, α v represents the system linear term coefficient, represents the uncertain nonlinear internal interference of system parameter perturbation, σ i (v i (t), t) represents the total interference suffered by the system at time t.
3. The distributed active disturbance rejection elastic control method for an empty rail train facing virtual formation according to claim 1, characterized in that The formation state error model is as follows: where, e i,j represents the formation state error between the train numbered i and the neighboring train numbered j, represents the expected relative state between the train numbered i and the neighboring train numbered j, d i,j represents the actual relative state between the train numbered i and the neighboring train numbered j, e i,0 represents the formation state error between the train numbered i and the leader train, represents the expected relative state between the train numbered i and the leader train, d i,0 represents the actual relative state between the train numbered i and the leader train, || represents the absolute value symbol, F represents the set of follower trains, F = {1, 2, 3, …, N - 1}; G represents the set of neighboring trains of the train numbered i, G = {i - 1, i + 1}, j = i - 1 or j = i + 1; In the formula, represents the expected interval between train No. i and the neighboring train No. j, h is the train length, and S t represents the safety margin at time t, S brake (v i ) represents the braking distance when the speed of train No. i is v i ; S brake (v j ) represents the braking distance when the speed of the neighboring train No. j is v j ; x i,j represents the actual interval between train No. i and the neighboring train No. j, represents the expected relative speed between train No. i and the neighboring train No. j, v i,j represents the actual relative speed between train No. i and the neighboring train No. j, x i represents the current position of train No. i, x j represents the current position of the neighboring train No. j, v i represents the current speed of train No. i, v j represents the current speed of the neighboring train No. j, and the superscript T represents the transpose; In the formula, represents the expected interval between the train numbered i and the leading train, and x i,0 represents the actual interval between the train numbered i and the leading train, represents the expected relative speed between the train numbered i and the leading train, and v i,0 represents the actual relative speed between the train numbered i and the leading train. x0 represents the current position of the leading train, and v0 represents the current speed of the leading train.
4. The distributed active disturbance rejection elastic control method for the suspended monorail train facing virtual formation according to claim 1, characterized in that According to the distributed dynamic model and the formation state error model, use the tanh function to smooth the formation state error between adjacent trains, and establish a cooperation and collision avoidance control protocol based on a two-way - leader communication topology structure, specifically including: Design a tracking differentiator to enable the train to accelerate or decelerate smoothly as follows: where d i,0 (t) represents the actual relative state of train No. i and the leading train at time t, and d i,0 (t)=[x i,0 (t), v i,0 (t)] T , x i,0 (t) represents the actual headway between train No. i and the leading train at time t, and v i,0 (t) represents the actual relative speed between train No. i and the leading train at time t; represents the step input signal when d i,0 (t) jumps to , represents the step input signal when x i,0 (t) jumps to , represents the desired headway between train No. i and the leading train at time t, represents the step input signal when v i,0 (t) jumps to , represents the desired relative speed between train No. i and the leading train at time t; represents the desired relative state between train No. i and the leading train at time t, represents transition signal, represents the running headway transition signal at time t, represents the relative speed transition signal at time t; and both represent differential signal, represents differential signal; r0 and r1 both represent the transition speed of the signal, and h0 and h1 both represent the terms that filter the noise of the signal; fhan() represents the fhan function; In order to observe all unknown terms including internal disturbances and external disturbances, design a nonlinear extended state observer based on a high-order sliding mode differentiator as follows: Wherein, ψ1(χ1) represents the estimated value of the train speed χ1, and χ1 = v i , v i represents the current speed of train No. i; ψ2(χ1) represents the estimated value of the total system disturbance χ2, and χ2 = σ i (v i ), σ i (v i ) represents the total disturbance suffered by the system; e1 represents the error of the train speed, and e2 represents the error of the total system disturbance; represents the differential of ψ1(χ1), represents the differential of ψ2(χ1); η1(χ1) represents the internal state of the nonlinear extended observer; f(v i ) represents the resistance suffered by train No. i; b represents the controller coefficient, and u i represents the traction / braking force of train No. i; a1, b1, q1, and p1 are all adjustable observer parameters, with a1 > 0, b1 > 0, 0 < q1 < p1; sgn() represents the sign function; || represents the absolute value symbol; Introduce an improved repulsive force field function as follows: In the formula, represents the improved repulsive force field of the train numbered j in the neighborhood relative to the train numbered i, and ρ j represents the radius of the repulsive force field of the train numbered j in the neighborhood, and r j represents the actual size radius of the train numbered j in the neighborhood, μ represents the repulsive force scale factor, and ρ i,j represents the actual interval between the train numbered i and the train numbered j in the neighborhood, and Δρ i,j represents ρ i,j and r j deviation, and Δρ represents ρ j and r j deviation, represents the expected interval between the train numbered i and the train numbered j in the neighborhood, represents ρ i,j and the absolute value of the deviation, F represents the set of follower trains, and G represents the set of neighboring trains of the train numbered i; According to the improved repulsive force field function, obtain the repulsive force exerted on each train by a neighboring train as follows: Wherein, represents the repulsive force generated by the train numbered j in the neighborhood on the train numbered i, i ∈ F, j ∈ G, n j→i represents the unit repulsive force vector from the train numbered j in the neighborhood to the train numbered i, n i→B represents the unit vector that causes the train numbered i to move in the direction B of the desired formation, represents the train numbered i, n j→i the repulsive force received in the direction, represents the train numbered i, n i→B the repulsive force received in the direction; According to the repulsive force exerted on each train by a neighboring train, the total repulsive force exerted on each train by the neighboring trains in front and behind is determined as follows: where F i rep represents the total repulsive force exerted on the train numbered i by the neighboring trains in front and behind, i ∈ F; represents the repulsive force exerted on the train numbered i by the neighboring train in front, represents the repulsive force exerted on the train numbered i by the neighboring train behind; Construct the collision avoidance control by combining the repulsive force exerted on each train by a neighboring train and the total repulsive force exerted on each train by the front and rear neighboring trains. The collision avoidance control, the tracking differentiator, and the nonlinear extended state observer together constitute a cooperation and collision avoidance control protocol based on a two-way - leader communication topology structure.
5. The distributed active disturbance rejection elastic control method for the suspended monorail train facing virtual formation according to claim 4, wherein The distributed active disturbance rejection controller is as follows: where, e i,0 represents the formation state error between the train numbered i and the leader train, and e i,j represents the formation state error between the train numbered i and the neighboring trains, represents the transition signal, represents the desired relative state between the train numbered i and the neighboring train in front, i ∈ F; ξ i represents the state vector of the train numbered i, and ξ i = [x i , v i T , where x i represents the current position of the train numbered i, and v i represents the current speed of the train numbered i; ξ0 represents the state vector of the leader train, and ξ j represents the state vector of the neighboring trains; represents the cooperative part of the train numbered i, and k, and both represent the adjustable coefficient matrices of the cooperative and collision avoidance control protocols, and ψ2 represents the estimated value of the total system disturbance; k = [k1, k2], k1 represents the component in the position direction of k, and k2 represents the component in the speed direction of k, represents the component in the position direction of represents the component in the speed direction of represents the component in the position direction of represents the component in the speed direction of i represents the traction / braking force of the train numbered i, and F i rep represents the total repulsive force exerted on the train numbered i by the neighboring trains in front and behind, t represents the moment of t, and tanh() represents the hyperbolic tangent function. 6. The distributed active disturbance rejection elastic control method for the aerial rail train facing virtual formation according to claim 1, characterized in that The adaptive mutation mechanism of mayflies based on Steffensen is specifically as follows: Let the ratio of the number of weight nodes adjacent to the current optimal solution of mayflies to the total number of weight nodes be used as the evaluation index of mayfly population diversity; If the diversity evaluation index is less than the preset threshold, the Steffensen method is adopted and, according to the formula mutate the population of mayflies; in the formula, X 1,k represents the k-th dimension of the individual X1 to be mutated, and X 1,k ' is the value of X 1,k after mutation; θ(X 1,k ) represents the deviation between the distance between X1 and Z1 and the expected distance between X 1,k ' and Z1, θ(X 1,k ) = L(X1, Z1) - l * , L(X1, Z1) represents the distance between the mayfly individual X1 to be mutated and the weight node Z1, and l * represents the expected distance between X 1,k ' and Z1, l * = α·L(X1, Z1), α ∈ [0, 1], and α represents a random number in the interval [0, 1].
7. The distributed active disturbance rejection elastic control method for the suspended monorail train facing virtual formation according to claim 6, characterized in that Use the adaptive mayfly optimization method to optimize the hyperparameters of the deep deterministic policy gradient algorithm, specifically including: Initialize the mayfly population and generate a set of weight nodes uniformly distributed in the target space; Use the initialized mayfly population to assign values to the hyperparameters to be optimized in the deep deterministic policy gradient algorithm; the hyperparameters to be optimized include the noise standard deviation, the noise standard deviation attenuation rate, the learning rate of the Actor network, and the learning rate of the Critic network; Calculate the diversity evaluation index of the current mayfly population; If the diversity evaluation index of the current mayfly population is less than the preset threshold, then based on the weight nodes, use the Steffensen method to mutate each mayfly in the current population, and select the better ones from the mutated mayfly population and the original mayfly population to recombine into a new mutated mayfly population with the same number as the original population; Divide the new mutated mayfly population into half females and half males, calculate the fitness values of the mayfly individuals included in the new mutated population according to the fitness function based on the ITAE index, and select the individual optimal mayfly and the global optimal mayfly from all male mayflies; According to the individual optimal mayfly and the global optimal mayfly selected so far, use the male mayfly state update formula to update the position and velocity of the male mayfly, and update the individual optimal mayfly and the global optimal mayfly; use the female mayfly state update formula to update the position and velocity of the female mayfly, and calculate the fitness values of all mayflies in the mayfly population; Each pair of female and male mayflies in the mayfly population mate, and two offspring are generated according to the mating and reproduction behavior formula. Select excellent individuals with the same number as the original population from the mated mayflies and retain them for the next generation; Detect whether the end condition is reached; if so, output the optimal hyperparameters; if not, replace the initialized mayfly population with the latest mutated population, and return to the step "Assign the hyperparameters to be optimized in the deep deterministic policy gradient algorithm using the initialized mayfly population".
8. The distributed active disturbance rejection elastic control method for the virtual formation-oriented sky train according to claim 3, characterized in that According to the expected running relative speed and expected running interval between adjacent follower trains and the leader train in the virtual formation, determine the formation state error model of the virtual formation. After that, it also includes: Make the following assumptions: Assumption 1: Assume that each train obtains the actual speed and position information of the front and rear trains and the leader train in the neighborhood in real time through vehicle-to-vehicle wireless communication; Hypothesis 2: Assume that the initial moment is \(t_0\), and the initial state deviation between trains in the neighborhood satisfies \(\|\xi i (t_0)-\xi j (t_0)\|\leq\delta_1, j\in G\), and the external disturbance on the system satisfies \(\|\sigma i (t)\|\leq\delta_2\); where \(\delta_1\) and \(\delta_2\) are both bounded positive constants, \(\xi i (t_0)\) represents the state vector of train numbered \(i\) at the initial moment \(t_0\), \(\xi j (t_0)\) represents the state vector of the neighboring train numbered \(j\) at the initial moment \(t_0\), and \(\sigma i (t)\) represents the external disturbance on the system at time \(t\); \(\|\ \|\) represents the norm. Assumption 3: Assume that the actual speed of the leader train is consistent with the target speed trajectory; The ultimate control objective of multi-train virtual formation cooperative tracking is expressed as: Among them, and are both extremely small positive values, and v0 represents the speed of train numbered i at the initial moment.
9. The distributed active disturbance rejection elastic control method for the suspended monorail train facing virtual formation according to claim 1, characterized in that Control the empty rail trains in the virtual formation so that multiple empty rail trains in the virtual formation maintain the expected formation under the action of interference. Specifically, it includes: Based on the state information of a said empty rail train, the front neighborhood train, the rear neighborhood train, and the leader train at the current moment, take the running interval transition signal and relative speed transition signal output by the tracking differentiator, the estimated value of the total system disturbance output by the nonlinear extended state observer, and the total repulsive force signal output by the collision avoidance control as the current input of the distributed active disturbance rejection controller, and calculate the output control quantity of the distributed active disturbance rejection controller for a said empty rail train at the next moment; the state information includes actual speed and position information; Use the output control quantity at the next moment to drive a said empty rail train to track and run collaboratively with the front neighborhood train, and obtain the actual speed and position information of a said empty rail train at the next moment; Update the current moment to the next moment, and return to the step "Obtain the actual speed and position information of the front neighborhood train, the rear neighborhood train, and the leader train of a said empty rail train at the current moment respectively, and calculate the actual interval between a said empty rail train and the front neighborhood train, the actual interval between a said empty rail train and the rear neighborhood train, and the actual interval and actual relative speed between a said empty rail train and the leader train at the current moment through deviation calculation".
Citation Information
Patent Citations
Virtual marshalling train reference curve calculation method based on improved reinforcement learning algorithm
CN116090336A
Virtual coupling high-speed train group cooperative operation control method
CN116279690A