Safety reinforcement learning control method and device for maglev train suspension system and medium

By constructing a suspension system dynamic model that considers multiple external interferences, and defining the candidate function of Lyapunov and stability constraints, the problems of instability and insufficient safety of the suspension system in the prior art are solved, and the progressive stability and safety guarantee of the suspension system are achieved.

CN120215280AActive Publication Date: 2025-06-27TONGJI UNIV

Patent Information

Application Number
CN202510667687.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-27
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing maglev train suspension system control method cannot theoretically ensure the stability of the control system, and does not consider the safety restrictions of suspended air gaps, which may lead to accidents such as "rail smashing".

Method used

By constructing a dynamic model of the suspension system that considers multiple external interference, time delay and orbital unevenness, defining the Liyapunov candidate function and stability constraints, modeling the control problem of the suspension system as a constraint optimization problem, and iteratively solving the objective function through the gradient descent method, and designing the cost function in combination with the barrier function to ensure the safety of the system.

Benefits of technology

The gradual stability of the suspension system in the face of complex disturbances is achieved, and the state of the system is effectively limited to the feasible area, ensuring the safety of the suspension system and the stability of the control system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215280A_ABST
    Figure CN120215280A_ABST
Patent Text Reader

Abstract

The invention discloses a safety reinforcement learning control method and device for a maglev train suspension system and a medium, and relates to the technical field of maglev control, and the method comprises the following steps: S1, constructing a suspension system dynamic model as a training environment; s2, modeling a control problem of the suspension system as a constraint optimization problem; s3, converting a constrained optimization problem into an unconstrained optimization problem, and performing iterative solution on the target function; s4, constructing a cost function, accelerating solution of an unconstrained optimization problem, and designing a penalty term to guide a system state to be away from a security boundary; and S5, in a reinforcement learning environment, combining a cost function, and carrying out iterative training solution on the unconstrained optimization problem. The method does not depend on an accurate suspension system model, an independent control strategy does not need to be designed for each specific challenge, the optimal control strategy can be adaptively learned through interaction with the suspension system, and the method has high robustness and flexibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of maglev control, and particularly to a safety reinforcement learning control method, device and medium for a maglev train suspension system. Background Art

[0002] Maglev trains achieve contactless high-speed operation through the electromagnetic force between the vehicle and the track, which can fill the speed gap between high-speed railways (300 km / h) and air transportation (800 km / h), and is of great significance for quickly connecting major urban agglomerations and promoting cross-regional integration and development. Due to its magnetic circuit characteristics, the suspension system is a typical strongly nonlinear and open-loop unstable system, and a controller needs to be designed to apply active control. However, the working environment of the suspension system is extremely complex, affected by external disturbances caused by track irregularities and impacts caused by aerodynamic loads, and there are bounded constraints on the suspension air gap, with the rated air gap being only about 10 mm. At the same time, the control current has unidirectional and saturation constraints. In addition, the induced electromotive force generated by the inductance of the suspension electromagnet coil hinders the rapid change of the control current, resulting in a large lag in the adjustment of the electromagnetic force, further exacerbating the control difficulty. Especially at speeds of 600 km / h or higher, the suspension system faces more stringent requirements, and the performance of the suspension controller directly determines the safety and stability of the train operation. Developing a suspension controller with high adaptive ability and strong robustness is an important requirement and a major challenge.

[0003] Chinese Patent Application with Publication No. CN118790056A discloses a deep reinforcement learning collaborative control method for a maglev train suspension system. This method first constructs a deep reinforcement learning environment for the unilateral suspension frame suspension system based on the system dynamics model, then designs a reinforcement learning algorithm and an initial reinforcement learning agent, and combines them with the deep reinforcement learning environment to form a deep reinforcement learning collaborative control algorithm. Then, the reinforcement learning agent is trained using the deep reinforcement learning environment and the collaborative control algorithm, and finally, the trained reinforcement learning agent controller is obtained to achieve effective control of the unilateral suspension frame suspension system under different disturbance conditions.

[0004] However, the reinforcement learning agent established by the above method directly optimizes the control algorithm by maximizing the cumulative reward, without considering the stability requirements of the control system, and cannot theoretically guarantee the stability of the control system, nor can it avoid the agent from making dangerous actions. In addition, this method does not consider the bounded constraints between the vehicle and the track, which may lead to accidents such as "hitting the track" and cannot ensure the safety of the suspension control system. Summary of the Invention

[0005] The present invention overcomes the deficiencies of the prior art and provides a safety reinforcement learning control method, device and medium for a maglev train suspension system.

[0006] To achieve the above object, the technical solution adopted by the present invention is: a safety reinforcement learning control method for a maglev train suspension system, comprising the following steps: S1. Construct a dynamic model of the suspension system considering various external disturbances, time delays, and track unevenness as the training environment for the reinforcement learning algorithm; S2. Model the control problem of the suspension system as a constrained optimization problem, and define a Lyapunov candidate function and stability constraint conditions; S3. Convert the constrained optimization problem in step S2 into an unconstrained optimization problem, and iteratively solve the objective function by the gradient descent method; S4. Construct a cost function based on the barrier function to accelerate the solution of the unconstrained optimization problem in step S3, and design a penalty term in combination with the physical boundary constraints of the suspension system to guide the system state away from the safety boundary; S5. In the reinforcement learning environment established in step S1, combine the cost function in step S4, and iteratively train and solve the unconstrained optimization problem in step S3 to obtain a reinforcement learning agent for the suspension system with stability guarantee and safety constraints.

[0007] In a preferred embodiment of the present invention, in the step S1, the dynamic model of the suspension system includes the following sub-models: a magnetic circuit equation established based on Maxwell's equations and Biot-Savart's law, a system dynamic model including an electromagnetic force equation, a voltage equation, and a vertical dynamic equation, a time-varying perturbation model for describing track unevenness interference, and an actuator model considering chopper characteristics and time delays.

[0008] In a preferred embodiment of the present invention, the magnetic circuit equation can be obtained based on Maxwell's equations and Biot-Savart's law: ; wherein, is the number of turns of the coil; is the coil current; is the magnetic flux; is the magnetic resistance; is the suspension air gap between the magnetic pole and the track; is the magnetic pole area; is the air permeability; The electromagnetic force equation is: ; wherein, is the controllable electromagnetic force; is the total energy stored in the magnetic field; is the partial derivative of with respect to; is the differential symbol; Based on Kirchhoff's voltage law, the voltage equation can be obtained as follows: ; where is the voltage of the electromagnet coil winding; is the coil resistance; is the coil inductance; Further, by performing a vertical dynamic analysis on the suspension system, the system equation can be obtained; ; where is the suspension mass; is the gravity acting on the suspended object; The influence of the track irregularity on the system can be described as: ; where is the distance between the electromagnet and the track reference plane; is the suspension air gap; is the track irregularity; Design a chopper equation to replace the voltage equation: ; where is the coil current; is the sign function; is the control current, which is given by the control algorithm; is the maximum rate of change of the current, which is determined by the coil inductance and the performance of the chopper; Further considering time delay and other disturbances, the system equation can be rewritten as: ; where is the disturbance force; is the actuator time delay; is the control strategy; is the state feedback time delay.

[0009] In a preferred embodiment of the present invention, in the step of S2, by defining the Lyapunov candidate function, the control problem of the suspension system is constructed as a constrained optimization problem in a constrained Markov decision process: ; ; ; where is to find; is the optimal strategy to be found; is to maximize; is the reward; is the reward at the moment; is constrained to be; and are the current state and the state at the next moment respectively; and are the current action and the action at the next moment respectively; is the experience replay pool; is the reward sampled from the experience replay pool and take the expectation; is the cost; is a bounded positive definite cost function; is a positive definite Lyapunov-like function, described by a neural network with parameters ; ; described; is a constant; The action taken by the agent is generated by an action network with parameters to obtain the mean of the Gaussian distribution and the variance ; through reparameterization: ; wherein, is the action of the agent; is a noise variable sampled from the standard Gaussian distribution ;

[0010] In a preferred embodiment of the present invention, in the step of S3, the objective function of the unconstrained optimization problem is: ; wherein, and are Lagrange multipliers; is the observation value received by the agent in the environment; is the reward sampled from the experience replay pool and take the expectation; is the reward at the moment; is the cost; Considering that ensuring stability can satisfy maximizing the reward, the above formula is simplified to: ; Update through the gradient descent method: ; wherein, is to take the gradient of ; is to calculate the gradient; The Lagrange multiplier is updated by maximizing the following objective: ; where is the expectation obtained from the experience replay pool ; and is restricted to positive values: ; where is the learning rate; is to calculate the gradient; The loss function of the Lyapunov critic is defined as: ; where is the Lyapunov candidate function, serving as the objective function for updating; Furthermore, is selected as the cost sum within the finite time window : ; where is the cost function; is the system state at time

[0011] In a preferred embodiment of the present invention, in the step of S4, the design of the barrier function includes: Define the safety set of the system and the function for measuring whether the system state is within , satisfying: ; ; ; where is the safety boundary; is the n-dimensional Euclidean space; is the set of real numbers; is for measuring the system state; Construct the barrier function, which satisfies that when the system approaches the safety boundary, the value of the barrier function tends to infinity, and the barrier function is: ; where is the safety boundary of the suspension system, and the above formula can meet the requirements of the barrier function: ; ; Among them, is for any system state; Combine the barrier function with the error-sensitive term to form the comprehensive cost function: Increase the cost term , and increase the cost value near the equilibrium point Designed as: ; Among them, is the base number for adjusting the error sensitivity; , , is an adjustable coefficient; Then the total cost is: ; Among them, , , , , , , , are all adjustable coefficients; is the suspension air gap; is the air gap velocity; Take , , , , , , , which is the constructed cost function.

[0012] In a preferred embodiment of the present invention, in the step of S5, the training parameters of the reinforcement learning agent include: The action network and the Lyapunov critic network adopt a multi-layer neural network structure; Set the minimum batch size, the capacity of the experience replay pool, and the length of the finite time window; Screen and save qualified agents according to the average cost threshold.

[0013] In a preferred embodiment of the present invention, the hidden layer structure of the action network and the Lyapunov critic network is two layers, each layer contains 256 neurons, and the activation function adopts the rectified linear unit function.

[0014] The present invention provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor; the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, the processor is enabled to execute the safety reinforcement learning control method of the maglev train levitation system as described in any one of the above.

[0015] The present invention provides a computer-readable storage medium storing computer instructions for enabling a processor to implement the safety reinforcement learning control method of the maglev train levitation system as described in any one of the above when executed.

[0016] The present invention solves the defects existing in the background art and has the following beneficial effects: The present invention provides a safety reinforcement learning control method, device and medium for a maglev train levitation system. By constructing a stability constraint optimization problem of the levitation system based on the Lyapunov function, it is transformed into an unconstrained optimization objective through the Lagrange multiplier method, and combined with the control objective of the levitation system, the unconstrained optimization problem is streamlined, reducing the computational amount while ensuring the stability constraint, not relying on an accurate levitation system model, without the need to design a separate control strategy for each specific challenge, and can adaptively learn the optimal control strategy through interaction with the levitation system, having strong robustness and flexibility.

[0017] In the present invention, based on the Lyapunov stability theory, the constraint conditions of the action network are constructed. Compared with the existing levitation control methods based on reinforcement learning, it avoids the action network from making dangerous actions, theoretically ensuring the asymptotic stability of the levitation system in the face of complex disturbances.

[0018] In the present invention, a barrier function is introduced in the design of the cost function, enabling the unconstrained optimization problem to quickly move away from the physical boundary of the levitation system during training and solution, accelerating the solution speed of the unconstrained optimization problem, while ensuring the safety of the levitation system, effectively limiting the system state within the feasible region. The designed cost function can provide a large gradient near the physical constraint boundary and the equilibrium point of the system, guiding the levitation system to quickly tend to stability.

[0019] In the present invention, by establishing a levitation system model considering various disturbances and time delays, a general chopper model is given, making the training environment of the algorithm closer to the real object, and reducing the difficulty of migrating the trained reinforcement learning agent with stability guarantee and safety constraints from simulation to real vehicle application. Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings; Figure 1 It is a block diagram of the safety reinforcement learning control method for a maglev train suspension system according to a preferred embodiment of the present invention; Figure 2 It is a model diagram of a single-point suspension system according to a preferred embodiment of the present invention; Figure 3 It is the air gap response diagram of the control method of the present invention and the PID control method during static levitation with time delay in working condition 1; Figure 4 It is a partial enlarged view of the air gap response diagram of the control method of the present invention and the PID control method during static levitation with time delay in working condition 1; Figure 5 It is the electromagnet current diagram of the control method of the present invention and the PID control method during static levitation with time delay in working condition 1; Figure 6 It is a partial enlarged view of the electromagnet current diagram of the control method of the present invention and the PID control method during static levitation with time delay in working condition 1; Figure 7 It is the air gap response diagram of the control method of the present invention and the PID control method when there are time delay and impact interference in working condition 2; Figure 8 It is the electromagnet current diagram of the control method of the present invention and the PID control method when there are time delay and impact interference in working condition 2; Figure 9 It is the air gap response diagram of the control method of the present invention and the PID control method when there are time delay and uneven interference in working condition 3; Figure 10 It is the electromagnet current diagram of the control method of the present invention and the PID control method when there are time delay and uneven interference in working condition 3. Detailed implementation manners

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0022] In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited by the specific embodiments disclosed below.

[0023] Overview of the Application: At present, in response to the challenges faced by the levitation system, such as strong nonlinearity, response lag, track irregularity interference, bounded constraints, etc., many scholars have proposed various solutions, such as fuzzy control, sliding mode control, model predictive control, etc. However, the drawback is that the existing methods are usually "patchwork" solutions designed for specific problems, relying on an accurate model of the controlled object, and a targeted control strategy needs to be designed for specific problems, unable to ensure sufficient adaptability.

[0024] Different from traditional control methods, reinforcement learning, as a data-driven method, directly learns the optimal control strategy through interaction with the environment, without the need for separate modeling and controller design for each specific problem. However, the existing reinforcement learning-based levitation control methods directly learn the control strategy through data, lacking strict stability analysis, and directly applying them to the levitation system cannot guarantee the stability of the control system.

[0025] In addition, the existing reinforcement learning-based levitation control strategies do not consider the safety limitations of the levitation air gap, which may lead to accidents such as "rail hitting", and cannot ensure the safety of the levitation control system.

[0026] Therefore, the present invention designs a safety reinforcement learning control method for a maglev train levitation system with stability guarantee and safety constraints. Based on the Lyapunov theory, stability constraints are constructed, and through strict mathematical analysis, it is proved that the designed method can guarantee the stability of the levitation system. In addition, considering physical constraints, a cost function is constructed and a barrier function is introduced to guide the levitation system away from the physical boundary, ensuring the safety of the levitation system.

[0027] As Figure 1 shown, a safety reinforcement learning control method for a maglev train levitation system includes the following steps: S1. Construct a dynamic model of the levitation system considering various external interferences, time delays, and track irregularities as the training environment for the reinforcement learning algorithm; S2. Model the control problem of the levitation system as a constrained optimization problem, and define the Lyapunov candidate function and stability constraint conditions; S3. Convert the constrained optimization problem in step S2 into an unconstrained optimization problem, and iteratively solve the objective function by the gradient descent method; S4. Construct a cost function based on the barrier function to accelerate the solution of the unconstrained optimization problem in step S3, and design a penalty term in combination with the physical boundary constraints of the suspension system to guide the system state away from the safety boundary. S5. In the reinforcement learning environment established in step S1, combined with the cost function in step S4, iteratively train and solve the unconstrained optimization problem in step S3 to obtain a reinforcement learning agent for the suspension system with stability guarantee and safety constraints.

[0028] It should be noted that the present invention constructs a stability constraint optimization problem for the suspension system based on the Lyapunov function, converts it into an unconstrained optimization objective through the Lagrange multiplier method, and simplifies the unconstrained optimization problem in combination with the control objective of the suspension system, reducing the computational complexity while ensuring the stability constraint.

[0029] As Figure 2 shown, in some specific embodiments, in step S1, the suspension system dynamics model includes the following sub-models: a magnetic circuit equation established based on Maxwell's equations and the Biot-Savart law, a system dynamics model including an electromagnetic force equation, a voltage equation, and a vertical dynamics equation, a time-varying perturbation model for describing track irregularity interference, and an actuator model considering chopper characteristics and time delay. Based on Maxwell's equations and the Biot-Savart law, the magnetic circuit equation can be obtained: ; Among them, is the number of coil turns; is the coil current; is the magnetic flux; is the magnetic reluctance; is the suspension air gap between the magnetic pole and the track; is the magnetic pole area; is the air permeability; The electromagnetic force equation is: ; Among them, is the controllable electromagnetic force; is the total energy stored in the magnetic field; is the partial derivative of with respect to; is the differential symbol; Based on Kirchhoff's voltage law, the voltage equation can be obtained: ; Among them, is the voltage of the electromagnet coil winding; is the coil resistance; is the coil inductance; Further vertical dynamic analysis of the suspension system yields the system equation; ; where, is the suspended mass; is the gravity acting on the suspended object; The influence of track irregularity on the system can be described as: ; where, is the distance between the electromagnet and the track reference plane; is the suspension air gap; is the track irregularity; Design the chopper equation to replace the voltage equation: ; where, is the coil current; is the sign function; is the control current, given by the control algorithm; is the maximum rate of change of current, determined by the coil inductance and the chopper performance; Further considering time delay and other disturbances, the system equation can be rewritten as: ; where, is the disturbance force; is the actuator time delay; is the control strategy; is the state feedback time delay.

[0030] It should be noted that by establishing a suspension system model considering various disturbances and time delays, a general chopper model is given, making the training environment of the algorithm closer to the real object and reducing the difficulty of transferring the trained reinforcement learning agent with stability guarantee and safety constraints from simulation to real vehicle application.

[0031] In some specific embodiments, in step S2, by defining a Lyapunov candidate function, the control problem of the suspension system is constructed as a constrained optimization problem in a constrained Markov decision process: ; ; ; where, is to find; is the optimal policy to be found; is to maximize; is the reward; is the reward at time is constrained as; and are the current state and the state at the next moment, respectively; and are the current action and the action at the next moment, respectively; is the experience replay pool; is the reward sampled from the experience replay pool and takes the expectation; is the cost; is a bounded positive definite cost function; is a positive definite Lyapunov-like function, described by a neural network with parameters ; described; is a constant; The action taken by the agent is generated by an action network with parameters to generate the mean of a Gaussian distribution and the variance , and is obtained by reparameterization as: ; where is the action of the agent; is a noise variable sampled from the standard Gaussian distribution ;

[0032] In some specific embodiments, in the step of S3, the objective function of the unconstrained optimization problem is: ; where and are Lagrange multipliers; is the observation received by the agent in the environment; is the reward sampled from the experience replay pool and takes the expectation; is the reward at time is the cost; Considering that ensuring stability can satisfy maximizing the reward, the above formula is simplified to: ; It is updated by the gradient descent method: ; where is the gradient of with respect to; is the gradient of with respect to; The Lagrange multiplier is updated by maximizing the following objective: ; where, is the expectation obtained from the experience replay pool ; and is restricted to be a positive value: ; where, is the learning rate; is to calculate the gradient; The loss function of the Lyapunov critic is defined as: ; where, is the Lyapunov candidate function, serving as the objective function for the update of ; Furthermore, is selected as the cost sum within the finite time window : ; where, is the cost function; is the system state at time

[0033] In some specific embodiments, in the step of S4, the design of the barrier function includes: defining the safety set of the system and the function for measuring whether the system state is within , satisfying: ; ; ; where, is the safety boundary; is the n-dimensional Euclidean space; is the set of real numbers; is to measure the system state; constructing a barrier function, which satisfies that when the system approaches the safety boundary, the value of the barrier function tends to infinity, and the barrier function is: ; where, is the safety boundary of the suspension system, and the above formula can meet the requirements of the barrier function: ; ; Among them, is for any system state; Combine the barrier function with the error-sensitive term to form a comprehensive cost function: Increase the cost term , and increase the cost value near the equilibrium point Designed as: ; Among them, is the base for adjusting the error sensitivity; , , is an adjustable coefficient; Then the total cost is: ; Among them, , , , , , , , are all adjustable coefficients; is the suspension air gap; is the air gap velocity; Take , , , , , , , which is the constructed cost function.

[0034] It should be noted that by constructing the safety barrier function of the suspension system, the unconstrained optimization problem can be quickly moved away from the physical boundary of the suspension system during training and solution, which speeds up the solution speed of the unconstrained optimization problem and at the same time ensures the safety of the suspension system.

[0035] In some specific implementation manners, in the step of S5, the training parameters of the reinforcement learning agent include: The action network and the Lyapunov critic network adopt a multi-layer neural network structure; Set the minimum batch size, the capacity of the experience replay pool, and the length of the finite time window; Screen and save qualified agents according to the average cost threshold; Among them, the hidden layer structure of the action network and the Lyapunov critic network is two layers, each layer contains 256 neurons, and the activation function uses the rectified linear unit.

[0036] Specifically, the maximum number of training steps per round is designed to be 1000, and an evaluation and saving is performed every 5000 steps. The simulation step size is 0.001 s, and the finite time window is 30. Both the action network and the Lyapunov critic network are designed with two hidden layers, each with 256 neurons, and the rectified linear activation function is used. The learning rate of the action network is , and the learning rate of the Lyapunov critic network is . The minimum batch size is designed to be 256, and the maximum experience replay pool is . Agents with an average cost less than 50 are saved, and a reinforcement learning agent for a maglev train suspension system with stability guarantee and safety constraints can be obtained.

[0037] The present invention provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor; the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, the processor can execute the safety reinforcement learning control method of the maglev train suspension system as described in any one of the above.

[0038] The present invention provides a computer-readable storage medium, which stores computer instructions for causing a processor to implement the safety reinforcement learning control method of the maglev train suspension system as described in any one of the above when executed.

[0039] To further illustrate the present invention, in a numerical simulation manner, the safety reinforcement learning control method of the present invention is compared with the proportional-integral-differential (PID) control algorithm currently used in actual vehicles under three different working conditions, and quantitative analysis is performed through the root mean square error (RMSE) and the maximum error. The parameters of the PID are selected as the proportional gain , the integral gain , and the derivative gain .

[0040] Specifically, the three working conditions are respectively: Working condition 1: Static levitation, with a state time delay of 2 ms and an actuator time delay of 3 ms, a maximum chopper adjustment time of 30 ms, and no interference.

[0041] Working condition 2: With a state time delay of 2 ms and an actuator time delay of 3 ms, a maximum chopper adjustment time of 30 ms, and respectively at being subjected to impact interference.

[0042] Working condition 3: With a state time delay of 2 ms and an actuator time delay of 3 ms, a maximum chopper adjustment time of 30 ms, and interference from track irregularities and random noise. The track irregularities are simulated by a sine wave, and the expression is: ; Among them, is the amplitude of long - wave irregularity; is the train running speed; is the bridge span.

[0043] In condition 1, the suspension air - gap response is as shown in Figure 3 and Figure 4 . In the adjustment stage of the first 0.3 s, affected by time - delay, the PID control shows obvious oscillation, and it basically reaches stability after 0.44 s. The proposed method is not significantly affected by time - delay and is more stable than the PID control. The comparison of the electromagnet current is as shown in Figure 5 and Figure 6 . The electromagnet current of the PID control shows a large - range fluctuation at about 0.05 s, and there is a gradually decaying oscillation within the subsequent 3 s. The proposed method is more stable, and the jitter can be completely eliminated after about 0.23 s, and the electromagnet current stabilizes at 21.8 A.

[0044] In condition 2, the suspension air - gap response is as shown in Figure 7 . When , the interference is small, and both the PID and the proposed method can better resist the impact. When , the interference increases. The PID control deviates more from the equilibrium point than the proposed method, reaching a maximum of 0.64 mm, while the proposed method is only 0.32 mm. And after the interference disappears, the PID control takes longer to return to the equilibrium state. When , the interference reaches the maximum. The maximum error of the suspension air - gap under the PID control reaches 1.24 mm, and the maximum error of the proposed method is only 0.46 mm, which is 63.2% less than that of the PID control. When the interference disappears, as shown in the electromagnet current in Figure 8 , the PID takes longer to return to the equilibrium state than the proposed method. In summary, in condition 2, the proposed method has stronger robustness than the PID control.

[0045] In condition 3, the influence of track irregularity on the suspension air - gap is as shown in Figure 9 . The suspension air - gap fluctuation of the proposed method is smaller than that of the PID control. The RMSE and the maximum error of the proposed method are reduced by 14.1% and 49.9% respectively compared with the PID, and it is less affected by the irregularity. It can be clearly seen from the electromagnet current shown in Figure 10 that the fluctuation of the proposed method is much smaller than that of the PID, and the response is more stable.

[0046] In summary, the safety reinforcement learning control method for the maglev train suspension system of the present invention does not rely on an accurate suspension system model, does not require designing a separate control strategy for each specific challenge, can adaptively learn the optimal control strategy through interaction with the suspension system, and has strong robustness and flexibility. Moreover, compared with the existing suspension control methods based on reinforcement learning, the present invention constructs the constraint conditions of the action network based on the Lyapunov stability theory, avoids the action network from making dangerous actions, and theoretically ensures the asymptotic stability of the suspension system in the face of complex disturbances. At the same time, the present invention introduces a barrier function in the design of the cost function, effectively limits the state of the system within the feasible region, and the designed cost function can provide a large gradient when the system is near the physical constraint boundary and the equilibrium point, guiding the suspension system to quickly tend to stability.

[0047] Based on the inspiration of the ideal embodiments of the present invention as described above, through the above description, for those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

[0048] In addition, it should be understood that although this specification is described according to the embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A safety reinforcement learning control method for a maglev train suspension system, characterized in that It includes the following steps: S1. Construct a suspension system dynamics model considering various external interferences, time delays, and track irregularities as the training environment for the reinforcement learning algorithm; S2. Model the control problem of the suspension system as a constrained optimization problem, and define the Lyapunov candidate function and stability constraint conditions; S3. Convert the constrained optimization problem in step S2 into an unconstrained optimization problem, and iteratively solve the objective function by the gradient descent method; S4. Construct a cost function based on the barrier function to accelerate the solution of the unconstrained optimization problem in step S3, and combine the physical boundary constraints of the suspension system to design a penalty term to guide the system state away from the safety boundary; S5. In the reinforcement learning environment established in step S1, combine the cost function in step S4, and iteratively train and solve the unconstrained optimization problem in step S3 to obtain a reinforcement learning agent for the suspension system with stability guarantee and safety constraints.

2. The safety reinforcement learning control method for a maglev train suspension system according to claim 1, wherein: In the step S1, the suspension system dynamics model includes the following sub-models: a magnetic circuit equation established based on Maxwell's equations and Biot-Savart's law, a system dynamics model including an electromagnetic force equation, a voltage equation, and a vertical dynamics equation, a time-varying perturbation model for describing track irregularity interference, and an actuator model considering chopper characteristics and time delays.

3. The safety reinforcement learning control method for a maglev train suspension system according to claim 2, characterized in that: The magnetic circuit equation can be obtained based on Maxwell's equations and Biot-Savart's law: ; Among them, is the number of turns of the coil; is the coil current; is the magnetic flux; is the magnetic reluctance; is the suspension air gap between the magnetic pole and the track; is the magnetic pole area; is the air permeability; The electromagnetic force equation is: ; Among them, is a controllable electromagnetic force; is the total energy stored in the magnetic field; is the partial derivative with respect to ; is the differential symbol; The voltage equation can be obtained based on Kirchhoff's voltage law: ; Among them, is the voltage of the electromagnet coil winding; is the coil resistance; is the coil inductance; Further perform vertical dynamics analysis on the suspension system to obtain the system equation; ; Among them, is the suspended mass; is the gravity acting on the suspended object; The influence of the track irregularity on the system can be described as: ; Among them, is the distance between the electromagnet and the track reference surface; is the suspension air gap; is the track irregularity; Design a chopper equation to replace the voltage equation: ; Among them, is the coil current; is the sign function; is the control current, which is given by the control algorithm; is the maximum rate of change of current, which is determined by the coil inductance and the performance of the chopper; Further considering time delays and other perturbations, the system equation can be rewritten as: ; Among them, is the disturbance force; is the actuator delay; is the control strategy; is the state feedback delay.

4. The safety reinforcement learning control method for a maglev train suspension system according to claim 1, characterized in that: In the step S2, by defining the Lyapunov candidate function, the control problem of the suspension system is constructed as a constrained optimization problem in a constrained Markov decision process: ; ; ; Among them, is to search for; is the optimal strategy to be searched for; is to maximize; is the reward; is the reward at time is constrained to be; and are the current state and the state at the next time respectively; and are the current action and the action at the next time respectively; is the experience replay pool; is the reward sampled from the experience replay pool ; to find the expectation; is the cost; is a bounded positive definite cost function; is a positive definite Lyapunov-like function, described by a neural network with parameters ; described by; is a constant; The action made by the agent passes through an action network with a parameter of to generate the mean of a Gaussian distribution and the variance , and through reparameterization, we get: ​ ; Among them, is the action of the agent; is the noise variable sampled from the standard Gaussian distribution among them.

5. The safety reinforcement learning control method for a maglev train suspension system according to claim 1, characterized in that: In the step S3, the objective function of the unconstrained optimization problem is: ; Among them, and are Lagrange multipliers; is the observation received by the agent in the environment; is the reward sampled from the experience replay pool and taking the expectation; is the reward at time and is the cost; Considering that ensuring stability can satisfy maximizing the reward, the above formula is simplified to: ; Update by the gradient descent method: ; Among them, is to calculate the gradient; is to calculate the gradient; The Lagrange multiplier is updated by maximizing the following objective: ; Among them, is the expectation obtained from the experience replay pool ; And is restricted to be positive: ; Among them, is the learning rate; is to calculate the gradient; The loss function of the Lyapunov critic is defined as: ; Among them, is the Lyapunov candidate function, serving as the updated objective function; Further, is selected as the cost sum within a finite time window: ; Among them, is the cost function; is the system state at time 6. The safety reinforcement learning control method for a maglev train suspension system according to claim 1, characterized in that: In the step S4, the design of the barrier function includes: Define the security set of the system and measure the system state Whether it is within the function within , satisfying: ; ; ; Among them, is the safety margin; is an n-dimensional Euclidean space; is the set of real numbers; is the metric system state; Construct the barrier function, which satisfies that when the system approaches the safety boundary, the value of the barrier function tends to infinity. The barrier function is as follows: ; Among them, is the safety boundary of the suspension system, and the above formula can meet the requirements of the barrier function: ; ; wherein, is for any system state; Combine the barrier function with the error-sensitive term to form the comprehensive cost function: Increase the cost item , and increase the cost value near the equilibrium point Designed as: ; Among them, is the base number for adjusting the error sensitivity; , , are adjustable coefficients; The total cost is as follows: ; Among them, , , , , , , , are all adjustable coefficients; is the suspension air gap; is the air gap velocity; Take , , , , , , , which is the constructed cost function.

7. The safety reinforcement learning control method for a maglev train suspension system according to claim 1, characterized in that: In the step S5, the training parameters of the reinforcement learning agent include: The action network and the Lyapunov critic network adopt a multi-layer neural network structure; Set the minimum batch size, the capacity of the experience replay pool, and the length of the finite time window; Screen and save qualified agents according to the average cost threshold.

8. The safety reinforcement learning control method for a maglev train suspension system according to claim 7, characterized in that: The hidden layer structure of the action network and the Lyapunov critic network is two layers, each layer contains 256 neurons, and the activation function adopts the rectified linear unit function.

9. An electronic device, characterized in that: The electronic device includes: at least one processor, and a memory communicatively connected to the at least one processor; the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, the processor is enabled to execute the safety reinforcement learning control method of the maglev train levitation system according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions for implementing the safety reinforcement learning control method of the maglev train levitation system according to any one of claims 1-8 when the computer instructions are executed by a processor.

Citation Information

Patent Citations

  • Levitation system control method for maglev train

    CN111806246A

  • Method for obtaining maglev train suspension system controller based on deep reinforcement learning

    CN116520721A

  • Tracking operation characteristic modeling method for unmanned driving of magnetically levitated train

    CN116880217A

  • Interference suppression suspension control method and system for magnetically levitated train

    CN118534771A

  • Deep reinforcement learning cooperative control method for EMS type maglev train suspension frame

    CN118790056A

Cited By

  • Multi-point collaborative reinforcement learning compensation control method, system and device for magnetic levitation system based on UAV-InSAR and medium

    CN122008890A