UAV-InSAR-based maglev system multi-point collaborative reinforcement learning compensation control method, system, device and medium
By using UAV mapping and a hierarchical control system, combined with a reinforcement learning compensator, the complex disturbance problem of the high-speed maglev train's suspension system was solved, achieving high-precision monitoring and global stable control, and improving the stability and anti-disturbance capability of the suspension system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TONGJI UNIV
- Filing Date
- 2026-04-10
- Publication Date
- 2026-06-23
AI Technical Summary
Faced with complex and ever-changing external disturbances and stringent constraints, traditional control methods struggle to achieve real-time monitoring with high spatiotemporal resolution and wide coverage, limiting the performance improvement of the suspension control system. Furthermore, reinforcement learning methods cannot guarantee stability and are prone to causing instability accidents such as track collisions.
A lightweight X-band interferometric synthetic aperture radar system equipped with an unmanned aerial vehicle (UAV) is used for high-precision mapping. A power spectrum model of track irregularities is constructed. Combined with a hierarchical control system, the lower layer is an amplitude saturation controller and the upper layer is a compensator based on reinforcement learning. The compensation control quantity is trained through the power spectrum model of track irregularities to generate the total electromagnetic force control law, thereby realizing the active collaborative compensation and global stability control of the suspension system.
It achieves quantitative modeling and spectral characterization of external disturbances, suppresses actuator saturation under strong disturbance conditions, ensures the global stability and self-learning ability of the suspension system, and improves the stability and disturbance rejection of the suspension system. Simulation verification shows that the error and control effect are better than traditional methods.
Smart Images

Figure CN122008890B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of high-speed maglev suspension technology, and in particular to a multi-point cooperative reinforcement learning compensation control method, system, equipment and medium for maglev systems based on UAV-InSAR. Background Technology
[0002] High-speed maglev trains rely on electromagnetic force to achieve contactless high-speed operation, combining speed advantages with environmental friendliness. They hold irreplaceable strategic significance for building efficient and convenient rapid transit networks in urban clusters and promoting regional coordinated development. The suspension system is the core subsystem of the maglev train. It can actively control and adjust the magnitude of the electromagnetic force in real time to maintain the suspension air gap stable near its rated value. Its performance directly determines the safety and stability of train operation. However, the suspension system is inherently a complex system with strong nonlinearity and open-loop instability, facing complex and ever-changing external disturbances and a series of stringent constraints during operation. In particular, the combined disturbances of multi-point coupling effects and track irregularities, the saturation constraint of electromagnet output, and the inherent hysteresis characteristics of the chopper become key factors restricting the performance improvement of the suspension control system.
[0003] The geometric irregularities of maglev lines are key physical quantities affecting vehicle dynamic response, levitation control robustness, and operational smoothness. High-precision, continuous, and spatially distributed characterization of these irregularities is crucial for modern control strategy design. While traditional point measurement methods offer high accuracy, their low coverage efficiency and reliance on on-site maintenance make them insufficient for meeting the demands of high spatiotemporal resolution and wide coverage in real-time monitoring of long-distance maglev systems. In recent years, unmanned aerial vehicles (UAVs) have gained widespread attention as a highly flexible and low-cost mobile sensing platform for monitoring track and railway infrastructure. UAV-InSAR research demonstrates that, under high-precision deformation monitoring conditions, interferometric imaging with related systems mounted on UAVs possesses the potential for continuous acquisition and high-resolution inversion of local track disturbances.
[0004] The inversion results of high-precision track irregularity spectra not only enable the establishment of realistic external disturbance models for maglev systems, but also provide input data for the training and verification of intelligent control algorithms. This irregularity modeling process based on UAV-InSAR constitutes a key link in environmental modeling within the reinforcement learning control framework, completing the important task of external disturbance identification. However, the aforementioned method of directly applying reinforcement learning cannot theoretically guarantee the stability of the levitation system, and in practical applications, it is difficult to avoid instability accidents such as track-vehicle collisions. Summary of the Invention
[0005] In view of this, the present invention provides a multi-point cooperative reinforcement learning compensation control method, system, device and medium for maglev systems based on UAV-InSAR to solve the above problems.
[0006] This invention provides a multi-point cooperative reinforcement learning compensation control method for a maglev system based on UAV-InSAR, comprising: using a UAV equipped with a lightweight X-band interferometric synthetic aperture radar system to perform high-precision mapping of the maglev line; constructing a track irregularity power spectrum model based on interferometric phase inversion; constructing a hierarchical control system, with an amplitude saturation controller set at each suspension point at the lower layer and a compensator based on reinforcement learning at the upper layer; the amplitude saturation controller generating a basic electromagnetic force control law based on the current state variables of the suspension point; the reinforcement learning-based compensator acting as an intelligent agent interacting with the environment formed by the suspension frame, track, and amplitude saturation controller, being trained using a standardized irregularity function obtained from the track irregularity power spectrum model, and outputting compensation control quantities for each suspension point; superimposing the basic electromagnetic force control law with the compensation control quantities to generate a total electromagnetic force control law for each suspension point, thereby performing active cooperative compensation and global stable suspension control of the suspension system.
[0007] In another implementation of the present invention, the step of using a UAV equipped with a lightweight X-band interferometric synthetic aperture radar system to perform high-precision mapping of the maglev line and constructing a track irregularity power spectrum model based on interferometric phase inversion includes: using a UAV equipped with a lightweight X-band interferometric synthetic aperture radar system to perform high-precision mapping of the maglev line; determining the interferometric phase according to the geometric relationship between the main track and the secondary track during the mapping process; performing inversion through the interferometric phase to obtain the elevation disturbance in the track normal direction; and establishing a track irregularity power spectrum model based on the elevation disturbance field.
[0008] In another implementation of the present invention, the expression for the elevation disturbance in the orbital normal direction is:
[0009]
[0010] in, These are the azimuth coordinates. For distance coordinates, The elevation change is obtained from interferometric phase inversion. The angle of incidence is denoted as .
[0011] In another implementation of the present invention, the expression for the orbital irregularity power spectrum model is:
[0012]
[0013] in, For wave number, The imaginary unit, These are the azimuth coordinates. The wavelength represents the non-uniform spatial wavelength; L represents the observation length. The normal elevation disturbance field of the benchmark orbit.
[0014] In another implementation of the present invention, the expression for the standardized non-roughness function is:
[0015]
[0016] in, For random phase, satisfying , For the first Discrete wavenumber points, For the first The interval of each wavenumber interval.
[0017] In another implementation of the present invention, the expression for the total electromagnetic force control law is:
[0018]
[0019] in, Based on the fundamental electromagnetic force control law; The compensation control values for each suspension point are learned by the reinforcement learning agent.
[0020] In another implementation of the present invention, the expression for the basic electromagnetic force control law is:
[0021]
[0022] in, , This is the gain coefficient; For the mass equivalent to a single levitation point, For the first The error of each floating point For the first The error derivative of each suspension point It is the arctangent function. This is the acceleration due to gravity.
[0023] Another aspect of the present invention provides a multi-point cooperative reinforcement learning compensation control system for a maglev system based on UAV-InSAR, comprising: a track irregularity power spectrum model construction module: used to perform high-precision mapping of the maglev line using a lightweight X-band interferometric synthetic aperture radar system mounted on an unmanned aerial vehicle, and construct a track irregularity power spectrum model based on interferometric phase inversion; a hierarchical control system construction module: used to construct a hierarchical control system, with an amplitude saturation controller set at each suspension point at the lower layer and a compensator based on reinforcement learning at the upper layer; a total electromagnetic force control law calculation module: the amplitude saturation controller generates a basic electromagnetic force control law based on the current state of the suspension point; the reinforcement learning-based compensator acts as an intelligent agent, interacting with the environment formed by the suspension frame, the track, and the amplitude saturation controller, and is trained using a standardized irregularity function obtained from the track irregularity power spectrum model to output compensation control quantities for each suspension point; the basic electromagnetic force control law is superimposed with the compensation control quantities to generate a total electromagnetic force control law for each suspension point, thereby performing active cooperative compensation and global stable suspension control of the suspension system.
[0024] In another aspect, the present invention provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of a multi-point cooperative reinforcement learning compensation control method for a UAV-InSAR-based maglev system as described in any of the preceding claims.
[0025] In another aspect, the present invention provides a computer storage medium storing a computer program, which, when executed by a processor, implements the steps of a multi-point cooperative reinforcement learning compensation control method for a UAV-InSAR-based maglev system as described in any of the preceding claims.
[0026] This invention presents a multi-point cooperative reinforcement learning compensation control method for maglev systems based on UAV-InSAR. It utilizes the UAV-InSAR system for high-precision mapping of the maglev line and constructs a track irregularity power spectrum model based on interferometric phase inversion, achieving quantitative modeling and spectral characterization of external disturbances. The lower layer configures amplitude saturation controllers for each suspension point to suppress actuator saturation under strong disturbance conditions. The upper layer relies on reinforcement learning to construct a compensator, actively coordinating each suspension point to achieve global stable suspension of the suspension frame. Based on Lyapunov stability theory, it is proven that even if the reinforcement learning model training has not reached convergence, the suspension system can still maintain consistent boundedness. As the reinforcement learning training gradually converges, the system can asymptotically achieve global asymptotic stability, enabling the suspension system to possess both self-learning capability and theoretical stability. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. By reading the detailed description of the embodiments below, the advantages and benefits of the solutions will become clear to those skilled in the art. The accompanying drawings are only for illustrating preferred embodiments and are not intended to limit the present invention. In the accompanying drawings:
[0028] Figure 1 This is a schematic diagram of a multi-point cooperative reinforcement learning compensation control method for a maglev system based on UAV-InSAR, according to an embodiment of the present invention.
[0029] Figure 2 This is a schematic diagram of the interferometric geometry model of a non-parallel spatial baseline according to an embodiment of the present invention.
[0030] Figure 3 This is a schematic diagram of SAR imaging fitted to the UAV track according to an embodiment of the present invention.
[0031] Figure 4 This is a schematic diagram of a suspension control system according to an embodiment of the present invention.
[0032] Figure 5 This is a schematic diagram of a single suspension training environment based on MuJoCo, according to an embodiment of the present invention.
[0033] Figure 6 This is a schematic diagram of track irregularity disturbance according to an embodiment of the present invention.
[0034] Figure 7 This is a schematic diagram of the four-point suspension gap at a speed of 300 km / h, according to an embodiment of the present invention.
[0035] Figure 8 This is a schematic diagram of four-point control current at a speed of 300 km / h, according to an embodiment of the present invention.
[0036] Figure 9 This is a schematic diagram of the four-point suspension gap at a speed of 600 km / h, according to an embodiment of the present invention.
[0037] Figure 10 This is a schematic diagram of four-point control current at a speed of 600 km / h, according to an embodiment of the present invention. Detailed Implementation
[0038] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art should fall within the protection scope of the present invention.
[0039] Figure 1 A schematic flowchart of a multi-point cooperative reinforcement learning compensation control method for a maglev system based on UAV-InSAR, provided as an embodiment of the present invention, is shown below. Figure 1 As shown, this embodiment mainly includes:
[0040] S101. A lightweight X-band interferometric synthetic aperture radar system equipped with an unmanned aerial vehicle is used to conduct high-precision mapping of the maglev line, and a power spectrum model of track irregularities is constructed based on interferometric phase inversion.
[0041] S102. Construct a hierarchical control system, with the lower layer being an amplitude saturation controller (ASC) set at each suspension point, and the upper layer being a compensator based on reinforcement learning.
[0042] S103. The amplitude saturation controller generates a basic electromagnetic force control law based on the current state of the suspension point.
[0043] S104. The reinforcement learning-based compensator acts as an intelligent agent, interacting with the environment formed by the suspension frame, the track, and the amplitude saturation controller. It is trained using the standardized irregularity function obtained from the track irregularity power spectrum model and outputs the compensation control quantity for each suspension point.
[0044] S105. The basic electromagnetic force control law is superimposed with the compensation control quantity to generate the total electromagnetic force control law for each suspension point, and the suspension system is actively coordinated to compensate and globally stabilized to control the suspension.
[0045] The present invention discloses a multi-point cooperative reinforcement learning compensation control method for a maglev system based on UAV-InSAR. The lower layer configures amplitude saturation controllers for each suspension point to suppress actuator saturation under strong disturbance conditions. The upper layer relies on reinforcement learning to construct a compensator, which actively coordinates each suspension point to achieve global stable suspension of the suspension frame. Based on Lyapunov stability theory, it is proved that even if the reinforcement learning model training has not reached a convergence state, the suspension system can still maintain consistent boundedness. As the reinforcement learning training gradually converges, the system can asymptotically achieve global asymptotic stability, so that the suspension system has both self-learning ability and theoretical stability.
[0046] In another implementation of the present invention, the step of using a UAV equipped with a lightweight X-band interferometric synthetic aperture radar system to perform high-precision mapping of the maglev line and constructing a track irregularity power spectrum model based on interferometric phase inversion includes: using a UAV equipped with a lightweight X-band interferometric synthetic aperture radar system to perform high-precision mapping of the maglev line; determining the interferometric phase according to the geometric relationship between the main track and the secondary track during the mapping process; performing inversion through the interferometric phase to obtain the elevation disturbance in the track normal direction; and establishing a track irregularity power spectrum model based on the elevation disturbance field.
[0047] In another implementation of the invention, a single suspension frame is considered, and a mathematical model of a four-point suspension system for the single suspension frame is established. The guidance system is ignored during modeling, meaning the system does not generate a directional system. Translation and rotation of the axis The shaft rotates. The displacements of the four suspension points are as follows:
[0048] (1)
[0049] in, , , , This represents the vertical displacement of the four suspended points relative to their initial positions. This represents the vertical displacement of the center of mass relative to its initial position. , They are respectively around shaft and The rotation angle of the shaft, and These represent the width and length distance between the four floating points, respectively.
[0050] The mechanical equations for translation and rotation are:
[0051] (2)
[0052] (3)
[0053] (4)
[0054] in, The total mass of the suspension frame, and They are respectively around shaft and moment of inertia of the shaft , , , The electromagnetic forces at the four suspension points are represented by the following electromagnetic equations:
[0055] (5)
[0056] in, air permeability, The number of coil turns. The area of the magnetic poles. For coil current, This is the initial suspended air gap.
[0057] Select the state variable as Control quantity Then the state equation of the system is:
[0058] (6)
[0059] In another implementation of the present invention, the expression for the elevation disturbance in the orbital normal direction is:
[0060]
[0061] in, These are the azimuth coordinates. For distance coordinates, The elevation change is obtained from interferometric phase inversion. The angle of incidence is denoted as .
[0062] For example, to achieve high-precision identification and spectral characterization of geometric irregularities in maglev lines, this paper introduces a lightweight X-band interferometric synthetic aperture radar (UAV-InSAR) system mounted on an unmanned aerial vehicle (UAV) to realize non-contact, high-resolution track elevation mapping. Compared with traditional ground-based surveying, UAV-InSAR has the advantages of flexible flight paths, high spatial resolution, and strong resistance to weather interference, and can construct the spatial spectral field of track irregularities over a large area.
[0063] like Figure 2 As shown, during the two flight imaging processes, there is a spatial baseline B between the main orbit and the secondary orbit. Let the incident angle be... Arbitrary scattering point The interference phase can be expressed as:
[0064] (7)
[0065] in, For radar wavelength, The straight-line distance from the radar antenna to the ground point. The elevation angle of the main track towards the secondary track. The angle of motion of the track relative to the main track. The angle between the master and slave trajectories.
[0066] If the trajectories are not parallel ( The second term introduces the azimuth phase component. This leads to interference fringe distortion and decreased coherence. The UAV attitude (pitch angle, roll angle, yaw angle) is acquired in real time using the POS system, and the trajectory equations for each flight are fitted to obtain the common heading angle.
[0067] (8)
[0068] Based on this, parallel baseline compensation is performed to eliminate the azimuth interference components caused by non-parallelism between tracks, making the trajectory equivalent to an ideal parallel baseline, and realizing a multi-temporal comparable interferometric geometry configuration.
[0069] The line-of-sight deformation (or elevation change) obtained by interferometric phase inversion is as follows:
[0070] (9)
[0071] The elevation disturbance in the orbital normal direction can be obtained further through geometric projection:
[0072] (10)
[0073] in The incident angle. Because the baseline of UAV imaging is relatively short, Geometric distortion can be accurately estimated and corrected using a POS system (inertial navigation and GPS).
[0074] like Figure 3 As shown, the UAV is affected by turbulence and attitude jitter during flight, resulting in nonlinear residual motion errors. To eliminate local geometric mismatches, a block registration model based on the Fourier transform coherence coefficient method is adopted. A binary quadratic correlation coefficient surface is constructed in the neighborhood of the control points of the main and auxiliary images:
[0075] (11)
[0076] Find the extreme value to obtain the control point offset:
[0077] (12)
[0078] The coefficients are solved by least squares fitting. This process involves completing local geometric mapping and subpixel-level resampling. This significantly reduces azimuth residual error, improves coherence, and provides high-quality interferometric data for subsequent spectral analysis.
[0079] In another implementation of the present invention, the expression for the orbital irregularity power spectrum model is:
[0080]
[0081] in, For wave number, The wavelength represents the non-uniform spatial wavelength; L represents the observation length. The normal elevation disturbance field of the benchmark orbit.
[0082] For example, the orbital normal elevation disturbance field is obtained. Subsequently, based on its spatial distribution, a power spectral density (PSD) model for orbital irregularities can be established:
[0083] (13)
[0084] in For wave number, Let L be the wavelength of the irregular spatial spectrum and L be the observation length. An empirical power-law model can be used to approximate the irregular spectrum:
[0085] (14)
[0086] In the formula For spectral intensity parameters, The spectral attenuation index is used to characterize the random geometric irregularities of maglev tracks. This spectrum describes the energy distribution of different wavelength components during track elevation disturbances.
[0087] In another implementation of the invention, a standardized non-roughness function can be defined:
[0088] (15)
[0089] in, For random phase, satisfying Equation (15) realizes the statistical simulation of the actual track irregularity field and can be used as the perturbation input function in the reinforcement learning training environment.
[0090] In another implementation of the present invention, the expression for the total electromagnetic force control law is:
[0091]
[0092] in, Based on the fundamental electromagnetic force control law; This is the compensation control value for each suspension point.
[0093] In another implementation of the present invention, the expression for the basic electromagnetic force control law is:
[0094]
[0095] in, , This is the gain coefficient; This is the mass equivalent to a single floating point.
[0096] For example, a common engineering practice is to treat each suspension point as completely independent, each controlled by a Proportional-Integral-Differential (PID) controller. As shown in (6), the suspension points are coupled together through the suspension electromagnet assembly and the suspension frame, and cannot be considered completely independent. Furthermore, PID control is a linear control method, which is difficult to handle the strong nonlinear characteristics of the suspension system. Therefore, the controller design proposed in this invention adopts a hierarchical control approach, such as... Figure 4 As shown, the bottom layer is an ASC controller, which treats the coupling effect of each suspension point and external disturbances as a unified disturbance term during the design. The upper layer adopts a reinforcement learning method to uniformly compensate for the coupling effect and disturbance of each point, thus forming a suspension control system.
[0097] Consider the first i Mathematical model of a floating point:
[0098] (16)
[0099] in, For the mass equivalent to a single levitation point, For suspended air gap, External disturbances such as track irregularities and multi-point coupling, , The electromagnetic force that actually acts is expressed as follows:
[0100] (17)
[0101] in, This represents the actual electromagnetic force coefficient. To control the current.
[0102] Define dimensionless error as:
[0103] (18)
[0104] in, For the target air gap, The unit length.
[0105] The electromagnetic force control law is designed as follows:
[0106] (19)
[0107] in, , This is the gain coefficient.
[0108] Combined with upper-level reinforcement learning compensation, the total electromagnetic force control law for a single point is:
[0109] (20)
[0110] in, For the upper-level reinforcement learning controller to the first i The compensation control quantity for each floating point.
[0111] The back-calculated control current is:
[0112] (twenty one)
[0113] in, To take into account the actual electromagnetic force coefficient It is difficult to accurately describe the defined nominal electromagnetic force coefficient.
[0114] From (17)(19)(21), the actual electromagnetic force can be obtained as:
[0115] (twenty two)
[0116] From (16) and (22), the dynamic equations of the closed-loop system can be obtained as follows:
[0117] (twenty three)
[0118] Depend on It can be known That is, the error equation of the system is:
[0119] (twenty four)
[0120] in, .
[0121] Define the compensation amount for each descent point generated by reinforcement learning. The compensator is an intelligent agent, and the suspension frame, track, and ASC controller together constitute the environment. During training, for each training segment, the suspension frame is reset to repeatedly run from the end of the track until the reinforcement learning method learns the optimal compensation amount. The system state is defined as follows:
[0122] (25)
[0123] The output action is:
[0124] (26)
[0125] The compensator is developed based on the SAC algorithm in deep reinforcement learning. Its core is the maximum entropy objective function, which introduces policy entropy on the basis of traditional reward maximization. To balance exploration and exploitation, and improve policy robustness. Its optimization objective is:
[0126] (27)
[0127] in, For instant rewards, As a discount factor, For the temperature parameter, the weights of reward and entropy are weighed. For strategy Induced state-action distribution.
[0128] To mitigate Q-value overestimation, SAC employs a dual Q-network. The target Q value is the minimum of the two.
[0129] (28)
[0130] in, For state transition distribution, The target Q network parameters are updated with a delay to maintain stability.
[0131] The loss function of the Q-network is:
[0132] (29)
[0133] in, This serves as a buffer for experience replay.
[0134] Policy Network Using a stochastic strategy, the goal is to maximize the soft value of state-action pairs:
[0135] (30)
[0136] in, These are the parameters of the policy network.
[0137] Policy network input state Output the mean and log-standard deviation of the Gaussian distribution:
[0138] (31)
[0139] in, The standard deviation is used. The original motion is generated through reparameterized sampling and then compressed to the target interval using tanh.
[0140] (32)
[0141] in, For standard normal noise, This is the scaling factor for the action range. This is element-wise multiplication.
[0142] SAC introduces an entropy-constrained objective and optimizes it through gradient descent. To match the preset target entropy :
[0143] (33)
[0144] Wherein, the target entropy is set as , For the action space dimension.
[0145] Example 1
[0146] Theorem 1: The designed hierarchical control method can still guarantee consistent boundedness even when the reinforcement learning training has not converged, and the system will satisfy asymptotic stability as the reinforcement learning gradually converges.
[0147] Proof 1: Choose the Lyapunov function as:
[0148] (34)
[0149] To verify the positive definiteness of (34), the constructor is:
[0150] (35)
[0151] Easy to obtain .right Differentiation yields:
[0152] (36)
[0153] when hour, ; hour, ; hour, Then we can obtain .because , , Then we can obtain:
[0154] (37)
[0155] If and only if The value can be 0. That is, (34) satisfies the conditions of the Lyapunov function.
[0156] Differentiating equation (34) with respect to time and substituting it into equation (24) yields:
[0157] (38)
[0158] From (32), we can see that the output of reinforcement learning is mapped to a bounded quantity through the tanh function, hence in (24) Use the following important inequalities:
[0159] (39)
[0160] (40)
[0161] From (38)(39)(40), we can obtain:
[0162] (41)
[0163] Define a set:
[0164] (42)
[0165] in, The parameters are selected to make .
[0166] Outside the set, i.e. Then, from (41), we can obtain:
[0167] (43)
[0168] That is, the derivative of the Lyapunov function of the system outside the set is less than 0, satisfying the uniform eventual boundedness.
[0169] When the system reaches steady state, that is Then, from (24), we can obtain:
[0170] (44)
[0171] The steady-state error then satisfies:
[0172] (45)
[0173] As reinforcement learning gradually converges, that is:
[0174] (46)
[0175] but From (41) and (45), we can obtain , The system gradually converges to asymptotic stability, thus proving the point.
[0176] The total reward function is obtained by designing a separate function for each point and summing the results:
[0177] (47)
[0178] Analysis of the above formula shows that the reward function reaches its global maximum when the error at each point reaches its minimum. Since the reward function implicitly includes the requirement of minimizing the error collaboratively among all points, the collaborative error term is not explicitly considered in the design of the reward function.
[0179] Example 2
[0180] A training environment for a single-suspension suspension system was built based on the MuJoCo physics simulation platform, mainly including the suspension frame and track system, such as... Figure 5 As shown in Table 1, the suspension system comprises the suspension frame body, suspension electromagnets, guide electromagnets, and auxiliary components. Each suspension electromagnet unit has a gap sensor at each end to collect real-time suspension gap data. The electromagnetic force is concentrated at both ends of the suspension electromagnet unit, resulting in a total of four gap sensors and four suspension points. The track system is constructed in segments along the forward direction, with left and right suspension guide rails configured synchronously. Specific parameter settings are shown in Table 1. This environment can effectively reproduce complex operating conditions such as track irregularities, multi-point coupled disturbances, and actuator saturation, providing a high-fidelity platform for training and verifying reinforcement learning collaborative feedforward compensation control strategies.
[0181] Table 1 Physical parameters of the single suspension system
[0182]
[0183] Due to the inherent characteristics of the electromagnet and the performance of the chopper, the variation of the control quantity is limited as follows:
[0184] (48)
[0185] in, For electromagnet current, This represents the maximum rate of change of current. and These are the upper and lower limits of the current, respectively.
[0186] The random tracks added during runtime are not smooth, such as Figure 6 As shown, the training can be obtained by sampling the orbital irregularity spectrum shown in (15).
[0187] The reinforcement learning agent is set to a total training step count of 1,200,000, with a maximum step count of 3,000 per segment. Training is performed every 2 steps, and evaluation and saving are performed every 5 segments. Specific parameters are shown in Table 2.
[0188] Table 2 Training parameter configuration
[0189]
[0190] In the simulation, the reinforcement learning compensation control method proposed in this invention is compared with PID (Proportional-Integral-Derivative) control, sliding mode control (SMC) and ASC control.
[0191] The expression for PID control is:
[0192] (49)
[0193] in, , and These are the proportional, integral, and differential coefficients, respectively.
[0194] The expression for sliding mode control is:
[0195] (50)
[0196] in, It is a velocity-type sliding surface. For air gap error, The sliding surface coefficient, To switch the gain, Boundary layer thickness, It is a saturation function.
[0197] The specific parameters for PID, SMC, and ASC are as follows: , , , , , , , .
[0198] In this verification, two different operating conditions were designed for comparative analysis. Operating condition 1 and operating condition 2 simulate train operation at speeds of 300 km / h and 600 km / h, respectively. The performance of the suspension control algorithm under long-wave irregularity interference was analyzed. The two operating conditions were quantitatively analyzed using root mean square error and maximum error.
[0199] Under operating condition 1, the 4-point suspended air gap is as follows: Figure 7As shown in the table, PID control exhibits significant fluctuations and oscillations, while SMC (Slip Mode Control) suffers from steady-state errors due to roughness disturbances, making it difficult to stabilize on the sliding surface. ASC (Automatic Slip Control) is similar to the proposed method, but the proposed method recovers faster after deviating from the equilibrium point due to disturbances. The root mean square error and maximum error of each suspension point are shown in Table 3, and the average error of the four suspension points is shown in Table 4. The average root mean square error of the proposed method is 0.5297 mm, which is approximately 19.49% lower than PID control, approximately 65.19% lower than SMC control, and approximately 28.51% lower than ASC control. The average maximum error of the proposed method is 1.9548 mm, which is approximately 2.43% lower than PID control, approximately 5.65% lower than SMC control, and approximately 4.60% lower than ASC control. All error indicators are optimal, and the method exhibits good suppression of roughness disturbances at 300 km / h.
[0200] Under operating condition 1, the control current at 4 points is as follows: Figure 8 As shown, PID control exhibits significant oscillations, repeatedly reaching current saturation, while ASC control shows smaller oscillations, consistently remaining near the equilibrium current. The proposed method incorporates current compensation based on ASC, thus increasing oscillations compared to ASC control, but less than PID and SMC control.
[0201] Table 3. Errors of each suspension point under operating condition 1
[0202]
[0203] Table 4 Average Error under Operating Condition 1
[0204]
[0205] Industrial control 2, 4-point suspended air gap, such as Figure 9As shown, compared to operating condition 1, the increased operating speed significantly increases the fluctuation of the suspension gap. The fluctuation under PID and SMC control is greatly amplified, reaching ±6mm, making it difficult to suppress the effects of roughness; under ASC control, the gap fluctuation range is approximately ±3mm, significantly affected by roughness. The proposed method shows relatively good control performance, with the fluctuation range of suspension points 1 and 2 approximately ±1.5mm, but the fluctuation range of suspension points 3 and 4 is relatively large due to coupling effects. The root mean square error and maximum error of each suspension point are shown in Table 5, and the average error of the four suspension points is shown in Table 6. The average root mean square error of the proposed method is 0.6065 mm, which is approximately 64.38% lower than PID control, approximately 65.39% lower than SMC control, and approximately 47.25% lower than ASC control. Its average maximum error is 2.4266 mm, which is approximately 55.85% lower than PID control, approximately 41.82% lower than SMC control, and approximately 21.29% lower than ASC control. Under high-speed, high-disturbance operating conditions, the proposed method shows significant improvement in error performance, demonstrating excellent disturbance suppression capability and control stability. The control current at four points in Industrial Control System 2 is as follows: Figure 10 As shown, both PID and SMC control reach the saturation current, making effective control difficult. ASC control and the proposed method can effectively control within the saturation range, but some oscillations exist.
[0206] Table 5. Errors of each suspension point under operating condition 2
[0207]
[0208] Table 6 Average Error under Operating Condition 2
[0209]
[0210] Therefore, it can be seen that:
[0211] (1) To address the problems of track irregularity interference, multi-point coupling disturbance, actuator saturation, and response lag during high-speed operation of maglev train suspension systems, this invention proposes a hierarchical collaborative feedforward compensation control strategy based on reinforcement learning, taking into account the characteristics of the train's reciprocating cyclic operation. A lightweight X-band interferometric synthetic aperture radar system mounted on an unmanned aerial vehicle (UAV) is introduced to provide disturbance modeling support. This strategy adopts a hierarchical architecture: the lower-level ASC controller addresses actuator saturation under strong disturbances; the upper-level reinforcement learning compensator based on the SAC algorithm utilizes the track irregularity power spectrum model information obtained from UAV-InSAR to coordinate with each suspension point to achieve global disturbance compensation, overcoming the limitations of traditional control.
[0212] (2) Based on Lyapunov stability theory, this invention rigorously proves the system stability, clarifying the uniform boundedness of the system when reinforcement learning training has not converged and the asymptotic stability after convergence. Simulation verification using a single suspension frame with four suspension points as the object shows that, under high-speed conditions of 600 km / h, the root mean square error and maximum error of the suspension air gap of the proposed method are reduced by more than 40% compared with the traditional ASC and SMC methods. It can effectively compensate for the multi-source interference and multi-constraint limitations faced by the suspension system, improve the system stability and disturbance rejection, highlight the advantages of UAV-InSAR and reinforcement learning synergistic compensation, and improve the comprehensive control performance of the suspension system.
[0213] (3) Future research will expand the system modeling dimensions, incorporate disturbances such as guidance system dynamics and multi-suspension coupling, and realize train-level collaborative control optimization; at the same time, optimize UAV-InSAR mapping accuracy and data processing efficiency, conduct physical experiments, refine the details of algorithm engineering implementation, and promote the control strategy from simulation to engineering application.
[0214] Another aspect of the present invention provides a multi-point cooperative reinforcement learning compensation control system for a maglev system based on UAV-InSAR, comprising:
[0215] Track irregularity power spectrum model construction module: used to conduct high-precision mapping of maglev lines using a lightweight X-band interferometric synthetic aperture radar system mounted on an UAV, and to construct a track irregularity power spectrum model based on interferometric phase inversion.
[0216] Layered Control System Construction Module: Used to build a layered control system, with the lower layer being an amplitude saturation controller set at each suspension point and the upper layer being a compensator based on reinforcement learning.
[0217] Total Electromagnetic Force Control Law Calculation Module: The amplitude saturation controller generates a basic electromagnetic force control law based on the current state variables of the suspension points; the reinforcement learning-based compensator acts as an intelligent agent, interacting with the environment formed by the suspension frame, track, and amplitude saturation controller, and is trained using a standardized non-compliance function obtained from the track non-compliance power spectrum model, outputting compensation control quantities for each suspension point; the basic electromagnetic force control law is superimposed with the compensation control quantities to generate the total electromagnetic force control law for each suspension point, performing active collaborative compensation and global stable suspension control of the suspension system.
[0218] The present invention relates to a multi-point cooperative reinforcement learning compensation control system for a maglev system based on UAV-InSAR. The lower layer is configured with amplitude saturation controllers for each suspension point to suppress actuator saturation under strong disturbance conditions. The upper layer relies on reinforcement learning to construct a compensator, which actively coordinates each suspension point to achieve global stable suspension of the suspension frame. Based on Lyapunov stability theory, it is proved that even if the reinforcement learning model training has not reached a convergence state, the suspension system can still maintain consistent boundedness. As the reinforcement learning training gradually converges, the system can asymptotically achieve global asymptotic stability, so that the suspension system has both self-learning ability and theoretical stability.
[0219] In another aspect of the present invention, the electronic device includes: a processor, a memory, and a communication bus and a communication interface.
[0220] in:
[0221] The processor, memory, and communication interface communicate with each other via a communication bus.
[0222] A communication interface is used to communicate with other electronic devices or servers.
[0223] The processor is used to execute programs, specifically, it can execute any of the steps of the multi-point cooperative reinforcement learning compensation control method for maglev systems based on UAV-InSAR in the above embodiments.
[0224] Specifically, the program may include program code, which includes computer operation instructions.
[0225] The processor may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.
[0226] Memory is used to store programs. Memory may include high-speed RAM, and may also include non-volatile memory, such as at least one disk drive.
[0227] Specifically, the program can be used to cause the processor to execute the steps of any of the UAV-InSAR-based multi-point cooperative reinforcement learning compensation control methods for maglev systems described in the embodiments. The specific implementation of each step in the program can be found in the corresponding descriptions of the steps and units executed in any of the UAV-InSAR-based multi-point cooperative reinforcement learning compensation control methods for maglev systems described above, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments.
[0228] An exemplary embodiment of this application also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods of various embodiments of this application.
[0229] The methods described above according to embodiments of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded via a network. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0230] Specific embodiments of the present invention have now been described. Other embodiments are within the scope of the appended claims. In some cases, the actions described in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result.
[0231] It should be noted that all directional indications (such as up, down, left, right, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship between the components in a certain order (as shown in the figure). If the specific order changes, the directional indication will also change accordingly.
[0232] In the description of this invention, the terms "first" and "second" are used only for convenience in describing different components or names, and should not be construed as indicating or implying a sequential relationship, relative importance, or implicitly specifying the number of technical features indicated. Thus, a feature defined with "first" and "second" may explicitly or implicitly include at least one of that feature.
[0233] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0234] It should be noted that although specific embodiments of the present invention have been described in detail with reference to the accompanying drawings, this should not be construed as limiting the scope of protection of the present invention. Various modifications and variations that can be made by those skilled in the art without inventive effort within the scope described in the claims still fall within the scope of protection of the present invention.
[0235] The examples of the embodiments of the present invention are intended to concisely illustrate the technical features of the embodiments of the present invention, so that those skilled in the art can intuitively understand the technical features of the embodiments of the present invention, and are not intended to be an improper limitation of the embodiments of the present invention.
[0236] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-point cooperative reinforcement learning compensation control method for a maglev system based on UAV-InSAR, characterized in that, include: A lightweight X-band interferometric synthetic aperture radar system mounted on an unmanned aerial vehicle (UAV) was used to perform high-precision mapping of the maglev line. Based on interferometric phase inversion, a track irregularity power spectrum model was constructed, including: performing high-precision mapping of the maglev line using a UAV equipped with a lightweight X-band interferometric synthetic aperture radar system; determining the interferometric phase based on the geometric relationship between the main track and the secondary track during the mapping process; performing inversion using the interferometric phase to obtain the elevation disturbance in the track normal direction; and establishing a track irregularity power spectrum model based on the elevation disturbance field. A hierarchical control system is constructed, with the lower layer being an amplitude saturation controller set at each suspension point, and the upper layer being a compensator based on reinforcement learning; The amplitude saturation controller generates a basic electromagnetic force control law based on the current state of the suspension point; The reinforcement learning-based compensator acts as an intelligent agent, interacting with the environment formed by the suspension frame, the track, and the amplitude saturation controller. It is trained using a standardized non-roughness function obtained from the track non-roughness power spectrum model and outputs the compensation control quantity for each suspension point. The basic electromagnetic force control law is superimposed with the compensation control quantity to generate the total electromagnetic force control law for each suspension point, thereby performing active collaborative compensation and global stable suspension control on the suspension system.
2. The method according to claim 1, characterized in that, The expression for the elevation disturbance in the normal direction of the orbit is: in, These are the azimuth coordinates. For distance coordinates, The elevation change is obtained from interferometric phase inversion. The angle of incidence is denoted as .
3. The method according to claim 2, characterized in that, The expression for the power spectrum model of track irregularities is: in, For wave number, The imaginary unit, These are the azimuth coordinates. The wavelength represents the non-uniform spatial wavelength; L represents the observation length. The normal elevation disturbance field of the benchmark orbit.
4. The method according to claim 3, characterized in that, The expression for the standardized non-smooth function is: in, For random phase, satisfying , For the first Discrete wavenumber points, For the first The interval of each wavenumber interval.
5. The method according to claim 1, characterized in that, The expression for the total electromagnetic force control law is: in, Based on the fundamental electromagnetic force control law; The compensation control values for each suspension point are learned by the reinforcement learning agent.
6. The method according to claim 5, characterized in that, The expression for the fundamental electromagnetic force control law is as follows: in, , This is the gain coefficient; For the mass equivalent to a single levitation point, For the first The error of each floating point For the first The error derivative of each suspension point It is the arctangent function. This is the acceleration due to gravity.
7. A multi-point cooperative reinforcement learning compensation control system for a maglev system based on UAV-InSAR, characterized in that, include: The track irregularity power spectrum model construction module is used to perform high-precision mapping of maglev lines using a UAV equipped with a lightweight X-band interferometric synthetic aperture radar system, and to construct a track irregularity power spectrum model based on interferometric phase inversion. This includes: performing high-precision mapping of the maglev line using a UAV equipped with a lightweight X-band interferometric synthetic aperture radar system; determining the interferometric phase based on the geometric relationship between the main track and the secondary track during the mapping process; performing inversion using the interferometric phase to obtain the elevation perturbation in the track normal direction; and establishing the track irregularity power spectrum model based on the elevation perturbation field. Hierarchical control system construction module: used to build a hierarchical control system, with the lower layer being an amplitude saturation controller set at each suspension point and the upper layer being a compensator based on reinforcement learning; Total Electromagnetic Force Control Law Calculation Module: The amplitude saturation controller generates a basic electromagnetic force control law based on the current state variables of the suspension points; the reinforcement learning-based compensator acts as an intelligent agent, interacting with the environment formed by the suspension frame, track, and amplitude saturation controller, and is trained using a standardized non-compliance function obtained from the track non-compliance power spectrum model, outputting compensation control quantities for each suspension point; the basic electromagnetic force control law is superimposed with the compensation control quantities to generate the total electromagnetic force control law for each suspension point, performing active collaborative compensation and global stable suspension control of the suspension system.
8. An electronic device, characterized in that, include: The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the multi-point cooperative reinforcement learning compensation control method for a maglev system based on UAV-InSAR as described in any one of claims 1 to 6.
9. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which, when executed by a processor, implements the steps in the multi-point cooperative reinforcement learning compensation control method for a maglev system based on UAV-InSAR as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Maglev train, levitation control system and method for improving operation stability
CN113619402A
Control method of maglev train suspension assembly
CN116674393A