An aircraft-trailer system coordination control method fusing reinforcement learning
By constructing a mathematical model of the aircraft-trailer system and designing adaptive and reinforcement learning controllers, the coordinated motion problem of the aircraft-trailer system under coupling and external disturbances was solved, achieving high-precision trajectory tracking and stable control.
Patent Information
- Application Number
- CN202411924960.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2044-12-25
AI Technical Summary
The coupling characteristics and external disturbances of existing aircraft-trailer systems make it difficult for automatic control to achieve coordinated movement. In particular, under uncertain conditions of system coupling and external disturbances, traditional control methods are unable to ensure coordinated movement of the aircraft and trailer.
By constructing a mathematical model of the aircraft-trailer system, a model-based adaptive controller and a reinforcement learning-based attitude compensation controller are designed. The adaptive controller and the attitude compensation controller are integrated to output control signals to achieve motion control of the aircraft and the trailer.
It achieves trajectory tracking control to overcome uncertainties caused by system coupling and external disturbances, ensuring the motion coordination and stability of the aircraft-trailer system, and improving control accuracy and system adaptability.
Smart Images

Figure CN119759018B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of aircraft traction control technology, and in particular to a coordinated control method for an aircraft-trailer system that incorporates reinforcement learning. Background Technology
[0002] Since aircraft lack reverse gear, towing operations are necessary to move them from their parking positions to designated locations. Traditional aircraft towing methods, such as towing and taxiing, have proven effective in reducing operating costs and pollutant emissions. Furthermore, the assistance of trailers allows for better route planning, thereby improving airport operational efficiency.
[0003] However, the coupling characteristics of aircraft-trailer systems present a series of challenges. To facilitate towing and steering, aircraft and towing trailers are mostly rigidly connected via landing gear. This connection method requires strong traction to ensure the movement of heavy aircraft. Due to the nonlinear coupling of aircraft-trailer systems and the potential scissor effect, automatic control of aircraft-trailer systems faces significant difficulties. In addition, operational coordination is affected by factors such as communication, air traffic control, and system commands.
[0004] In summary, under conditions of system coupling and uncertainty of external disturbances, existing control methods are insufficient to ensure coordinated movement of the aircraft and the trailer. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a coordinated control method for aircraft-trailer systems that integrates reinforcement learning. This method can accurately track and control the trajectory of uncertainties caused by system coupling and external disturbances, thereby achieving coordinated motion control of the aircraft-trailer system.
[0006] The objective of this invention can be achieved through the following technical solution: a coordinated control method for an aircraft-trailer system incorporating reinforcement learning, comprising the following steps:
[0007] S1. Construct a mathematical model of the aircraft-trailer system and determine the coupling relationship between control variables and state variables. The control variables include trailer movement control variables and system steering control variables.
[0008] S2. Design a model-based adaptive controller for the trailer movement control variables;
[0009] For the system steering control variables, an attitude compensation controller based on reinforcement learning is designed;
[0010] S3 integrates the adaptive controller and attitude compensation controller, and outputs corresponding control signals to the aircraft and trailer to complete the motion control of the aircraft and trailer.
[0011] Furthermore, the specific process of step S1 is as follows:
[0012] Define ΣOa, ΣOm, and ΣOt as frames fixed to the aircraft, steering mechanism, and trailer, respectively. Frame ΣOa is used to describe the aircraft's position p. a The frame ΣOm is used to describe the position p of the steering mechanism. m The frame ΣOt is used to describe the trailer position p. t ;
[0013] Define s a =[x a ,y a ,γ a ] T Represents the aircraft's attitude and s in the global frame ΣOg. t =[x t ,y t ,γ t ] T This indicates the attitude of the trailer within the global framework ΣOg;
[0014] Considering the aircraft's motion state, s a and s t The relationship between them is,
[0015]
[0016] sγ a =sin(γ) a )
[0017] sγ m =sin(γ) m )
[0018] cγ a =cos(γ) a )
[0019] cγ m =cos(γ) m )
[0020] Where, γ t γ is the rotation angle of the trailer. m γ is the rotation angle of the steering mechanism. a L is the aircraft's rotation angle. ma and L tm These are the distances between the steering mechanism and the aircraft, and the distances between the trailer and the steering mechanism, respectively.
[0021] For s a and s t Differentiate the relationship between them:
[0022]
[0023] Where, q s Let S(s) represent the joint space state of the aircraft-trailer system, and S(s) be the state transition matrix.
[0024] Further, step S2 includes the following process:
[0025] S21. Regarding the trailer movement control variable x t and y t Introducing dummy variable μ t To achieve x t and y t After decoupling, an adaptive control law was designed to perform closed-loop control of the trailer's moving speed;
[0026] S22, Regarding the system steering control variable γ t and γ m We construct a reinforcement learning agent to achieve compliant steering.
[0027] Furthermore, in step S21, the dummy variable μ t By the turning variable γ t and γ m The decision is made jointly, and the following conditions are met: lim t→∞ |e γ,t ,e γ,m |=0, the virtual variable μ1 is specifically:
[0028]
[0029] Where, x e and y e These are the x-coordinates of the aircraft-trailer system in the world coordinate system. g and y g The error in direction, c s =[c s,k1 ,c s,d1 ,c s,k2 ,c s,d2 ] are control parameters, used to correspond to the adjustment error x respectively. e , y e , Gain, and It is the expected speed of the trailer, e γ,t e represents the trailer rotation angle error. γ,m This represents the rotation angle error of the rotating mechanism.
[0030] Furthermore, the specific process of step S21 is as follows:
[0031] For trailer speed control, and As control variables, these three control variables are coupled, and the choice... and As an independent variable, and the error is defined as x. e,t and y e,t Then, a virtual input μ is introduced. t ,get:
[0032]
[0033] Define Lyapunov functions And differentiate:
[0034]
[0035] Design position control law:
[0036]
[0037] Among them, c x,k and c y,k These are control parameters, which control the speed control law v. t And adaptive law Set them respectively to,
[0038]
[0039] in, It is an auxiliary signal, λ k It is a positive real number, c x,d It is a normal number.
[0040] Furthermore, step S22 employs a two-layer reinforcement learning (RL) architecture to synthesize the dummy variable μ. t and steering angular velocity ω m Attitude compensation controller.
[0041] Furthermore, the two-layer reinforcement learning architecture includes a basic policy module π and an adaptive module ρ, wherein the basic policy module π uses the current state, previous actions, different desired trajectories, and an adaptive learning rate. As input, the output yields a steering motion control strategy, and the adaptive module ρ is trained to improve the system's adaptability to uncertainties, including internal parameters and external changes.
[0042] Furthermore, the operation of the basic strategy module π includes:
[0043] In each step, the input state is defined as s s =[s a ,st ,s m [This refers to the current global attitude of the trailer, aircraft, and steering mechanism.] It is s s The first derivative, s a,d For the aircraft's motion reference trajectory;
[0044] Action represented as a t =[ω m ,μ t ] T ω m It is the action sequence for updating the angular velocity of the steering mechanism, and the designed virtual input μ t Used to adjust the angular velocity of the system;
[0045] Transform the objective-conditional reinforcement learning problem into a Markov decision process: The elements in the equation are state, action, transition dynamics, reward function, and a discount factor γ, for any state-action sequence. t-1 ,s t ,r t The RL agent is trained to construct a policy π that maximizes the expected reward. The agent then solves a problem based on this policy, which is defined as follows:
[0046]
[0047] Where τ={(s1,a1,r1),(s2,a2,r2),…,(s t ,a t ,r t P = [P1, P2, ..., P] is the sequence of actions taken by the agent when executing policy π. t ] is the possible sequence of actions under π.
[0048] To guide the agent in coordinated motion within the coupled and nonlinear conditions of the aircraft-trailer system, a reward function is designed that includes task-oriented and performance-enhancing rewards. The task-oriented reward is used to encourage the aircraft-trailer system to track the desired trajectory s in different environments. a,d Performance improvement rewards focus on dynamic adjustments; in addition, the reward function also considers real-time and predicted states.
[0049] Furthermore, the task-oriented reward is defined as:
[0050]
[0051] Where r0 is the initial reward, e s =[e a ,e t ,e m ] is the state error matrix, e s =e s,t -e s,d ξ0 is an adjustable parameter, Δ e It is the threshold, f(s,a) is the defined reward rate function, P1, P2 and P3 are positive definite matrices, and e s T P1e s The impact of state error on the reward function is described;
[0052] The performance improvement reward is defined as:
[0053]
[0054] in, It is the difference between the action values at times t and t-1. e tm =e t -e m It is the relative error between the trailer and the steering mechanism, P a P e and P tm It is a positive definite matrix. and These are used to adjust the smoothness of actions and states, respectively. To improve coordination in relative steering motion, during training, the scaling factor of each reward item is adjusted to maximize the reward.
[0055] Furthermore, the operation of the adaptive module ρ includes:
[0056] For the nominal inertia matrix and uncertain external friction To provide compensation, define for:
[0057]
[0058] Where s = [s t [f] is the input vector. It is the Gaussian function, W M It involves learning network weights; the adaptive learning principle is...
[0059] For the adaptive module, a radial basis function-based neural network is used to approximate the uncertain dynamics of the system. This neural network consists of three feedforward layers. The input layer includes the system state and the defined nonlinear friction force. The hidden layer is a radial basis vector with 10 nodes. The output layer is the adaptive term for the internal system parameters and the external uncertain friction. Finally, the... As input to the basic policy π.
[0060] Compared with the prior art, the present invention has the following advantages:
[0061] This invention first describes the aircraft-trailer system using a model and analyzes the coupling relationship between control states. Then, for the control variables describing the trailer's movement, a model-based adaptive controller is designed. For the control variables describing the aircraft-trailer system's steering, a reinforcement learning agent is constructed, and a reinforcement learning-based attitude compensation controller is designed. Finally, the adaptive controller and the attitude compensation controller are fused to achieve coordinated control of the aircraft-trailer system and to accurately track and control the trajectory despite uncertainties caused by model coupling and external disturbances.
[0062] This invention designs an adaptive controller by introducing a virtual variable through kinematic decoupling. This virtual variable is jointly determined by the trailer rotation angle and the rotation angle of the rotating mechanism, which enables system position decoupling and ensures the trailer's state [x] t ,y t and linear velocity v t Tracking performance.
[0063] This invention addresses system parameter variations and external uncertainties by designing a reinforcement learning-based compensation controller. It employs a two-layer reinforcement learning architecture to synthesize the dummy variable μ. t and steering angle ω m The agile controller, with its two-layer reinforcement learning architecture, consists of two layers: a basic policy π and an adaptive module ρ. The basic policy π is used to ensure coordinated steering motion and designs task-oriented and performance-enhancing rewards, which can guide the agent to coordinate motion under the coupling and nonlinearity of the aircraft-trailer system. The adaptive module ρ considers the time-varying characteristics of system parameters and external uncertainties, and can effectively compensate for inertia matrix and uncertain external friction, thereby improving motion coordination. Attached Figure Description
[0064] Figure 1 This is a schematic diagram of the method flow of the present invention;
[0065] Figure 2 This is a schematic diagram illustrating the application process of an example.
[0066] Figure 3 This invention relates to a model of an aircraft-trailer system;
[0067] Figure 4 A coordinated control diagram for an aircraft-trailer system incorporating reinforcement learning;
[0068] Figure 5 To reinforce the learning control strategy graph;
[0069] Figure 6Simulation scenarios under different masses and friction parameters (Simulation Scenario 1);
[0070] Figure 7 This is a reward graph for a simulated scenario.
[0071] Figure 8 This is a trajectory tracking simulation scenario (Simulation Scenario 2);
[0072] Figure 9 The graph shows the change in control law in simulation scenario two.
[0073] Figure 10 The trajectory tracking error in simulation scenario two;
[0074] Figure 11 This is the reward chart for simulation scenario two. Detailed Implementation
[0075] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0076] Example
[0077] like Figure 1 As shown, a coordinated control method for an aircraft-trailer system incorporating reinforcement learning includes the following steps:
[0078] S1. Construct a mathematical model of the aircraft-trailer system and determine the coupling relationship between control variables and state variables. The control variables include trailer movement control variables and system steering control variables.
[0079] S2. Design a model-based adaptive controller for the trailer movement control variables;
[0080] For the system steering control variables, an attitude compensation controller based on reinforcement learning is designed;
[0081] S3 integrates the adaptive controller and attitude compensation controller, and outputs corresponding control signals to the aircraft and trailer to complete the motion control of the aircraft and trailer.
[0082] This embodiment applies the above solution, and the process is as follows: Figure 2 As shown, the aircraft-trailer system is first modeled and the coupling relationships between control states are analyzed. The control variable x describing the trailer's movement is then... t and y t A model-based adaptive control strategy is designed. For the control variable γ describing the steering of the aircraft-trailer system... t and γ m We construct a reinforcement learning agent to achieve compliant steering.
[0083] Control variable x for trailer movement t and y tIntroducing dummy variable μ t To achieve x t and y t After decoupling, an adaptive control rate was designed to perform closed-loop control of the trailer's moving speed.
[0084] Dummy variable μ t The variable γ is changed by the aircraft trailer system. t and γ m A joint decision should satisfy the following conditions: lim t→∞ |e γ,t ,e γ,m |=0.
[0085] To ensure the dummy variable μ t Convergence, design a reinforcement learning agent, and define the state as The behavior of an intelligent agent is defined as a t =[ω m ,μ t ] T By interacting with the aircraft-trailer model environment, attitude stability control is ensured.
[0086] By fusing model-based adaptive controllers and reinforcement learning-based controllers, coordinated control of the aircraft-trailer system can be achieved.
[0087] Specifically, it includes:
[0088] I. Aircraft-Trailer System Model Construction
[0089] like Figure 3 As shown in the system description section, the system consists of an aircraft and a trailer, which are connected by a designed steering mechanism. To effectively describe the modeling process, the preparations and relevant definitions are as follows: ΣOg is used as the global frame. Let ΣOa, ΣOm, and ΣOt represent the frames fixed to the aircraft, steering mechanism, and trailer, respectively. Frame ΣOa is used to describe the aircraft's position p. a The frame ΣOm is used to describe the position p of the steering mechanism. m The frame ΣOt is used to describe the trailer position p. t . s a =[x a ,y a ,γ a ] T This represents the aircraft's attitude within the frame ΣOg, s t =[x t ,y t ,γ t ] T This represents the trailer's attitude within the global frame ΣOg. Similarly, other states in the world coordinate system are also labeled with superscripts. γ a,mIt is the relative angle between the aircraft and the steering mechanism, γ m,t It is the relative angle between the steering mechanism and the trailer. D a and D t These are the wheelbases of the aircraft and the trailer, respectively. L ma and L tm They are points p a p m and p t The distance between them.
[0090] Unlike independent Ackerman kinematics models, this system needs to consider the aircraft's motion state, and the control objective is to achieve the desired motion of the aircraft. a and s t The relationship between them can be deduced as follows:
[0091]
[0092] In the formula, sγ a =sin(γ) a ), sγ m =sin(γ) m ), cγ a =cos(γ) a ), cγ m =cos(γ) m ), γ t γ is the rotation angle of the trailer. m γ is the rotation angle of the steering mechanism. a L is the aircraft's rotation angle. ma and L tm These are the distances between the steering mechanism and the aircraft, and the distances between the trailer and the steering mechanism, respectively.
[0093] Differentiating equation (1), we get:
[0094]
[0095] In the formula, S(s) is the state transition matrix, defined as follows:
[0096]
[0097] As can be seen from equations (1) and (2), the aircraft-trailer system is in a coupled state. Specifically, the angular velocity of the trailer depends on its own attitude and its relative state with the aircraft. Furthermore, equations (1) and (2) describe the kinematic characteristics of the system. If factors such as joint uncertainties and kinematic friction are considered, the dynamic characteristics of the system will become more complex. For such coupled variables, this design uses a reinforcement learning method to handle them.
[0098] II. Adaptive Compensation Controller Design
[0099] Based on the above analysis, a coordinated control scheme integrating reinforcement learning (RL) and closed-loop feedback is proposed. For the system's linear velocity (system variable v)... t An adaptive control method was employed. For the coupled angular velocity (system variable γ), t and γ m It uses an intelligent agent to better make steering decisions under different conditions and uncertainties. The control framework is as follows: Figure 4 As shown.
[0100] For trailer speed control, it can be and These three variables are coupled and are used as control variables. and As an independent variable, and the error is defined as x. e,t and y e,t Therefore, it can be considered as point p. t Position control, and then a virtual input μ is introduced. t ,get:
[0101]
[0102] Define Lyapunov functions And by taking the derivative, we can obtain,
[0103]
[0104] in, and This is the expected speed of the trailer.
[0105] Design position control law,
[0106]
[0107] Among them, c x,k and c y,k These are control parameters.
[0108] Therefore, the speed control law v t And adaptive law Set them to:
[0109]
[0110] In the formula, It is an auxiliary signal, λ k It is a positive real number, c x,d It is a normal number.
[0111] After the airplane and trailer form a solid, the designed virtual input μt The system's attitude needs to be considered, which allows us to obtain...
[0112]
[0113] In the formula, x e and y e These are the x-coordinates of the aircraft-trailer system in the world coordinate system. g and y g Error in direction. s =[c s,k1 ,c s,d1 ,c s,k2 ,c s,d2 ] is a control parameter.
[0114] Formula (7) is based on Lyapunov stability design and can ensure the trailer's condition [x t ,y t and linear velocity v t The tracking performance is good. However, it cannot guarantee the tracking performance of the dummy variable μ. t The convergence of μ in the aircraft-trailer system. t By γ t and γ m A joint decision should satisfy the following conditions: lim t→∞ |e γ,t ,e γ,m |=0.
[0115] The parameter matrix c in equation (9) s Only the convergence rate can be adjusted, but stability cannot be guaranteed. Therefore, the coupled directional variables and uncertainties are intended to be controlled through the designed reinforcement learning (RL) agent.
[0116] III. Strengthening Learning Coordination Strategies
[0117] The goal of this approach is to enable an agent to control the directional variables of an aircraft-trailer system, and then train a policy that can ensure more flexible steering under constraints of large size and inertia. To achieve this, this approach proposes a two-layer reinforcement learning (RL) architecture to synthesize the dummy variable μ. t and steering angular velocity ω m Agile controller. An overview of the learning model is as follows: Figure 5 As shown.
[0118] The RL architecture consists of two layers: a base policy π and an adaptive module ρ. The base policy π is based on the current state, previous actions, different desired trajectories, and an adaptive learning rate. As input, the adaptive module ρ is trained to improve the system's adaptability to uncertainties, including internal parameters (such as mass variations) and external variations (friction due to different terrains). They work together to achieve robust adaptability to diverse environments and strongly coupled uncertainties within the system.
[0119] (1) Basic Strategy π
[0120] In each step, the input state is defined as Here s s =[s a ,s t ,s m [This refers to the current global attitude of the trailer, aircraft, and steering mechanism.] It is s s The first derivative. Furthermore, unlike traditional trailer systems, this solution directly defines the aircraft's trajectory as the reference trajectory s. a,d This ensures the aircraft's maneuverability. Then The analysis at the velocity level is based on Equation (2), where linear velocity is achieved through a model-based control law (see Equation (7)), and the designed strategy π is only concerned with how to more coordinately ensure steering motion.
[0121] Action represented as a t =[ω m ,μ t ] T , here, ω m It is the action sequence for updating the angular velocity of the steering mechanism, and the designed virtual input μ t Used to adjust the system's angular velocity. Unlike differential drive trailers, the variable μ t The coupling effect between the aircraft and the trailer needs to be considered. RL provides an adaptive learning framework that can be applied to control decisions for coupled and nonlinear models. This embodiment uses the dominant actor-commentator (A2C) approach to control the system and formulates the problem as a goal-conditional reinforcement learning problem. Of course, other RL algorithms are also applicable.
[0122] This problem is then transformed into a Markov decision process, represented as: The elements are state, action, transition dynamics, reward function, and a discount factor γ. For any state-action sequence... t-1 ,s t ,r t The RL agent is trained to construct a policy π that maximizes the expected reward, and then the agent solves a problem based on this policy. Policy π is defined as follows:
[0123]
[0124] In the formula, τ={(s1,a1,r1),(s2,a2,r2),…,(s t ,a t ,r t P = [P1, P2, ..., P] is the sequence of actions taken by the agent when executing policy π. t ] is the possible sequence of actions under π.
[0125] To guide the agent in coordinated motion within the coupled and nonlinear environment of the aircraft-trailer system, the reward is designed with two components: task-oriented and performance-enhancing. The task-oriented reward encourages the aircraft-trailer system to follow the desired trajectory in different environments. a,d The performance enhancement reward focuses on dynamic adjustments, such as the crawling problem caused by the aircraft's massive mass, and instability caused by external uncertainties. Furthermore, the reward function considers both the real-time state and the predicted state to ensure the convergence of the RL agent.
[0126] The task-oriented part is defined as follows:
[0127]
[0128] In the formula, r0 is the initial reward, e s =[e a ,e t ,e m ] is the state error matrix, e s =e s,t -e s,d ξ0 is an adjustable parameter, Δ e It is the threshold. f(s,a) is the defined reward rate function.
[0129]
[0130] In the formula, P1, P2, and P3 are positive definite matrices. s T P1e s The effect of state error on the reward function is described.
[0131] The performance improvement is defined as follows:
[0132]
[0133] In the formula, It is the difference between the action values at times t and t-1, similarly. e tm =e t -e m It is the relative error between the trailer and the steering mechanism. P a P e and P tmIt is a positive definite matrix.
[0134] Equation (12) describes the effect of state error and its derivative on the agent. The terms in Equation (3) and Adjust the smoothness of the movements and states separately. This improves the coordination of relative steering movements. During training, the scaling factor for each reward is adjusted to maximize the reward.
[0135] (2) Adaptive Module
[0136] Based on state feature analysis and a fundamental policy model of the desired motion trajectory, the agent is able to make decisions at the kinematic level. The basic policy π enables motion decisions at the kinematic level. However, the time-varying characteristics of system parameters and external uncertainties may affect the coordination of motion. To address this issue, this scheme designs an adaptive module ρ, which is activated by a set of learned adaptive weights to ensure stability and convergence.
[0137] The aircraft has a very large and variable mass, which leads to changes in system parameters and affects the system's dynamic characteristics. Furthermore, changes in mass also affect friction parameters, causing problems such as low-speed crawling or steady-state oscillations during system motion. Additionally, since the aircraft-trailer system is in a low-speed motion state during taxiing, this design requires consideration of the nominal inertia matrix. and uncertain external friction Compensation. Definition. for,
[0138]
[0139] Where s = [s t [f] is the input vector. It is the Gaussian function, W M It involves learning network weights; the adaptive learning principle is...
[0140] For the adaptive learning module, a radial basis function-based neural network is used to approximate the uncertain dynamics of the system. The neural network consists of three feedforward layers. The input layer includes the system state and the defined nonlinear friction force. The hidden layer is a radial basis vector with 10 nodes. The output layer is an adaptive term for the internal system parameters and the external uncertain friction. Finally, It is used as input to the basic policy π.
[0141] To verify the effectiveness of this solution, this embodiment conducts simulation verification. Two scenarios are considered in the simulation environment: First, a sudden change in the quality of the transported load causes a sudden change in the load parameters of the coordinated transport system, verifying the compliant adjustment capability of the coordinated impedance model under this condition. Second, trajectory tracking motion; by designing different motion trajectories, the control accuracy of the proposed control method is verified.
[0142] In the above simulation scenario, the aircraft-trailer system is implemented in a simulation platform based on dynamics engine interaction, and the model scale is scaled 25 times proportionally to the real aircraft-trailer system.
[0143] The following are two case studies described in detail:
[0144] Case 1: Simulation Verification of Mass and Friction Variation Scenarios
[0145] Simulation scenario setup: An aircraft-trailer system moves under varying ground friction conditions, aiming to follow a sinusoidal trajectory. During this motion, the aircraft's mass changes; this simulation verifies the robustness of the controller under abrupt changes in mass and friction parameters. In this training scenario, the angle between the aircraft and the trailer is arbitrary. The agent's goal is to reduce the angle to the desired tracking state.
[0146] like Figure 6 As shown, the agent was trained for 2000 epochs, with 50 iterations per epoch, and the average reward was output. Under different mass and friction conditions, the agent was able to converge well under the designed reward policy. As the number of training iterations increased, the fluctuation of the reward gradually decreased, indicating that the agent possesses memory and adaptive capabilities. Furthermore, the reward also showed different fluctuations under different internal conditions. Particularly in the fourth case (random variation in friction), the step reward also showed some fluctuation, but the average reward, which determines the agent's learning ability, showed a higher smoothness. When the aircraft mass was 25 kg, the agent needed approximately 150 epochs to stabilize, while when the mass was 15 kg, it only needed 80 epochs to adapt. This indicates that internal parameters have a certain influence on the training and learning of the agent.
[0147] Figure 7 The simulation examines the reward variations in a scenario under different terrain conditions. It shows that under the proposed fusion control, the reward converges and stabilizes quickly, indicating that the proposed method has a fast convergence speed. Furthermore, although the step reward fluctuates somewhat under different terrains, the average reward remains generally stable, demonstrating the algorithm's high robustness.
[0148] Case Study 2: Trajectory Tracking Simulation Verification
[0149] To better observe the control performance under different speeds and trajectories, dual-circle trajectory tracking was chosen as the trajectory tracking task. The diameters of the two circles are D1 = 9 meters and D2 = 6 meters, respectively. The tangents of the two circles can be used to verify the accuracy of repeatability. In this simulation, the initial position of the trailer is set to [ g x t,0 , g y t,0 , g γ t,0 ] = [-0.2m, -0.3m, 12°], the desired initial position is [ g x t,des , g y t,des , g γ t,des [] = [0m, 0m, 0°]. This is to verify the convergence and dynamic response capabilities of the proposed algorithm.
[0150] like Figure 8 As shown, the aircraft-trailer system should achieve the desired trajectory tracking within a given time t = 400 seconds. At approximately t = 15 seconds, the trailer pulls the aircraft into the desired trajectory. The green and red solid lines represent the trajectories of the trailer and aircraft, respectively. At t = 100 seconds, half of the trajectory tracking of a circle with a diameter D1 = 9m has been completed. The black dashed line represents the desired trajectory. At t = 200 seconds, the system has completed half of the trajectory tracking. At t = 200 seconds, the system has completed trajectory tracking. The proposed controller achieves effective trajectory tracking and a high level of repeatability accuracy.
[0151] The control rules of trailers, such as Figure 9 As shown, this also reflects the relevant characteristics. In the first step (approximately 10 seconds), rapid linear and angular velocities are used to correct the initial state error. The second and third steps correspond to the trajectory tracking of the large and small circles, respectively. It can be observed that the control law is very smooth at the switching point (t = 200 seconds). Figure 10 This represents the tracking error of the aircraft and trailer in Cartesian space. Position and attitude errors are distinguished using solid and dashed lines, with different colors used to differentiate different dimensions. Despite the aircraft's large mass, high inertia, and lack of trajectory correction capabilities, the proposed controller remains highly effective in controlling the tracking error. Particularly in the initial stages, the aircraft's tracking error is smaller than the trailer's.
[0152] Figure 11 The reward for the designed reinforcement learning (RL) agent demonstrates better convergence in dynamically coupled environments. From Figure 11As shown in the reward convergence diagram, during the two laps of the aircraft, the reinforcement learning algorithm mainly aims to achieve stable control at different angular velocities. This requires the angular velocity to converge quickly to the desired angle. The reward convergence speed along the way can reflect the error tracking of the angular velocity. Figure 11 The relatively fast convergence speed ensured the accuracy of trajectory tracking. Despite various instabilities in this simulation, the cumulative reward demonstrated good convergence and stability. This is because the agent was designed with the system's coupling states fully in mind, and further adjustments were made during the controller design process.
[0153] In summary, this solution addresses the trajectory tracking control problem of an aircraft-trailer system by proposing a method combining a model-based controller and a reinforcement learning algorithm. This method first designs an adaptive velocity controller based on dummy variables through kinematic decoupling. Then, it designs a reinforcement learning agent to handle system coupling and uncertainties for attitude control. This enables the controller to accurately track and control the trajectory despite uncertainties caused by model coupling and external disturbances, thereby achieving motion coordination of the aircraft-trailer system.
[0154] This approach achieves precise control of system coupling characteristics by combining model-based and data-driven techniques. First, a mathematical model is established and coupling characteristics are analyzed. Then, dummy variables are designed to decouple the system position. Next, a velocity controller is constructed based on Lyapunov stability, and intelligent directional control is achieved through a reinforcement learning agent. Finally, by integrating the closed-loop control law and the RL agent output, the control strategy is optimized, control performance is improved, and the complexity and training time of the RL agent are reduced. Simulation tests verify that the aircraft can achieve precise trajectory tracking using this method, demonstrating significant advantages in improving control accuracy, enhancing system stability, and optimizing control strategies.
Claims
1. A method for coordinating control of an aircraft-trailer system using reinforcement learning, the method comprising: The method comprises the following steps: S1, constructing a mathematical model of the aircraft-trailer system to determine the coupling relationship between control variables and state variables, wherein the control variables include trailer movement control variables and system steering control variables; S2, designing a model-based adaptive controller for the trailer movement control variables; designing a reinforcement learning-based attitude compensation controller for the system steering control variables; S3, fusing the adaptive controller and the attitude compensation controller to output corresponding control signals to the aircraft and the trailer to complete the motion control of the aircraft and the trailer; Step S2 includes the following processes: S21, for trailer moving control variable x t and y t , introduce virtual variable μ t , realize decoupling of x t and y t , then design adaptive control rate, close loop control trailer moving speed; S22, for the system steering control variable γ t and γ m , build reinforcement learning agent to achieve compliant operation of steering; The virtual variable μ in step S21 t is determined by the steering variable γ t and γ m together and satisfies: lim t→∞ |e γ,t ,e γ,m |=0, the virtual variable μ t is specifically: where x e and y e are the errors of the airplane-trailer system in the x g and y g directions in the world coordinate system, c s = [c s,k1 , c s,d1 , c s,k2 , c s,d2 ] are control parameters for adjusting the gains of the errors x e , y e , , and is the desired speed of the trailer, e γ,t is the trailer rotation angle error, and e γ,m is the rotation mechanism rotation angle error. The specific process of step S21 is: For trailer speed control, we have and As control variables, these three control variables are coupled, we choose and as independent variables, and define the error as x e,t and y e,t Then introduce a virtual input μ t , we get: Defining the Lyapunov function And derive: Design a position control law: where c x,k and c y,k are control parameters, the velocity control law v t and the adaptive law are set as, respectively, wherein is an auxiliary signal, λ k is a positive real number, c x,d is a normal number; Step S22 employs a double-layer reinforcement learning architecture for synthesizing the virtual variable μ t and a pose compensation controller of the steering angular velocity ω m , the double-layer reinforcement learning architecture comprising a base policy module π and an adaptive module ρ, the base policy module π taking as input the current state, the previous action, different desired trajectories, and an adaptive learning rate and outputting a steering motion control policy, the adaptive module ρ being trained to improve the adaptability of the system to uncertainties, including internal parameters and external variations.
2. The method of claim 1, wherein The specific process of step S1 is: wherein the definitions of∑Oa,∑Om and∑Ot are frames fixed to the aircraft, the steering mechanism and the trailer respectively, the frame∑Oa being used to describe the aircraft position p a , the frame∑Om being used to describe the steering mechanism position p m , the frame∑Ot being used to describe the trailer position p t ; Definition s α = [x a , y a , γ a ] T denotes the attitude of the airplane in the global frame ΣOg, s t = [x t , y t , γ t ] T denotes the attitude of the trailer in the global frame ΣOg; Taking into account the motion state of the aircraft, s a and the relationship between s t is sγ a = sin(γ a ) sγ m = sin(γ m ) cγ a = cos(γ a ) cγ m = cos(γ m ) where γ t is the rotation angle of the trailer, γ m is the rotation angle of the steering mechanism, γ a is the rotation angle of the airplane, L ma and L tm are the distances between the steering mechanism and the airplane, and between the trailer and the steering mechanism, respectively. For s a and s t between s and s is derived: where q s is the joint space state of the aircraft-trailer system, S(s) is the state transition matrix.
3. The method of claim 1, wherein, The working process of the basic strategy module π includes: In each step, the input state is defined as s s = [s a , s t , s m ] is the current global pose of the trailer, the airplane and the steering mechanism, is the first derivative of s s , s a,d is the motion reference trajectory of the airplane; Action is denoted as a t = [ω m , μ t ] T , ω m is the action sequence that updates the steering mechanism angular velocity, the designed virtual input μ t is used to adjust the angular velocity of the system; The goal-conditioned reinforcement learning problem is translated into a Markov Decision Process: where the elements are state, action, transition state, reward function and a discount factor γ, for any state-action sequence t-1 s t ,r t >, the RL agent is trained to build a policy π that maximizes the expected return, then the agent solves the problem based on the above policy, the policy π is defined as, where τ = {(s1, a1, r1), (s2, a2, r2), …, (s t , a t , r t )} is the sequence when the agent executes the policy π, and P = [P1, P2, …, P t ] is the sequence of possibilities of actions under π. To guide the agents to coordinate motion under the coupling and nonlinearity of the aircraft-trailer system, the reward function is designed to include: task-oriented and performance improvement, wherein the task-oriented reward is used to encourage the aircraft-trailer system to track the expected trajectory s a,d in different environments, and the performance improvement reward focuses on dynamic adjustment, in addition, the reward function also considers real-time state and predicted state.
4. The method of claim 3, wherein, The task-oriented reward is defined as: where r0is the initial reward, e s = [e a < e t < e m ] is the state error matrix, e s = e s,t - e s,d , 0 is a tunable parameter, A e is a threshold value, f(s, a) is a defined reward rate function, P1, P2, and P3 are positive definite matrices, e s T P1e s describes the impact of the state error on the reward function; The performance improvement reward is defined as: where, is the difference between the action values at t and t-1, e tm = e t - e m is the relative error between the trailer and the steering mechanism, P a , P e and P tm are positive definite matrices, and are used to adjust the smoothness of the actions and states, respectively, are used to improve the coordination of the relative steering motion, during training, by adjusting the scaling factors of each reward term to maximize the reward.
5. The method of claim 3, wherein, The working process of the adaptive module ρ includes: to the nominal inertia matrix and uncertain external friction is compensated, defined as: where s = [s t ,f] is the input vector, is the Gaussian kernel, W M are the learning network weights, and the adaptive learning rule is For the adaptive module, a neural network based on radial basis functions is used to approximate the uncertain dynamics of the system, which consists of three layers of feedforward layers, the input layer includes the system states and the defined nonlinear friction force, the hidden layer is a radial basis vector with 10 nodes, and the output layer is the adaptive term of the internal system parameters and the external uncertain friction. Finally, the input of the base policy π.