Path-following control method for intelligent electric vehicles based on Q-learning genetic algorithm
Through the intelligent electric vehicle path tracking control method based on Q-learning genetic algorithm, the system output is redefined and the model is decomposed. Combined with adaptive generalized sliding mode control, the controller parameters are optimized, which solves the problem of tire nonlinearity under high-speed driving and achieves improved path tracking accuracy and dynamic stability.
Patent Information
- Application Number
- CN202410971315.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-07-19
AI Technical Summary
When driving at high speeds, the nonlinear factors of the lateral factors of smart car tires are enhanced, and the path tracking accuracy under steering control is affected. Especially on high-speed and low-adhesion roads, the car is prone to dangerous conditions such as skidding and tailspin. Existing control methods are difficult to ensure dynamic stability during path tracking.
A path tracking control method for intelligent electric vehicles based on Q-learning genetic algorithm is adopted. By redefining the system output, the path tracking model is decomposed into an input-output subsystem and a zero dynamic subsystem. Input-output linearization and adaptive generalized sliding mode control are used to optimize the controller parameters and achieve coordinated control of four-wheel steering and driving torque.
It improves the path tracking capability and dynamic stability of smart electric vehicles under extreme working conditions, prevents excessive wheel slip, and achieves the asymptotic stability of the system and good tracking control effect.
Smart Images

Figure CN118928401B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent electric vehicle path tracking control, and mainly relates to an intelligent electric vehicle path tracking control method based on a Q-learning genetic algorithm. Background Art
[0002] Path-following control is a key technology for achieving intelligent vehicles. It uses vehicle status information and a pre-planned desired path to control the front wheel steering angle in real time, aiming to minimize lateral offset and heading errors between the vehicle and the target path, ensuring the vehicle follows the intended path. Currently, intelligent vehicles can achieve good path-following performance under most operating conditions. However, when the vehicle is traveling at high speeds and with large curvatures, the enhanced lateral nonlinearity of the tires significantly affects the path-following accuracy under steering control, making it difficult for the front wheel steering angle determined by the path-following controller to meet control requirements. On high-speed, low-adhesion roads, the vehicle is prone to dangerous conditions such as skidding and tailspin when cornering. With only the front wheel steering angle as a control variable, the steering controller struggles to maintain both path tracking and vehicle dynamic stability under extreme conditions.
[0003] The steering system directly generates lateral acceleration through tire lateral force, thereby directly affecting lateral displacement. Active yaw moment control (DYC) indirectly affects the vehicle's lateral response by affecting the yaw rate. Coordinated control of DYC and active steering can effectively improve the vehicle's dynamic performance during path tracking. Four-wheel independent drive electric vehicles (FWD) have independently controllable driving / braking torques on all four wheels. This allows for precise and rapid control of individual wheel driving / braking torques and inter-axle and inter-wheel torque distribution within the motor's capabilities. This provides additional degrees of freedom for lateral control and significantly improves the vehicle's maneuverability and stability. Commonly used control methods for coordinating active steering and DYC in FWD intelligent electric vehicles include optimal control, sliding mode control, adaptive inversion control, and model predictive control. Reference 1 [Liang J, Lu Y, Yin G, et al. A Distributed Integrated Control Architecture of AFS and DYC based on MAS for Distributed Drive Electric Vehicles [J]. IEEE Transactions on Vehicular Technology, 2021, (99): 1-1.] regards the AFS system and the DYC system as multi-agents, and proposes a new distributed framework for AFS+DYC system integrated control based on the MAS idea. The collaborative control strategy of the two agents is obtained through the Pareto optimal theory to improve the lateral stability of the vehicle and reduce the driver's computational workload during path tracking. Reference 2 [Hang P, Chen X, Luo F. LPV / H-infinity Controller Design for Path Tracking of Autonomous Ground Vehicles Through Four-WheelSteering and Direct Yaw-Moment Control [J]. International Journal of Automotive Technology, 2019 (4): 20.] A four-wheel steering (4WS) and direct yaw moment control (DYC) system is used to design the path tracking control of an autonomous ground vehicle (AGV). The linear variable parameter (LPV) / H∞ control is used as the upper-level controller, and the weighted least squares (WLS) allocation algorithm is used in the lower layer for torque distribution.Reference 3 [Mashadi, B., Ahmadizadeh, P., Majidi, M. et al. Integrated robust controller for vehicle path following. Multibody Syst Dyn 33, 207–228 (2015).] For an integrated 4WS+DYC control system, a robust controller was designed based on the μ-synthesis method to achieve path following for an AGV. This method has the powerful ability to enable the vehicle to track the desired path in the presence of parameter uncertainty. Both the active steering and DYC coordinated control schemes mentioned above utilize the redundant control degrees of freedom provided by active yaw torque to enhance the path following capability of four-wheel-drive smart electric vehicles while also maintaining dynamic stability. However, these methods rarely consider the effective trade-offs between tire slip, controller parameters, and system dynamic performance during implementation.
[0004] The main idea of sliding mode control is to design a dynamic, continuous or discrete control strategy so that the state of the system can slide on a preset sliding mode surface to achieve the desired performance. Therefore, the present invention takes into account the slip state, trajectory tracking error and yaw stability of the four wheels, redefines the output of the path tracking system, and decomposes the path tracking system into an input-output subsystem and a zero-dynamic subsystem through input-output linearization; for the input-output subsystem, an adaptive generalized sliding mode control method is proposed to make the state of the input-output subsystem quickly follow its ideal value; for the zero-dynamic subsystem, the stability constraint conditions of the zero-dynamic subsystem are obtained through stability analysis, and on this basis, a controller parameter design method based on Q-learning genetic algorithm optimization is proposed to achieve asymptotic stability of the intelligent electric vehicle path tracking control system near the equilibrium point, thereby improving the path tracking capability of the intelligent electric vehicle and ensuring its dynamic stability under extreme working conditions. Summary of the Invention
[0005] In order to improve the path tracking capability of intelligent electric vehicles and ensure their dynamic stability under extreme working conditions, the present invention adopts a method of redefining the system output and proposes an intelligent electric vehicle path tracking control method based on Q-learning genetic algorithm. While ensuring good tracking control effects under different working conditions, it improves the dynamic stability of the system.
[0006] The technical solutions adopted by the present invention to solve the technical problems are as follows:
[0007] A path tracking control method for an intelligent electric vehicle based on a Q-learning genetic algorithm comprises the following steps:
[0008] Step 1: Based on the kinematics and dynamics of the vehicle, a path tracking model for the four-wheel steering and four-wheel independent drive smart electric vehicle is established;
[0009] Step 2: Based on the path tracking model obtained in step 1, redefine the output of the system and decompose the path tracking model into input-output subsystem and zero dynamic subsystem through input-output linearization;
[0010] Step 3: Use the CarSim car model to obtain the vehicle's real-time parameters: Y-axis coordinate, X-axis coordinate, heading angle, lateral velocity and acceleration, longitudinal velocity and acceleration, heading angular velocity, rolling angular velocity of each wheel, vertical force of each wheel, etc.
[0011] Step 4: Based on the Y-axis coordinates, X-axis coordinates, heading angle, lateral velocity and acceleration, longitudinal velocity and acceleration, heading angular velocity, rolling angular velocity of each wheel, vertical force of each wheel, etc. output by the CarSim vehicle model in step 3, a controller parameter design method based on Q-learning genetic algorithm optimization is proposed with the goal of rapid convergence of the zero dynamic subsystem state to obtain the optimized controller parameters;
[0012] Step 5. Based on the given reference trajectory, the input-output subsystem obtained in step 2, the Y-axis coordinate, X-axis coordinate, heading angle, lateral velocity and acceleration, longitudinal velocity and acceleration, heading angular velocity, rolling angular velocity of each wheel, vertical force of each wheel, etc. output by the CarSim vehicle model in step 3, and the controller parameters obtained in step 4, the four-wheel steering angle and four-wheel drive torque are obtained through the input-output subsystem adaptive generalized sliding mode control module and input into the CarSim vehicle model.
[0013] The beneficial effects of the present invention are as follows:
[0014] 1) Based on the dynamic motion mechanism of the vehicle, the present invention establishes a mathematical model of the path tracking system, considers the slip state of the four wheels, trajectory tracking error and yaw stability, redefines the system output, and uses input-output linearization to decompose the path tracking system into input-output subsystems and a zero dynamic subsystem, thus achieving dimensionality reduction and decoupling of the model.
[0015] 2) The present invention proposes an adaptive generalized sliding mode control strategy for the input and output subsystem, so that the state of the input and output subsystem: the four-wheel slip state Δω σ , vehicle lateral motion Δγ and Δv y It converges to zero in a finite time, thus preventing excessive wheel slip while controlling the vehicle's handling stability.
[0016] 4) The present invention aims to achieve rapid convergence of the heading angle following error and lateral displacement following error of the zero dynamic subsystem near the equilibrium point, and proposes a controller parameter design method based on Q-learning genetic algorithm optimization, thereby ensuring the asymptotic stability of the intelligent electric vehicle trajectory tracking control system.
[0017] 5) The method proposed in the present invention is simple, easy to implement, and has a wide range of applications, and is suitable for widespread promotion and application. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a principle block diagram of the intelligent electric vehicle path tracking control method based on the Q learning genetic algorithm of the present invention.
[0019] Figure 2 This is the principle block diagram of the controller parameter design method based on Q learning genetic algorithm optimization of the present invention DETAILED DESCRIPTION
[0020] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] like Figure 1 As shown, the intelligent electric vehicle path tracking control method based on the Q learning genetic algorithm of the present invention includes: step 1, according to the kinematic and dynamic mechanism of the vehicle, establishing a path tracking model of the four-wheel steering and four-wheel independent drive intelligent electric vehicle; step 2, based on the path tracking model obtained in step 1, redefining the output of the system, and decomposing the path tracking model into an input-output subsystem and a zero dynamic subsystem through input-output linearization; step 3, using the CarSim car model to obtain the real-time parameters of the vehicle: Y-axis coordinate, X-axis coordinate, heading angle, lateral speed and acceleration, longitudinal speed and acceleration, heading angular velocity, rolling angular velocity of each wheel, vertical force of each wheel, etc.; step 4, based on the Y-axis coordinate, X-axis coordinate, heading angle output by the CarSim car model in step 3 Angle, lateral velocity and acceleration, longitudinal velocity and acceleration, heading angular velocity, rolling angular velocity of each wheel, vertical force of each wheel, etc., with the goal of rapid convergence of the zero dynamic subsystem state, a controller parameter design method based on Q learning genetic algorithm optimization is proposed to obtain the optimized controller parameters; Step 5, according to the given reference trajectory, the input and output subsystem obtained in Step 2, the Y-axis coordinate, X-axis coordinate, heading angle, lateral velocity and acceleration, longitudinal velocity and acceleration, heading angular velocity, rolling angular velocity of each wheel, vertical force of each wheel, etc. output by the CarSim car model in Step 3, and the controller parameters obtained in Step 4, the four-wheel steering angle and four-wheel drive torque are obtained through the input and output subsystem adaptive generalized sliding mode control module, and they are input into the CarSim car model.
[0022] like Figure 2 As shown, the genetic algorithm optimization steps based on Q learning of the present invention include: step 1, population and parameter initialization; step 2, penalty factor optimization based on Q learning; step 3, termination criterion 1 judgment; step 4, selection, adaptive crossover and mutation operations; step 5, generation of a new generation population; step 6, termination criterion 2 judgment.
[0023] The present invention is based on the intelligent electric vehicle path tracking control method of the Q learning genetic algorithm, and the specific implementation steps are as follows:
[0024] Step 1: Based on the vehicle's dynamics, a path tracking model for a four-wheel drive smart electric vehicle is established. The specific implementation of the path tracking model for a four-wheel drive smart electric vehicle is as follows:
[0025] 1.1 Vehicle dynamics model
[0026] In the controller design, the handling dynamics model of the four-wheel independent drive smart electric vehicle can be described as:
[0027]
[0028] Among them, v x and v y are the longitudinal and lateral velocities of the vehicle, respectively. r is the yaw rate. m and I z is the vehicle mass and yaw inertia. F yf and F yr is the lateral force of the front and rear tires. f and δ r is the front and rear wheel steering angle. f and l r ΔM is the distance between the center of mass of the vehicle and the front and rear axles respectively. z The additional yaw moment generated by the longitudinal force difference between the left and right tires is:
[0029]
[0030] Among them, T qσ is the driving torque of each wheel, σ=fl, fr, rl, rr, representing the left front wheel, right front wheel, left rear wheel, right rear wheel, t wσ Satisfy t wfl =t wfr =t wf , t wrl =t wrr =t wr , t wf , t wr are the wheelbases of the front and rear wheels of the vehicle, R e is the wheel rolling radius.
[0031] When the lateral acceleration is less than 0.4g (g is the acceleration due to gravity), the slip angle does not exceed 4° to 5°. The relationship between the tire lateral force and the slip angle can be expressed as a linear relationship, which can be written as:
[0032] F yf =-k f α f , F yr=-k r α r (4)
[0033] Among them, k f and k r are the cornering stiffness of the front and rear tires respectively. f and α r is the front and rear tire slip angle, defined as:
[0034]
[0035] Assuming that the front and rear wheel steering angles are small, cosδ in equations (1) and (2) is f ≈cosδ r ≈1. Substituting equations (4) and (5), the vehicle model described by equations (1) and (2) can be simplified to the following linear dynamic equations:
[0036]
[0037] Wheel dynamics equation:
[0038]
[0039] Where, ω σ is the rolling angular velocity of each wheel, I ω is the wheel moment of inertia, F zσ The vertical force on each wheel when the vehicle rolls and pitches is not considered, and μ is the road friction coefficient.
[0040]
[0041] Where m ω is the sprung mass, and h is the height of the sprung center of mass.
[0042] The desired wheel rolling angular velocity ω is calculated from the wheel center velocity dσ :
[0043]
[0044] Where η wfl =η wrl =-1,η wfr =η wrr =1.
[0045] Desired wheel rolling angular velocity ω dσ Derivative:
[0046]
[0047] In order to control the tire slip rate, the desired wheel angular velocity is set to the angular velocity of the wheel when it is in pure rolling, so the desired tire slip state is Δω σ =ω σ -ω dσ =0, indicating that the tire has no slip.
[0048]
[0049] 1.2 Path tracking model
[0050] In order to enable the vehicle to accurately track the target path, a path tracking control strategy is designed to reduce the heading angle error The lateral position error Δy is minimized. Δy is the lateral deviation between the center of mass of the car and the nearest point d on the reference path, R is the arc length along the reference path, and are the desired heading angle on the reference path and the actual heading angle of the car, respectively.
[0051] Actual heading angle and target heading angle The difference is the heading angle error, which is
[0052]
[0053] Based on the Serret-Frenet coordinate system and linearized by the small angle assumption, we can obtain:
[0054]
[0055] Where ρ is the curvature corresponding to the reference path point.
[0056] The lateral position error satisfies the following equation:
[0057]
[0058] Under ideal steady-state conditions, Therefore, according to formulas (13) and (14), γ d =v x ρ; v yd = 0, let Δγ = γ - γ d , Δv y =v y -v yd is the desired yaw rate γ d Tracking deviation, expected lateral velocity v yd tracking deviation.
[0059]
[0060]
[0061] make The path tracking model obtained from equations (11), (13)-(16) is:
[0062]
[0063] u z =[δ β ,δ γ ,T fl ,T fr ,T rl ,T rr ] T ,
[0064]
[0065]
[0066]
[0067] Step 2: Based on the path tracking model obtained in step 1, redefine the output of the system and decompose the path tracking model into an input-output subsystem and a zero dynamic subsystem through input-output linearization.
[0068] The specific implementations of the input and output subsystem and the zero dynamic subsystem are as follows:
[0069] In order to control the wheel slip rate while controlling the vehicle's handling stability and prevent excessive wheel slip, the wheel slip state Δω is calculated. σ , vehicle lateral motion Δγ and Δv y As a state variable, let x z1 =Δv y ,x z2 =Δγ,x z3 =Δω fl ,x z4 =Δω fr ,x z5 =Δω rl ,x z6 =Δω rr , considering the influence of the slip state of the four wheels on the trajectory tracking error and yaw stability, the output of the system is redefined as:
[0070]
[0071] In the formula, the design parameter k i1 、k i2 、k i3is an undetermined constant, i=1,2,…,6. By selecting appropriate design parameters k i1 、k i2 、k i3 , which can ensure the internal dynamic stability of the system.
[0072] By taking the derivative of equation (18), we can obtain the input and output subsystem according to the path tracking model (17), heading angle error (12) and lateral position error (13):
[0073]
[0074] Where,
[0075]
[0076] K z3 =[k 12 ,k 22 ,k 32 ,k 42 ,k 52 ,k 62 ] T
[0077]
[0078] Since the trajectory tracking model (17) is 8-dimensional and the input-output subsystem (19) is only 6-dimensional, the remaining 2-dimensional system states, namely the heading angle following error (13) and the lateral displacement following error (14), constitute the internal subsystem of the trajectory tracking control system. It can be seen that by adopting the method of redefining the system output (18), the path tracking control system (17) is decomposed into the input-output subsystem (19) and the internal subsystems (13)-(14).
[0079] When a specific control input u z , so that when the output z of the input-output subsystem (19) is zero, the internal subsystems (13) and (14) are zero dynamic subsystems. Ensuring the local stability of the zero dynamic subsystem can ensure the local asymptotic stability of the entire closed-loop system. Therefore, the present invention proposes an input-output subsystem control strategy based on adaptive piecewise sliding mode for the input-output subsystem, so that the input-output subsystem converges to zero in a finite time; for the internal subsystem, when a specific control input u z When the output of the input-output subsystem is zero, the internal subsystem is the zero dynamic subsystem. By selecting reasonable controller design parameters, the zero dynamic subsystem is made asymptotically stable near the equilibrium point, thereby ensuring the asymptotic stability of the intelligent electric vehicle trajectory tracking control system.
[0080] Step 3: Use the CarSim car model to obtain the vehicle's real-time parameters: Y-axis coordinate, X-axis coordinate, heading angle, lateral speed and acceleration, longitudinal speed and acceleration, heading angular velocity, rolling angular velocity of each wheel, vertical force of each wheel, etc.
[0081] Step 4: Based on the Y-axis coordinates, X-axis coordinates, heading angle, lateral velocity and acceleration, longitudinal velocity and acceleration, heading angular velocity, rolling angular velocity of each wheel, vertical force of each wheel, etc. output by the CarSim vehicle model, with the goal of rapid convergence of the zero dynamic subsystem state, a controller parameter design method based on Q-learning genetic algorithm optimization is proposed. The optimized controller parameters are then sent to the adaptive generalized sliding mode control module of the input and output subsystems. The specific implementation method is as follows:
[0082] The 2D state variables of the internal subsystem (heading angle following error The lateral displacement following error Δy) is derivatized twice so that the control u z Appear:
[0083]
[0084]
[0085] General Substituting (20) and (21), we can get
[0086]
[0087] Where,
[0088]
[0089] From the input and output subsystem, we can see that the system state converges to zero in a finite time, that is, When u z =-(K z1 B z ) -1 [(K z1 A z +K z2 )x z +D z ], and bring it into the internal subsystem (22), we can get the zero dynamic subsystem:
[0090]
[0091] make:
[0092]
[0093] We can get:
[0094]
[0095] Substituting equation (24) into equation (23), we can obtain
[0096]
[0097] From formula (25), we can see that if the matrix [A ο -B ο (K z1 B z ) -1 K N2 K N3 ] is strictly guaranteed to be in the left half plane of the complex plane, then the equilibrium point of the original nonlinear system is asymptotically stable. Therefore, select a suitable redefinition of the system output parameter k i1 、k i2 、k i3 , i=1,2,…,6, so that the matrix [A ο -B ο (K z1 B z ) -1 K N2 K N3 ] are all negative, then the zero dynamic subsystem (25) will be at the equilibrium point The path tracking control system is asymptotically stable.
[0098] In order to better achieve rapid stabilization of the path tracking control system, the present invention takes the rapid convergence of the zero dynamic subsystem state as the goal and proposes a controller parameter design method based on Q-learning genetic algorithm optimization to optimize the set system performance indicators.
[0099] Step 1 Population and parameter initialization
[0100] The parameter k to be optimized i1 、k i2 、k i3 (i=1,2,…,6) represents the gene of the genetic algorithm. All genes are connected in series to form individuals, which are coded with real numbers. Multiple individuals form a group. The initial population is generated randomly, the population size is Size, and the number of generations is m. G , maximum evolutionary generation m Gmax , the number of Q learning iterations is m Q , maximum number of iterations m Qmax , let m G =0, m Q =0, current action a k =0.2, current state s k =x o T, k is the current moment, and other related parameters are initialized.
[0101] Step 2: Penalty factor optimization based on Q-learning
[0102] The quadratic performance index related to the heading control error and its rate of change is used as the optimization objective function:
[0103]
[0104] Where Θ∈R 4×4 is a positive symmetric matrix.
[0105] At the same time, each individual in the population must satisfy [A ο -B ο (K z1 B z ) -1 K N2 K N3 ](25)The real part of the eigenvalue is less than zero:
[0106] I Γ =Re[λ Γ (A o )]<0 (27)
[0107] Where Γ is an integer, 1≤Γ≤4. Γ (A o ) represents the matrix A o The Γth eigenvalue of .
[0108] According to the optimization objective function (26) and the inequality constraint (27), the following penalty function (28) is selected:
[0109]
[0110] Where, β J (m G ) is the penalty factor, β J (m G )=a k , m G For evolutionary algebra.
[0111] In order to strike a balance between the objective function and the constraints, the present invention adopts the Q learning method to optimize the penalty factor in equation (28) in real time. The specific implementation steps are as follows:
[0112] Step 2.1. Calculate s k Execute action a in state k Q k (s k ,a k )value
[0113] Let the fitness function:
[0114] F(m G )=1 / J β (m G )
[0115] According to the current state k and action a k Calculate the fitness values of all individuals in the current population, and then use formula (29) to calculate Q k (s k ,a k ), Q k (s k ,a k ) is in s k Execute action a in state k Q value.
[0116]
[0117] Where, ∑F feasible Represents the sum of the fitness function values of all feasible solutions in the population, num feasible represents the number of feasible solutions in the population, max(F Ξ ) represents the maximum fitness function value of individuals in the population, Ξ is an integer, 0<Ξ<size, size is the population size, ∑sum viol The total number of individuals violating the constraint, ∑mum viol The total number of individuals violating the constraint.
[0118] Step 2.2, calculate the next moment state s k+1 and immediate return value R k
[0119] According to the current state k , bring the design parameters corresponding to each individual in the population into the zero dynamic subsystem (25), and calculate the state s at the next moment k+1 .
[0120] Execute k+1 All possible actions corresponding to k a is the number of all possible actions, Υ=1,2,...,k a , calculate the immediate reward value R according to formula (30) k ;
[0121]
[0122] Where, ε J is a small positive number, r (k+1,Υ) For s k+1 Execute A in state k+1A single reward value obtained when taking any action in .
[0123] Step 2.3, Update Q k (s k ,a k ) value, and obtain the optimal action a k *
[0124] Order s k+1 Any action among all possible actions corresponding to the state is a′, and s k+1 and a′ into equations (28) and (29) to calculate Q k (s k+1 ,a′). Combined with Q k (s k ,a k ), update Q using formula (31) k (s k ,a k ).
[0125]
[0126] Where ε1 is the learning factor and ε2 is the discount factor.
[0127] Use greedy strategy to select the optimal action:
[0128]
[0129] Order s k =s k+1 , a k =a k * , m Q =m Q +1.
[0130] Step 3 Termination Criteria 1
[0131] If the number of Q learning iterations is m Q Equal to the maximum number of Q learning iterations m Qmax , then the iteration ends and outputs a k , jump to step 4; if the number of Q learning iterations is m Q Not equal to the maximum number of Q learning iterations m Qmax , return to step 2.1.
[0132] Step 4: Selection, adaptive crossover, and mutation operations
[0133] will a kSubstitute into formula (28) to calculate the fitness function value of each individual, arrange the population in order from best to worst according to the fitness function value, and determine the selection probability of the individual according to the roulette method to form the father.
[0134] Using the average, maximum, and minimum fitness function values F of the current population avg (m G ), F max (m G ), F min (m G ), estimate the individual diversity of the current population:
[0135] f d (m G )=F avg (m G ) / [ε+F max (m G )-F min (m G )] 0≤m G ≤m Gmax
[0136] Where m G is the evolutionary algebra; ε is a small positive number to ensure that no singularity occurs.
[0137] Design adaptive crossover and mutation operators P c 、P m , so that it can be adaptively adjusted according to population diversity and evolutionary generations:
[0138]
[0139]
[0140] Where m Gmax is the maximum evolutionary generation; P m0 ,P c0 ∈[0,1],b1,b2∈R + , can take any value. Thus, when the population distribution is relatively concentrated (ie tends to local convergence, f d (m G ) is larger), p m will increase, and p c Decrease, otherwise the opposite; in the overall trend, p m and p c Both decrease slowly with the genetic process.
[0141] The population is updated by performing selection, crossover, and mutation operations on the population.
[0142] Step 5: Generate a new generation of population
[0143] The original population and the updated population after selection, crossover and mutation operations are sorted according to fitness values, and the top 50% are selected as the new generation population.
[0144] Let m G =m G +1,m Q =0.
[0145] Step 6 Termination Criteria 2
[0146] If the genetic evolution generation number m G Equal to the maximum evolutionary generation m Gmax , then the iteration ends and the controller parameter k is output i1 、k i2 、k i3 (i=1,2,…,6); if the genetic evolution generation m G Not equal to the maximum evolutionary generation m Gmax , skip to step 2.1.
[0147] Step 5: Based on the given reference trajectory, the input-output subsystem obtained in step 2, the Y-axis coordinate, X-axis coordinate, heading angle, lateral velocity and acceleration, longitudinal velocity and acceleration, heading angular velocity, rolling angular velocity of each wheel, vertical force of each wheel, etc. output by the CarSim vehicle model in step 3, and the controller parameters obtained in step 4, the four-wheel steering angle and four-wheel drive torque are obtained through the input-output subsystem adaptive generalized sliding mode control module and input into the CarSim vehicle model. The specific implementation is as follows:
[0148] For the input-output subsystem (19), the following generalized sliding mode surface is adopted:
[0149]
[0150] Where s g =[s g1 ,s g2 ,s g3 ,s g4 ,s g5 ,s g6 ] T , c s =diag(c s1 ,c s2 ,c s3 ,c s4 ,c s5 ,c s6 ), Γ=diag(G1,G2,G3,G4,G5,G6),
[0151] is the differential operator, csi >0 is the design parameter, i=1,2,...,6.
[0152] Select a complementary sliding surface in space that is orthogonal to equation (34):
[0153]
[0154] Where s h =[s h1 ,s h2 ,s h3 ,s h4 ,s h5 ,s h6 ] T .
[0155] Define s=s g +s h , from (34) and (35) we can get:
[0156]
[0157] For the input-output subsystem (19), the following adaptive generalized sliding mode control strategy is proposed:
[0158]
[0159] In the formula, sgn(s)=[sgn(s1),sgn(s2),sgn(s3),sgn(s4),sgn(s5),sgn(s6)] T , D z The estimated upper bounds of the elements in the matrix, The adaptive rate is designed as follows:
[0160]
[0161] Among them, α D is the design parameter, α D =diag(α D1 ,α D2 ,α D3 ,α D4 ,α D5 ,α D6 ), α Di >0, s|=diag(|s1|,|s2|,|s3|,|s4|,|s5|,|s6|).
Claims
1. A path tracking control method for intelligent electric vehicles based on Q-learning genetic algorithm, characterized in that: The method comprises the following steps: Step 1: Based on the kinematics and dynamics of the vehicle, a path tracking model for the four-wheel steering and four-wheel independent drive smart electric vehicle is established; Step 2: Based on the path tracking model obtained in step 1, redefine the output of the system and decompose the path tracking model into input-output subsystem and zero dynamic subsystem through input-output linearization; Step 3: Use the CarSim car model to obtain the vehicle's real-time parameters: Y-axis coordinate, X-axis coordinate, heading angle, lateral velocity and acceleration, longitudinal velocity and acceleration, heading angular velocity, rolling angular velocity of each wheel, and vertical force of each wheel; Step 4: Based on the Y-axis coordinates, X-axis coordinates, heading angle, lateral velocity and acceleration, longitudinal velocity and acceleration, heading angular velocity, rolling angular velocity of each wheel, vertical force of each wheel, etc. output by the CarSim vehicle model in step 3, a controller parameter design method based on Q-learning genetic algorithm optimization is proposed with the goal of rapid convergence of the zero dynamic subsystem state to obtain the optimized controller parameters; Step 5. Based on the given reference trajectory, the input-output subsystem obtained in step 2, the Y-axis coordinate, X-axis coordinate, heading angle, lateral velocity and acceleration, longitudinal velocity and acceleration, heading angular velocity, rolling angular velocity of each wheel, and vertical force of each wheel output by the CarSim vehicle model in step 3, and the controller parameters obtained in step 4, the four-wheel steering angle and four-wheel drive torque are obtained through the input-output subsystem adaptive generalized sliding mode control module and input into the CarSim vehicle model. The implementation process of step 5 above is as follows: For the input-output subsystem (19), the following generalized sliding mode surface is adopted: Where s g =[s g1 ,s g2 ,s g3 ,s g4 ,s g5 ,s g6 ] T , c s =diag(c s1 ,c s2 ,c s3 ,c s4 ,c s5 ,c s6 ), Γ=diag(G1,G2,G3,G4,G5,G6), is the differential operator, c si >0 is the design parameter, i=1,2,...,6; Select a complementary sliding surface in space that is orthogonal to equation (34): Where s h =[s h1 ,s h2 ,s h3 ,s h4 ,s h5 ,s h6 ] T ; Define s=s g +s h , from (34) and (35) we can get: For the input-output subsystem (19), the following adaptive generalized sliding mode control strategy is proposed: In the formula, sgn(s)=[sgn(s1),sgn(s2),sgn(s3),sgn(s4),sgn(s5),sgn(s6)] T , D z The estimated upper bounds of the elements in the matrix, The adaptive rate is designed as follows: Among them, a D For design parameters, a D =diag(a D1 ,a D2 ,a D3 ,a D4 ,a D5 ,a D6 ),a Di >0,|s|=diag(|s1|,|s2|,|s3|,|s4|,|s5|,|s6|)。 2. The intelligent electric vehicle path tracking control method based on Q learning genetic algorithm as claimed in claim 1, characterized in that: The process of establishing the path tracking model of the four-wheel steering and four-wheel independent drive smart electric vehicle described in step 1 is as follows: 1.1 Vehicle dynamics model In the controller design, the handling dynamics model of the four-wheel independent drive intelligent electric vehicle is: Among them, v x and v y are the longitudinal and lateral velocities of the vehicle, respectively; r is the yaw rate; m and I z is the vehicle mass and yaw inertia; F yf and F yr is the lateral force of the front and rear tires; δ f and δ r is the front and rear wheel steering angle; l f and l r are the distances from the vehicle's center of mass to the front and rear axles, respectively; ΔM z The additional yaw moment generated by the longitudinal force difference between the left and right tires is: Among them, T qσ is the driving torque of each wheel, σ=fl, fr, rl, rr, representing the left front wheel, right front wheel, left rear wheel, right rear wheel, t wσ Satisfy t wfl =t wfr =t wf , t wrl =t wrr =t wr , t wf , t wr are the wheelbases of the front and rear wheels of the vehicle, R e is the wheel rolling radius; When the lateral acceleration is less than 0.4g (g is the acceleration of gravity), the side slip angle does not exceed 4 0 5 0 , the relationship between tire lateral force and sideslip angle can be expressed as a linear relationship, which can be written as: F yf =-k f a f ,F yr =-k r a r (4) Among them, k f and k r are the cornering stiffness of the front and rear tires respectively; α f and α r is the front and rear tire slip angle, defined as: Assuming that the front and rear wheel steering angles are small, cosδ in equations (1) and (2) is f ≈cosδ r ≈1; Substituting equations (4) and (5), the vehicle model described by equations (1) and (2) is simplified to the following linear dynamic equation: Wheel dynamics equation: Where, ω σ is the rolling angular velocity of each wheel, I ω is the wheel moment of inertia, F zσ is the vertical force on each wheel when the vehicle roll and pitch motion are not considered, μ is the road friction coefficient; Where m ω is the sprung mass, h is the height of the sprung center of mass; The desired wheel rolling angular velocity ω is calculated from the wheel center velocity dσ : where η wfl = η wrl = -1, η wfr = η wrr = 1; Desired wheel rolling angular velocity ω dσ Derivative: The desired tire slip state is Δω σ =ω σ -ω dσ =0, indicating that the tire has no slip; 1.2 Path tracking model Design a path tracking control strategy to make its heading angle error The lateral position error Δy is minimized; Δy is the lateral deviation between the center of mass of the vehicle and the nearest point d on the reference path, R is the arc length along the reference path, and are the desired heading angle on the reference path and the actual heading angle of the car; Actual heading angle and target heading angle The difference is the heading angle error, which is Based on the Serret-Frenet coordinate system and linearized by the small angle assumption, we can obtain: Where ρ is the curvature corresponding to the reference path point; The lateral position error satisfies the following equation: Under ideal steady-state conditions, Therefore, according to formulas (13) and (14), γ d =v x ρ; v yd = 0, let Δγ = γ - γ d , Δv y =v y -v yd is the desired yaw rate γ d Tracking deviation, expected lateral velocity v yd tracking deviation; make The path tracking model obtained from equations (11), (13)-(16) is: u z =[δ β ,δ γ ,T fl ,T fr ,T rl ,T rr ] T , 3. The intelligent electric vehicle path tracking control method based on Q learning genetic algorithm as claimed in claim 1, characterized in that: In step 2, the path tracking model is decomposed into input and output subsystems and zero dynamic subsystems by redefining the system output. The implementation process is as follows: The slip state of each wheel Δω σ , vehicle lateral motion Δγ and Δv y Control as a state variable; let x z1 =Δv y ,x z2 =Δγ,x z3 =Δω fl ,x z4 =Δω fr ,x z5 =Δω rl ,x z6 =Δω rr , considering the influence of the slip state of the four wheels on the trajectory tracking error and yaw stability, the output of the system is redefined as: Where, the design parameter k i1 、k i2 、k i3 is an undetermined constant, i=1,2,…,6; by selecting a suitable design parameter k i1 、k i2 、k i3 , which can ensure the internal dynamic stability of the system; By taking the derivative of equation (18), we can obtain the input and output subsystem according to the path tracking model (17), heading angle error (12) and lateral position error (13): Where, K z3 =[k 12 ,k 22 ,k 32 ,k 42 ,k 52 ,k 62 ] T 4. The intelligent electric vehicle path tracking control method based on Q learning genetic algorithm as claimed in claim 1, characterized in that: The controller parameter design method based on Q-learning genetic algorithm optimization described in step 4 is implemented as follows: The 2D state variables of the internal subsystem (heading angle following error The lateral displacement following error Δy) is derivatized twice so that the control u z Appear: General Substituting (20) and (21), we can get Where, From the input and output subsystem, we can see that the system state converges to zero in a finite time, that is, When u z =-(K z1 B z ) -1 [(K z1 A z +K z2 )x z +D z ], and bring it into the internal subsystem (22), we can get the zero dynamic subsystem: make: We can get: Substituting equation (24) into equation (23), we can obtain From formula (25), we can see that if the matrix [A ο -B ο (K z1 B z ) -1 K N2 K N3 ] is strictly guaranteed to be in the left half plane of the complex plane, then the equilibrium point of the original nonlinear system is asymptotically stable; therefore, select an appropriate redefine the system output parameter k i1 、k i2 、k i3 , i=1,2,…,6, so that the matrix [A ο -B ο (K z1 B z ) -1 K N2 K N3 ] are all negative, then the zero dynamic subsystem (25) will be at the equilibrium point It is asymptotically stable at the point where the path tracking control system is asymptotically stable. Aiming at the rapid convergence of the zero-dynamic subsystem state, a controller parameter design method based on Q-learning genetic algorithm optimization is proposed to optimize the set system performance indicators. Step 1 Population and parameter initialization The parameter k to be optimized i1 、k i2 、k i3 (i=1,2,…,6) represents the gene of the genetic algorithm. All genes are connected in series to form individuals, which are coded with real numbers. Multiple individuals form a group. The initial population is generated randomly, the population size is Size, and the number of evolution generations is m. G , maximum evolutionary generation m Gmax , the number of Q learning iterations is m Q , maximum number of iterations m Qmax , let m G =0, m Q =0, current action a k =0.2, current state s k =x o T , k is the current moment, initialize other related parameters; Step 2: Penalty factor optimization based on Q-learning The quadratic performance index related to the heading control error and its rate of change is used as the optimization objective function: min f J =exp(s k Θs k T ) (26) Where Θ∈R 4×4 is a positive definite symmetric matrix; At the same time, each individual in the population must satisfy [A ο -B ο (K z1 B z ) -1 K N2 K N3 ](25)The real part of the eigenvalue is less than zero: I Γ =Re[λ Γ (A o )]<0 (27) Where Γ is an integer, 1≤Γ≤4; λ Γ (A o ) represents the matrix A o The Γth eigenvalue of ; According to the optimization objective function (26) and the inequality constraint (27), the following penalty function (28) is selected: Where, β J (m G ) is the penalty factor, β J (m G )=a k , m G For evolutionary algebra; In order to strike a balance between the objective function and the constraints, the Q-learning method is used to optimize the penalty factor in Equation (28) in real time. The specific implementation steps are as follows: Step 2.
1. Calculate s k Execute action a in state k Q k (s k ,a k )value Let the fitness function: F(m G )=1 / J β (m G ) According to the current state k and action a k Calculate the fitness values of all individuals in the current population, and then use formula (29) to calculate Q k (s k ,a k ), Q k (s k ,a k ) is in s k Execute action a in state k Q value; Where, ∑F feasible Represents the sum of the fitness function values of all feasible solutions in the population, num feasible represents the number of feasible solutions in the population, max(F Ξ ) represents the maximum fitness function value of individuals in the population, Ξ is an integer, 0<Ξ<size, size is the population size, ∑sum viol The total number of individuals violating the constraint, ∑mum viol The total number of individuals violating the constraint; Step 2.2, calculate the next moment state s k+1 and immediate return value R k According to the current state k , bring the design parameters corresponding to each individual in the population into the zero dynamic subsystem (25), and calculate the state s at the next moment k+1 ; Execute k+1 All possible actions corresponding to k a is the number of all possible actions, Υ=1,2,...,k a , calculate the immediate reward value R according to formula (30) k ; Where, ε J is a small positive number, r (k+1,Υ) For s k+1 Execute A in state k+1 A single reward value obtained when taking any action in Step 2.3, Update Q k (s k ,a k ) value, and obtain the optimal action a k * Order s k+1 Any action among all possible actions corresponding to the state is a′, and s k+1 and a′ into equations (28) and (29) to calculate Q k (s k+1 ,a′); combined with Q k (s k ,a k ), update Q using formula (31) k (s k ,a k ); Where ε1 is the learning factor and ε2 is the discount factor; Use greedy strategy to select the optimal action: Let s k = s k+1 , a k = a k * , m Q = m Q + 1; Step 3 Termination Criteria 1 If the number of Q learning iterations is m Q Equal to the maximum number of Q learning iterations m Qmax , then the iteration ends and outputs a k , skip to step 4; If the number of Q learning iterations is m Q Not equal to the maximum number of Q learning iterations m Qmax , return to step 2.1; Step 4: Selection, adaptive crossover, and mutation operations will a k Substitute into formula (28) to calculate the fitness function value of each individual, arrange the population in order from best to worst according to the fitness function value, and determine the selection probability of the individual according to the roulette method to form the father. Using the average, maximum, and minimum fitness function values F of the current population avg (m G ), F max (m G ), F min (m G ), estimate the individual diversity of the current population: f d (m G )=F avg (m G ) / [ε+F max (m G )-F min (m G )] 0≤m G ≤m Gmax Where m G is the evolutionary algebra; ε is a small positive number to ensure that no singularity occurs; Design adaptive crossover and mutation operators P c 、P m , so that it can be adaptively adjusted according to population diversity and evolutionary generations: Where m Gmax is the maximum evolutionary generation; P m0 ,P c0 ∈[0,1],b1,b2∈R + , can be any value; thus, when the population distribution is relatively concentrated (ie tends to local convergence, f d (m G ) is larger), p m will increase, and p c Decrease, otherwise the opposite; in the overall trend, p m and p c All of them decrease slowly with the genetic process; The population is updated by performing selection, crossover, and mutation operations on the population; Step 5: Generate a new generation of population Sort the original population and the updated population after selection, crossover, and mutation operations according to fitness values, and select the top 50% as the new generation population; Let m G =m G +1,m Q =0; Step 6 Termination Criteria 2 If the genetic evolution generation number m G Equal to the maximum evolutionary generation m Gmax , then the iteration ends and the controller parameter k is output i1 、k i2 、k i3 (i=1,2,…,6); if the genetic evolution generation m G Not equal to the maximum evolutionary generation m Gmax , skip to step 2.1.
Citation Information
Patent Citations
Multi-agent-based track tracing control method of reconfigurable modular flexible manipulator
CN109240092A
And inputting saturated automatic driving automovable path tracking control method
CN111176302A