Independent drive electric vehicle stability undisturbed switching control method
By combining model predictive control with a perturbation-free switching method based on deep reinforcement learning, the stability control problem of independently driven electric vehicles under complex operating conditions is solved, achieving high-precision and smooth stability control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGAN UNIV
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-05
AI Technical Summary
Existing stability control methods for independently driven electric vehicles rely on precise vehicle models, which are difficult to adapt to complex operating conditions. Furthermore, data-driven methods face challenges in terms of interpretability and generalized safety.
Combining model predictive control and deep reinforcement learning, a disturbance-free switching control method is designed using a two-degree-of-freedom vehicle dynamics model and the Actor-Critic deep reinforcement learning framework. This method integrates the prior knowledge of the physical model with the adaptive capabilities of data-driven approaches, and utilizes a long short-term memory network to process time-series parameters to optimize yaw moment control.
It improves control accuracy under complex working conditions, avoids abrupt changes in control quantities and vehicle impacts, and enhances vehicle stability and ride comfort.
Smart Images

Figure CN121973757A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of electric vehicle technology, specifically relating to a method for stability-uninterrupted switching control of independently driven electric vehicles. Background Technology
[0002] Independent drive electric vehicles allow for independent control of each wheel, resulting in rapid response and representing a crucial direction for the development of new energy vehicles. Vehicle stability control aims to ensure a vehicle's ability to resist external disturbances and maintain stable driving. Existing stability control strategies mainly include direct yaw moment control and active front-wheel steering, with control methods largely based on vehicle models or preset rules.
[0003] However, model-based methods heavily rely on the accuracy of the model. Simplifying the model reduces control precision, while complex, high-precision models are difficult to model and computationally intensive, making them unsuitable for real-time control. Furthermore, vehicle parameters are time-varying and operating conditions are complex, resulting in insufficient robustness of methods based on fixed models or parameters. In recent years, data-driven control methods, such as deep learning, do not rely on precise physical models and achieve control through self-learning capabilities, but they face challenges in interpretability and generalization safety. Summary of the Invention
[0004] To address the aforementioned problems, the purpose of this invention is to provide a stability-uninterrupted switching control method for independently driven electric vehicles, solving the problem of how to organically combine model-based control and data-driven control, leveraging their respective strengths to design a stability control method that is adaptable to complex operating conditions, highly accurate, and stable.
[0005] To achieve the above objectives, the technical solution adopted by the present invention includes:
[0006] A method for stability-unobstructed switching control of an independently driven electric vehicle includes the following steps: Step 1: Establish a two-degree-of-freedom vehicle dynamics model, which includes the vehicle's lateral dynamics equations and vehicle yaw dynamics equations. Step 2: Based on the vehicle's motion state and road adhesion conditions, obtain the ideal state information for stability control. The ideal state information includes the ideal center of gravity sideslip angle and the ideal yaw rate. Step 3: Based on the two-degree-of-freedom vehicle dynamics model in Step 1, design a model predictive stability control method based on the physical model. The model predictive stability control method includes three parts: prediction model, optimization objective, and constraint conditions. The input is the actual state information of the vehicle at the current moment and the ideal state information obtained in Step 2. The output is the additional yaw moment control quantity calculated based on the model predictive control. Step 4: Build the Actor-Critic deep reinforcement learning framework, taking the actual state information of the vehicle and the additional yaw moment control quantity obtained in Step 3 as input, and outputting the additional yaw moment control quantity based on deep reinforcement learning; the actual state information of the vehicle includes the sideslip angle and yaw rate of the center of mass at the current moment and the past T time steps. Step 5: The additional yaw moment control quantity obtained in Step 3 and the additional yaw moment control quantity based on deep reinforcement learning obtained in Step 4 are used as inputs to the model predictive controller and the reinforcement learning controller, respectively; the outputs of the model predictive controller and the reinforcement learning controller are the final additional yaw moment control quantities after perturbationless switching; the model predictive controller adopts the model predictive stability control method, and the reinforcement learning controller adopts the reinforcement learning stability control method; a switching control method is designed to switch between the model predictive controller and the reinforcement learning controller; Step six: Based on the tire adhesion utilization rate, allocate the final additional yaw moment control amount obtained in step five after the seamless switching to each drive wheel.
[0007] Preferably, the vehicle's lateral dynamics equations in step one are as follows:
[0008] The vehicle yaw dynamics equations are as follows:
[0009] in, , Indicates the lateral stiffness of the front and rear axles; The value represents the sideslip angle of the center of mass; u represents the longitudinal velocity; a and b represent the distances from the center of mass to the front and rear axes, respectively. Indicates yaw rate. This represents the rate of change of yaw rate; Indicates the front wheel steering angle; Indicates the quality of the car; Indicates lateral velocity. Indicates the rate of change of lateral velocity; Indicates that the car is going around Moment of inertia of the shaft.
[0010] Preferably, the ideal centroid sideslip angle in step two and ideal yaw rate The formula for calculation is:
[0011] in, Indicates the wheelbase of the car; Represents the stability factor. ; μThis represents the road adhesion coefficient under different road types; denoted by , g represents the mass of the car; g represents the acceleration due to gravity; sgn represents the sign function.
[0012] Preferably, the model prediction stability control method in step three specifically includes: Based on the two-degree-of-freedom vehicle dynamics model from step one, a nonlinear state-space prediction model for the control system is established, and its nonlinear state-space equations are as follows:
[0013] in, , , Indicates the change in state; The constraints are ,in and These are the minimum and maximum allowable values for the additional yaw moment, respectively. The optimization goal is , representing the actual state Tracking the ideal state The cost, where Q is the weight matrix of the target being tracked. Let W represent the control input, and let W represent the weight matrix of the control input.
[0014] Preferably, in step four, the Actor-Critic deep reinforcement learning framework includes an Actor network and a Critic network; the Actor network includes a first input layer, a first long short-term memory network, a first fully connected layer, a second fully connected layer, and a first output layer connected in series; the Critic network includes a second input layer, a first long short-term memory network, a third fully connected layer, a fourth fully connected layer, and a second output layer connected in series, with the first output layer also connected in series with the third fully connected layer; the actual state information of the vehicle is used as the input to the first and second input layers, and the additional yaw moment control quantity is used as the input to the third fully connected layer.
[0015] Preferably, the Long Short-Term Memory (LSTM) network comprises an LSTM input layer, an LSTM hidden layer, an LSTM fully connected layer, and an LSTM output layer connected in sequence; the LSTM input layer is used to receive the actual state information of the vehicle; the LSTM hidden layer is used to extract the time-series features of the actual state information of the vehicle; the LSTM fully connected layer uses ReLU as the activation function to achieve feature mapping; and the LSTM output layer is used to output the additional yaw moment control quantity or the corresponding value function.
[0016] Preferably, step five, the switching control method, specifically includes: defining a control flag (Flag); when Flag=1, selecting the reinforcement learning controller; and when Flag=0, selecting the model prediction controller.
[0017] in, To enhance the instantaneous reward under learning stability control, To predict the instantaneous return under stability control for the model.
[0018] Preferably, step six specifically includes: Step 601: Calculate the adhesion utilization rate of a single tire based on the ratio of the utilized adhesion of a single wheel to the maximum adhesion that the ground can provide.
[0019] Where i=1,2,3,4 represent the left front wheel, right front wheel, left rear wheel, and right rear wheel, respectively; This represents the longitudinal force of the i-th tire; This represents the lateral force of the i-th tire; This represents the vertical force of the i-th tire; Indicates the road surface adhesion coefficient; Step 602, using the sum of the squares of the adhesion utilization rates of the four tires as the optimization objective function:
[0020] Where i=1,2,3,4 represent the left front wheel, right front wheel, left rear wheel, and right rear wheel, respectively; This represents the longitudinal force of the i-th tire; This represents the lateral force of the i-th tire; This represents the vertical force of the i-th tire; Indicates the road surface adhesion coefficient; Step 603 satisfies the following equality constraints:
[0021] Among them, F x1 F x2 F x3 and F x4 F represents the final additional yaw moment control amount for the left front wheel, right front wheel, left rear wheel, and right rear wheel of the car; d Represents the total driving force of the car; d represents the wheelbase. This represents the final additional yaw moment control amount after a bumpless switch.
[0022] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the independent drive electric vehicle stability-uninterrupted switching control method of this application.
[0023] A computer program product includes a computer program / instructions that, when executed by a processor, implement the independent drive electric vehicle stability-uninterrupted switching control method of this application.
[0024] Compared with the prior art, the advantages of the present invention are: (1) The independent drive electric vehicle stability disturbance-free switching control method of the present invention integrates the advantages of model predictive control and deep reinforcement learning control. It utilizes the prior knowledge of the model and leverages the adaptive capability of data-driven nonlinear and time-varying characteristics, thereby improving the control accuracy under complex working conditions and avoiding the negative impact of tire lateral stiffness and load transfer changes on control under extreme working conditions.
[0025] (2) The independent drive electric vehicle stability disturbance-free switching control method of the present invention is different from the traditional reinforcement learning method based on fully connected networks. The present invention uses a long short time memory network (LSTM) to build an Actor-Critic architecture, thereby better processing time series parameters and improving the network learning effect.
[0026] (3) The independent drive electric vehicle stability disturbance-free switching control method of the present invention has designed a disturbance-free switching mechanism to avoid the control quantity jump and vehicle impact caused by direct switching of different control strategies, and ensure the smoothness of the vehicle stability control process and the ride comfort. Attached Figure Description
[0027] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the following detailed description to explain the invention, but do not constitute a limitation thereof. In the drawings: Figure 1 It is a two-degree-of-freedom dynamic model of a vehicle.
[0028] Figure 2 This is a model predictive control flowchart.
[0029] Figure 3 This is a diagram of the Actor-Critic architecture under an LSTM network.
[0030] Figure 4 This is the stability switching control diagram of the present invention.
[0031] Figure 5 This is the logic diagram for the non-disruptive switching of the present invention.
[0032] Figure 6 This invention compares the yaw rate under low-adhesion road conditions, with the steering wheel in a sinusoidal state and the vehicle speed at 108 km / h, with the yaw rate under no-control conditions.
[0033] Figure 7This invention compares the center of gravity sideslip angle under low-adhesion road conditions, with the steering wheel in a sinusoidal state and the vehicle speed at 108 km / h, with that under no-control conditions.
[0034] Figure 8 This invention compares the lateral acceleration under low-adhesion road conditions, with the steering wheel in a sinusoidal state and the vehicle speed at 108 km / h, with that under no-control conditions. Detailed Implementation
[0035] The invention is not limited to the specific embodiments described below. All equivalent modifications made based on the technical solutions of this application fall within the protection scope of this invention. Unless otherwise specified, all components and devices in this invention utilize components and devices known in the prior art.
[0036] This application discloses a method for stability-unobstructed switching control of independently driven electric vehicles, including the following steps: Step one: This method only studies the lateral motion of the vehicle and assumes that the longitudinal velocity of the vehicle remains constant. Therefore, a two-degree-of-freedom vehicle dynamics model is established, such as... Figure 1 As shown, the two-degree-of-freedom vehicle dynamics model includes the vehicle lateral dynamics equation and the vehicle yaw dynamics equation; The microequations of the two-degree-of-freedom vehicle dynamics model in this embodiment are:
[0037]
[0038] in, , Indicates the lateral stiffness of the front and rear axles; The value represents the sideslip angle of the center of mass; u represents the longitudinal velocity; a and b represent the distances from the center of mass to the front and rear axes, respectively. Indicates yaw rate. This represents the rate of change of yaw rate; Indicates the front wheel steering angle; Indicates the quality of the car; Indicates lateral velocity. Indicates the rate of change of lateral velocity; Indicates that the car is going around Moment of inertia of the shaft.
[0039] Step two: Based on the vehicle's motion state and road adhesion conditions, obtain the ideal state information for stability control. The ideal state information includes the ideal centroid sideslip angle. and ideal yaw rate The specific formula is as follows:
[0040] in, Indicates the wheelbase of the car; Represents the stability factor. ; μ This represents the road adhesion coefficient under different road types; denoted by , g represents the mass of the car; g represents the acceleration due to gravity; sgn represents the sign function.
[0041] Step 3, as Figure 2 As shown, a model predictive stability control method based on a physical model is designed. The model predictive stability control method consists of three parts: a prediction model, an optimization objective, and constraints. The inputs are the actual state information of the vehicle at the current moment and the ideal state information obtained in step two. The output is the additional yaw moment control quantity calculated based on the model predictive control.
[0042] The model prediction stability control method in this embodiment specifically includes: Based on the two-degree-of-freedom vehicle dynamics model from step one, a nonlinear state-space prediction model for the control system is established, and its nonlinear state-space equations are as follows:
[0043] in, , , Indicates the change in state; The constraints are ,in and These are the minimum and maximum allowable values for the additional yaw moment, respectively. The optimization goal is , representing the actual state Tracking the ideal state The cost, where Q is the weight matrix of the target being tracked. Let W represent the control input, and let W represent the weight matrix of the control input.
[0044] Step 4: Build the Actor-Critic deep reinforcement learning framework, taking the actual state information of the vehicle and the additional yaw moment control quantity obtained in Step 3 as input, and outputting the additional yaw moment control quantity based on deep reinforcement learning; the actual state information of the vehicle includes the sideslip angle and yaw rate of the center of mass at the current time and the past T time steps.
[0045] The Actor-Critic deep reinforcement learning framework of this embodiment includes an Actor network and a Critic network. The Actor network includes a first input layer, a first long short-term memory network, a first fully connected layer, a second fully connected layer, and a first output layer, which are connected in series. The Critic network includes a second input layer, a first long short-term memory network, a third fully connected layer, a fourth fully connected layer, and a second output layer, which are connected in series. The first output layer is also connected in series with the third fully connected layer. The actual state information of the vehicle is used as the input to the first input layer and the second input layer, and the additional yaw moment control quantity is used as the input to the third fully connected layer.
[0046] The Long Short-Term Memory network in this embodiment includes an LSTM input layer, an LSTM hidden layer, an LSTM fully connected layer, and an LSTM output layer connected in sequence. The LSTM input layer is used to receive the actual state information of the vehicle; The LSTM hidden layer is used to extract time-series features of the vehicle's actual state information; The LSTM fully connected layer uses ReLU as the activation function to achieve feature mapping; The LSTM output layer is used to output additional yaw moment control quantities or the corresponding value functions.
[0047] The Actor-Critic deep reinforcement learning framework built in this embodiment is as follows: Figure 3 As shown, the specific steps include: Step 401: Build the Actor-Critic network framework. The input of the Actor network is the vehicle state, and the output is the control action; the input of the Critic network is the vehicle state, and the output is the value function.
[0048] Step 402: Use a Long Short-Term Memory (LSTM) neural network to record the current state of the vehicle's center of gravity sideslip angle and yaw rate, as well as the historical state information of the past 9 time steps. This information serves as the trend information of the vehicle's state data. The data is input into the LSTM neural network, and the output of the network is passed to the fully connected network as shown in the following formula.
[0049] in, t For time, s t Let t represent the vehicle state at time t. h t 1. Historical status information h t = s t T , s t T+1 To integrate the current vehicle status with historical status information, w These are network parameters. f The mapping relationship represented by the LSTM network. o t For network output; Step 403: Select RuLU as the activation function for the fully connected network and output the action and value functions.
[0050] In this embodiment, the Actor-Critic deep reinforcement learning framework constructs a reinforcement learning training sample set based on a vehicle dynamics simulation environment during the training phase. The vehicle state sequence is used as the network input, and the additional yaw moment is used as the network output. The control effect is evaluated through a reward function, and a deep reinforcement learning method based on policy gradients is used to iteratively update the network parameters until the network output control policy meets the preset stability control requirements. The specific interactive environment design is as follows: State space design: Using the ideal state information obtained in step two as the control objective, construct the state space. ,in The yaw rate is angular velocity. It is the centroid sideslip angle.
[0051] Motion space design: Setting up motion space ,in To add yaw moment control, this motion space The output from the Actor network will directly affect the vehicle's dynamic state; Reward function design: A reward function is established based on the degree to which the vehicle's actual state information tracks the ideal state information. R t .
[0052] Step 5: The additional yaw moment control quantity obtained in Step 3 and the additional yaw moment control quantity based on deep reinforcement learning obtained in Step 4 are used as inputs to the model predictive controller and the reinforcement learning controller, respectively. The outputs of the model predictive controller and the reinforcement learning controller are the final additional yaw moment control quantities after perturbationless switching. The model predictive controller adopts the model predictive stability control method, and the reinforcement learning controller adopts the reinforcement learning stability control method. A switching control method is designed to switch between the model predictive controller and the reinforcement learning controller to eliminate the jumps and shocks caused by hard switching between different controls.
[0053] The switching control method in this embodiment specifically includes: Define a control flag (Flag) based on the relationship between the instantaneous reward RRL of reinforcement learning stability control and the instantaneous reward RMPC of model prediction stability control. When Flag=1, the reinforcement learning controller is selected; when Flag=0, the model prediction controller is selected.
[0054] in, To enhance the instantaneous reward under learning stability control, To predict the instantaneous return under stability control for the model.
[0055] The control logic for seamless switching between the two control methods is as follows: Figure 5 As shown.
[0056] Step six: Based on the tire adhesion utilization rate, allocate the final additional yaw moment control amount obtained in step five after the seamless switching to each drive wheel, specifically including: Step 601: Calculate the utilization rate of a single tire based on the ratio of the utilized adhesion of a single wheel to the maximum adhesion that the ground can provide.
[0057] Where i=1,2,3,4 represent the left front wheel, right front wheel, left rear wheel, and right rear wheel, respectively; This represents the longitudinal force of the i-th tire; This represents the lateral force of the i-th tire; This represents the vertical force of the i-th tire; Indicates the road surface adhesion coefficient; Step 602, using the sum of squares of the utilization rates of the four tires as the optimization objective function:
[0058] Step 603: In order for the vehicle to drive normally, the resultant longitudinal force of each wheel must meet the total driving force requirement of the vehicle, and at the same time, the yaw moment caused by the different longitudinal forces of each wheel must meet the stability control requirements, satisfying the following equation constraints:
[0059] Among them, F x1 F x2 F x3 F x4 Fd represents the final additional yaw moment control amount for the left front wheel, right front wheel, left rear wheel, and right rear wheel of the car; Fd represents the total driving force of the car; d represents the track width. This represents the final additional yaw moment control amount after a bumpless switch.
[0060] This application also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the independent drive electric vehicle stability-uninterrupted switching control method of this application.
[0061] This application also discloses a computer program product, including a computer program / instructions, which, when executed by a processor, implements the independent drive electric vehicle stability-uninterrupted switching control method of this application.
[0062] Example This embodiment presents a stability-unobstructed switching control method for independently driven electric vehicles based on the disclosure of this application. The effectiveness of the invented stability-unobstructed switching control is verified using the MATLAB software environment. Key parameters of the test vehicle are shown in Table 1. The front wheel steering angle sinusoidal signal is selected as the input, the simulated vehicle speed is 108 km / h, and the simulated road surface with adhesion coefficient μ = 0.4 is compared with the model predictive stability control (NMPC), reinforcement learning stability control (RL), unobstructed switching stability control (BT), and no control method.
[0063] Table 1 Vehicle Parameters
[0064] Depend on Figure 6 As shown, the reference range for yaw rate is -0.0854 to 0.0849 rad / s, the range for yaw rate without control is -0.3214 to 0.2602 rad / s, the range for yaw rate using the method of this invention is -0.0993 to 0.1077 rad / s, the range for yaw rate based on reinforcement learning is -0.1100 to 0.1021 rad / s, and the range for yaw rate based on model predictive control is -0.1426 to 0.1360 rad / s. Using the strategy proposed in this invention, the yaw rate can be controlled within a stable range, and compared to no control, the peak yaw rate decreases by 69.10%. Compared to single model predictive control and reinforcement learning stability control, the peak yaw rate decreases by 30.36% and 9.72%, respectively. When the yaw rate changes around 7 seconds, the yaw rate under this invention converges to the vicinity of the reference value better than other control methods.
[0065] like Figure 7As shown, the reference value of the centroid sideslip angle varies from -0.0104 to 0.0104 rad, the range of the centroid sideslip angle without control is -0.1150 to 0.1498 rad, the range of the centroid sideslip angle using the method of this invention is -0.0145 to 0.0103 rad, the range of the centroid sideslip angle based on reinforcement learning is -0.0122 to 0.0147 rad, and the range of the centroid sideslip angle based on model predictive control is -0.0245 to 0.0250 rad. When using the control strategy proposed in this invention, the peak value of the centroid sideslip angle decreases by 93.12% compared to no control. Compared to single model predictive control and reinforcement learning stability control, the peak value of the centroid sideslip angle decreases by 58.80% and 29.93% respectively, and the peak value of the centroid sideslip angle decreases by 42%. When the centroid sideslip angle changes around 7 seconds, the centroid sideslip angle under this invention is closer to the reference value of the centroid sideslip angle than other control strategies.
[0066] like Figure 8 As shown, the range of lateral acceleration without control is -3.1107 to 3.0839 m / s². 2 The lateral acceleration of the method of the present invention varies in the range of -2.1473 to 2.2122 m / s². 2 The lateral acceleration under model predictive control varies from -2.2316 to 2.4126 m / s². 2 The lateral acceleration variation range based on reinforcement learning is -2.2803 to 2.1456 m / s². 2 When using the control strategy proposed in this invention, the peak lateral acceleration decreased by 30.97% compared to no control; compared to single model predictive control and reinforcement learning stability control, the peak lateral acceleration decreased by 28.26% and 26.69%, respectively.
[0067] Existing vehicle stability methods, while improving vehicle stability to some extent, are highly dependent on model parameters. Vehicles possess complex nonlinear characteristics; tire lateral stiffness and axle load are typical time-varying parameters. Therefore, it is difficult to characterize the dynamic characteristics under all operating conditions using a simple physical model. Data-driven control methods, on the other hand, do not require a precise system model and can fit vehicle dynamic characteristics through deep learning using neural networks, which helps improve control performance. However, end-to-end data-driven control lacks interpretability. This invention addresses the characteristics of model-based and data-driven stability control methods by proposing a disturbance-free switching control method for independently driven electric vehicles. This method can select the optimal additional yaw moment in different time domains to track the vehicle's optimal stable state.
[0068] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.
[0069] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0070] Furthermore, the various implementation methods disclosed in this solution can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content invented by this disclosure.
Claims
1. A method for stability-uninterrupted switching control of an independently driven electric vehicle, characterized in that, Includes the following steps: Step 1: Establish a two-degree-of-freedom vehicle dynamics model, which includes the vehicle lateral dynamics equation and the vehicle yaw dynamics equation. Step 2: Based on the vehicle's motion state and road adhesion conditions, obtain the ideal state information for stability control, which includes the ideal center of gravity sideslip angle and the ideal yaw rate. Step 3: Based on the two-degree-of-freedom vehicle dynamics model in Step 1, design a model predictive stability control method based on the physical model. The model predictive stability control method includes three parts: prediction model, optimization objective, and constraint conditions. The input is the actual state information of the vehicle at the current moment and the ideal state information obtained in Step 2. The output is the additional yaw moment control quantity calculated based on the model predictive control. Step 4: Build the Actor-Critic deep reinforcement learning framework, take the actual state information of the vehicle and the additional yaw moment control quantity obtained in Step 3 as input, and output the additional yaw moment control quantity based on deep reinforcement learning. The actual state information of the vehicle includes the sideslip angle and yaw rate at the current moment and the past T time steps; Step 5: The additional yaw moment control quantity obtained in Step 3 and the additional yaw moment control quantity based on deep reinforcement learning obtained in Step 4 are used as inputs to the model prediction controller and the reinforcement learning controller, respectively; the outputs of the model prediction controller and the reinforcement learning controller are the final additional yaw moment control quantities after perturbationless switching; the model prediction controller adopts the model prediction stability control method, and the reinforcement learning controller adopts the reinforcement learning stability control method. Design a switching control method to switch between the model prediction controller and the reinforcement learning controller; Step six: Based on the tire adhesion utilization rate, allocate the final additional yaw moment control amount obtained in step five after the seamless switching to each drive wheel.
2. The stability-uninterrupted switching control method for independently driven electric vehicles as described in claim 1, characterized in that, The lateral dynamics equations of the vehicle in step one are as follows: The vehicle yaw dynamics equations are as follows: in, , Indicates the lateral stiffness of the front and rear axles; The value represents the sideslip angle of the center of mass; u represents the longitudinal velocity; a and b represent the distances from the center of mass to the front and rear axes, respectively. Indicates yaw rate. This represents the rate of change of yaw rate; Indicates the front wheel steering angle; Indicates the quality of the car; Indicates lateral velocity. Indicates the rate of change of lateral velocity; Indicates that the car is going around Moment of inertia of the shaft.
3. The stability-uninterrupted switching control method for independently driven electric vehicles as described in claim 1, characterized in that, The ideal centroid sideslip angle mentioned in step two and ideal yaw rate The formula for calculation is: in, Indicates the wheelbase of the car; Represents the stability factor. ; μ This represents the road adhesion coefficient under different road types; denoted by , g represents the mass of the car; g represents the acceleration due to gravity; sgn represents the sign function.
4. The stability-uninterrupted switching control method for independently driven electric vehicles as described in claim 1, characterized in that, The model prediction stability control method described in step three specifically includes: Based on the two-degree-of-freedom vehicle dynamics model from step one, a nonlinear state-space prediction model for the control system is established, and its nonlinear state-space equations are as follows: in, , , Indicates the change in state; The constraints are ,in and These are the minimum and maximum allowable values for the additional yaw moment, respectively. The optimization goal is , representing the actual state Tracking the ideal state The cost, where Q is the weight matrix of the target being tracked. Let W represent the control input, and let W represent the weight matrix of the control input.
5. The stability-uninterrupted switching control method for independently driven electric vehicles as described in claim 1, characterized in that, In step four, the Actor-Critic deep reinforcement learning framework includes an Actor network and a Critic network; The Actor network comprises a first input layer, a first long short-term memory network, a first fully connected layer, a second fully connected layer, and a first output layer, which are connected in series. The Critic network includes a second input layer, a long short-term memory network, a third fully connected layer, a fourth fully connected layer, and a second output layer connected in series. The first output layer is also connected in series with the third fully connected layer. The actual state information of the vehicle is used as the input to the first and second input layers, and the additional yaw moment control quantity is used as the input to the third fully connected layer.
6. The stability-uninterrupted switching control method for independently driven electric vehicles as described in claim 5, characterized in that, The Long Short-Term Memory network comprises an LSTM input layer, an LSTM hidden layer, an LSTM fully connected layer, and an LSTM output layer connected in sequence. The LSTM input layer is used to receive the actual state information of the vehicle; The LSTM hidden layer is used to extract time-series features of the vehicle's actual state information; The LSTM fully connected layer uses ReLU as the activation function to achieve feature mapping; The LSTM output layer is used to output additional yaw moment control quantities or the corresponding value functions.
7. The stability-uninterrupted switching control method for independently driven electric vehicles as described in claim 1, characterized in that, The switching control method described in step five specifically includes: Define a control flag (Flag). When Flag=1, the reinforcement learning controller is selected; when Flag=0, the model prediction controller is selected. in, To enhance the instantaneous reward under learning stability control, To predict the instantaneous return under stability control for the model.
8. The stability-uninterrupted switching control method for independently driven electric vehicles as described in claim 1, characterized in that, Step six specifically includes: Step 601: Calculate the adhesion utilization rate of a single tire based on the ratio of the utilized adhesion of a single wheel to the maximum adhesion that the ground can provide. Where i=1,2,3,4 represent the left front wheel, right front wheel, left rear wheel, and right rear wheel, respectively; This represents the longitudinal force of the i-th tire; This represents the lateral force of the i-th tire; This represents the vertical force of the i-th tire; Indicates the road surface adhesion coefficient; Step 602, using the sum of the squares of the adhesion utilization rates of the four tires as the optimization objective function: Where i=1,2,3,4 represent the left front wheel, right front wheel, left rear wheel, and right rear wheel, respectively; This represents the longitudinal force of the i-th tire; This represents the lateral force of the i-th tire; This represents the vertical force of the i-th tire; Indicates the road surface adhesion coefficient; Step 603 satisfies the following equality constraints: Among them, F x1 F x2 F x3 and F x4 F represents the final additional yaw moment control amount for the left front wheel, right front wheel, left rear wheel, and right rear wheel of the car; d Represents the total driving force of the car; d represents the wheelbase. This represents the final additional yaw moment control amount after a bumpless switch.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the independent drive electric vehicle stability-uninterrupted switching control method according to any one of claims 1-8.
10. A computer program product, characterized in that, Includes a computer program / instruction, which, when executed by a processor, implements the independent drive electric vehicle stability-uninterrupted switching control method according to any one of claims 1-8.