Electric drive system gear shifting process control method based on reinforcement learning

Through the optimized EMT shift process control based on reinforcement learning, the problem of difficult to accurately model friction resistance in traditional modeling methods is solved, efficient and adaptive shift control is achieved, and driving comfort and economy are improved.

CN120274059APending Publication Date: 2025-07-08SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510429810.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional dynamic modeling methods have friction resistance during EMT shifting, which leads to interruption of power or shifting impact during shifting, affecting driving comfort and economy.

Method used

Using a method based on reinforcement learning, a reinforcement learning agent is constructed, and the gear shift process control strategy is optimized through the Q-value function, combined with the simulation model of the servo motor, shift actuator and gearbox, and the DQN algorithm is used for training and optimization.

Benefits of technology

It improves the accuracy and efficiency of gear shift process control, reduces R&D costs, has adaptability, adapts to different models and transmission systems, and meets the power and economic needs of gear shifting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120274059A_ABST
    Figure CN120274059A_ABST
Patent Text Reader

Abstract

The invention discloses an electric drive system gear shifting process control method based on reinforcement learning, and relates to the technical field of electric drive system gear shifting control. The method comprises the steps that a gear shifting strategy model is built, and a target gear in the gear shifting process is obtained; establishing a simulation model of the two-gear EMT transmission system according to the kinetic equation, and receiving an angular displacement parameter of a gear shifting motor and a rotating speed parameter of a driving motor; and establishing a reinforcement learning environment, and selecting a state input variable and an action output variable of the reinforcement learning environment according to the simulation model and the target gear. A reinforcement learning method is adopted, simulation models of a servo motor, a gear shifting execution mechanism and a gearbox are established on Matlab / Simulink, a reinforcement learning agent (RL Agent) is constructed, training is carried out in the simulation models by utilizing a DQN algorithm, and finally verification and optimization are carried out in combination with a hardware test bed, so that the development efficiency of a gear shifting process control system is improved, and the development cost of the gear shifting process control system is reduced. And the research and development cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of shift control for electric drive systems, and specifically to a method for controlling the shift process of an electric drive system based on reinforcement learning. Background Art

[0002] An automatic transmission is a key component for an electric vehicle to achieve automatic shifting, and its control effect directly affects driving comfort and power performance. EMT refers to a mechanical automatic transmission driven by an electric motor (electric drive system), which is similar in structure to a traditional AMT and consists of a mechanical gearbox and a shift actuator. The difference is that EMT relies on an electric motor for driving while AMT relies on an engine for driving. Compared with a traditional engine, an electric motor has more excellent rotational speed dynamic regulation and torque dynamic regulation characteristics, and can achieve power unloading during the shift process by controlling the electric motor. Thanks to the excellent speed regulation characteristics of the electric motor, EMT can achieve rapid synchronization of the rotational speeds of the engaging sleeve and the engaging gear ring during the rotational speed synchronization stage by controlling the drive motor instead of using a synchronizer ring for rotational speed synchronization during the shift process, thereby achieving fast and smooth shifting. This can not only reduce the mass and volume of the transmission, but also reduce the manufacturing cost and improve the economy of the entire vehicle.

[0003] The solution of actively synchronizing rotational speeds during the shift process by controlling the electric motor poses higher requirements for the shift process control method of the electric motor transmission integrated system. If there is a long-time power interruption or a large shift shock during the shift process of EMT, it will significantly affect driving comfort. Therefore, the shift process control method for a clutchless EMT has become a key research focus at present.

[0004] When designing the shift process control method for EMT, there are certain errors in traditional dynamic modeling methods. For example, during the shift process of an automobile, frictional resistances such as dry friction between clutch plates, friction between synchronizer cone surfaces, and friction at gear meshing points have different natures and laws, and the calculation of their frictional resistances is more complex. Therefore, on the premise that it is difficult to accurately model some variables of the EMT variable speed system, the obtained shift process control method has significant deficiencies. For this reason, the present invention proposes a method for controlling the shift process of an electric drive system based on reinforcement learning. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for controlling the shift process of an electric drive system based on reinforcement learning, apply the theory and algorithm of reinforcement learning to the shift process control of EMT, construct a reinforcement learning agent and its input variables, output variables, and reward function after comprehensively considering shift power performance and economic factors, and train the agent to obtain a shift process control strategy with higher accuracy.

[0006] To achieve the above object, the present invention provides the following technical solutions: A control method for the shifting process of an electric drive system based on reinforcement learning, comprising the following steps:

[0007] Establish a simulation model of a two-speed EMT transmission system according to the dynamic equation, and receive the angular displacement parameter of the shifting motor and the rotational speed parameter of the driving motor as inputs;

[0008] Establish a reinforcement learning environment, and select the state input variables and action output variables of the reinforcement learning environment according to the simulation model and the target gear;

[0009] Based on the rotational speed synchronization situation, shifting time, and engagement sleeve displacement during the shifting process, establish a reward function for the reinforcement learning environment;

[0010] Construct a reinforcement learning agent and connect the reinforcement learning agent to the reinforcement learning environment;

[0011] Taking the maximum Q-value function of the reinforcement learning agent as the optimization objective, train the reinforcement learning agent to obtain the optimal control strategy for the shifting process.

[0012] Furthermore, the rotational speed and angular displacement parameters are respectively collected by a rotational speed sensor and an angular displacement sensor.

[0013] Furthermore, the dynamic simulation model of the two-speed EMT transmission system includes a shifting motor model, a driving motor model, a shifting actuator model, and a two-speed EMT gearbox model, specifically as follows:

[0014] (31) The shifting motor model is a brushed DC motor, taking the rotational speed command issued by the shifting motor servo driver as the input, calculating the electromagnetic torque of the shifting motor, and outputting the real-time rotational speed parameter:

[0015] The electromechanical coupling equation of the shifting motor is:

[0016]

[0017] In the formula, J is the motor inertia; w is the motor output angular velocity; is the first derivative of w with respect to time; k d is the motor torque coefficient; i a is the armature current; is i a the first derivative with respect to time; f is the damping coefficient; T L is the load torque; L a is the armature inductance; u i is the excitation voltage; k b is the back electromotive force constant; R a is the armature resistance;

[0018] The electromagnetic torque Te output by the shift motor is obtained by multiplying the torque coefficient k d by the armature current i a ; the angular displacement output by the shift motor is obtained by integrating the angular velocity w output by the motor; the output speed n is obtained according to the conversion formula between angular velocity and speed;

[0019] (32) The drive motor model is a permanent magnet synchronous motor, which takes the speed / torque command issued by the drive motor servo driver as the input and is directly connected to the input shaft of the two-speed EMT transmission model;

[0020] In the speed synchronization stage, the speed of the output shaft is read through the servo driver at the load end, and the speed output by the drive motor is calculated according to the conversion of the speed ratio. This speed command is input into the drive motor servo driver, and the servo driver adjusts the drive motor to run at a given speed through the internally integrated PID control module to achieve the speed synchronization between the engaging sleeve and the engaging gear ring;

[0021] (33) The shift actuator model is a rocker-type shift actuator driven by a shift motor. The torque output by the shift motor is amplified by the gear transmission ratio and finally acts on the engaging sleeve to move the engaging sleeve;

[0022] The calculation formula for the shift force of the shift actuator is as follows:

[0023]

[0024] In the formula, Te is the electromagnetic torque output by the shift motor; R1 is the radius of the first-stage transmission gear of the shift actuator model; R21 is the ratio of the radius of the second-stage transmission gear to the first-stage transmission gear of the shift actuator model; R32 is the ratio of the radius of the third-stage transmission gear to the second-stage transmission gear of the shift actuator model;

[0025] According to the shift force, the current actual displacement of the engaging sleeve can be obtained. The specific calculation formula is as follows:

[0026]

[0027] In the formula, m slv is the mass of the engaging sleeve; x slv is the actual displacement of the engaging sleeve; is the second derivative of x slv with respect to time; F s is the shift force; F f is the resistance; integrating twice continuously with respect to can obtain the actual displacement of the engaging sleeve;

[0028] (34) Two-speed EMT transmission model

[0029] The drive motor of the two-speed EMT transmission model is directly connected to the input shaft, and then coupled to the first and second gear through the intermediate shaft. The load uses a load motor of the same model as the drive motor;

[0030] When modeling the transmission, the calculation formula for the moment of inertia on the input side is as follows:

[0031]

[0032] In the formula, J in is the moment of inertia on the input side; J md is the moment of inertia of the drive motor; J ish is the moment of inertia of the input shaft; J csh is the moment of inertia of the intermediate shaft; J gr1 is the moment of inertia of the first gear; J gr2 is the moment of inertia of the second gear; ig0 is the transmission ratio between the input shaft and the intermediate shaft; i g1 is the transmission ratio between the intermediate shaft and the first gear; i g2 is the transmission ratio between the intermediate shaft and the second gear;

[0033] The output side uses a load motor of the same model as the drive motor to simulate the real load of the vehicle. The calculation formula for the moment of inertia on the output side is as follows:

[0034] J out = J ml + J osh + J slv

[0035] In the formula, J out is the moment of inertia on the output side; J ml is the moment of inertia of the load motor; J osh is the moment of inertia of the output shaft; J slv is the moment of inertia of the engagement sleeve.

[0036] Furthermore, according to the simulation model and the target gear, the state input variables and action output variables of the reinforcement learning environment are selected as follows:

[0037] State input variables: During the gear engagement stage, it is the difference between the expected displacement and the real-time displacement of the engagement sleeve at the target gear; during the rotational speed synchronization stage, it is the difference between the rotational speed of the engaged gear ring and the output shaft.

[0038] Action output variables: The rotational speed command of the EMT shift motor; the rotational speed command of the drive motor.

[0039] Furthermore, based on the rotational speed synchronization situation, shift time, and engagement sleeve displacement during the shift process, a reward function for the reinforcement learning environment is established. The reward function r is calculated by the following formula:

[0040] r = r1 + r2 + r3

[0041] Wherein, r1 is the shift time reward function, r2 is the shift accuracy reward function, and r3 is the rotational speed synchronization reward function.

[0042] Further, the calculation method of the shift time reward function r1 is as follows:

[0043]

[0044] Wherein, delta_t is the time interval from issuing the shift command to the completion of gear engagement.

[0045] Further, the calculation method of the shift accuracy reward function r2 is as follows:

[0046]

[0047] Wherein, x0 is the displacement of the synchronizer sleeve at the expected gear position, and x is the real-time displacement of the synchronizer sleeve.

[0048] Further, the calculation method of the rotational speed synchronization reward function r3 is as follows:

[0049]

[0050] Wherein, n0 is the target rotational speed of the input shaft during the rotational speed synchronization stage, and n is the real-time rotational speed of the input shaft.

[0051] Further, the construction of the reinforcement learning agent and the connection of the reinforcement learning agent to the reinforcement learning environment are as follows:

[0052] (91) Selection of the reinforcement learning agent policy

[0053] The DQN agent is adopted, and the used policy is the ε-greedy policy. The ε-greedy policy randomly selects actions with probability ε to explore different states and actions in the environment; selects the action with the maximum Q value with probability 1-ε: The formula expression of the ε-greedy policy is:

[0054]

[0055] Wherein, a t is the action selected by the DQN agent at time step t; Q(s,a;θ) represents the Q value estimation of the "state-action" pair (s,a) by the policy network under the parameter θ;

[0056] (92) Establishment of the reinforcement learning agent critic network

[0057] The used neural network is a feedforward neural network, including a feature input layer, two fully connected layers, two activation function layers and an output layer;

[0058] Fit a feedforward neural network to the critic network to approximate the Q-value function;

[0059] Let θ be the parameters of the critic network, and Q(s,a;θ) represent the Q-value estimation of the "state-action" pair (s,a) by the network under the parameters θ;

[0060] The DQN updates the network parameters θ by minimizing the loss function. Specifically, the mean squared error loss function is used, and the specific expression is:

[0061] L(θ) = E (s,a,r,s')~U(D) [(r + γmax a' Q(s',a';θ - ) - Q(s,a;θ)) 2

[0062] In the formula, U(D) represents the "state-action-reward-next state" quadruple (s,a,r,s') uniformly sampled from the experience replay buffer D; γ is the discount factor; θ - is the parameter of the target network;

[0063] The network parameters θ are updated using the gradient descent method. The specific formula is:

[0064]

[0065] Among them, θ' is the updated network parameter; α is the learning rate, set to α = 0.001; L(θ) is the gradient of the mean squared error loss function L(θ) with respect to θ, and the gradient clipping threshold is set to 1.

[0066] Furthermore, taking the maximum of the action value function of the reinforcement learning agent as the optimization goal, the reinforcement learning agent is trained to obtain the optimal shift process control strategy, specifically as follows:

[0067] (101): Use the difference between the expected displacement of the shift sleeve and the real-time displacement of the shift sleeve during the upshift stage, and the difference between the rotational speed of the engagement gear ring and the output shaft during the rotational speed synchronization stage as the state input variables of the reinforcement learning agent; use the rotational speed command of the EMT shift motor and the rotational speed command of the drive motor as the output variables of the reinforcement learning agent;

[0068] (102): The critic network collects the values of the reward function and the state input variables, minimizes the loss function with the help of the DQN algorithm, and optimizes the current shift process control method strategy. Subsequently, the optimized control strategy is output to the reinforcement learning environment through the actuator;

[0069] ​(103): Calculate the reward function for the shifting process under the new shifting method, calculate the Q value and feedback it to the critic network; the reinforcement learning environment passes the new state input variables to the reinforcement learning agent;

[0070] (104): Repeat steps (102) and (103) to maximize the Q value in the Q function. At this time, the shifting process control strategy is the current optimal control strategy, which is used to enable the reinforcement learning agent to achieve the optimal decision-making during the shifting process.

[0071] The present invention has at least the following beneficial effects:

[0072] (1) By adopting the reinforcement learning method, the present invention establishes simulation models of a servo motor, a shifting actuator, and a gearbox on Matlab / Simulink, constructs a reinforcement learning agent (RL Agent), trains it in the simulation model using the DQN algorithm, and finally combines it with a hardware test bench for verification and optimization, improving the development efficiency of the shifting process control system and reducing the R & D cost.

[0073] (2) By adopting the reinforcement learning method, the present invention has strong adaptability and generalization ability. It can continuously interact with the environment and adjust its shifting strategy in real time according to the feedback rewards, and can adapt to different vehicle models and transmission systems. The traditional PID algorithm needs to change parameters to maintain good shifting control effects, which is relatively difficult in practical applications.

[0074] (3) By adopting the reinforcement learning method, while meeting the requirements of vehicle shifting economy and power performance, the present invention can dynamically coordinate the working states of multiple modules such as a drive motor and a shifting motor, achieve the optimal completion of each shifting task, and meet the power performance and economic requirements during vehicle shifting. The traditional shifting method needs to be specifically designed for each variable, and there are large errors in modeling, making it difficult to obtain the optimal shifting strategy.

[0075] Of course, it is not necessary for any product implementing the present invention to simultaneously achieve all the above-mentioned advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 is a schematic flowchart of the control method of the present invention;

[0077] Figure 2 is a schematic structural diagram of a two-speed EMT gearbox in the present invention;

[0078] Figure 3 is a schematic framework principle diagram of the control method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0079] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0080] Please refer to Figure 1 , the present invention provides a technical solution: a method for controlling the shifting process of an electric drive system based on reinforcement learning, including:

[0081] S1. Construct a shifting strategy model to obtain the target gear during the shifting process;

[0082] S2. Establish a simulation model of a two-speed EMT transmission system according to the dynamic equation, and receive the angular displacement parameter of the shifting motor and the rotational speed parameter of the driving motor as inputs;

[0083] The rotational speed and angular displacement parameters are respectively collected by a rotational speed sensor and an angular displacement sensor;

[0084] The dynamic simulation model of the two-speed EMT transmission system includes a shifting motor model, a driving motor model, a shifting actuator model, and a two-speed EMT transmission model, specifically as follows:

[0085] (S21) The shifting motor model is a brushed DC motor. Taking the rotational speed command issued by the shifting motor servo driver as the input, calculate the electromagnetic torque of the shifting motor, and output the real-time rotational speed parameter:

[0086] The electromechanical coupling equation of the shifting motor is:

[0087]

[0088] In the formula, J is the motor inertia; w is the motor output angular velocity; is the first derivative of w with respect to time; k d is the motor torque coefficient; i a is the armature current; is i a about the first derivative of time; f is the damping coefficient; T L is the load torque; L a is the armature inductance; u i is the excitation voltage; k b is the back electromotive force constant; R a is the armature resistance;

[0089] The electromagnetic torque Te output by the shifting motor is multiplied by the torque coefficient k d by the armature current i aObtained; the angular displacement output by the shift motor is obtained by integrating the motor output angular velocity w; the output rotational speed n is obtained according to the conversion formula between angular velocity and rotational speed; the load torque TL of the motor is difficult to accurately represent and can be replaced by a constant first. When the reinforcement learning agent completes training, it will optimize the shift strategy in real time according to the true value of TL;

[0090] (S22) The drive motor model is a permanent magnet synchronous motor, taking the speed / torque command issued by the drive motor servo driver as the input and directly connected to the input shaft of the two-speed EMT transmission model;

[0091] In the speed synchronization stage, the rotational speed of the output shaft is read through the servo driver at the load end, and the rotational speed output by the drive motor is calculated according to the conversion of the speed ratio. This rotational speed command is input into the drive motor servo driver, and the servo driver adjusts the drive motor to run at a given speed through the internally integrated PID control module to achieve the speed synchronization of the engagement sleeve and the engagement gear ring;

[0092] (S23) The shift actuator model is a rocker-type shift actuator driven by a shift motor. The torque output by the shift motor is amplified by the gear transmission ratio and finally acts on the engagement sleeve to move the engagement sleeve;

[0093] The calculation formula for the shift force of the shift actuator is as follows:

[0094]

[0095] In the formula, Te is the electromagnetic torque output by the shift motor; R1 is the radius of the first-stage transmission gear of the shift actuator model; R21 is the ratio of the radius of the second-stage transmission gear to the first-stage transmission gear of the shift actuator model; R32 is the ratio of the radius of the third-stage transmission gear to the second-stage transmission gear of the shift actuator model;

[0096] According to the shift force, the current actual displacement of the engagement sleeve can be obtained, and the specific calculation formula is as follows:

[0097]

[0098] In the formula, m slv is the mass of the engagement sleeve; x slv is the actual displacement of the engagement sleeve; is the second derivative of x slv with respect to time; F s is the shift force; F f is the resistance; Integrating twice continuously can obtain the actual displacement of the engagement sleeve; when conducting physical verification, in this embodiment, an angular displacement sensor is used to measure the rotation angle of the input shaft of the shift actuator, and then the axial displacement of the engagement sleeve is obtained. The resistance F in the above formula fIt cannot be accurately represented and can be replaced by a constant first. After the reinforcement learning agent completes training, it will optimize the shifting strategy in real time according to the true value of F f ;

[0099] (S24) Two-speed EMT transmission model

[0100] The structure of the two-speed EMT transmission model is as Figure 2 shown. The driving motor of the two-speed EMT transmission model is directly connected to the input shaft, and then coupled to the first and second gear through the intermediate shaft. The load is replaced by a load motor of the same model as the driving motor, and a torque is output to simulate the load of the vehicle in the real working condition;

[0101] When modeling the transmission, the calculation formula for the moment of inertia on the input side is as follows:

[0102]

[0103] In the formula, J in is the moment of inertia on the input side; J md is the moment of inertia of the driving motor; J ish is the moment of inertia of the input shaft; J csh is the moment of inertia of the intermediate shaft; J gr1 is the moment of inertia of the first gear; J gr2 is the moment of inertia of the second gear; ig0 is the transmission ratio between the input shaft and the intermediate shaft; i g1 is the transmission ratio between the intermediate shaft and the first gear; i g2 is the transmission ratio between the intermediate shaft and the second gear;

[0104] On the output side, a load motor of the same model as the driving motor is used to simulate the real load of the vehicle. The calculation formula for the moment of inertia on the output side is as follows:

[0105] J out = J ml + J osh + J slv

[0106] In the formula, J out is the moment of inertia on the output side; J ml is the moment of inertia of the load motor; J osh is the moment of inertia of the output shaft; J slv is the moment of inertia of the engaging sleeve;

[0107] For the shifting process of the two-speed EMT transmission system without a clutch, the upshift from the first gear to the second gear is used as an example here:

[0108] Before the start of gear shifting, the system is in the first gear engaged state; when the system receives a gear shifting command, the torque of the driving motor is adjusted to a value close to 0 to reduce the contact force between the first gear engaging ring gear and the engaging sleeve during gear shifting; subsequently, the gear shifting motor outputs torque and drives the engaging sleeve to disengage from the first gear engaging ring gear and enter the neutral gear; when the system is in the neutral gear state, the input end and the output end of the EMT are decoupled, and at this time, the rotational speed of the driving motor needs to be adjusted so that the rotational speed of the second gear engaging ring gear is consistent with the rotational speed of the output shaft, that is, the rotational speed is synchronized; when the rotational speed synchronization stage is completed, the gear shifting motor outputs torque and drives the engaging sleeve to couple with the second gear engaging ring gear to complete the final gear shifting operation, and the gear shifting process ends;

[0109] Regarding the technical solution of this embodiment, in the two-speed EMT transmission system model, the following parameters are set as follows:

[0110] Motor moment of inertia: J = 0.002;

[0111] Motor torque coefficient: kd = 0.06277;

[0112] Damping coefficient: f = 0.000006;

[0113] Armature inductance: La = 0.002;

[0114] Back electromotive force constant: kb = 0.063;

[0115] Armature resistance: Ra = 0.5;

[0116] Radius of the first-stage transmission gear of the gear shifting actuator: R1 = 0.008;

[0117] Ratio of the radius of the second-stage transmission gear to the first-stage transmission gear of the gear shifting actuator: R21 = 1.5;

[0118] Ratio of the radius of the third-stage transmission gear to the second-stage transmission gear of the gear shifting actuator: R32 = 0.057;

[0119] Mass of the engaging sleeve: mslv = 0.4;

[0120] Moment of inertia on the input side: Jin = 0.77;

[0121] Moment of inertia on the output side: Jout = 7.86;

[0122] Transmission ratio between the input shaft and the intermediate shaft: ig0 = 2;

[0123] Transmission ratio between the intermediate shaft and the first gear: ig1 = 2;

[0124] Transmission ratio between the intermediate shaft and the second gear: ig2 = 1.25;

[0125] S3. Establish a reinforcement learning environment, and select the state input variables and action output variables of the reinforcement learning environment according to the simulation model and the target gear position;

[0126] State input variables: During the gear engagement stage, it is the difference between the expected displacement of the sliding sleeve at the target gear position and the real-time displacement of the sliding sleeve; during the rotational speed synchronization stage, it is the difference between the rotational speed of the engaging gear ring and the output shaft;

[0127] Action output variables: The rotational speed command of the EMT shift motor; the rotational speed command of the drive motor;

[0128] S4. Based on the rotational speed synchronization situation, shift time, and sliding sleeve displacement during the shifting process, establish the reward function of the reinforcement learning environment;

[0129] (S41) The reward function r is calculated by the following formula:

[0130] r = r1 + r2 + r3

[0131] In the formula, r1 is the shift time reward function, r2 is the shift accuracy reward function, and r3 is the rotational speed synchronization reward function;

[0132] (S42) Shift time reward function

[0133] The shift time delta_t of the two-speed EMT is an important indicator of the research content of this embodiment. In the simulation, 0.5s is used as the expected value of the shift time, and this expected value can be adjusted. When delta_t < 0.5s, the reward function is an inverse proportional function of delta_t to encourage the agent to minimize the value of delta_t; when delta_t ≥ 0.5s, the reward function is a linear function -20delta_t of delta_t, and the shift time reward function is:

[0134]

[0135] where delta_t is the time interval from issuing the shift command to the completion of gear engagement;

[0136] (S43) Shift accuracy reward function

[0137] The axial displacement error of the engaging sleeve is the difference between the displacement x0 of the engaging sleeve at the expected gear position and the real-time displacement x of the engaging sleeve. If |x - x0| < 0.01m, it can be approximately considered that the engaging sleeve has reached the target gear position, and the reward function value is 10. If |x - x0| ≥ 0.01m, it is considered that the engaging sleeve has not reached the target gear position, and the reward function value at this time is set as a linear function -10|x - x0| of the error |x - x0|. Considering the specific mechanical structure of the gearbox, the error |x - x0| cannot be greater than certain specific values, otherwise it will cause wear between the engaging sleeve and the inner wall of the gearbox, and even cause gear shifting failure, endangering the safety of the vehicle. Therefore, the reward function value here needs to be set as -100 to avoid such situations.

[0138] The reward function for shifting accuracy is:

[0139]

[0140] where x0 is the displacement of the engaging sleeve at the expected gear position, and x is the real-time displacement of the engaging sleeve.

[0141] (S44) Reward function for speed synchronization

[0142] During the gear shifting process of the two-speed EMT, when the engaging sleeve is in the neutral position, it is necessary to wait for the rotational speed of the driving motor to be consistent with the output shaft rotational speed of the engaging gear ring before performing the upshift operation. Otherwise, the phenomenon of "gear clash" will occur, exacerbating the damage to the transmission system. Therefore, during the speed synchronization stage, when the real-time rotational speed n of the input shaft reaches the target rotational speed n0, the reward function value is positive, otherwise it is negative. When the transmission system is in other stages, the reward function r3 has no effect and the value is 0. The reward function for speed synchronization is:

[0143]

[0144] where n0 is the target rotational speed of the input shaft during the speed synchronization stage, and n is the real-time rotational speed of the input shaft.

[0145] S5. Construct a reinforcement learning agent and connect the reinforcement learning agent to the reinforcement learning environment, as follows:

[0146] (S51) Selection of the reinforcement learning agent strategy

[0147] The DQN agent is adopted, and the strategy used is the ε-greedy strategy. The ε-greedy strategy randomly selects actions with probability ε to explore different states and actions in the environment; it selects the action with the largest Q value with probability 1 - ε to utilize the knowledge already learned. As the training progresses, ε usually gradually decreases, enabling the DQN agent to transition from exploration-based to utilization-based.

[0148] The formula expression of the ε-greedy strategy is:

[0149]

[0150] where a t is the action selected by the DQN agent at time step t; Q(s, a; θ) represents the Q-value estimate of the "state-action" pair (s, a) by the policy network with parameters θ;

[0151] (S52) Establish the critic network of the reinforcement learning agent

[0152] The neural network used in this embodiment is a feedforward neural network, which includes a feature input layer (the input dimension is determined by the input variables), two fully connected layers (each layer contains 200 neurons), two activation function layers, and an output layer (the output dimension is equal to the number of actions in the action space);

[0153] In Matlab, the established feedforward neural network is fitted into the critic network to approximate the Q-value function;

[0154] Let θ be the parameter of the critic network, and Q(s, a; θ) represents the Q-value estimate of the "state-action" pair (s, a) by the network with parameters θ;

[0155] The DQN updates the network parameter θ by minimizing the loss function. Specifically, the mean squared error loss function (MSE) is used, and the specific expression is:

[0156] L(θ) = E (s,a,r,s')~U(D) [(r + γ max a' Q(s', a'; θ - ) - Q(s, a; θ)) 2

[0157] where U(D) represents the "state-action-reward-next state" quadruple (s, a, r, s') uniformly sampled from the experience replay buffer D; γ is the discount factor; θ - is the parameter of the target network;

[0158] In Matlab, the gradient descent method is used to update the network parameter θ, and the specific formula is:

[0159]

[0160] where θ' is the updated network parameter; α is the learning rate, set to α = 0.001; L(θ) is the gradient of the mean squared error loss function L(θ) with respect to θ, and the gradient clipping threshold is set to 1 to prevent gradient explosion;

[0161] ​S6. With the maximization of the Q-value function of the reinforcement learning agent as the optimization goal, train the reinforcement learning agent to obtain the optimal shift process control strategy;

[0162] Use the Simulink simulation model to train the RL Agent, and then optimize the shift process control strategy (method) obtained after training in combination with the EMT physical test bench;

[0163] As Figure 3 shown, the shift process control strategy (method) of this embodiment relies on the continuous interaction between the reinforcement learning agent and the reinforcement learning environment, and its general training process is as follows:

[0164] (S61): Use the difference between the expected displacement of the synchronizer sleeve and the real-time displacement of the synchronizer sleeve during the upshift stage, and the difference between the rotational speed of the engaging gear ring and the output shaft during the rotational speed synchronization stage as the state input variables of the reinforcement learning agent; Use the rotational speed command of the EMT shift motor and the rotational speed command for driving the motor as the output variables of the reinforcement learning agent;

[0165] (S62): The critic network collects the values of the reward function and the state input variables, minimizes the loss function with the help of the DQN algorithm, and optimizes the current shift process control strategy. Then, output the optimized control strategy to the reinforcement learning environment through the actuator;

[0166] (S63): Calculate the reward function of the shift process under the new shift method, calculate the Q-value and feedback it to the critic network; The reinforcement learning environment transmits the new state input variables to the reinforcement learning agent;

[0167] (S64): Repeat steps (S62) and (S63) to maximize the Q-value in the Q function. At this time, the shift process control strategy is the current best control strategy, which is used to enable the reinforcement learning agent to achieve the optimal decision-making during the shift process;

[0168] S7. Monitor the shift process in real time during vehicle driving, and regularly update and optimize the shift process control strategy.

[0169] In summary, by adopting the reinforcement learning method, the present invention establishes simulation models of a servo motor, a shift actuator, and a gearbox on Matlab / Simulink, constructs a reinforcement learning agent (RL Agent), uses the DQN algorithm to train in the simulation model, and finally combines with a hardware test bench for verification and optimization, improving the development efficiency of the shift process control system and reducing the R & D cost; by adopting the reinforcement learning method, it has strong adaptability and generalization ability, can continuously interact with the environment, and adjust its shift strategy in real time according to the feedback reward, and can adapt to different vehicle models and transmission systems; by adopting the reinforcement learning method, while meeting the requirements of vehicle shift economy and power performance, it can dynamically coordinate the working states of multiple modules such as a drive motor and a shift motor, achieve the optimal completion of each shift task, and meet the power performance and economic requirements during vehicle shifting.

[0170] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0171] For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. When an element is referred to as "assembled on", "mounted on", "fixed to" or "disposed on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are only for the purpose of illustration and do not represent the only implementation.

[0172] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

[0173] In the description of this specification, the description referring to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

Claims

1. A control method for the shifting process of an electric drive system based on reinforcement learning, characterized in that It includes the following steps: Establish a simulation model of a two-speed EMT transmission system according to the dynamic equation, and receive the angular displacement parameter of the shift motor and the rotational speed parameter of the drive motor as inputs; Establish a reinforcement learning environment, and select the state input variables and action output variables of the reinforcement learning environment according to the simulation model and the target gear; Based on the rotational speed synchronization situation, shift time, and engagement sleeve displacement during the shifting process, establish a reward function for the reinforcement learning environment; Construct a reinforcement learning agent and connect the reinforcement learning agent to the reinforcement learning environment; Taking the maximum Q-value function of the reinforcement learning agent as the optimization goal, train the reinforcement learning agent to obtain the optimal shift process control strategy.

2. The control method for the shifting process of the electric drive system based on reinforcement learning according to claim 1, wherein: The rotational speed and angular displacement parameters are respectively collected by a rotational speed sensor and an angular displacement sensor.

3. The control method for the shifting process of the electric drive system based on reinforcement learning according to claim 2, wherein: The dynamic simulation model of the two-speed EMT transmission system includes a shift motor model, a drive motor model, a shift actuator model, and a two-speed EMT transmission model, specifically as follows: (31) The shift motor model is a brushed DC motor. Taking the rotational speed command issued by the shift motor servo driver as the input, calculate the electromagnetic torque of the shift motor, and output the real-time rotational speed parameters: The electromechanical coupling equation of the shift motor is: Where, J is the moment of inertia of the motor; w is the output angular velocity of the motor; is the first derivative of w with respect to time; k d is the motor torque coefficient; i a is the armature current; is the first derivative of i a with respect to time; f is the damping coefficient; T L is the load torque; L a is the armature inductance; u i is the excitation voltage; k b is the back electromotive force constant; R a is the armature resistance; The electromagnetic torque Te output by the shift motor is obtained by multiplying the torque coefficient k d by the armature current i a ; the angular displacement output by the shift motor is obtained by integrating the angular velocity w output by the motor; The output rotational speed n is obtained according to the conversion formula between angular velocity and rotational speed; (32) The drive motor model is a permanent magnet synchronous motor. Taking the rotational speed / torque command issued by the drive motor servo driver as the input and directly connecting to the input shaft of the two-speed EMT transmission model; During the rotational speed synchronization stage, read the rotational speed of the output shaft through the servo driver at the load end, calculate the rotational speed output by the drive motor according to the conversion of the speed ratio, input this rotational speed command into the drive motor servo driver, and the servo driver adjusts the drive motor to run at a given speed through the internally integrated PID control module to achieve the rotational speed synchronization of the engagement sleeve and the engagement gear ring; (33) The shift actuator model is a rocker-type shift actuator driven by a shift motor. The torque output by the shift motor is amplified by the gear transmission ratio and finally acts on the engagement sleeve to move the engagement sleeve; The shift force calculation formula of the shift actuator is as follows: In the formula, Te is the electromagnetic torque output by the shift motor; R1 is the radius of the first-stage transmission gear of the shift actuator model; R21 is the ratio of the radius of the second-stage transmission gear to the first-stage transmission gear of the shift actuator model; R32 is the ratio of the radius of the third-stage transmission gear to the second-stage transmission gear of the shift actuator model; According to the shift force, the current actual displacement of the engagement sleeve can be obtained, and the specific calculation formula is as follows: Where, m slv is the mass of the engaging sleeve; x slv is the actual displacement of the engaging sleeve; is the second derivative of x slv with respect to time; F s is the shifting force; F f is the resistance; Integrating twice continuously with respect to can obtain the actual displacement of the engaging sleeve; (34) Two-speed EMT transmission model The drive motor of the two-speed EMT transmission model is directly connected to the input shaft, and then coupled to the first and second gear through the intermediate shaft. The load uses a load motor of the same model as the drive motor; When modeling the transmission, the calculation formula of the moment of inertia on the input side is as follows: where, J in is the moment of inertia of the input side; J md is the moment of inertia of the drive motor; J ish is the moment of inertia of the input shaft; J csh is the moment of inertia of the intermediate shaft; J gr1 is the moment of inertia of the first gear; J gr2 is the moment of inertia of the second gear; ig0 is the transmission ratio between the input shaft and the intermediate shaft; i g1 is the transmission ratio between the intermediate shaft and the first gear; i g2 is the transmission ratio between the intermediate shaft and the second gear; On the output side, a load motor of the same model as the drive motor is used to simulate the real load of the vehicle. The calculation formula of the moment of inertia on the output side is as follows: J out = J ml + J osh + J slv where J out is the moment of inertia on the output side; J ml is the moment of inertia of the load motor; J osh is the moment of inertia of the output shaft; J slv is the moment of inertia of the sliding sleeve.

4. The method for controlling the shifting process of an electric drive system based on reinforcement learning according to claim 3, characterized in that: According to the simulation model and the target gear, select the state input variables and action output variables of the reinforcement learning environment, specifically as follows: State input variables: During the gear engagement stage, it is the difference between the expected displacement of the synchronizer sleeve at the target gear and the real-time displacement of the synchronizer sleeve; during the rotational speed synchronization stage, it is the difference between the rotational speed of the engaging gear ring and the output shaft. Action output variables: The rotational speed command of the EMT shift motor; the rotational speed command of the drive motor.

5. The control method for the shifting process of an electric drive system based on reinforcement learning according to claim 4, characterized in that: Based on the rotational speed synchronization situation, shift time, and synchronizer sleeve displacement during the shifting process, a reward function for the reinforcement learning environment is established. The reward function r is calculated by the following formula: r=r1+r2+r3 In the formula, r1 is the shift time reward function, r2 is the shift accuracy reward function, and r3 is the rotational speed synchronization reward function.

6. The control method for the shifting process of the electric drive system based on reinforcement learning according to claim 5, wherein: The calculation method of the shift time reward function r1 is: Where, delta_t is the time interval from issuing the shift command to the completion of gear engagement.

7. The control method for the gear shifting process of the electric drive system based on reinforcement learning according to claim 5, characterized in that: The calculation method of the shift accuracy reward function r2 is: Where, x0 is the displacement of the synchronizer sleeve at the expected gear, and x is the real-time displacement of the synchronizer sleeve.

8. The control method for the shifting process of the electric drive system based on reinforcement learning according to claim 5, characterized in that: The calculation method of the rotational speed synchronization reward function r3 is: Where, n0 is the target rotational speed of the input shaft during the rotational speed synchronization stage, and n is the real-time rotational speed of the input shaft.

9. The control method for the shifting process of the electric drive system based on reinforcement learning according to claim 1, wherein Construct a reinforcement learning agent and connect the reinforcement learning agent to the reinforcement learning environment, specifically as follows: (91) Selection of the reinforcement learning agent strategy The DQN agent is adopted, and the strategy used is the ε-greedy strategy. The ε-greedy strategy randomly selects actions with probability ε to explore different states and actions in the environment; selects the action with the maximum Q value with probability 1-ε. The formula expression of the ε-greedy strategy is: where a t is the action selected by the DQN agent at time step t; Q(s, a; θ) represents the Q-value estimation of the "state-action" pair (s, a) by the policy network under the parameter θ; (92) Establish the critic network of the reinforcement learning agent The neural network used is a feedforward neural network, including a feature input layer, two fully connected layers, two activation function layers, and an output layer; Fit the feedforward neural network into the critic network to approximate the Q value function; Let θ be the parameter of the critic network, and Q(s,a;θ) represents the Q value estimation of the "state-action" pair (s,a) by the network under the parameter θ; DQN updates the network parameter θ by minimizing the loss function. Specifically, the mean squared error loss function is used, and the specific expression is: L(θ) = E (s,a,r,s')~U(D) [(r + γmax a' Q(s', a'; θ - ) - Q(s, a; θ)) 2 ​ where \(U(D)\) represents the "state-action-reward-next state" quadruple \((s, a, r, s')\) uniformly sampled from the experience replay buffer \(D\); \(\gamma\) is the discount factor; \(\theta\) - is the parameter of the target network; Use the gradient descent method to update the network parameter θ, and the specific formula is: Among them, θ' is the updated network parameter; α is the learning rate, set as α = 0.001; L(θ) is the gradient of the mean squared error loss function L(θ) with respect to θ, and the gradient clipping threshold is set to 1.

10. The control method for the shifting process of the electric drive system based on reinforcement learning according to claim 1, wherein: Taking the maximum of the action value function of the reinforcement learning agent as the optimization goal, train the reinforcement learning agent to obtain the optimal shift process control strategy, specifically as follows: (101): Take the difference between the expected displacement of the synchronizer sleeve and the real-time displacement of the synchronizer sleeve during the gear engagement stage, and the difference between the rotational speed of the engaging gear ring and the output shaft during the rotational speed synchronization stage as the state input variables of the reinforcement learning agent; take the rotational speed command of the EMT shift motor and the rotational speed command of the drive motor as the output variables of the reinforcement learning agent; (102): The critic network collects the values of the reward function and the state input variables, minimizes the loss function with the help of the DQN algorithm, and optimizes the current shift process control strategy. Then, output the optimized control strategy to the reinforcement learning environment through the actuator; (103): Calculate the reward function of the shift process under the new shift method, calculate the Q value and feedback it to the critic network; the reinforcement learning environment transfers the new state input variables to the reinforcement learning agent; (104): Repeat steps (102) and (103) to maximize the Q-value in the Q function. The shift process control strategy at this time is the current optimal control strategy, which is used to enable the reinforcement learning agent to achieve an optimal decision during the shift process.