Automatic gearbox gear shifting control system and control method

By introducing intelligent gear decision module, shift timing optimization module and smooth shift control module into the automatic transmission control system, the problem of difficulty in taking into account fuel economy and shift smoothness in the existing system is solved, and a better traffic environment and driver style adaptability is achieved.

CN120100898APending Publication Date: 2025-06-06JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510277123.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing automatic transmission control system is difficult to balance fuel economy and shift smoothness, and it fails to adapt to different traffic environments and driver styles.

Method used

The intelligent gear decision module, shift timing optimization module and smooth shift control module are adopted to optimize shift decisions and execution through logical rules, simulated annealing optimization algorithm and depth deterministic strategy gradient methods.

Benefits of technology

It improves the fuel economy and smoothness of gear shifts, enhances adaptability to different traffic environments and driver styles, and reduces the calibration and tuning costs of automatic transmissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120100898A_ABST
    Figure CN120100898A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic gearbox gear shifting control system and method, and belongs to the field of automobile dynamic control. A gear shifting control system adopted by the method comprises an intelligent gear decision module, a gear shifting opportunity optimization module and a smooth gear shifting control module. The intelligent gear decision-making module sends a gear instruction through a logic rule based on the traffic information and the vehicle information; the gear shifting opportunity optimization module decides the optimal gear shifting opportunity based on a simulated annealing optimization algorithm and corrects the gear shifting opportunity according to the driving style; the smooth shift control module adjusts a transmission actuator control current based on the depth deterministic policy gradient. The economical efficiency, the dynamic property and the smoothness of the automatic gearbox facing different working conditions can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of automobile dynamic control, and specifically provides an automatic transmission shift control system and a control method thereof. Background Art

[0002] National policies and market demand have driven the development of transmission technology towards multi-gear, integrated and intelligent, improving the power, smoothness and economy of the transmission. In recent years, with the rapid development of the automotive industry and the growing market demand for high-performance, low-energy vehicles, automatic transmissions have played an important role in improving vehicle power, economy and driving comfort. Automatic transmission control technology is not only the core of the automotive power transmission field, but also the key to promoting technological progress in the automotive industry and meeting market demand. The current automatic transmission shift strategy is usually based on fixed shift logic or simple working condition adaptation rules, which fails to fully consider the traffic environment and driver style, resulting in difficulty in balancing vehicle fuel economy and shifting smoothness.

[0003] As traditional automatic transmission control systems and methods are unable to balance fuel economy and smoothness, and are unable to adapt to different traffic environments and driver styles, it is particularly important to develop an automatic transmission shift control system that can be used in complex traffic environments and for a variety of driving styles by combining advanced intelligent control algorithms and optimization methods. Summary of the invention

[0004] The purpose of the embodiments of the present invention is to provide an automatic transmission shift control system and a control method thereof, aiming to solve the problems raised in the above-mentioned background technology.

[0005] The embodiment of the present invention is implemented as follows: an automatic transmission shift control system and a control method thereof, comprising an intelligent gear decision module, a gear shift timing optimization module and a smooth gear shift control module. The intelligent gear decision module sends a gear command through logical rules based on traffic information and vehicle information; the gear shift timing optimization module determines the best gear shift timing based on a simulated annealing optimization algorithm, and modifies the gear shift timing according to the driving style; and the smooth gear shift control module adjusts the gearbox actuator control current based on a deep deterministic strategy gradient.

[0006] Compared with the prior art, the beneficial effects achieved by the present invention include:

[0007] (1) An intelligent shift decision module in an automatic transmission shift control system can predictively give upshift or downshift suggestions based on traffic information and vehicle information, saving time and improving safety when passing through traffic lights.

[0008] (2) A shift timing optimization module in an automatic transmission shift control system performs optimal shifting based on a simulated annealing algorithm and then corrects the shift timing based on driving style, thereby improving shifting fuel economy and expanding driver adaptability.

[0009] (3) A smooth shift control module in an automatic transmission shift control system improves the smoothness of the shift and can reduce the calibration and tuning costs of the automatic transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a schematic diagram of the automatic transmission shift control system;

[0011] Figure 2 It is a 9-speed automatic transmission structure diagram DETAILED DESCRIPTION

[0012] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0013] [First embodiment]

[0014] like Figure 1 The figure shows an automatic transmission shift control system provided by the first embodiment of the present invention, wherein the system comprises an intelligent gear position decision module, a gear shift timing optimization module and a smooth gear shift control module.

[0015] The intelligent gear decision module sends gear instructions through logical rules based on traffic information and vehicle information;

[0016] The shift timing optimization module determines the best shift timing based on the simulated annealing optimization algorithm and modifies the shift timing according to the driving style;

[0017] The smooth shift control module adjusts a transmission actuator control current based on a deep deterministic policy gradient.

[0018] To further illustrate the automatic transmission shift control system proposed in the first embodiment of the present invention, a 9-speed automatic transmission is selected as a research object, and its structural schematic diagram is shown in FIG. Figure 2 As shown,

[0019] like Figure 2 The figure shows the automatic transmission shifting power transmission system of the present invention, which is mainly composed of a torque converter, four parallel planetary gear mechanisms, four wet clutches and two brake clutches.

[0020] like Figure 1 As shown, the intelligent gear position decision module in this embodiment obtains the signal light cycle information and the distance from the signal light through the vehicle communication system, and triggers the intelligent gear position decision module based on the distance from the signal light. When the distance d is less than the distance threshold dy When the intelligent gear decision module is triggered, otherwise it is not triggered. 1 , the time for the red / yellow light to turn green is t 2 ) and road speed limit v max To calculate the distance threshold d from the traffic light y , d y =v max *max{t 1 , t 2}, max{} is the maximum value. max is the upper speed limit, then the time range for passing the intersection can be expressed as t min ∈[d y / v max , d y / v 0 ]. Among them, v 0 is the initial velocity. 1 Greater than or equal to d y / v 0 When t 1 Less than or equal to d y / v max When t 2 Greater than or equal to d y / v 0 When t 2 Less than or equal to d y / v max When the gear is in the gear position, the driver is sent a command to maintain the current gear or to accelerate upshift. The above command can be rejected by pressing the opposite command pedal, and accepted by pressing the same command pedal. In other cases, no command is sent.

[0021] In the shift timing optimization module, the engineering dual-parameter shift benchmark mapping table can use vehicle speed and throttle opening as table lookup mapping table inputs, and the gear position signal as the table lookup mapping table output. On this basis, taking fuel economy as the optimization target, the shift time and gear position as the optimization target variable, and the optimization range can be defined as [0, d y / v 0 ] and {1, 2, ... D], D is the highest gear of the gearbox, and each gear corresponds to a corresponding transmission ratio. The optimization model can be expressed as:

[0022]

[0023] In the formula,

[0024] min means the minimum value,

[0025] J is the target optimization function,

[0026] F(t h ,D h ) is the optimal shift time t h and the best gear D h The nonlinear fuel consumption function is

[0027] st represents a constraint,

[0028] t h represents the optimal shift time,

[0029] D h Indicates the best gear.

[0030] The optimal shift time t is obtained based on the simulated annealing algorithm. h and the best gear D h , then we can get the corresponding t at this moment h Vehicle speed v, engine speed ω e and throttle opening α, v, ω e and α and D h The corresponding combination is the optimal shift three-parameter mapping chart. On this basis, the speed, acceleration and pedal opening signals are collected, and the abnormal values ​​are eliminated, and the data is smoothed and filtered:

[0031]

[0032] In the formula,

[0033] After smoothing the filtered data,

[0034] x i is the original data variable, i takes 1, 2 and 3 to represent speed, acceleration and pedal opening data respectively,

[0035] t is time,

[0036] N is the sliding window size, which can generally be between 10 and 20.

[0037] Furthermore, the input variables are normalized to avoid the influence of different dimensional features on model training:

[0038]

[0039] In the formula,

[0040] is the normalized data,

[0041] max(x i ) and min(x i ) are xi The maximum and minimum values ​​of .

[0042] The above data are subjected to feature extraction, including mean value, standard deviation and rate of change:

[0043]

[0044] In the formula,

[0045] μ i is the average value,

[0046] σ i is the standard deviation,

[0047] is the rate of change of the variable,

[0048] Expressed as The next moment value,

[0049] Δt is the time rate of change,

[0050] n is the number of data.

[0051] The data are labeled as three driving styles: adventurous, normal, and conservative. The radial basis function is selected as the support vector machine kernel function. The kernel function can be expressed as:

[0052]

[0053] In the formula,

[0054] K is the kernel function,

[0055] λ i ,λ j They are all sample characteristic variables, which can be μ i , σ i and

[0056] γ is the kernel function width parameter,

[0057] exp() is the exponential function.

[0058] The one-to-many strategy is used to decompose the three different driving styles into three binary classification problems: adventurous versus non-adventurous, normal versus abnormal, and conservative versus non-conservative. The support vector machine optimization model can be expressed as:

[0059]

[0060] sty i (w T λ i +b)≥1-ζ i ,ζi ≥0

[0061] In the formula,

[0062] min is the minimum value,

[0063] st represents a constraint,

[0064] w is the normal vector of the decision hyperplane,

[0065] w T is the transpose of w,

[0066] C is the penalty coefficient,

[0067] ζ i is the slack variable,

[0068] y i is a hyperplane, and y i =w T λ i +b,y i Take 1 or -1,

[0069] b is the distance between the hyperplane and the origin.

[0070] n is the number of data.

[0071] 70% of the data is used for training and 30% of the data is used for verification. The final driver style recognition model based on support vector machine can be obtained, and the shift timing is corrected according to the driver style. When in the adventurous type, the shift delay time is increased. As the gear position increases, the delay time increases. The normal type shift time is not corrected. When in the conservative type, the gear is shifted in advance. As the gear position increases, the advance shift time increases. The specific delay time can be obtained according to the actual vehicle model calibration, and the correction time mapping chart under different driving styles is further obtained. Finally, the optimal shift three-parameter mapping chart and the shift timing correction mapping chart are combined into the shift timing optimization mapping chart.

[0072] In the smooth shift control module, the shift behavior is established into a Markov decision process model consisting of a state space, an action space, a transition probability and a reward function:

[0073] MDP=(S,A,T,r)

[0074] In the formula,

[0075] MDP is a Markov decision process model.

[0076] S is the state space, and includes the shift time, shift phase, control current of the solenoid hydraulic valve and oil pressure filling signal,

[0077] A is the action space and is the differential control current,

[0078] T is the transition probability,

[0079] r is the reward function, which is a weighted function of shift shock, shift time and clutch slip loss.

[0080] The deep deterministic policy gradient method is adopted, that is, the actor-critic framework containing two deep neural networks. The deterministic policy function is selected, and data is generated through the vehicle dynamics model and stored in the experience pool. The random sampling method is used to form experience replay, and the neural network policy parameters are updated online. After obtaining the online neural network parameters, the parameters in the target network are updated according to the soft update principle. In order to improve the learning ability of the agent, action noise is added based on the action output of the actor network. The specific implementation steps can be expressed as:

[0081] (1) Initialize the network parameters and environment, and initialize the actor network and critic network parameters. Assume that the parameters of the actor network are θ actor , the parameter of the critic network is θ critic . Build a vehicle dynamics model including an automatic transmission to generate shift data and establish a simulation environment.

[0082] (2) Initialize the experience pool, set the action noise, and create an experience pool R to store the agent’s experience (S) at each time t. t ,A t ,r t ,S t+1 ), where S t is the current state, A t is the selected action, r t For reward, S t+1 is the next state. Set the action noise Z t To increase exploratory potential, noise can be defined using

[0083] Z t =Z t-1 +θ(η-Z t-1 )+W t

[0084] In the formula,

[0085] Z t represents the action noise,

[0086] Z t-1 Indicates the action noise of the last moment,

[0087] θ represents the regression rate,

[0088] η represents the intensity of noise,

[0089] W t represents a Gaussian noise,

[0090] t represents time.

[0091] (3) Select actions based on the current strategy and exploration noise, based on the current actor network πθ actor (S t ) selects an action, which is to increase the differential current and add noise Z t ,This ensures that the agent can both utilize known strategies and conduct sufficient exploration during the training process.

[0092] A t =πθ actor (S t )+Z t

[0093] In the formula,

[0094] A t Indicates the execution of an action.

[0095] Z t Indicates motion noise

[0096] πθ actor (S t ) represents the actor network in state S t Next, select the action.

[0097] (4) Execute the action, observe the reward, observe the new state, and execute the selected action A t , and obtain the immediate reward r of the environment t , and observe the new state S t+1 .

[0098] (5) Store the data in the experience pool and store the current state S t 、Action A t , Reward t and the new state S t+1 Stored in the experience pool R:

[0099] R=R∪{(S t ,A t ,r t ,S t+1 )}

[0100] (6) Randomly sample data from the experience pool to update the critic network. Randomly sample a batch of data from the experience pool R and update the critic network according to the Bellman equation:

[0101]

[0102] In the formula,

[0103] j represents random sampling data,

[0104] S j represents the state of random sampling,

[0105] A j represents the random sampling action,

[0106] r j represents the reward of random sampling,

[0107] S j+1 represents the state of the next random sampling,

[0108] L(θ critic ) represents the updated value of the critic network,

[0109] U represents the batch size of random sampling in the experience pool,

[0110] θ c ' ritic represents the target critic network parameters,

[0111] Represented as the target actor network in state S j+1 Next, select the action.

[0112] Indicates that in state S j and action A j The evaluation function of the critic network under

[0113] is the discount factor,

[0114] Indicates that in state S j+1 Next, according to the target network's policy The maximum future reward that can be obtained from the chosen action.

[0115] (7) Use the sampled gradient to update the policy network and use the gradient information of the critic network to update the parameters of the actor network:

[0116]

[0117] In the formula,

[0118] ▽θ actor J represents the actor network parameters θ actor The gradient objective function is

[0119] U represents the batch size of random sampling in the experience pool,

[0120] represents the evaluation function of the critic network at state Sj and action Aj,

[0121] ▽θ actor Denotes the actor network parameter θ actor gradient.

[0122] (8) Update the target network and use the soft update rule to update the parameters of the target network. The parameters of the target network θ c ' ritic and θ a ' ctor According to the current network parameters θ critic and θ actor renew:

[0123] θ c ' ritic ←τθ critic +(1-τ)θ critic

[0124] θ a ' ctor ←τθ actor +(1-τ)θ actor

[0125] In the formula,

[0126] θ a ' ctor represents the target actor network parameters,

[0127] θ c ' ritic represents the target critic network parameters,

[0128] θ critic represents the critic network parameters,

[0129] θ actor represents the actor network parameters,

[0130] τ represents the soft update ratio.

[0131] (9) Iterate until the set requirements are met. Repeat the above steps and iterate the training until the agent's shifting strategy meets the predetermined performance goals or task requirements.

[0132] [Second embodiment]

[0133] The second embodiment of the present invention further provides an automatic transmission shift control, comprising the following steps:

[0134] 1) Intelligent shift trigger: First, the distance threshold is calculated based on the traffic light cycle information and the distance to the traffic light obtained by the vehicle communication system; secondly, the maximum speed limit at the road end is used as the upper speed limit to calculate the time range for successfully passing the traffic light intersection; finally, the driver is given an acceleration or deceleration command through a logical control command, and the command can be rejected by pressing the opposite command pedal and accepted by pressing the same command pedal.

[0135] 2) Perform optimal shifting: When a shifting command is received, the gear is shifted through the optimal shifting three-parameter mapping chart. The optimal shifting three-parameter mapping chart is obtained as follows: first, automatic shifting is performed according to the engineering two-parameter shifting benchmark mapping chart; second, the shifting time and gear position are used as optimization target variables, and fuel economy is used as the optimization target. The optimal shifting time and the optimal gear position are optimized based on the simulated annealing algorithm; finally, the vehicle speed, engine speed and throttle opening corresponding to the optimal shifting time are solved, and the vehicle speed, engine speed and throttle opening are combined with the optimal gear position to form the optimal shifting three-parameter mapping chart.

[0136] 3) Correction of shift timing: After shifting through the optimal shift three-parameter mapping chart, the shift timing is corrected based on the driving style and organized into a shift timing correction mapping chart. The process of obtaining the shift timing correction mapping chart is as follows: first, the speed, acceleration and pedal opening signals are preprocessed and normalized; second, the data is labeled as three driving styles of adventurous, normal and conservative, and the radial basis function is selected to build a support vector machine model. Then, the processed speed, acceleration and pedal opening signals are used as inputs. After training and verification, a driver style recognition model based on the support vector machine can be obtained; then, for the adventurous driving style, its shift delay time is increased. As the gear position increases, the delay time increases. The normal shift time is not corrected. For the conservative driving style, the gear is shifted in advance. As the gear position increases, the advance shift time increases; finally, the correction time under different driving styles is obtained according to the actual vehicle model calibration, and it is organized into a shift time correction mapping chart according to the driving style.

[0137] 4) Smooth shift control: When the shift command is received, the shift current is controlled by the trained deep deterministic policy gradient method. The shift control process based on deep deterministic policy gradient: First, define the state space including shift time, shift stage, control current of electromagnetic hydraulic valve and oil pressure filling signal, define differential control current as action, and define reward function as weighted function of shift shock, shift time and clutch slip loss. Combined with the above definition, the shift control problem is modeled as a Markov decision process; secondly, based on the actor-critic architecture, a deep deterministic policy gradient method including actor network, critic network, target network and experience replay is established. Finally, after training and verification until the policy converges or the training termination condition is reached, the transmission actuator control current is trained.

Claims

1. An automatic transmission shift control system, characterized in that: It includes an intelligent gear decision module, a gear shift timing optimization module and a smooth gear shift control module; the intelligent gear decision module sends gear instructions through logical rules based on traffic information and vehicle information; the gear shift timing optimization module decides the best gear shift timing based on the simulated annealing optimization algorithm, and corrects the gear shift timing according to the driving style; the smooth gear shift control module adjusts the transmission actuator control current based on the deep deterministic strategy gradient.

2. A system according to claim 1, characterized in that: In the intelligent gear position decision module, the traffic light cycle information and the distance from the traffic light are obtained through the vehicle communication system, and the intelligent gear position decision module is triggered based on the distance from the traffic light. When the distance d is less than the distance threshold d y When the intelligent gear decision module is triggered, otherwise it is not triggered; according to the signal light cycle time and the road speed limit v max To calculate the distance threshold d from the traffic light y , d y =v max *max{t1, t2}, where t1 is the time it takes for a green light to turn red / yellow, t2 is the time it takes for a red / yellow light to turn green, and max{} is the maximum value; the maximum speed limit v max is the upper speed limit, then the time range for passing the intersection can be expressed as t min ∈[d y / v max , d y / v0], where v0 is the initial velocity; when t1 is greater than or equal to d y / v0, the driver is sent a command to maintain the current gear or accelerate upshift; when t1 is less than or equal to d y / v max When t2 is greater than or equal to d y / v0, the driver is sent a command to maintain the next gear or decelerate and downshift; when t2 is less than or equal to d y / v max When the gear is in the gear position, a command to maintain the current gear or to accelerate upshift is sent to the driver; the above command can be rejected by pressing the opposite command pedal and accepted by pressing the same command pedal; in other cases, no command is sent.

3. A system according to claim 1, characterized in that: In the shift timing optimization module, automatic shifting is performed based on the engineering dual-parameter shift benchmark mapping chart. On this basis, fuel economy is taken as the optimization target, and the shift time and gear position are taken as the optimization target variable. The optimization range can be defined as [0, d y / v0] and {1, 2, ... D], D is the highest gear of the gearbox, and each gear corresponds to a corresponding transmission ratio; the optimal shift time t is obtained based on the simulated annealing algorithm. h and the best gear D h , we can further get t h Corresponding vehicle speed v, engine speed ω e and throttle opening α, v, ω e and α and D h The corresponding combination is the optimal shifting three-parameter mapping chart; on this basis, the speed, acceleration and pedal opening signals are preprocessed and normalized, and the data are labeled as adventurous, normal and conservative. The radial basis function is selected as the support vector machine kernel function, and the processed data is used as the input of the support vector machine model. After training and verification, the driver style recognition model is obtained, and the shifting timing is corrected according to the driver style; when in the adventurous type, the shifting delay time is increased. As the gear position increases, the delay time increases. The normal type shifting time is not corrected. When in the conservative type, the gear is shifted in advance. As the gear position increases, the advance shifting time increases. The specific delay time can be obtained according to the actual vehicle model calibration, and the correction time mapping chart under different driving styles is further obtained; finally, the optimal shifting three-parameter mapping chart and the shifting timing correction mapping chart are combined into a shifting timing optimization mapping chart.

4. A system according to claim 1, characterized in that: In the smooth shift control module, the shift behavior is established into a Markov decision process model consisting of a state space, an action space, a transfer probability and a reward function; the state space is defined to include the shift time, the shift stage and the control current signal of the electromagnetic hydraulic valve; the action space is defined as the differential control current; the reward function is defined as a weighted function of the shift shock, the shift time and the clutch slip loss index; a deep deterministic policy gradient method is used, after initializing the network and the environment, continuously interacting with the environment to generate experience, randomly sampling data from the replay pool, updating the critic network and the actor network, and softly updating the target network until the strategy converges or the training termination condition is reached, and the transmission actuator control current is trained.

5. A control method for an automatic transmission shift control system according to any one of claims 1 to 4, characterized in that: The following steps are involved: (1) Intelligent shift triggering: First, the distance threshold is calculated based on the traffic light cycle information and the distance to the traffic light obtained by the vehicle communication system. Second, the maximum speed limit at the road end is used as the upper speed limit to calculate the time range for successfully passing the traffic light intersection. Finally, the driver is given an acceleration or deceleration command through a logic control command. The command can be rejected by pressing the opposite command pedal and accepted by pressing the same command pedal. (2) Performing optimal gear shifting: When a gear shifting instruction is received, the gear shifting is performed using the optimal gear shifting three-parameter mapping chart; wherein the optimal gear shifting three-parameter mapping chart is obtained in the following process: first, automatic gear shifting is performed according to the engineering two-parameter gear shifting benchmark mapping chart; second, the gear shifting time and gear position are used as optimization target variables, and fuel economy is used as the optimization target, and the optimal gear shifting time and optimal gear position are obtained based on the simulated annealing algorithm; finally, the vehicle speed, engine speed and throttle opening corresponding to the optimal gear shifting time are solved, and the vehicle speed, engine speed and throttle opening are combined with the optimal gear position to form the optimal gear shifting three-parameter mapping chart; (3) Correcting the gear shift timing: After the gear shift is performed through the optimal gear shift three-parameter mapping chart, the gear shift timing is corrected based on the driving style and organized into a gear shift timing correction mapping chart; wherein, the process of obtaining the gear shift timing correction mapping chart is as follows: first, the speed, acceleration and pedal opening signals are preprocessed and normalized; second, the data is labeled into three driving styles: adventurous, normal and conservative, and the radial basis function is selected to build a support vector machine model; then, the processed speed, acceleration and pedal opening signals are used as inputs, and after training and verification, a driver style recognition model based on the support vector machine is obtained; then, for the adventurous driving style, its gear shift delay time is increased, and the delay time increases with the increase of gear position, and the normal gear shift time is not corrected. For the conservative driving style, the gear is shifted in advance, and the advance gear shift time increases with the increase of gear position; finally, the correction time under different driving styles is obtained according to the actual vehicle model calibration, and it is organized into a gear shift time correction mapping chart according to the driving style; (4) Smooth shift control: When a shift command is received, the shift current is controlled by the trained deep deterministic policy gradient method. The shift control process based on deep deterministic policy gradient is as follows: First, the state space including the shift time, shift phase and control current signal of the electromagnetic hydraulic valve is defined, the differential control current is defined as the action, and the reward function is defined as a weighted function of the shift shock, shift time and clutch slip loss index. Combining the above definitions, the shift control problem is modeled as a Markov decision process. Secondly, based on the actor-critic architecture, a deep deterministic policy gradient method including an actor network, a critic network, a target network and experience replay is established. Finally, after training and verification until the strategy converges or the training termination condition is reached, the transmission actuator control current is trained.