Series-parallel hybrid electric vehicle mode switching self-learning system and method

By combining a self-learning system with multiple algorithm optimization and control strategies, the problems of long calibration cycle and poor adaptability of mode switching control in series-parallel hybrid electric vehicles have been solved, achieving efficient and stable mode switching control and improving the driving experience.

CN121989906APending Publication Date: 2026-05-08CHANGZHOU INST OF MECHATRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGZHOU INST OF MECHATRONIC TECH
Filing Date
2026-03-24
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing hybrid electric vehicle mode switching control relies on manual experience for calibration, which has a long calibration cycle and high testing costs. Furthermore, the fixed parameters of traditional control strategies are difficult to adapt to different acceleration and vehicle speed conditions, resulting in a poor driving experience and issues such as speed deviation and vehicle impact.

Method used

Employing a self-learning system, combined with an engine operation optimization module, a basic mode switching control module, and a self-learning compensation control module, the system utilizes online dynamic programming algorithms, proportional-integral closed-loop control, particle swarm optimization algorithms, and soft actor-commentator reinforcement learning algorithms to achieve dynamic differential equation optimization and motor torque coordinated control, adapting to multiple operating conditions.

Benefits of technology

It improves driving smoothness and adaptability during mode switching, reduces calibration cycle and testing costs, and enhances the consistency and stability of control strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121989906A_ABST
    Figure CN121989906A_ABST
Patent Text Reader

Abstract

The invention provides a series-parallel hybrid electric vehicle mode switching self-learning system and method, and relates to the field of hybrid electric vehicle power control. Comprising an engine work optimization module, a basic mode switching control module and a self-learning compensation control module. Optimal starting and pre-working rotating speed curves of an engine are obtained through optimization of an online dynamic programming algorithm, dual-motor optimal basic output torque is obtained through combination of PI closed-loop control and a particle swarm optimization algorithm, and dual-motor compensation torque at the moment t is generated based on an SAC reinforcement learning algorithm. And linearly combining the compensation torque and the basic torque into an actual output torque to complete mode switching. The mode switching smoothness is effectively improved, the energy consumption is reduced, accurate compensation torque output under multiple working conditions is achieved, the working condition self-adaptive capacity and robustness of mode switching control are enhanced, and the real-time control requirement of the whole vehicle is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hybrid electric vehicle power control, and specifically to a self-learning system and method for mode switching in a series-parallel hybrid electric vehicle. Background Technology

[0002] Hybrid electric vehicle technology and pure electric vehicle technology have long coexisted and complemented each other strategically. Among them, the series-parallel hybrid electric vehicle combines multiple power sources and coupling mechanisms such as planetary gears, and has the technical characteristics of both series and parallel hybrid systems. It not only has excellent power performance and no range anxiety, but also has significant energy-saving and emission-reduction effects, making it the main research configuration in the current field of hybrid electric vehicles.

[0003] The overall vehicle control of a series-parallel hybrid electric vehicle (SPEV) not only needs to improve fuel economy and power performance under steady-state conditions through power coupling and power source complementarity, but also requires precise coordination of transient torque to ensure driving smoothness. During actual driving, influenced by driving environment and operating conditions, SPEVs need to frequently switch between pure electric mode and hybrid mode. This transient mode switching process requires operations such as engine starting and torque matching of multiple power sources, placing extremely high demands on the control strategy.

[0004] Existing hybrid electric vehicles rely heavily on manual calibration for mode switching, resulting in long calibration cycles, high testing costs, and inconsistent calibration quality due to the influence of engineers' expertise. Furthermore, traditional control strategies use fixed parameters, making it difficult to ensure smooth switching under different acceleration and vehicle speed conditions. The mode switching control also has weak adaptive capabilities, leading to issues such as large speed deviations, vehicle shocks, and jerking, which degrade the driving experience.

[0005] In recent years, integrated control methods combining optimization algorithms, feedback control, and intelligent algorithms have provided a new direction for mode switching control of series-parallel hybrid electric vehicles. How to achieve precise control of the mode switching process through algorithm fusion, while enabling the control strategy to have self-learning capabilities to adapt to a wide range of operating conditions, has become a research focus for those skilled in the art. Summary of the Invention

[0006] The purpose of this invention is to overcome at least one technical problem existing in the prior art and to provide a self-learning system and method for mode switching of a series-parallel hybrid electric vehicle.

[0007] On one hand, this invention provides a self-learning system for mode switching in a series-parallel hybrid electric vehicle (SPEV). The SPEV adapted to this system is a power-split SPEV, whose transmission structure includes a dual planetary gear set mechanism, a torsional damper, a lock-up clutch, an engine, a first motor, a second motor, and an output shaft. The dual planetary gear set mechanism comprises a first planetary gear set and a second planetary gear set, both consisting of a sun gear, a planet carrier, and a ring gear. The engine is connected to the planet carrier of the first planetary gear set via the torsional damper, and the ring gear of the second planetary gear set is connected to the output shaft. The first motor and the second motor are respectively connected to the sun gears of the first and second planetary gear sets. The ring gear of the first planetary gear set is connected to the ring gear of the second planetary gear set, and the planet carrier of the second planetary gear set is locked to the frame via a lock-up clutch. The self-learning system includes an engine operation optimization module, a basic mode switching control module, and a self-learning supplementation module. The system comprises a compensation control module and an engine operation optimization module. The engine operation optimization module uses mode switching time, power source energy consumption, engine speed tracking error, and engine terminal speed tracking error as four-element optimization objectives, and solves for the optimal engine start-up and pre-operation speed curves based on an online dynamic programming algorithm. The basic mode switching control module uses the optimal start-up and pre-operation speed curves output by the engine operation optimization module as the speed tracking benchmark, and combines proportional-integral closed-loop control and particle swarm optimization algorithm to determine the optimal basic output torque of the dual motors to adapt to the mode switching process. The self-learning compensation control module generates dual motor compensation torque based on a soft actor-critic reinforcement learning algorithm for different acceleration conditions. The dual motor compensation torque and the optimal basic output torque of the dual motors are linearly combined to obtain the actual output torque of the dual motors. The actual output torque of the dual motors is used as a motor torque command to control the torque coordination control of the hybrid electric vehicle to complete the mode switching.

[0008] Furthermore, in the engine operation optimization module, the engine speed of the transmission structure, the planet carrier speed of the first planetary gear set, the angular error between the engine and the first planetary gear set, the ring gear speed of the second planetary gear set, the output shaft speed, and the angular error between the second planetary gear set and the output shaft are used as state variables. The output torque of the first motor and the second motor are used as system control input variables, and the output shaft speed and the vehicle speed are used as output variables. The dynamic differential equation of the mode switching process of the series-parallel hybrid electric vehicle is established as follows: ; ; ; ; In the formula, x is the state variable of the mode switching process. This is expressed as engine speed. This indicates the rotational speed of the planet carrier in the first planetary gear set. This indicates the angular error between the engine and the first planetary gear set. This indicates the rotational speed of the ring gear of the second planetary gear set. Indicates the output shaft speed. This represents the angular error between the second planetary gear set and the output shaft; u is the motor output torque; and The output torques of the first and second motors are represented respectively; w is the known input variable disturbance of the system. and These represent the engine output torque and the load on the output shaft, respectively. , and These represent the coefficient matrices for the state variables, control input variables, and known output variables, respectively.

[0009] Furthermore, the engine operation optimization module, based on the dynamic differential equation of the mode switching process, uses a weighted function of mode switching time, power source energy consumption, engine speed tracking error, and engine terminal speed tracking error as the optimization objective function, and uses the engine start-up and pre-operation curves as the optimization targets. Through online optimization, it obtains the optimal engine start-up and pre-operation curves. The mathematical expression of the online optimization model is as follows: ; ; In the formula, min represents minimization; This represents the overall optimization objective function of the engine operation optimization module; Indicates the mode switching time function; Represents the energy consumption function of the power source; This represents the engine speed tracking error function; The function representing the engine terminal speed tracking error is denoted by t; t represents time. The weighting coefficients of the mode switching time function; The weighting coefficients represent the energy consumption function of the power source; The weighting coefficients represent the engine speed tracking error function; The weighting coefficients represent the engine terminal speed tracking error function; and These represent the end and start times of mode switching, respectively; u represents the motor output torque; and These represent the optimal engine start-up and pre-operation target speed and the actual speed, respectively. This represents the engine terminal speed tracking error; the engine operation optimization module iteratively solves the optimization model using an online dynamic programming algorithm to obtain the optimal speed curves for the engine start-up and pre-operation stages during mode switching; wherein, the solution process uses the dynamic differential equation of the hybrid electric vehicle mode switching process as a constraint condition, combining dynamic characteristics with optimization objectives to achieve optimal planning of the engine speed curve.

[0010] Furthermore, the basic mode switching control module is used to: merge the optimal output shaft speed curve calculated based on the target vehicle speed and the optimal engine start-up and pre-operation speed curves output by the engine operation optimization module to obtain a multi-dimensional total speed tracking target. ; Calculate the deviation between the total rotational speed tracking target and the state variables during the mode switching process. ; to deviation As the input state variable for proportional-integral control, the output torque of the two motors is calculated based on proportional-integral control. The calculation formula is as follows: ; ; ; In the formula, This represents the optimal engine start-up and pre-operation target speed. This indicates the target rotational speed of the output shaft, calculated from the target vehicle speed. This indicates the real-time rotational speed of the output shaft. represents the engine speed, and x represents the state variable during the mode switching process; This is the proportional coefficient, used to suppress the speed deviation at the current moment; is the integral coefficient, used to eliminate the steady-state speed difference during the entire mode switching process; u is the motor output torque.

[0011] Furthermore, the basic mode switching control module is used to adjust and optimize the objective function using a weighted function of mode switching impact, dual-motor energy consumption, and speed tracking error as parameters. , To optimize parameters and set upper and lower limits based on the hardware performance of the dual motors and the ride comfort requirements of the vehicle, a constrained parameter optimization model is constructed. The parameter optimization model is as follows: ; ; In the formula, min represents minimization; This represents the optimization objective function in the basic mode switching control module; Represents the impact function; Represents the energy consumption function of the motor; This represents the speed tracking error function; The weighting coefficients of the impact function are represented. The weighting coefficients represent the energy consumption function of the motor. The weighting coefficients represent the speed tracking error function; This is the proportionality coefficient; The integral coefficient; , They represent The minimum and maximum values; , They represent The minimum and maximum values; 'a' is the vehicle acceleration; The target engine speed; This refers to the actual engine speed. The basic mode switching control module uses a particle swarm optimization algorithm to... , By performing iterative optimization, the optimal proportional coefficient can be obtained. and optimal integral coefficients Substituting this into the proportional-integral control calculation formula yields the optimal basic output torque of the dual motors. .

[0012] Furthermore, the self-learning compensation control module uses a soft actor-critic reinforcement learning algorithm to train the motor compensation amount. It uses vehicle speed, acceleration, engine speed, vehicle speed tracking error, motor output torque, impact, and mode switching time as the observation space, motor compensation torque as the action space, and a weighted function of vehicle speed tracking error, impact, and mode switching time as the reward space. Random acceleration curves are generated in the reset function for tracking training to obtain a motor compensation torque suitable for multiple operating conditions. The motor compensation torque is combined with the optimal motor output torque obtained by the basic mode switching control module to become the actual motor output torque.

[0013] Furthermore, the observation space is the set of state parameters at time t. ,in, For the vehicle speed, For the acceleration of the whole vehicle, For engine speed, For vehicle speed tracking error, For motor output torque, For the impact of mode switching, The time elapsed for mode switching; the action space is the motor compensation torque at time t. The reward space is the reward function at time t. , This is a weighted function of vehicle speed tracking error, impact, and mode switching time, characterizing the condition adaptation effect of the current motor compensation amount. The value is positively correlated with the mode switching control effect.

[0014] Furthermore, the training process of the soft actor-critic reinforcement learning algorithm includes initialization, environment interaction, experience storage, sample sampling, target Q-value calculation, network update, target network soft update, and iterative convergence steps, specifically: S1, Initialization: Initialize the policy network. Two commentator networks , Two target commentator networks Entropy temperature coefficient Experience replay pool R; S2, Environmental interaction: In the vehicle control environment, based on the current state Explore by sampling actions using a random policy; the action expression is: Where t represents time t; For policy networks; It is a random noise variable; For the action generation function based on policy network and random noise, the observed state and noise are mapped to specific motor torque compensation quantities; S3, storing experience: storing the experience data group obtained from environmental interaction ( , , , ) Stored into the experience replay pool R, where Let be the reward function value at time t. S4, Sampling data: Randomly sample empirical data samples of batch size N from the empirical replay pool R. , , , ), where i is the sample number, Let i be the observed state of the i-th sample. For the motor compensation torque of the i-th sample, Let i be the reward function value for the i-th sample. S5. Calculate the target Q value for the (i+1)th sample: [This is the state observed at the next time step.] Sample the optimal action from the current policy network. The target Q value of the i-th sample is calculated by combining the entropy term. ,include: Where i is the i-th data sample; In order to be in The following actions are sampled from the policy network; As a discount factor, and ; S6. Update the critic network: Train two critic networks using mean squared error loss. The loss function expression is: ;in, The mean squared error of the two critic networks; for and Q-value estimation; N is the batch size; S7, updating the policy network: updating the policy network by minimizing the policy loss. parameters The policy loss expression is: ;in, S8, Target Network Soft Update: Iteratively update the parameters of the two target commentator networks using a soft update method; S9, Iterative Convergence: Repeat steps S2-S8, continuously interacting with the vehicle control environment and updating network parameters until the network loss tends to stabilize or reaches the preset maximum number of iterations, at which point training convergence is determined.

[0015] Furthermore, the motor compensation torque obtained after training convergence, applicable to multiple operating conditions, is converted into a lookup table mode for storage. The input of the lookup table mode is the real-time operating parameters of the vehicle, and the output is the matched motor torque compensation torque. The motor compensation torque retrieved from the lookup table is linearly combined with the optimal basic output torque of the dual motors to obtain the actual output torque of the dual motors. The actual output torque of the motors serves as the final motor control command in the mode switching process of the series-parallel hybrid electric vehicle.

[0016] Secondly, embodiments of the present invention also provide a self-learning method for mode switching in a series-parallel hybrid electric vehicle, applied to the aforementioned self-learning system for mode switching in a series-parallel hybrid electric vehicle. The method includes: when the engine operation optimization module receives a mode switching command, using mode switching time, power source energy consumption, engine speed tracking error, and engine terminal speed tracking error as four-element optimization objectives, and solving for the optimal start-up and pre-operation speed curves of the engine based on an online dynamic programming algorithm; the basic mode switching control module uses the optimal start-up and pre-operation speed curves output by the engine operation optimization module as the speed tracking benchmark, and combines proportional-integral closed-loop control and particle swarm optimization algorithm to determine the optimal basic output torque of the dual motors to adapt to the mode switching process; the self-learning compensation control module generates dual-motor compensation torque based on a soft actor-critic reinforcement learning algorithm for different acceleration conditions, and linearly combines the dual-motor compensation torque with the optimal basic output torque of the dual motors to obtain the actual output torque of the dual motors, and uses the actual output torque of the dual motors as a motor torque command to control the series-parallel hybrid electric vehicle to complete the mode switching.

[0017] Thirdly, embodiments of the present invention also provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the above-described self-learning method for mode switching of a series-parallel hybrid electric vehicle.

[0018] Fourthly, embodiments of the present invention also provide a readable storage medium, which, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to execute the above-described hybrid electric vehicle mode switching self-learning method.

[0019] The advantages of this invention compared to the prior art are: (1) From engine curve planning to motor torque calculation, the multi-objective weighted optimization criterion is adopted, taking into account indicators such as switching time, energy consumption, speed tracking accuracy, and impact. The mode switching process is accurately described by dynamic differential equations, and the closed-loop suppression of speed deviation is achieved by combining PI control, which effectively reduces the vehicle impact and jerking during the mode switching process and improves driving smoothness. (2) The self-learning compensation control module is based on the SAC reinforcement learning algorithm. It is trained by simulating different acceleration and vehicle speed conditions through random acceleration curves. The obtained motor compensation amount can be adapted to multiple conditions, which solves the problem that the fixed parameters of the traditional control strategy cannot be adapted to complex conditions, and greatly improves the working condition adaptation capability of the mode switching control. (3) The integration of "dynamic programming algorithm + particle swarm optimization algorithm + reinforcement learning algorithm" replaces the traditional manual experience calibration, which shortens the calibration cycle, reduces the testing cost, and avoids the influence of the engineer's level on the calibration quality, thereby improving the consistency and stability of the control strategy. (4) All algorithms of the present invention are designed for engineering implementation. The motor compensation amount obtained by training can be converted into a lookup table mode for real-time calling. The entire control process is executed online in real time. The control strategy can be dynamically adjusted according to the actual working conditions of the vehicle, adapting to the actual driving scenarios of hybrid electric vehicles. It can also be extended to various hybrid configurations such as power split type and series-parallel type. Attached Figure Description

[0020] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0021] Figure 1 This is a schematic diagram of the self-learning mode switching system for a series-parallel hybrid electric vehicle provided in Embodiment 1 of the present invention.

[0022] Figure 2 This is a simplified structural diagram of a power-split hybrid electric vehicle provided in Embodiment 1 of the present invention.

[0023] Figure 3 This is a flowchart of a self-learning method for mode switching in a series-parallel hybrid electric vehicle provided in Embodiment 2 of the present invention.

[0024] Figure 4 This is a partial block diagram of the electronic device provided in Embodiment 3 of the present invention. Detailed Implementation

[0025] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subroutine, etc.

[0026] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0027] The present invention will now be described in detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0028] Example 1 The specific implementation method is as follows: like Figure 1 The diagram shown is a schematic of the self-learning mode switching system for a series-parallel hybrid electric vehicle provided by the present invention.

[0029] As an example, the self-learning system includes an engine operation optimization module, a basic mode switching control module, and a self-learning compensation control module. The engine operation optimization module uses mode switching time, power source energy consumption, engine speed tracking error, and engine terminal speed tracking error as four-element optimization objectives, and solves for the optimal engine start-up and pre-operation speed curves based on an online dynamic programming algorithm. The basic mode switching control module uses the optimal start-up and pre-operation speed curves output by the engine operation optimization module as the speed tracking benchmark, and combines proportional-integral closed-loop control and particle swarm optimization algorithm to determine the optimal basic output torque of the dual motors to adapt to the mode switching process. The self-learning compensation control module generates dual-motor compensation torque based on a soft actor-critic reinforcement learning algorithm for different acceleration conditions, and linearly combines the dual-motor compensation torque and the optimal basic output torque of the dual motors to obtain the actual output torque of the dual motors. The actual output torque of the dual motors is used as a motor torque command to control the torque coordination control of the hybrid electric vehicle to complete the mode switching.

[0030] Preferred, combined Figure 2 As shown, the hybrid electric vehicle adapted to this system is a power-split hybrid electric vehicle. Its transmission structure includes a dual planetary gear set mechanism, a torsional damper 1, a lock-up clutch 2, an engine 3, a first motor 4, a second motor 5, and an output shaft 6. The dual planetary gear set mechanism includes a first planetary gear set 7 and a second planetary gear set 8. Both the first planetary gear set 7 and the second planetary gear set 8 are composed of a sun gear 70, a planet carrier 71, and a ring gear 72. The engine 3 is connected to the planet carrier 71 of the first planetary gear set 7 through the torsional damper 1. The ring gear of the second planetary gear set 8 is connected to the output shaft 6. The first motor 4 and the second motor 5 are respectively connected to the sun gear 70 of the first planetary gear set and the sun gear of the second planetary gear set 8. The ring gear 72 of the first planetary gear set 7 is connected to the ring gear of the second planetary gear set 8. The planet carrier of the second planetary gear set 8 is locked to the frame through the lock-up clutch 2.

[0031] In some feasible implementations, in the engine operation optimization module, the engine speed of the transmission structure, the speed of the planet carrier 71 of the first planetary gear set 7, the angular error between the engine 3 and the first planetary gear set 7, the speed of the ring gear of the second planetary gear set 8, the speed of the output shaft, and the angular error between the second planetary gear set 8 and the output shaft 6 are used as state variables. The output torque of the first motor 4 and the second motor 5 are used as system control input variables, and the speed of the output shaft 6 and the vehicle speed are used as output variables. The dynamic differential equation of the mode switching process of the series-parallel hybrid electric vehicle is established as follows: ; ; ; ; In the formula, x is the state variable of the mode switching process. This is expressed as engine speed. This indicates the rotational speed of the planet carrier in the first planetary gear set. This indicates the angular error between the engine and the first planetary gear set. This indicates the rotational speed of the ring gear of the second planetary gear set. Indicates the output shaft speed. This represents the rotational angle error between the second planetary gear set and the output shaft; u here serves as the system control input variable, which is essentially the motor output torque. and The output torques of the first and second motors are represented respectively; w is the known input variable disturbance of the system. and These represent the engine output torque and the load on the output shaft, respectively. , and These represent the coefficient matrices for the state variables, control input variables, and known output variables, respectively.

[0032] Specifically, ; , ; In the formula, and These are the equivalent rotational inertia of the engine and the output shaft, respectively. and These are the damping and stiffness coefficients between the engine and the planetary carrier of the first planetary set, respectively. and These are the damping and stiffness coefficients between the gear ring of the second planetary gear set and the output shaft, respectively. , , and All are equivalent inertia coefficients, and , , , ; , , and All are equivalent inertia coefficients, and , , , ;in, , , and These are the rotational inertia coefficients of the planet carrier, sun gear, ring gear, and first motor of the first planetary gear set, respectively. , , and These are the rotational inertia coefficients of the planet carrier, sun gear, ring gear, and second motor of the second planetary set, respectively. and These are the characteristic parameters of the first and second planetary arrays, respectively. =3.8, =2.5.

[0033] In some feasible implementations, the engine operation optimization module, based on the dynamic differential equation of the mode switching process, uses a weighted function of mode switching time, power source energy consumption, engine speed tracking error, and engine terminal speed tracking error as the optimization objective function, and uses the engine start-up and pre-operation curves as the optimization targets. Through online optimization, it obtains the optimal engine start-up and pre-operation curves. The mathematical expression of the online optimization model is as follows: ; ; In the formula, min represents minimization; This represents the overall optimization objective function of the engine operation optimization module; Indicates the mode switching time function; Represents the energy consumption function of the power source; This represents the engine speed tracking error function; The function representing the engine terminal speed tracking error is denoted by t; t represents time. The weighting coefficients of the mode switching time function; The weighting coefficients represent the energy consumption function of the power source; The weighting coefficients represent the engine speed tracking error function; The weighting coefficients represent the engine terminal speed tracking error function; and These represent the end and start times of mode switching, respectively; u represents the motor output torque; and These represent the optimal engine start-up and pre-operation target speed and the actual speed, respectively. This represents the engine terminal speed tracking error; the engine operation optimization module iteratively solves the optimization model using an online dynamic programming algorithm to obtain the optimal speed curves for the engine start-up and pre-operation stages during mode switching; wherein, the solution process uses the dynamic differential equation of the hybrid electric vehicle mode switching process as a constraint condition, combining dynamic characteristics with optimization objectives to achieve optimal planning of the engine speed curve.

[0034] Specifically, the core logic of the engine operation optimization module is: to establish dynamic differential equations, construct a four-element multi-objective optimization model, and obtain the optimal engine speed curve by solving the dynamic equations through online dynamic programming iteration.

[0035] For a specific example, the vehicle is initially in pure electric mode, driving at low speed in the city, with a real-time speed v=25km / h and a real-time acceleration of... When the driver presses the accelerator pedal to 55%, the vehicle control unit (ECU) determines that it needs to switch to hybrid mode, issues a switching command, and sets the start time of the switch. Target switching end time ≤0.8s, the switching process prioritizes smoothness and speed while also considering low energy consumption. Core parameters of the vehicle and transmission structure: Transmission structure: Dual planetary gear set mechanism (characteristic parameters of the first planetary gear set) =3.8, =2.5), one 1.5T engine (idle speed 800r / min, rated speed 5500r / min, maximum output torque) ), First motor (main drive, ), second motor (generator / drive, Torsional shock absorber, lock-up clutch (locked to frame), output shaft (gear ratio 4.1); Vehicle parameters: curb weight 1650kg, tire rolling radius 0.33m, mode switching impact threshold. The core parameters of the algorithm are as follows: online dynamic programming solution step size 0.02s; particle swarm optimization particle number 30, iteration count 100, learning factor 1.8; SAC algorithm discount factor γ=0.95, soft update coefficient α=0.005, experience replay pool capacity 5×10⁵, Gaussian noise variance 0.05, sample batch size 128. (All examples below are based on this.)

[0036] Taking low-speed, slow-acceleration operation as an example, since smoothness and speed are prioritized in low-speed, slow-acceleration operation, the value is: =0.25、 =0.2、 =0.3、 =0.25. Using the dynamic differential equation as a hard constraint, the optimization model is iteratively solved (step size 0.02s) to obtain the optimal engine start-up and pre-operation speed curves. The online dynamic programming (DP) is solved in reverse order: iterating from the terminal stage k=35 to the initial stage k=0, calculating the optimal cost function for each stage, and finally backtracking from k=0 to obtain the optimal decision sequence. This includes: discretizing the continuous dynamic equations and optimization objective, transforming them into a multi-stage decision problem; iterating in reverse order from the terminal stage, calculating the optimal cost and optimal decision for each state using the Bellman equation; and backtracking forward to obtain the optimal decision sequence to obtain the optimal engine speed curve, including: Start-up phase (0s~0.3s): The engine smoothly starts from 0r / min to idle speed of 800r / min, with speed tracking error. No speed fluctuation; Pre-operation phase (0.3s~0.7s): The engine speed increases linearly from 800r / min to 1600r / min (calculated by the vehicle's upper-level energy management strategy, which is the optimal operating speed of the engine for slow acceleration of 25km / h). Terminal stability ( =0.7s): The rotational speed stabilized at 1600 r / min.

[0037] In some feasible implementations, the basic mode switching control module is used to: merge the output shaft optimal speed curve calculated based on the target vehicle speed and the engine optimal start-up and pre-operation speed curves output by the engine operation optimization module to obtain a multi-dimensional total speed tracking target. ; Calculate the deviation between the total rotational speed tracking target and the state variables during the mode switching process. ; to deviation As the input state variable for proportional-integral control, the output torque of the two motors is calculated based on proportional-integral control. The calculation formula is as follows: ; ; ; In the formula, This represents the optimal engine start-up and pre-operation target speed. This indicates the target rotational speed of the output shaft, calculated from the target vehicle speed. This indicates the real-time rotational speed of the output shaft. represents the engine speed, and x represents the state variable during the mode switching process; This is the proportional coefficient, used to suppress the speed deviation at the current moment; is the integral coefficient, used to eliminate the steady-state speed difference during the entire mode switching process; u is the motor output torque.

[0038] Preferably, the basic mode switching control module is used to adjust and optimize the objective function using a weighted function of mode switching impact, dual-motor energy consumption, and speed tracking error as parameters. , To optimize parameters and set upper and lower limits based on the hardware performance of the dual motors and the ride comfort requirements of the vehicle, a constrained parameter optimization model is constructed. The parameter optimization model is as follows: ; ; In the formula, min represents minimization; This represents the optimization objective function in the basic mode switching control module; Represents the impact function; Represents the energy consumption function of the motor; This represents the speed tracking error function; The weighting coefficients of the impact function are represented. The weighting coefficients represent the energy consumption function of the motor. The weighting coefficients represent the speed tracking error function; This is the proportionality coefficient; The integral coefficient; , They represent The minimum and maximum values; , They represent The minimum and maximum values; 'a' is the vehicle acceleration; The target engine speed; This represents the actual engine speed. The basic mode switching control module uses a particle swarm optimization algorithm to... , By performing iterative optimization, the optimal proportional coefficient can be obtained. and optimal integral coefficients Substituting this into the proportional-integral control calculation formula yields the optimal basic output torque of the dual motors. .

[0039] Specifically, the core logic of the basic mode switching control module is to merge the multi-dimensional total speed tracking target, calculate the speed deviation and substitute it into the PI control to initially calculate the torque, optimize the PI parameters by particle swarm optimization, and solve for the optimal basic torque.

[0040] Specific examples, , , Set the upper and lower limits of the PI parameter: , The optimal PI parameters were obtained through particle swarm optimization (30 particles, 100 iterations). , The above , Substitute to The optimal basic output torque (two-dimensional vector) of the dual motors is calculated as follows: .

[0041] In some feasible implementations, the self-learning compensation control module uses a soft actor-critic reinforcement learning algorithm to train the motor compensation amount. The observation space is based on vehicle speed, acceleration, engine speed, vehicle speed tracking error, motor output torque, impact, and mode switching time. The action space is based on the motor compensation torque. The reward space is based on a weighted function of vehicle speed tracking error, impact, and mode switching time. Random acceleration curves are generated in the reset function for tracking training to obtain a motor compensation torque suitable for multiple operating conditions. The motor compensation torque is combined with the optimal motor output torque obtained by the basic mode switching control module to become the actual motor output torque.

[0042] Preferably, the observation space is the set of state parameters at time t. ,in, For the vehicle speed, For the acceleration of the whole vehicle, For engine speed, For vehicle speed tracking error, For motor output torque, For the impact of mode switching, The time elapsed for mode switching; the action space is the motor compensation torque at time t. The reward space is the reward function at time t. , This is a weighted function of vehicle speed tracking error, impact, and mode switching time, characterizing the condition adaptation effect of the current motor compensation amount. The value is positively correlated with the mode switching control effect.

[0043] Preferably, the training process of the soft actor-critic reinforcement learning algorithm includes initialization, environment interaction, experience storage, sample sampling, target Q-value calculation, network update, target network soft update, and iterative convergence steps, specifically: S1. Initialization: Initialize the policy network. Two commentator networks , Two target commentator networks Entropy temperature coefficient and experience replay pool R; S2, Environmental Interaction: In the vehicle control environment, based on the current state Explore by sampling actions using a random policy; the action expression is: ; Where t represents time t; For policy networks; It is a random noise variable; The action generation function is based on a policy network and random noise, which maps the observed state and noise to specific motor torque compensation quantities. S3, Storing Experience: This involves storing experience data sets obtained from environmental interactions. , , , ) Stored into the experience replay pool R, where Let be the reward function value at time t. The observed state at time t+1; S4. Sampling data: Randomly sample empirical data samples of batch size N from the empirical replay pool R. , , , ), where i is the sample number, Let i be the observed state of the i-th sample. For the motor compensation torque of the i-th sample, Let i be the reward function value for the i-th sample. This represents the observation state of the (i+1)th sample; S5. Calculate the target Q value: for the state at the next time step. Sample the optimal action from the current policy network. The target Q value of the i-th sample is calculated by combining the entropy term. ,include: ; Where i is the i-th data sample; In order to be in The following actions are sampled from the policy network; As a discount factor, and ; The probability is the logarithmic value of the action. S6. Update the critic network: Train two critic networks using mean squared error loss. The loss function expression is as follows: ; in, The mean squared error of the two critic networks; for and Q-value estimation; N is the batch size; S7. Update the policy network: Update the policy network by minimizing the policy loss. parameters The policy loss expression is: ; in, For policy network loss; S8. Target Network Soft Update: The parameters of the two target reviewer networks are iteratively updated using a soft update method. S9. Iterative convergence: Repeat steps S2-S8, continuously interact with the vehicle control environment and update network parameters until the network loss tends to stabilize or the preset maximum number of iterations is reached, and the training is determined to be converged.

[0044] Specifically, the core logic of the self-learning compensation control module is to construct the three spaces of the SAC algorithm, obtain the torque compensation torque through online training, merge the compensation torque with the basic torque by looking up the table, and coordinate the clutch and torque control to complete the switching.

[0045] For a specific example, based on the SAC algorithm training process described above, and combined with the mixed training data from the bench and real vehicles in the experience replay pool, online rapid training was performed for this operating condition. After 300 generations of network training, the loss tended to stabilize, and convergence was determined, thus obtaining the optimal compensation torque of the dual motors suitable for this low-speed, slow-acceleration operating condition: That is, the compensation amount of the first motor (Positive compensation, matching engine starting torque), second motor compensation amount (Positive compensation ensures smooth torque coupling of the dual planetary gear set). The optimal compensation torque is converted into a two-dimensional lookup table and stored in the ECU. The lookup input is "vehicle speed and acceleration".

[0046] The optimal compensation torque retrieved from the lookup table is linearly combined with the optimal base torque from the base mode switching control module to obtain the actual output torque of the dual motors: The vehicle's ECU will The torque is simultaneously sent to the first and second motor actuators. The torque from the dual motors and the engine is coupled through the dual planetary gear mechanism to drive the output shaft, thus completing the mode switching.

[0047] It should be noted that the above embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the design concept of the present invention should fall within the scope of protection of the present invention.

[0048] Example 2 Please see Figure 3 This embodiment provides a flowchart of a self-learning method for mode switching in a series-parallel hybrid electric vehicle.

[0049] As an example, the method is applied to the self-learning mode switching system for a parallel-hybrid electric vehicle described in Example 1, and the method includes: When the engine operation optimization module receives the mode switching command, it uses the mode switching time, power source energy consumption, engine speed tracking error, and engine terminal speed tracking error as the four-element optimization objectives, and solves the optimal engine start-up and pre-operation speed curves based on the online dynamic programming algorithm. The basic mode switching control module uses the optimal start-up and pre-operation speed curves output by the engine operation optimization module as the speed tracking benchmark. Combining proportional-integral closed-loop control and particle swarm parameter optimization algorithm, it determines the optimal basic output torque of the dual motors to adapt to the mode switching process. The self-learning compensation control module is based on the soft actor-critic reinforcement learning algorithm. It generates dual-motor compensation torque for different acceleration conditions. The dual-motor compensation torque is linearly combined with the optimal basic output torque of the dual motors to obtain the actual output torque of the dual motors. The actual output torque of the dual motors is used as a motor torque command to control the hybrid electric vehicle to complete the mode switching.

[0050] It is not difficult to see that this embodiment is a method embodiment corresponding to the first embodiment, and this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment.

[0051] Example 3 Please see Figure 4 The present invention also provides an electronic device, including: a memory and a processor; the memory stores at least one program instruction; the processor loads and executes the at least one program instruction to implement the hybrid electric vehicle mode switching self-learning method provided in Embodiment 2.

[0052] The memory 402 and processor 401 are connected via a bus, which may include any number of interconnecting buses and bridges, connecting various circuits of one or more processors 401 and memory 402 together. The bus may also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. A bus interface provides an interface between the bus and the transceiver. The transceiver may be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 401 is transmitted over a wireless medium via an antenna, which further receives data and transmits it to processor 401.

[0053] Processor 401 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory 402 can be used to store data used by processor 401 during operation.

[0054] Example 4 This invention also proposes a storage medium storing a self-learning method for mode switching in a series-parallel hybrid electric vehicle (SPEV). When the SPEV mode switching self-learning program is executed, it implements the steps of the above-described SPEV mode switching self-learning method. Since this storage medium employs all the technical solutions of the above embodiments, it possesses at least all the beneficial effects brought about by the technical solutions of the above embodiments, which will not be elaborated upon further here.

[0055] The above descriptions are merely embodiments of the present invention. Commonly known structures and characteristics are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, based on the guidance provided in this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A self-learning system for mode switching of a series-parallel hybrid electric vehicle, wherein the series-parallel hybrid electric vehicle adapted to the system is a power-split series-parallel hybrid electric vehicle, and its transmission structure includes a dual planetary gear set mechanism, a torsional damper (1), a lock-up clutch (2), an engine (3), a first motor (4), a second motor (5), and an output shaft (6); the dual planetary gear set mechanism includes a first planetary gear set (7) and a second planetary gear set (8), both the first planetary gear set (7) and the second planetary gear set (8) consisting of a sun gear (70) and a planet carrier (70). 1) and gear ring (72); the engine (3) is connected to the planet carrier (71) of the first planetary gear set (7) through a torsional damper (1), the gear ring of the second planetary gear set (8) is connected to the output shaft (6), the first motor (4) and the second motor (5) are respectively connected to the sun gear (70) of the first planetary gear set and the sun gear of the second planetary gear set (8), the gear ring (72) of the first planetary gear set (7) is connected to the gear ring of the second planetary gear set (8), and the planet carrier of the second planetary gear set (8) is locked to the frame through a locking clutch (2); characterized in that, The self-learning system includes an engine operation optimization module, a basic mode switching control module, and a self-learning compensation control module. The engine operation optimization module uses mode switching time, power source energy consumption, engine speed tracking error, and engine terminal speed tracking error as four optimization objectives, and solves the optimal engine start-up and pre-operation speed curves based on online dynamic programming algorithm. The basic mode switching control module uses the optimal start-up and pre-operation speed curves output by the engine operation optimization module as the speed tracking benchmark, and combines proportional-integral closed-loop control and particle swarm parameter optimization algorithm to determine the optimal basic output torque of the dual motors to adapt to the mode switching process. The self-learning compensation control module is based on the soft actor-critic reinforcement learning algorithm. It generates dual-motor compensation torque for different acceleration conditions, and linearly combines the dual-motor compensation torque with the optimal basic output torque of the dual motors to obtain the actual output torque of the dual motors. The actual output torque of the dual motors is used as the motor torque command to control the torque coordination control of the hybrid electric vehicle to complete the mode switching.

2. The self-learning mode switching system for a series-parallel hybrid electric vehicle according to claim 1, characterized in that, In the engine operation optimization module, the engine (3) speed, the planet carrier (71) speed of the first planetary gear set (7), the rotational angle error between the engine (3) and the first planetary gear set (7), the gear ring speed of the second planetary gear set (8), the output shaft speed, and the rotational angle error between the second planetary gear set (8) and the output shaft (6) are used as state variables. The output torque of the first motor (4) and the second motor (5) are used as system control input variables. The output shaft speed (6) and the vehicle speed are used as output variables. The dynamic differential equation of the hybrid electric vehicle mode switching process is established: ; ; ; ; In the formula, x is the state variable of the mode switching process. This is expressed as engine speed. This indicates the rotational speed of the planet carrier in the first planetary gear set. This indicates the angular error between the engine and the first planetary gear set. This indicates the rotational speed of the ring gear of the second planetary gear set. Indicates the output shaft speed. This represents the angular error between the second planetary gear set and the output shaft; u is the motor output torque; and The output torques of the first and second motors are represented respectively; w is the known input variable disturbance of the system. and These represent the engine output torque and the load on the output shaft, respectively. , and These represent the coefficient matrices for the state variables, control input variables, and known output variables, respectively.

3. The self-learning system for mode switching of a series-parallel hybrid electric vehicle according to claim 2, characterized in that, The engine operation optimization module, based on the dynamic differential equations of the mode switching process, uses a weighted function of mode switching time, power source energy consumption, engine speed tracking error, and engine terminal speed tracking error as the optimization objective function, and the engine start-up and pre-operation curves as the optimization targets. Through online optimization, it obtains the optimal engine start-up and pre-operation curves. The mathematical expression of the online optimization model is as follows: ; ; In the formula, min represents minimization; This represents the overall optimization objective function of the engine operation optimization module; Indicates the mode switching time function; Represents the energy consumption function of the power source; This represents the engine speed tracking error function; The function representing the engine terminal speed tracking error is denoted by t; t represents time. The weighting coefficients of the mode switching time function; The weighting coefficients represent the energy consumption function of the power source; The weighting coefficients represent the engine speed tracking error function; The weighting coefficients represent the engine terminal speed tracking error function; and These represent the end and start times of mode switching, respectively; u represents the motor output torque; and These represent the optimal engine start-up and pre-operation target speed and the actual speed, respectively. This indicates the engine terminal speed tracking error; The engine operation optimization module iteratively solves the optimization model using an online dynamic programming algorithm to obtain the optimal speed curves for the engine start-up and pre-operation stages during mode switching. The solution process uses the dynamic differential equation of the mode switching process of the series-parallel hybrid electric vehicle as a constraint condition, combining dynamic characteristics with optimization objectives to achieve optimal planning of the engine speed curve.

4. The self-learning mode switching system for a series-parallel hybrid electric vehicle according to claim 3, characterized in that, The basic mode switching control module is used for: By merging the optimal output shaft speed curve calculated based on the target vehicle speed and the optimal engine start-up and pre-operation speed curves output by the engine operation optimization module, a multi-dimensional total speed tracking target is obtained. ; Calculate the deviation between the total rotational speed tracking target and the state variables during the mode switching process. ; Deviation As the input state variable for proportional-integral control, the output torque of the two motors is calculated based on proportional-integral control. The calculation formula is as follows: ; ; ; In the formula, This represents the optimal engine start-up and pre-operation target speed. This indicates the target rotational speed of the output shaft, calculated from the target vehicle speed. This indicates the real-time rotational speed of the output shaft. represents the engine speed, and x represents the state variable during the mode switching process; This is the proportional coefficient, used to suppress the speed deviation at the current moment; This is the integral coefficient, used to eliminate the steady-state speed difference throughout the mode switching process; u represents the motor's output torque.

5. The self-learning system for mode switching of a series-parallel hybrid electric vehicle according to claim 4, characterized in that, The basic mode switching control module is used to adjust and optimize the objective function using a weighted function of mode switching impact, dual-motor energy consumption, and speed tracking error as parameters. , To optimize parameters and set upper and lower limits based on the hardware performance of the dual motors and the ride comfort requirements of the vehicle, a constrained parameter optimization model is constructed. The parameter optimization model is as follows: ; ; In the formula, min represents minimization; This represents the optimization objective function in the basic mode switching control module; Represents the impact function; Represents the energy consumption function of the motor; This represents the speed tracking error function; The weighting coefficients of the impact function are represented. The weighting coefficients represent the energy consumption function of the motor. The weighting coefficients represent the speed tracking error function; This is the proportionality coefficient; The integral coefficient; , They represent The minimum and maximum values; , They represent The minimum and maximum values; 'a' is the vehicle acceleration; The target engine speed; This refers to the actual engine speed. The basic mode switching control module uses a particle swarm optimization algorithm to... , By performing iterative optimization, the optimal proportional coefficient can be obtained. and optimal integral coefficients Substituting this into the proportional-integral control calculation formula yields the optimal basic output torque of the dual motors. .

6. The self-learning mode switching system for a series-parallel hybrid electric vehicle according to claim 1, characterized in that, The self-learning compensation control module uses a soft actor-critic reinforcement learning algorithm to train the motor compensation amount. It uses vehicle speed, acceleration, engine speed, vehicle speed tracking error, motor output torque, impact, and mode switching time as the observation space, motor compensation torque as the action space, and a weighted function of vehicle speed tracking error, impact, and mode switching time as the reward space. Random acceleration curves are generated in the reset function for tracking training to obtain motor compensation torque applicable to multiple operating conditions. The motor compensation torque is combined with the optimal motor output torque obtained by the basic mode switching control module to become the actual motor output torque.

7. The self-learning mode switching system for a series-parallel hybrid electric vehicle according to claim 6, characterized in that, The observation space is the set of state parameters at time t. ,in, For the vehicle speed, For the acceleration of the whole vehicle, For engine speed, For vehicle speed tracking error, For motor output torque, For the impact of mode switching, Time elapsed for mode switching; The action space is the motor compensation torque at time t. ; The reward space is the reward function at time t. , This is a weighted function of vehicle speed tracking error, impact, and mode switching time, characterizing the condition adaptation effect of the current motor compensation amount. The value is positively correlated with the mode switching control effect.

8. The self-learning mode switching system for a series-parallel hybrid electric vehicle according to claim 7, characterized in that, The training process of the soft actor-critic reinforcement learning algorithm includes initialization, environment interaction, experience storage, sample sampling, target Q-value calculation, network update, target network soft update, and iterative convergence steps, specifically: S1. Initialization: Initialize the policy network. Two commentator networks , Two target commentator networks Entropy temperature coefficient and experience replay pool R; S2, Environmental Interaction: In the vehicle control environment, based on the current state Explore by sampling actions using a random policy; the action expression is: ; Where t represents time t; For policy networks; It is a random noise variable; The action generation function is based on a policy network and random noise, which maps the observed state and noise to specific motor torque compensation quantities. S3, Storing Experience: This involves storing experience data sets obtained from environmental interactions. , , , ) Stored into the experience replay pool R, where Let be the reward function value at time t. The observed state at time t+1; S4. Sampling data: Randomly sample empirical data samples of batch size N from the empirical replay pool R. , , , ), where i is the sample number, Let i be the observed state of the i-th sample. For the motor compensation torque of the i-th sample, Let i be the reward function value for the i-th sample. This represents the observation state of the (i+1)th sample; S5. Calculate the target Q value: for the state at the next time step. Sample the optimal action from the current policy network. The target Q value of the i-th sample is calculated by combining the entropy term. ,include: ; Where i is the i-th data sample; In order to be in The following actions are sampled from the policy network; As a discount factor, and ; The probability is the logarithmic value of the action. S6. Update the critic network: Train two critic networks using mean squared error loss. The loss function expression is as follows: ; in, The mean squared error of the two critic networks; for and Q-value estimation; N is the batch size; S7. Update the policy network: Update the policy network by minimizing the policy loss. parameters The policy loss expression is: ; in, For policy network loss; S8. Target Network Soft Update: The parameters of the two target reviewer networks are iteratively updated using a soft update method. S9. Iterative convergence: Repeat steps S2-S8, continuously interact with the vehicle control environment and update network parameters until the network loss tends to stabilize or the preset maximum number of iterations is reached, and the training is determined to be converged.

9. The self-learning mode switching system for a series-parallel hybrid electric vehicle according to claim 8, characterized in that, After training convergence, the motor compensation torque applicable to multiple operating conditions is converted into a lookup table mode for storage. The input of the lookup table mode is the real-time operating parameters of the whole vehicle, and the output is the matched motor torque compensation torque. The motor compensation torque retrieved from the lookup table is linearly combined with the optimal basic output torque of the dual motors to obtain the actual output torque of the dual motors. The actual output torque of the motors serves as the final motor control command in the mode switching process of the series-parallel hybrid electric vehicle.

10. A self-learning method for mode switching in a series-parallel hybrid electric vehicle, applied to the self-learning system for mode switching of a series-parallel hybrid electric vehicle as described in any one of claims 1-9, characterized in that, The method includes: When the engine operation optimization module receives the mode switching command, it uses the mode switching time, power source energy consumption, engine speed tracking error, and engine terminal speed tracking error as the four-element optimization objectives, and solves the optimal engine start-up and pre-operation speed curves based on the online dynamic programming algorithm. The basic mode switching control module uses the optimal start-up and pre-operation speed curves output by the engine operation optimization module as the speed tracking benchmark. Combining proportional-integral closed-loop control and particle swarm parameter optimization algorithm, it determines the optimal basic output torque of the dual motors to adapt to the mode switching process. The self-learning compensation control module is based on the soft actor-critic reinforcement learning algorithm. It generates dual-motor compensation torque for different acceleration conditions. The dual-motor compensation torque is linearly combined with the optimal basic output torque of the dual motors to obtain the actual output torque of the dual motors. The actual output torque of the dual motors is used as a motor torque command to control the hybrid electric vehicle to complete the mode switching.