Engine torque control method and device based on reinforcement learning and vehicle-mounted controller
By using a reinforcement learning-based engine torque control method and a multi-objective optimized engine torque control model, the problems of poor generalization, lag in dynamic response, and high calibration cost in existing technologies are solved, achieving full-condition adaptability and efficient torque control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA FAW CO LTD
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-15
AI Technical Summary
Existing engine torque control methods have poor generalization ability, are difficult to adapt to complex operating conditions, have lag in dynamic response, low transient control accuracy, high calibration cost, and long iteration cycle.
An engine torque control method based on reinforcement learning is adopted. By acquiring real-time operating status information, a multi-objective optimized engine torque control model is used for control. The model is trained by combining reward functions with total energy consumption penalty, SOC deviation penalty and ineffective energy consumption penalty to achieve full operating condition adaptation and lightweight deployment to the vehicle controller.
It achieves full-condition adaptability across different ambient temperatures and speed ranges, reduces fuel consumption, improves torque tracking accuracy, shortens calibration cycles, reduces maintenance costs, and enhances product iteration efficiency.
Smart Images

Figure CN122040436A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of automotive engine electronic control technology, and more specifically, to an engine torque control method, device, and on-board controller based on reinforcement learning. Background Technology
[0002] Currently, torque control in EMS mainly relies on empirical control strategies (such as lookup table methods and PID control) or simplified model control (such as torque estimation based on mean models). Its core logic is to achieve torque output through pre-calibrated operating parameters (such as throttle opening and injection pulse width corresponding to speed-load). Mainstream EMS in the industry also generally adopt the "lookup table + feedback correction" mode, relying on a large amount of bench calibration data to cover limited operating conditions.
[0003] The existing technology has the following problems: poor generalization, making it difficult to adapt to complex working conditions; lag in dynamic response and low transient control accuracy; insufficient multi-objective optimization coordination; high calibration cost and long iteration cycle. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide an engine torque control method, device and vehicle controller based on reinforcement learning, which has good generalization, timely dynamic response, realizes multi-objective optimization and cooperation, and has low calibration cost.
[0005] This application provides an engine torque control method based on reinforcement learning, the method comprising the following steps: The system acquires real-time operating status information of the target vehicle, including: wheel-end power demand, vehicle speed, vehicle acceleration, real-time battery charge, target remaining battery charge, clutch engagement / disengagement status, and real-time engine speed. An engine torque control model is deployed in the target vehicle's onboard controller. This engine torque control model is matched to the target vehicle. The engine torque control model is trained in a simulation environment based on multi-objective optimization of the target torque action space, which includes the operating status information. The system inputs the real-time operating status information of the target vehicle into the engine torque control model, processes this information, and determines the target torque command for the target vehicle in real-time. The engine actuator is controlled to operate according to the target torque command.
[0006] In some embodiments, in the reinforcement learning-based engine torque control method, the reward function of the engine torque control model is optimized based on multi-objective optimization of total energy consumption penalty, SOC deviation penalty and ineffective energy consumption penalty; The weight of the total energy consumption penalty is greater than the weight of the SOC deviation penalty; the weight of the SOC deviation penalty is greater than the weight of the ineffective energy consumption penalty.
[0007] In some embodiments, the engine torque control model in the reinforcement learning-based engine torque control method is trained based on the following steps: Define the state space, action space, and reward function of the engine torque control model; the state space includes: wheel-end power demand, vehicle speed, vehicle acceleration, real-time battery charge, target remaining battery charge, clutch engagement / disengagement status, and real-time engine speed; the action space is the target engine torque; A simulation environment is set up; the simulation environment includes a hybrid electric vehicle simulation model, which includes an engine module, a power battery module, a motor module, a clutch module, a transmission system module, and a vehicle driving resistance sub-module; Based on the reward function, the engine torque control model is trained in a simulation environment through the interaction between the engine torque control model and the simulation environment until a preset training stop condition is met, thus obtaining a trained engine torque control model.
[0008] In some embodiments, the engine torque control model in the reinforcement learning-based engine torque control method is a lightweight model that has undergone precision quantization.
[0009] In some embodiments, the reinforcement learning-based engine torque control method further includes the following steps before deploying the engine torque control model to the vehicle controller: Using a calibration dataset that includes predetermined operating conditions, the accuracy of the trained engine torque control model is quantized to generate a lightweight engine torque control model. The performance indicators of the lightweight engine torque control model are then verified in a simulation environment, including total energy consumption, state of charge (SOC) control, and wheel-end power satisfaction rate.
[0010] In some embodiments, in the reinforcement learning-based engine torque control method, the engine torque control model is deployed to the vehicle controller through the following steps: The engine torque control model is converted into embedded code supported by the vehicle controller of the target vehicle; the signal interface parameters in the embedded code are configured according to the hardware I / O definition of the vehicle controller of the target vehicle. After configuring the signal interface parameters, the target flashing tool of the vehicle controller is used to securely transfer the binary file compiled based on the embedded code to the program memory of the vehicle controller.
[0011] In some embodiments, the reinforcement learning-based engine torque control method further includes the following steps after deploying the engine torque control model to the vehicle controller: After the vehicle controller is powered on, the vehicle bus messages are monitored to verify the vehicle controller's ability to receive the operating status information and send the target torque command, in order to perform communication verification. Functional testing is conducted by driving the target vehicle under various test conditions on a real vehicle or a rotary drum test bench, and collecting operational data; the operational data is used to determine the control effect evaluation results of the engine torque control model. The various test conditions include low-speed congestion conditions, high-speed cruising conditions, and rapid acceleration or hill climbing conditions.
[0012] In some embodiments, the reinforcement learning-based engine torque control method further includes: If the control effect evaluation result of the engine torque control model does not meet the preset compliance conditions, the engine torque control model is corrected based on the operating data until the control effect evaluation result meets the preset compliance conditions.
[0013] In some embodiments, a reinforcement learning-based engine torque control device is also provided, the device comprising: An acquisition module is used to acquire real-time operating status information of the target vehicle, including: wheel-end power demand, vehicle speed, vehicle acceleration, real-time battery charge, target remaining battery charge, clutch engagement / disengagement status, and real-time engine speed. An engine torque control model is deployed in the target vehicle's onboard controller. The engine torque control model is matched with the target vehicle. The engine torque control model is trained in a simulation environment based on multi-objective optimization of the target torque action space including the operating status information. A processing module is used to input the real-time operating status information of the target vehicle into the engine torque control model, process the real-time operating status information of the target vehicle through the engine torque control model, and determine the real-time target torque command of the target vehicle. The control module is used to control the operation of the engine actuator according to the target torque command. In some embodiments, an on-board controller is also provided, which performs the steps of the reinforcement learning-based engine torque control method.
[0014] This application provides an engine torque control method, device, and on-board controller based on reinforcement learning. The method acquires real-time operating status information of a target vehicle, including: wheel-end power demand, vehicle speed, vehicle acceleration, real-time battery charge, target remaining battery charge, clutch engagement / disengagement status, and real-time engine speed. A pre-trained engine torque control model is deployed in the on-board controller of the target vehicle. The real-time operating status information of the target vehicle is input into the pre-trained engine torque control model. The engine torque control model is matched to the target vehicle. The reinforcement learning model is trained in a simulation environment based on multi-objective optimization of the target torque action space including the operating status information. The engine torque control model processes the real-time operating status information of the target vehicle to determine the real-time target torque command. Based on the target torque command, the engine actuator is controlled to operate. The method leverages the real-time... Interactive learning capabilities enable full-condition adaptation to different ambient temperatures, speed ranges, and load rates, eliminating the need for extensive pre-calibration data and addressing the poor generalization issues of traditional strategies. Reinforcement learning algorithms in continuous action space optimize control command output, improving dynamic response lag while effectively controlling overshoot, thus resolving the dynamic response lag problem. A multi-dimensional reward function is designed to dynamically balance multiple performance objectives, achieving multi-objective collaborative optimization and avoiding performance sacrifices due to prioritizing a single objective, thus addressing insufficient multi-objective optimization collaboration. While ensuring torque accuracy meets standards, fuel consumption is reduced, avoiding the conflicting trade-offs of multiple objectives in traditional strategies. Utilizing the transfer learning capabilities of reinforcement learning models, strategy iteration can be completed with only a small amount of supplementary operating condition data after hardware changes, shortening the calibration cycle, solving the problem of high calibration costs, and effectively improving product iteration efficiency. It can also automatically adapt to non-calibration scenarios such as engine aging and high altitudes without manual parameter adjustments, reducing later maintenance costs and improving torque tracking accuracy under all operating conditions. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A flowchart of the engine torque control method based on reinforcement learning described in an embodiment of this application is shown; Figure 2 A flowchart of the method for training an engine torque control model according to an embodiment of this application is shown; Figure 3 A flowchart of the method for precision quantization of a model according to an embodiment of this application is shown; Figure 4 A flowchart illustrating the method for deploying the engine torque control model to an on-board controller according to an embodiment of this application is shown. Figure 5 A flowchart of another engine torque control method based on reinforcement learning, as described in an embodiment of this application, is shown. Figure 6 A schematic diagram of the structure of the engine torque control device based on reinforcement learning described in an embodiment of this application is shown. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0018] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0019] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0020] Currently, torque control in EMS mainly relies on empirical control strategies (such as lookup table methods and PID control) or simplified model control (such as torque estimation based on mean models). Its core logic is to achieve torque output through pre-calibrated operating parameters (such as throttle opening and injection pulse width corresponding to speed-load). Mainstream EMS in the industry also generally adopt the "lookup table + feedback correction" mode, relying on a large amount of bench calibration data to cover limited operating conditions.
[0021] The existing technology has the following problems: poor generalization, making it difficult to adapt to complex working conditions; lag in dynamic response and low transient control accuracy; insufficient multi-objective optimization coordination; high calibration cost and long iteration cycle.
[0022] Based on this, this application provides an engine torque control method, device, and on-board controller based on reinforcement learning. The method acquires real-time operating status information of a target vehicle, including: wheel-end power demand, vehicle speed, vehicle acceleration, real-time battery charge, target remaining battery charge, clutch engagement / disengagement status, and real-time engine speed. A pre-trained engine torque control model is deployed in the on-board controller of the target vehicle. The real-time operating status information of the target vehicle is input into the pre-trained engine torque control model. The engine torque control model is matched with the target vehicle. The reinforcement learning model is trained in a simulation environment based on multi-objective optimization of the target torque action space including the operating status information. The engine torque control model processes the real-time operating status information of the target vehicle to determine the real-time target torque command. Based on the target torque command, the engine actuator is controlled to operate. The method leverages reinforcement learning... Real-time interactive learning capabilities enable full-condition adaptation to different ambient temperatures, speed ranges, and load rates, eliminating the need for extensive pre-calibration data and addressing the poor generalization issues of traditional strategies. Reinforcement learning algorithms in continuous action space optimize control command output, improving dynamic response lag while effectively controlling overshoot, thus resolving the dynamic response lag problem. A multi-dimensional reward function is designed to dynamically balance multiple performance objectives, achieving multi-objective collaborative optimization and avoiding performance sacrifices due to prioritizing a single objective, thus addressing insufficient multi-objective optimization collaboration. While ensuring torque accuracy meets standards, fuel consumption is reduced, avoiding the conflicting trade-offs of multiple objectives in traditional strategies. Utilizing the transfer learning capabilities of reinforcement learning models, strategy iteration can be completed with only a small amount of supplementary operating condition data after hardware changes, shortening the calibration cycle, solving the problem of high calibration costs, and effectively improving product iteration efficiency. It can also automatically adapt to non-calibration scenarios such as engine aging and high altitudes without manual parameter adjustments, reducing later maintenance costs and improving torque tracking accuracy under all operating conditions.
[0023] Please refer to Figure 1, Figure 1 A flowchart of the engine torque control method based on reinforcement learning described in an embodiment of this application is shown; as follows: Figure 1 As shown, the method includes the following steps S101-S103: S101. Obtain the real-time operating status information of the target vehicle, including: wheel-end power demand, vehicle speed, vehicle acceleration, real-time power battery charge, target remaining power battery charge, clutch engagement / disengagement status, and real-time engine speed; the on-board controller of the target vehicle is equipped with a pre-trained engine torque control model; the engine torque control model is matched with the target vehicle; the engine torque control model is trained in a simulation environment based on multi-objective optimization of the target torque action space including the operating status information; S102. Input the real-time operating status information of the target vehicle into the engine torque control model, process the real-time operating status information of the target vehicle through the engine torque control model, and determine the real-time target torque command of the target vehicle; S103. Control the engine actuator to operate according to the target torque command. In step S101, the real-time operating status information of the target vehicle is obtained. The operating status information includes: wheel-end power demand, vehicle speed, vehicle acceleration, real-time power battery charge, target remaining power battery charge, clutch engagement / disengagement status, and real-time engine speed. The on-board controller of the target vehicle is equipped with a pre-trained engine torque control model.
[0024] In some embodiments, please refer to Figure 2 , Figure 2 A flowchart illustrating the method for training an engine torque control model according to an embodiment of this application is shown; as follows: Figure 2 As shown, the engine torque control model is trained based on the following steps S201-S203: S201. Define the state space, action space, and reward function of the engine torque control model; the state space includes: wheel-end power demand, vehicle speed, vehicle acceleration, real-time battery charge, target remaining battery charge, clutch engagement / disengagement status, and real-time engine speed; the action space is the target engine torque. S202. Set up a simulation environment; the simulation environment includes a hybrid electric vehicle simulation model, which includes an engine module, a power battery module, a motor module, a clutch module, a transmission system module, and a vehicle driving resistance sub-module. S203. Based on the reward function, in the simulation environment, the engine torque control model is trained by the interaction between the engine torque control model and the simulation environment until the preset training stop condition is met, and the trained engine torque control model is obtained.
[0025] The reward function of the engine torque control model is optimized based on multi-objective optimization of total energy consumption penalty, SOC deviation penalty and ineffective energy consumption penalty; The weight of the total energy consumption penalty is greater than the weight of the SOC deviation penalty; the weight of the SOC deviation penalty is greater than the weight of the ineffective energy consumption penalty.
[0026] The engine torque control model performs deep reinforcement learning of energy management strategies in a simulation environment. This stage is the core of training continuous energy management strategies. By remodeling the energy management problem of hybrid electric vehicles into a Markov decision process (MDP), optimizing the definition of state space and action space, and realizing interactive learning between strategy and environment in the simulation environment, the system collects state, action, and reward information and iteratively updates the energy management strategy, laying the foundation for subsequent lightweighting and deployment.
[0027] The core elements of the Markov Decision Process (MDP) are adjusted and defined as follows.
[0028] The state space S of the engine torque control model is used for multi-dimensional vehicle dynamics and energy status perception. Based on the actual needs of hybrid electric vehicle energy management, the relevant dimensions are determined as "speed, acceleration, real-time SOC, target SOC, clutch (open / closed) status, vehicle required torque, and engine real-time speed", which comprehensively covers vehicle driving dynamics, power battery status, power coupling status, and power demand.
[0029] Here, speed refers to vehicle speed, acceleration refers to vehicle acceleration, real-time SOC refers to the real-time charge of the power battery, target SOC refers to the target remaining charge of the power battery, and vehicle required torque refers to the target torque of the engine.
[0030] The state space S is defined as follows: S={v,a,SOC_real,SOC_target ,Clutch,P_req ,n_eng} The meanings and technical functions of each parameter are as follows: v: Vehicle speed (km / h), representing the basic operating conditions of vehicle operation (such as low-speed congestion, high-speed cruising). a: Vehicle acceleration (m / s²) reflects the dynamic changes in driving conditions (such as rapid acceleration, deceleration and coasting), and provides a basis for transient energy distribution; SOC_real: Real-time remaining power capacity (%) of the power battery, which is directly related to the priority of power utilization (prioritize power use when SOC is high, and conserve power when SOC is low). SOC_target: Target remaining battery charge (%), issued by the vehicle control strategy (e.g., setting a high target SOC for long-distance driving and a low target SOC for short-distance driving) to prevent the SOC from deviating from the reasonable range; Clutch: Clutch engagement / disengagement state (0 = disengaged, 1 = engaged), reflecting the coupling relationship between the engine and the drive system (when disengaged, the engine does not participate in driving; when engaged, the engine provides torque), ensuring that the strategy adapts to the power coupling mode; P_wheel,req: Wheel end power demand (kW), calculated by the vehicle controller (VCU) based on the driver's pedal signal, vehicle speed and driving resistance, directly reflects the actual power demand of the vehicle and adapts to the power output logic under different transmission ratios; n_eng: Real-time engine speed (rpm), associated with engine fuel consumption characteristics (the engine's optimal torque varies at different speeds), ensuring the engine operates in its high-efficiency range.
[0031] Compared to traditional simplified state spaces or the use of "vehicle demand torque", "wheel-end demand power" is more directly related to the vehicle's energy consumption (power × time = energy), which can avoid torque conversion errors caused by changes in transmission ratio (such as different power demands corresponding to the same wheel-end torque in different gears); at the same time, it can more accurately match the power coupling logic of "engine + motor" in hybrid systems, and reduce energy distribution deviations when multiple power sources work together (such as when the wheel-end power demand increases sharply under climbing conditions, the strategy can quickly determine the power distribution ratio between the engine and the motor).
[0032] Action space A is for precise engine torque control. The action space is defined as the engine target torque (a continuous variable), which directly corresponds to the engine's output control command, and is defined as follows: A = T_eng, target; T_eng,target ∈[0,T_eng,max]; T_eng,target represents the engine's target torque; T_eng,max represents the engine's maximum output torque, determined based on the engine's hardware parameters, and indicates the engine's real-time target output torque for strategy decisions.
[0033] "Engine torque" aligns with the underlying logic of engine control (engine controllers typically use torque as their core control objective), and can be calculated using the relationship between wheel-end power demand, gear ratio, and vehicle speed (P_wheel,req=T_wheel×ω_wheel, where P_wheel is the wheel-end power demand, T_wheel...). (where ω_wheel is the wheel angular velocity), the torque that the engine needs to output is derived in reverse to avoid accumulated errors in the power-torque conversion process; at the same time, the continuous motion space is adapted to the smoothness of torque output to avoid power shocks and energy consumption fluctuations caused by discrete motions.
[0034] The reward function R is guided by a multi-objective principle with the goal of minimizing total energy consumption.
[0035] The core of the reward function design is to guide the strategy to achieve multi-objective optimization of "minimum total energy consumption, stable SOC, and no ineffective energy consumption". The specific logic is divided into three parts according to weight priority, and the weights can be adjusted through offline simulation.
[0036] Total energy consumption penalty (highest weight, e.g., 70%): A comprehensive penalty is imposed on the fuel consumption and electricity consumption of hybrid vehicles. Fuel consumption is calculated by combining the engine's real-time speed, target torque, and fuel consumption rate graph. Electricity consumption is calculated by combining the difference between the power demanded at the wheel ends and the engine's output power (i.e., the power that the motor needs to supplement or recover) with the working time. After converting fuel consumption to oil price and electricity consumption to electricity price into cost units, the sum of the two is taken as the total energy consumption index. This index is penalized (the higher the total energy consumption, the lower the reward value), prioritizing the optimization of total energy consumption.
[0037] SOC deviation penalty (secondary weight, e.g., 20%): Calculate the absolute value of the difference between the real-time remaining power capacity (SOC) of the power battery and the target remaining power capacity, and penalize this difference (the larger the difference, the lower the reward value) to avoid the SOC deviating too much from the target range (e.g., excessive discharge at high SOC leads to no power available later, and excessive power preservation at low SOC leads to frequent high-load operation of the engine).
[0038] Ineffective energy consumption penalty (lowest weight, e.g., 10%): When the clutch is disengaged (the engine is not coupled to the drive system and does not participate in power output) but the engine is still outputting the target torque, an additional penalty (reduced reward value) is triggered to avoid ineffective fuel consumption caused by engine idling.
[0039] The specific process for training the engine torque control model is as follows: Simulation environment setup: A hybrid electric vehicle simulation model is built based on MATLAB / Simulink, integrating sub-modules for engine, power battery, motor, clutch, transmission system, and vehicle driving resistance; typical operating condition data (such as NEDC and WLTC conditions) are input, and the wheel-end power demand is calculated through "speed-acceleration-driving resistance" to simulate the dynamic changes of "wheel-end power-power source coordination", and the data required for training, such as fuel consumption, SOC change, and engine speed, are output; MDP initialization: Load simulation environment parameters, initialize the deep reinforcement learning agent, and set the experience replay pool capacity, training batch size, and learning rate; Interactive Training: The agent reads state space parameters in real time in the simulation environment (wheel-end power demand, vehicle speed, vehicle acceleration, real-time power battery charge, target remaining power battery charge, clutch engagement / disengagement status, and real-time engine speed), calculates the power demand corresponding to the wheel-end power based on the transmission ratio, and outputs the target engine torque (action space); the environment calculates the real-time reward value according to the above reward function and feeds back the state at the next moment, storing the "current state - action - reward - next state" sample into the experience replay pool, and randomly selecting samples to update the agent network parameters; Convergence verification: When the total energy consumption fluctuation over 1000 consecutive training steps is less than 3%, the SOC deviation (the absolute value of the difference between real-time SOC and target SOC) is less than 5%, and the wheel-end power satisfaction rate (actual wheel-end power / required power) is greater than 95%, the training converges, and the trained energy management strategy network (including state-action mapping parameters) is saved.
[0040] Once the engine torque control model is trained, the controller deployment phase begins.
[0041] In some embodiments, the engine torque control model is a lightweight model that has undergone precision quantization.
[0042] In other words, during the controller deployment phase, the lightweight quantization model is deployed to the vehicle controller to enable the real-time online operation of the energy management strategy and ensure the feasibility of the strategy implementation.
[0043] In some embodiments, please refer to Figure 3 , Figure 3 A flowchart of the method for precision quantization of the model according to an embodiment of this application is shown; before deploying the engine torque control model to the vehicle controller, the method further includes the following steps S301-S302: S301. Using a calibration dataset including predetermined operating conditions, the accuracy of the trained engine torque control model is quantified to generate a lightweight engine torque control model; S302. The performance indicators of the lightweight engine torque control model are verified in a simulation environment, including total energy consumption, SOC control, and wheel-end power satisfaction rate.
[0044] In other words, before deploying the engine torque control model to the vehicle controller, the engine torque control model needs to be quantized for accuracy. First, model quantization is performed on the trained policy network. The quantization scaling factor and zero-point offset are determined by calibrating the dataset to achieve model lightweighting. Then, accuracy verification is performed. The total energy consumption, SOC control, wheel-end power satisfaction rate, and other indicators of the quantized model are verified in a simulation environment to ensure that they meet the requirements of real vehicle applications.
[0045] After completing the lightweighting process of the model, vehicle deployment and testing were carried out.
[0046] First, model integration is performed by transforming the engine torque control model S-Function and integrating it into the energy management strategy model to replace the traditional rule-based torque control. In other words, the trained and quantized reinforcement learning model is encapsulated and embedded into the software framework of the vehicle energy management strategy through interface technologies such as SFunction, thereby replacing the traditional rule-based control logic at the algorithm level.
[0047] Then, controller code generation and adaptation are performed. In some embodiments, please refer to... Figure 4 , Figure 4 A flowchart illustrating the method for deploying the engine torque control model to the vehicle controller is shown; as follows: Figure 4 As shown, the engine torque control model is deployed to the vehicle controller through the following steps S401-S403: S401. Convert the engine torque control model into embedded code supported by the vehicle controller of the target vehicle; S402. Configure the signal interface parameters in the embedded code according to the hardware I / O definition of the vehicle controller of the target vehicle. S403. After configuring the signal interface parameters, use the target flashing tool of the vehicle controller to securely transfer the binary file compiled based on the embedded code to the program memory of the vehicle controller.
[0048] Here, code generation and adaptation are performed: the quantization model is converted into C code executable by the vehicle controller, and the hardware interface and communication protocol of the VCU (vehicle controller) are adapted. This mainly includes model to code conversion, interface parameter configuration, and firmware flashing.
[0049] Specifically, the model-to-code process is as follows: using MATLAB / Simulink, the quantized energy management strategy model is converted into C language code executable by the vehicle controller; during the code generation process, code optimization options are enabled, while retaining the conversion interface between wheel-end power demand and engine torque, adapting to the computing power characteristics of the vehicle controller architecture.
[0050] The interface parameter configuration process is as follows: Based on the hardware I / O definition of the vehicle controller, set the signal interface parameters in the code—input interface and output interface—to ensure signal interaction compatibility; The firmware flashing process is as follows: Using the vehicle controller target flashing tool, the compiled binary file is securely transferred to the VCU's program memory via the OBD interface or CAN bus; a verification mechanism is enabled during the flashing process to prevent firmware corruption from causing controller failure.
[0051] OBD (On-Board Diagnostics) refers to an onboard automatic diagnostic system for vehicles.
[0052] Please refer to Figure 5 In some embodiments, after deploying the engine torque control model to the vehicle controller, the method further includes the following steps S501-S502: S501. After the vehicle controller is powered on, the vehicle controller's ability to receive the operating status information and send the target torque command is verified by monitoring the vehicle bus messages, in order to perform communication verification. S502. Perform functional testing by driving the target vehicle under various test conditions on a real vehicle or a rotary drum test bench and collecting operating data; the operating data is used to determine the control effect evaluation result of the engine torque control model. The various test conditions include low-speed congestion conditions, high-speed cruising conditions, and rapid acceleration or hill climbing conditions.
[0053] This section is used for communication establishment and functional testing.
[0054] Communication link verification: After the VCU is powered on, the CAN bus messages are monitored by the host computer to confirm that the state space parameter messages can be received normally and the engine torque command messages can be sent normally.
[0055] Real-vehicle functional testing: Vehicles equipped with the deployed VCU will be placed on actual roads or a rotary test bench to test the effectiveness of the strategy under different operating conditions, such as: Low-speed congested conditions: Verify SOC maintenance capability and avoid increased energy consumption caused by frequent engine start-stop; High-speed cruise condition: Verify the engine torque control accuracy and the engine's high-efficiency range operating rate; Rapid acceleration / climbing conditions: Verify the real-time response of the strategy and the synergistic effect of multiple power sources.
[0056] If problems are found during testing, the controller logs are read through the host computer to locate the cause of the problem. After correction, the software is updated and the functions are tested again until the requirements of the actual vehicle application are met.
[0057] In other words, if the control effect evaluation result of the engine torque control model does not meet the preset compliance conditions, the engine torque control model is corrected based on the operating data until the control effect evaluation result meets the preset compliance conditions.
[0058] The control effect evaluation results are a set or more sets of quantitative index results used to quantitatively evaluate the actual performance of the reinforcement learning-based engine torque control system. For example, they include: energy consumption economy indexes, such as the equivalent fuel consumption per 100 kilometers under a specific test cycle (such as WLTC); power battery power management indexes, such as the absolute value of the deviation between the actual remaining power battery power and the target remaining power at the end of the test; torque tracking performance indexes, such as the tracking error (such as root mean square error RMSE) or response delay time of the actual output torque of the engine to the target torque; and drivability indexes, such as the wheel-end power demand satisfaction rate or power response time under a specific rapid acceleration condition.
[0059] The preset compliance conditions are technical indicator thresholds or ranges that match the quantitative indicator results in the control effect evaluation results.
[0060] The operational data refers to the time-series data actually collected and recorded by the vehicle bus network, sensors, and controllers during functional testing on a real vehicle or drum test bench. It is usually CAN bus message decoding data with timestamps or data logs recorded internally by the controller. By analyzing this data, problem scenarios can be reproduced and the specific reasons for improper model decisions can be located.
[0061] The engine torque control model is modified, specifically, based on the substandard "control effect evaluation results" and their corresponding "operational data", the parameters of the "engine torque control model" are adjusted to improve its performance through a retraining process.
[0062] Specifically, state-action-reward sequences can be extracted from "running data" to construct a new training dataset; usually, poorly performing work condition data are weighted or the sampling frequency is increased to enhance the model's learning under such work conditions.
[0063] Using the currently deployed model (i.e. the model whose performance has not met the target) as a pre-trained model, continue training iterations on the new dataset.
[0064] The reward function can also be adjusted. If the analysis finds that the multi-objective balance is improper, the weights of the multi-dimensional reward function defined in the claims (such as the penalty weights for total energy consumption, SOC bias, and ineffective energy consumption) can be adjusted in a targeted manner, and then training can be carried out based on the new reward function.
[0065] The revised model is then quantified, code generated, and deployed to the VCU, followed by a new round of functional testing and evaluation, forming a closed-loop engineering iteration of "testing-evaluation-correction-redeployment" until all "preset compliance conditions" are met.
[0066] In step S102, the real-time operating status information of the target vehicle is input into the engine torque control model, and the engine torque control model processes the real-time operating status information of the target vehicle to determine the real-time target torque command of the target vehicle.
[0067] The onboard sensor network and controller collect and aggregate all parameters that constitute the state space S in real time, forming the state vector S at the current moment. t This data is then input into the engine torque control model already deployed in the VCU.
[0068] After receiving the state vector, the engine torque control model (i.e., the quantized reinforcement learning policy network) executes the forward computation logic stored internally. This computation process performs nonlinear transformation and mapping on the input state based on the optimal parameters learned by the model during the training phase; for any input state, its output action can maximize the long-term cumulative reward (i.e., achieve a balance between multiple objectives such as minimizing total energy consumption).
[0069] The result of the model calculation is a continuous scalar value, which is directly defined as the engine target torque command at the current moment and sent to the engine controller through the VCU output interface.
[0070] In step S103, the engine actuator is controlled to operate according to the target torque command.
[0071] The target torque command is generated by the reinforcement learning decision module in the vehicle control unit (VCU) and sent to the engine controller in a standardized message format via the vehicle's internal communication network (such as the CAN bus). After receiving the torque command, the engine controller's internal low-level control logic (such as fuel injection control, throttle control, ignition timing control, variable valve timing control, and wastegate valve control sub-modules) will work together to execute the target torque command, ultimately ensuring that the engine's actual output torque tracks the aforementioned "target torque command".
[0072] The engine torque control method based on reinforcement learning described in this application has the following beneficial effects.
[0073] Improved energy consumption optimization: The introduction of "wheel-end power demand" in the state space enables the strategy to be more directly related to the energy demand of the whole vehicle, avoiding energy distribution deviation caused by transmission ratio conversion error; combined with the continuous engine torque action space and multi-objective reward function, the engine can be precisely controlled to operate in the high-efficiency range, and the total energy consumption is reduced compared with the traditional strategy, giving full play to the energy-saving potential of hybrid vehicles.
[0074] Enhanced strategy generalization: The multi-dimensional state space based on "wheel-end power demand" covers all operating conditions. The strategy can adapt to different vehicle models or different driving scenarios without relying on a large amount of pre-calibrated data. Model quantization retains the original strategy's adaptability to operating conditions. The quantized model can still cope with non-calibrated scenarios in real vehicles, and its generalization is improved compared to traditional calibration strategies. Based on the same inventive concept, this application also provides a task dispatching device corresponding to the reinforcement learning-based engine torque control method. Since the principle of the device in this application is similar to the reinforcement learning-based engine torque control method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0075] Please refer to Figure 6 , Figure 6 A schematic diagram of the structure of the reinforcement learning-based engine torque control device according to an embodiment of this application is shown. The device includes: The acquisition module 601 is used to acquire real-time operating status information of the target vehicle, including: wheel-end power demand, vehicle speed, vehicle acceleration, real-time power battery charge, target remaining power battery charge, clutch engagement / disengagement status, and real-time engine speed. An engine torque control model is deployed in the onboard controller of the target vehicle. The engine torque control model is matched with the target vehicle. The engine torque control model is trained in a simulation environment based on multi-objective optimization of the target torque action space including the operating status information. The processing module 602 is used to input the real-time operating status information of the target vehicle into the engine torque control model, process the real-time operating status information of the target vehicle through the engine torque control model, and determine the real-time target torque command of the target vehicle. The control module 603 is used to control the operation of the engine actuator according to the target torque command. In some embodiments, in the reinforcement learning-based engine torque control device, the reward function of the engine torque control model is optimized based on a multi-objective approach, including total energy consumption penalty, SOC deviation penalty, and ineffective energy consumption penalty. The weight of the total energy consumption penalty is greater than the weight of the SOC deviation penalty; the weight of the SOC deviation penalty is greater than the weight of the ineffective energy consumption penalty.
[0076] In some embodiments, in the reinforcement learning-based engine torque control device, the engine torque control model is trained based on the following steps: Define the state space, action space, and reward function of the engine torque control model; the state space includes: wheel-end power demand, vehicle speed, vehicle acceleration, real-time battery charge, target remaining battery charge, clutch engagement / disengagement status, and real-time engine speed; the action space is the target engine torque; A simulation environment is set up; the simulation environment includes a hybrid electric vehicle simulation model, which includes an engine module, a power battery module, a motor module, a clutch module, a transmission system module, and a vehicle driving resistance sub-module; Based on the reward function, the engine torque control model is trained in a simulation environment through the interaction between the engine torque control model and the simulation environment until a preset training stop condition is met, thus obtaining a trained engine torque control model.
[0077] In some embodiments, the engine torque control model in the reinforcement learning-based engine torque control device is a lightweight model that has undergone precision quantization.
[0078] In some embodiments, in the reinforcement learning-based engine torque control device, before deploying the engine torque control model to the vehicle controller, the accuracy of the trained engine torque control model is quantized using a calibration dataset including predetermined operating conditions to generate a lightweight engine torque control model; the performance indicators of the lightweight engine torque control model are verified in a simulation environment, the performance indicators including total energy consumption, SOC control, and wheel-end power satisfaction rate.
[0079] In some embodiments, in the reinforcement learning-based engine torque control device, the engine torque control model is deployed to the vehicle controller through the following steps: The engine torque control model is converted into embedded code supported by the vehicle controller of the target vehicle; the signal interface parameters in the embedded code are configured according to the hardware I / O definition of the vehicle controller of the target vehicle. After configuring the signal interface parameters, the target flashing tool of the vehicle controller is used to securely transfer the binary file compiled based on the embedded code to the program memory of the vehicle controller.
[0080] In some embodiments, in the reinforcement learning-based engine torque control device, after the engine torque control model is deployed to the vehicle controller, after the vehicle controller is powered on, the vehicle bus messages are monitored to verify the vehicle controller's ability to receive the operating status information and send the target torque command, so as to perform communication verification. Functional testing is conducted by driving the target vehicle under various test conditions on a real vehicle or a rotary drum test bench, and collecting operational data; the operational data is used to determine the control effect evaluation results of the engine torque control model. The various test conditions include low-speed congestion conditions, high-speed cruising conditions, and rapid acceleration or hill climbing conditions.
[0081] In some embodiments, in the reinforcement learning-based engine torque control device, if the control effect evaluation result of the engine torque control model does not meet the preset compliance conditions, the engine torque control model is corrected based on the operating data until the control effect evaluation result meets the preset compliance conditions.
[0082] Based on the same inventive concept, this application also provides an on-board controller corresponding to the reinforcement learning-based engine torque control method. Since the principle of the on-board controller in this application is similar to the reinforcement learning-based engine torque control method described above, the implementation of the on-board controller can refer to the implementation of the method, and the repeated parts will not be described again.
[0083] This application also provides an on-board controller that executes the steps of the reinforcement learning-based engine torque control method.
[0084] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.
[0085] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0086] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0087] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a platform server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0088] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An engine torque control method based on reinforcement learning, characterized in that, The method includes the following steps: The system acquires real-time operating status information of the target vehicle, including: wheel-end power demand, vehicle speed, vehicle acceleration, real-time battery charge, target remaining battery charge, clutch engagement / disengagement status, and real-time engine speed. An engine torque control model is deployed in the target vehicle's onboard controller. This engine torque control model is matched to the target vehicle. The engine torque control model is trained in a simulation environment based on multi-objective optimization of the target torque action space, which includes the operating status information. The system inputs the real-time operating status information of the target vehicle into the engine torque control model, processes this information, and determines the target torque command for the target vehicle in real-time. The engine actuator is controlled to operate according to the target torque command.
2. The engine torque control method based on reinforcement learning according to claim 1, characterized in that, The reward function of the engine torque control model is optimized based on multi-objective optimization of total energy consumption penalty, SOC deviation penalty and ineffective energy consumption penalty; The weight of the total energy consumption penalty is greater than the weight of the SOC deviation penalty; The weight of the SOC deviation penalty is greater than the weight of the ineffective energy consumption penalty.
3. The engine torque control method based on reinforcement learning according to claim 2, characterized in that, The engine torque control model is trained based on the following steps: Define the state space, action space, and reward function of the engine torque control model; the state space includes: wheel-end power demand, vehicle speed, vehicle acceleration, real-time battery charge, target remaining battery charge, clutch engagement / disengagement status, and real-time engine speed; the action space is the target engine torque; A simulation environment is set up; the simulation environment includes a hybrid electric vehicle simulation model, which includes an engine module, a power battery module, a motor module, a clutch module, a transmission system module, and a vehicle driving resistance sub-module; Based on the reward function, the engine torque control model is trained in a simulation environment through the interaction between the engine torque control model and the simulation environment until a preset training stop condition is met, thus obtaining a trained engine torque control model.
4. The engine torque control method based on reinforcement learning according to claim 1, characterized in that, The engine torque control model is a lightweight model that has undergone precision quantization.
5. The engine torque control method based on reinforcement learning according to claim 4, characterized in that, Before deploying the engine torque control model to the on-board controller, the method further includes the following steps: Using a calibration dataset that includes predetermined operating conditions, the accuracy of the trained engine torque control model is quantized to generate a lightweight engine torque control model. The performance indicators of the lightweight engine torque control model are then verified in a simulation environment, including total energy consumption, state of charge (SOC) control, and wheel-end power satisfaction rate.
6. The engine torque control method based on reinforcement learning according to claim 1 or 4, characterized in that, The engine torque control model is deployed to the on-board controller through the following steps: The engine torque control model is converted into embedded code supported by the vehicle controller of the target vehicle; the signal interface parameters in the embedded code are configured according to the hardware I / O definition of the vehicle controller of the target vehicle. After configuring the signal interface parameters, the target flashing tool of the vehicle controller is used to securely transfer the binary file compiled based on the embedded code to the program memory of the vehicle controller.
7. The engine torque control method based on reinforcement learning according to claim 1 or 4, characterized in that, After deploying the engine torque control model to the on-board controller, the method further includes the following steps: After the vehicle controller is powered on, the vehicle bus messages are monitored to verify the vehicle controller's ability to receive the operating status information and send the target torque command, in order to perform communication verification. Functional testing is conducted by driving the target vehicle under various test conditions on a real vehicle or a rotary drum test bench, and collecting operational data; the operational data is used to determine the control effect evaluation results of the engine torque control model. The various test conditions include low-speed congestion conditions, high-speed cruising conditions, and rapid acceleration or hill climbing conditions.
8. The engine torque control method based on reinforcement learning according to claim 6, characterized in that, The method further includes: If the control effect evaluation result of the engine torque control model does not meet the preset compliance conditions, the engine torque control model is corrected based on the operating data until the control effect evaluation result meets the preset compliance conditions.
9. An engine torque control device based on reinforcement learning, characterized in that, The device includes: An acquisition module is used to acquire real-time operating status information of the target vehicle, including: wheel-end power demand, vehicle speed, vehicle acceleration, real-time battery charge, target remaining battery charge, clutch engagement / disengagement status, and real-time engine speed. An engine torque control model is deployed in the target vehicle's onboard controller. The engine torque control model is matched with the target vehicle. The engine torque control model is trained in a simulation environment based on multi-objective optimization of the target torque action space including the operating status information. A processing module is used to input the real-time operating status information of the target vehicle into the engine torque control model, process the real-time operating status information of the target vehicle through the engine torque control model, and determine the real-time target torque command of the target vehicle. The control module is used to control the operation of the engine actuator according to the target torque command.
10. A vehicle-mounted controller, characterized in that, The on-board controller performs the steps of the engine torque control method based on reinforcement learning as described in any one of claims 1 to 8.