Energy management method and device based on reinforcement learning, vehicle and storage medium
Through the energy management method based on reinforcement learning, the lightweight model trained by deep learning is used to optimize the engine power demand, which solves the problems of hybrid vehicles in energy consumption optimization and generalization, and achieves more efficient energy saving and adaptability.
Patent Information
- Application Number
- CN202510805806.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-08-19
AI Technical Summary
The energy management strategy of hybrid vehicles is poor in energy consumption optimization and generalization, and it is difficult to adapt to complex driving conditions.
Using a reinforcement learning-based energy management method, by obtaining the current speed, acceleration and power battery state of the vehicle, the lightweight energy management model trained by deep learning is used to predict the engine power demand, and the engine operation is controlled to optimize energy distribution.
It improves the energy-saving effect and applicability of hybrid vehicles, can better adapt to complex driving conditions, and fully utilize energy-saving potential.
Smart Images

Figure CN120503768A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle technology, and in particular to an energy management method, device, vehicle, and storage medium based on reinforcement learning. Background Art
[0002] With the intensifying global energy crisis and rising environmental awareness, hybrid vehicles (HEVs), a mode of transportation that combines the advantages of both fuel-powered engines and electric motors, have garnered widespread attention. By rationally distributing power output between the engine and electric motor, HEVs can improve fuel economy while reducing exhaust emissions. However, optimizing HEV energy management strategies to achieve optimal energy efficiency has long been a key research topic in this field.
[0003] In related technologies, hybrid electric vehicle energy management strategies are primarily based on pre-set rules. For example, energy management rules are designed based on the vehicle's operating conditions and driving patterns, and the power distribution between the engine and electric motor is controlled using the pre-set rules.
[0004] However, related technologies suffer from poor energy optimization and generalizability. Rule-based approaches, while simple and easy to implement, struggle to adapt to complex driving conditions and have limited optimization effectiveness. Summary of the Invention
[0005] The present application provides an energy management method, device, vehicle and storage medium based on reinforcement learning to solve the problem that energy management technology has poor energy consumption optimization effect and generalization, and is difficult to adapt to complex driving conditions. The present application can fully tap the energy-saving potential of hybrid vehicles, improve energy-saving effects and has strong applicability.
[0006] The first embodiment of the present application provides an energy management method based on reinforcement learning, comprising the following steps: Obtain the current speed, current acceleration, current power battery status, and target power battery status of the target vehicle; Inputting the current speed, the current acceleration, the current power battery state, and the target power battery state into a preset lightweight energy management model to obtain the engine required power, wherein the preset lightweight energy management model is trained by a deep learning training framework; The engine is controlled to operate based on the engine power demand.
[0007] Optionally, in some embodiments, before inputting the current speed, the current acceleration, the current power battery state, and the target power battery state into the preset lightweight energy management model, the method further includes: Build simulation environment and initial energy management model; The initial energy management model is trained based on the simulation environment to obtain a target energy management model. Based on a preset loss function and according to the target energy management model, a preset small-scale neural network is trained to generate a preset lightweight energy management model.
[0008] Optionally, in some embodiments, the training the initial energy management model based on the simulation environment includes: Construct a deep learning training framework, wherein the training framework includes: state space, action space and reward function, The state space is: ; in, is the state space, is the vehicle acceleration, is the vehicle speed, ; The action space is: engine required power ; The reward function is:
[0009] in, is the first reward weight, is the second reward weight, is the third reward weight, For fuel prices, For fuel consumption, For electricity price, For power consumption, The real-time state of charge. is the target state of charge, It is the penalty item for low SOC.
[0010] Optionally, in some embodiments, the preset loss function is: ; in, is the preset loss function, Networking for Teachers in State Next, according to the parameters The output, Networking for Students in State Next, according to the parameters Output.
[0011] Optionally, in some embodiments, before deploying the lightweight energy management model to the target vehicle, the process includes: Deploy the lightweight energy management model to a test device, perform a functional test on the test device to obtain a test result, and determine whether the test result meets a preset application condition; If the test result meets the preset application conditions, the lightweight energy management model is applied to the target vehicle.
[0012] A second embodiment of the present application provides an energy management device based on reinforcement learning, including: An acquisition module is used to obtain the current speed, current acceleration, current power battery status and target power battery status of the target vehicle; a generation module, configured to input the current speed, the current acceleration, the current power battery state, and the target power battery state into a preset lightweight energy management model to obtain a required engine power, wherein the preset lightweight energy management model is trained using a deep learning training framework; A control module is configured to control the engine to operate based on the engine power demand.
[0013] Optionally, in some embodiments, before inputting the current speed, the current acceleration, the current power battery state, and the target power battery state into the preset lightweight energy management model, the generating module includes: A construction unit, used for constructing a simulation environment and an initial energy management model; A training unit is used to train the initial energy management model based on the simulation environment to obtain a target energy management model, and based on a preset loss function, train a preset small-scale neural network according to the target energy management model to generate a preset lightweight energy management model.
[0014] Optionally, in some embodiments, the training unit is specifically configured to: Construct a deep learning training framework, wherein the training framework includes: state space, action space and reward function, The state space is: ; in, is the state space, is the vehicle acceleration, is the vehicle speed, ; The action space is: engine required power ; The reward function is:
[0015] in, is the first reward weight, is the second reward weight, is the third reward weight, For fuel prices, For fuel consumption, For electricity price, For power consumption, The real-time state of charge. is the target state of charge, It is the penalty item for low SOC.
[0016] Optionally, in some embodiments, the preset loss function is: ; in, is the preset loss function, Networking for Teachers in State Next, according to the parameters The output, Networking for Students in State Next, according to the parameters Output.
[0017] Optionally, in some embodiments, after generating the preset lightweight energy management model, the training unit includes: A testing subunit is used to deploy the lightweight energy management model to a test device, perform a functional test on the test device to obtain a test result, and determine whether the test result meets a preset application condition; A deployment subunit is used to deploy the lightweight energy management model to the target vehicle when the test result meets the preset application conditions.
[0018] A third aspect of the present application provides a vehicle, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the energy management method based on reinforcement learning as described in the above embodiment.
[0019] The fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the energy management method based on reinforcement learning as described in the above embodiment.
[0020] Thus, by obtaining the target vehicle's current speed, current acceleration, current battery state, and target battery state, and inputting these into a preset lightweight energy management model, the engine power demand is determined. The preset lightweight energy management model is trained using a deep learning training framework and controls the engine's operation based on the engine power demand. This solves the problem of energy management technologies having poor energy consumption optimization and generalizability, making them difficult to adapt to complex driving conditions. This application can fully tap the energy-saving potential of hybrid electric vehicles, improve energy-saving effects, and has strong applicability.
[0021] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart of an energy management method based on reinforcement learning according to an embodiment of the present application; Figure 2 A schematic diagram of the principle of energy management model training in a simulation environment according to one embodiment of the present application; Figure 3 A schematic diagram of the principle of an energy management method based on deep reinforcement learning according to one embodiment of the present application; Figure 4 Schematic diagram of a block diagram of an energy management device based on reinforcement learning according to an embodiment of the present application; Figure 5 A schematic structural diagram of a vehicle provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0024] The following describes the energy management method, device, vehicle and storage medium based on reinforcement learning of the embodiment of the present application with reference to the accompanying drawings. In response to the problem that the energy management technology mentioned in the above background technology has poor energy consumption optimization effect and generalization, and is difficult to adapt to complex driving conditions, the present application provides an energy management method based on reinforcement learning. In this method, by obtaining the current speed, current acceleration, current power battery state and target power battery state of the target vehicle, and inputting the current speed, current acceleration, current power battery state and target power battery state into a preset lightweight energy management model, the engine demand power is obtained, wherein the preset lightweight energy management model is trained by a deep learning training framework, and controls the engine to operate based on the engine demand power. Thus, the problem that the energy management technology has poor energy consumption optimization effect and generalization, and is difficult to adapt to complex driving conditions is solved. The present application can give full play to the energy-saving potential of hybrid electric vehicles, improve energy-saving effects and has strong applicability.
[0025] Specifically, Figure 1 A flowchart of an energy management method based on reinforcement learning provided in an embodiment of the present application.
[0026] like Figure 1 As shown, the energy management method based on reinforcement learning includes the following steps: In step S101 , the current speed, current acceleration, current power battery state, and target power battery state of the target vehicle are acquired.
[0027] The target power battery state may be preset by a user, obtained through a limited number of experiments, or obtained through a limited number of computer simulations, which is not specifically limited here.
[0028] In step S102, the current speed, current acceleration, current power battery state and target power battery state are input into a preset lightweight energy management model to obtain the required power of the engine, wherein the preset lightweight energy management model is trained by a deep learning training framework.
[0029] Specifically, the embodiment of the present application predicts the engine's required power based on the current vehicle state (speed, acceleration, power battery state) and the target power battery state through a preset lightweight energy management model, wherein the preset lightweight energy management model predicts the engine's required power based on the input vehicle state and the target power battery state. This power value is used to control the engine's output to meet the vehicle's power requirements, while optimizing the use of the power battery to ensure that the vehicle maintains efficient operation while reaching the target power battery state.
[0030] Optionally, in some embodiments, before the current speed, current acceleration, power battery status and target power battery status are input, it includes: constructing a simulation environment and an initial energy management model; training the initial energy management model based on the simulation environment to obtain a target energy management model, based on a preset loss function, according to the target energy management model, training a preset small-scale neural network to generate a preset lightweight energy management model.
[0031] It should be noted that, in view of the poor energy consumption optimization effect and generalization of energy management technology, the embodiment of the present application generates a data-driven intelligent energy management strategy based on four-dimensional typical working conditions and hybrid vehicle energy consumption models, and uses deep reinforcement learning technology to fully tap the energy-saving potential of hybrid vehicles and optimize energy-saving effects. In addition, in order to address the difficulties in the actual vehicle engineering application of energy management technology based on deep reinforcement learning, knowledge distillation technology is used to simplify the neural network and realize its application in the vehicle controller.
[0032] like Figure 2 As shown, the embodiment of the present application is based on the energy management technology of deep reinforcement learning. It uses deep reinforcement learning technology to model the energy management problem as a Markov decision process, define the state space, action space and reward function, perform deep reinforcement learning of the energy management strategy in a simulation environment, and collect and update the strategy. In addition, the embodiment of the present application draws on the knowledge distillation model compression method to use the deterministic samples generated by the original policy network to train a small-scale neural network to achieve policy extraction.
[0033] Specifically, the energy management method principle based on deep reinforcement learning in the embodiment of the present application is as follows: Figure 3 As shown, the implementation of the embodiment of the present application is divided into three stages.
[0034] Phase 1: Deep reinforcement learning of energy management strategies in a simulation environment. This phase is the training phase of the continuous energy management strategy, which collects information such as state, action, and reward generated by the strategy-environment interaction and updates the energy management strategy.
[0035] Specifically, the energy management problem of hybrid vehicles is modeled as a Markov decision process, and its state space, action space, and reward function are preliminarily defined as follows: The state space is: ; in, is the state space, is the vehicle acceleration, is the vehicle speed, ; The action space is: Engine required power ; The reward function is: Hybrid electric vehicle energy management is a long-term sequential decision-making process with the goal of minimizing total energy consumption. For plug-in hybrid electric vehicles with large-capacity batteries, the reward function design should consider both fuel consumption and electricity consumption. In addition, the optimal energy allocation trajectory based on DP developed in Research Content 2 can be used as a reference. Trajectory , to guide the training of reinforcement learning. The reward function is set as follows:
[0036] in, is the first reward weight, is the second reward weight, is the third reward weight, For fuel prices, For fuel consumption, For electricity price, For power consumption, The real-time state of charge. is the target state of charge, It is the penalty item for low SOC.
[0037] in, and It can be pre-set by the user, obtained through a limited number of experiments, or obtained through a limited number of computer simulations, and is not specifically limited here.
[0038] Phase 2: Downloading the energy management strategy: After the strategy training is completed, the structure and parameters of the strategy network are downloaded and saved as the learned energy management strategy.
[0039] To meet the operational cycle requirements of each functional module of the vehicle controller, strategy extraction is required. This process consists of two steps: extracting and migrating strategy knowledge; and preserving the strategy parameters and network structure after migration.
[0040] In this embodiment, we draw on the knowledge distillation model compression method to achieve energy management policy extraction and migration. However, unlike knowledge distillation, because energy management uses a deterministic strategy, we directly use the deterministic samples generated by the original policy network as hard-labeled data to train a small-scale network to complete policy extraction.
[0041] First, the policy network obtained in the first stage serves as the "teacher" network, and a small neural network is constructed as the "student" network. The input layer, output layer, number of hidden layers, and activation function of Net-Student are the same as those of Net-Teacher, but the number of neurons in each hidden layer is reduced, meaning that Net-Student has fewer parameters. Next, sample policy decision data generated by Net-Teacher is collected and used to train Net-Student with the goal of minimizing the loss function that represents the difference in the outputs of the two policy networks, as shown in the following equation.
[0042] Among them, the preset loss function is: ; in, is the preset loss function, Networking for Teachers in State Next, according to the parameters The output, Networking for Students in State Next, according to the parameters Output.
[0043] Phase 3: Online application of energy management strategies. The downloaded energy management strategies only need to map states to actions in real time to achieve online application of the strategies.
[0044] Optionally, in some embodiments, before deploying the lightweight energy management model to the target vehicle, the process includes: deploying the lightweight energy management model to a test device, performing a functional test on the test device to obtain a test result, and determining whether the test result meets a preset application condition; if the test result meets the preset application condition, deploying the lightweight energy management model to the target vehicle.
[0045] The preset application conditions may be preset by the user and are not specifically limited here.
[0046] Specifically, after completing the parameterization extraction of the strategy, it is further applied to the hardware-in-the-loop platform. The first step is to deploy the controller code. (1) Use Simulink's code generation tools to convert the model into C / C++ code that can be executed by the controller. (2) Compile the generated code and perform necessary optimizations to adapt it to the characteristics of the target hardware platform.
[0047] The second step is to update the VCU software. (1) According to the I / O configuration of the target hardware, set the corresponding interface parameters in the code to ensure correct interaction between the software and the hardware. (2) Select the CAN communication protocol for data exchange between the controller and the host computer or other ECU (electronic control unit). (3) Securely transfer the compiled executable file to the VCU via the selected communication protocol. This usually requires a specialized flashing tool or software. (4) Use the flashing tool to download the new firmware to the VCU, updating its program memory.
[0048] The third step is to establish and test the communication. (1) Confirm that the communication link between the VCU and the host computer is normal. This can be verified by sending a simple diagnostic request. (2) Perform functional testing to ensure that the new energy management strategy works as expected and conduct problem analysis.
[0049] Therefore, hardware-in-the-loop testing of the strategy is realized through the hardware-in-the-loop testing platform to meet the conditions for the actual vehicle deployment of deep reinforcement learning intelligent energy management strategy and challenge the actual vehicle deployment of deep reinforcement learning intelligent agent.
[0050] In step S103 , the engine is controlled to operate based on the engine required power.
[0051] Specifically, the embodiment of the present application generates engine power demand operation based on a preset lightweight energy management model, so that the vehicle optimizes energy management while meeting power requirements.
[0052] The embodiment of the present application conducts deep reinforcement learning of energy management strategies in a simulation environment. This stage is the training stage of the continuous energy management strategy, which collects information such as states, actions, rewards, etc. generated by the interaction between the strategy and the environment, and updates the energy management strategy. Afterwards, the energy management strategy is downloaded. After the strategy training is completed, the structure and parameters of the strategy network are downloaded and saved as the learned energy management strategy. Finally, the energy management strategy is applied online. The downloaded energy management strategy only needs to map the state to the action in real time to realize the online application of the strategy. After completing the parameter extraction of the strategy, it is further applied to the hardware-in-the-loop platform.
[0053] According to the reinforcement learning-based energy management method proposed in the embodiments of this application, the engine power requirement is obtained by obtaining the current speed, current acceleration, current battery state, and target battery state of the target vehicle and inputting these into a preset lightweight energy management model. The preset lightweight energy management model is trained using a deep learning training framework and controls the engine to operate based on the engine power requirement. This solves the problem that energy management technologies have poor energy consumption optimization and generalization, making them difficult to adapt to complex driving conditions. This application can fully utilize the energy-saving potential of hybrid vehicles, improve energy-saving effects, and has strong applicability.
[0054] Next, the energy management device based on reinforcement learning proposed in accordance with an embodiment of the present application will be described with reference to the accompanying drawings.
[0055] Figure 4 4 is a block diagram of an energy management device based on reinforcement learning according to an embodiment of the present application.
[0056] like Figure 4 As shown, the energy management device 10 based on reinforcement learning includes: an acquisition module 100, a generation module 200 and a control module 300.
[0057] The acquisition module 100 is used to acquire the current speed, current acceleration, current power battery state and target power battery state of the target vehicle.
[0058] The generation module 200 is used to input the current speed, current acceleration, current power battery status and target power battery status into a preset lightweight energy management model to obtain the engine required power, wherein the preset lightweight energy management model is trained by a deep learning training framework.
[0059] The control module 300 is configured to control the engine to operate based on the engine's required power.
[0060] Optionally, in some embodiments, before the current speed, current acceleration, current power battery state and target power battery state are input into the preset lightweight energy management model, the generation module 200 includes: a construction unit and a training unit.
[0061] The construction unit is used to construct a simulation environment and an initial energy management model.
[0062] The training unit is used to train the initial energy management model based on the simulation environment to obtain the target energy management model. Based on the preset loss function and the target energy management model, the preset small-scale neural network is trained to generate a preset lightweight energy management model.
[0063] Optionally, in some embodiments, the training unit is specifically configured to: Build a deep learning training framework, which includes: state space, action space and reward function, The state space is: ; in, is the state space, is the vehicle acceleration, is the vehicle speed, ; The action space is: Engine required power ; The reward function is:
[0064] in, is the first reward weight, is the second reward weight, is the third reward weight, For fuel prices, For fuel consumption, For electricity price, For power consumption, The real-time state of charge. is the target state of charge, It is the penalty item for low SOC.
[0065] Optionally, in some embodiments, the preset loss function is: ; in, is the parameter of network S The loss function is For the network In state Next, the parameters are The expected value when For the network In state Next, the parameters are Expected value when .
[0066] Optionally, in some embodiments, after generating the preset lightweight energy management model, the training unit includes: a testing subunit and a deployment subunit.
[0067] Among them, the testing subunit is used to deploy the lightweight energy management model to the test equipment, perform functional testing on the test equipment to obtain test results, and determine whether the test results meet the preset application conditions.
[0068] The deployment subunit is used to deploy the lightweight energy management model to the target vehicle when the test results meet the preset application conditions.
[0069] It should be noted that the above explanation of the embodiment of the energy management method based on reinforcement learning is also applicable to the energy management device based on reinforcement learning in this embodiment, and will not be repeated here.
[0070] According to the reinforcement learning-based energy management device proposed in the embodiment of this application, the engine power demand is obtained by obtaining the current speed, current acceleration, current battery state, and target battery state of the target vehicle and inputting these into a preset lightweight energy management model. The preset lightweight energy management model is trained using a deep learning training framework and controls the engine to operate based on the engine power demand. This solves the problem that energy management technologies have poor energy consumption optimization and generalization, making them difficult to adapt to complex driving conditions. This application can fully utilize the energy-saving potential of hybrid vehicles, improve energy-saving effects, and has strong applicability.
[0071] Figure 5 A schematic diagram of the structure of a vehicle provided in an embodiment of the present application. The vehicle may include: Memory 501 , processor 502 , and computer programs stored in the memory 501 and executable on the processor 502 .
[0072] When the processor 502 executes the program, the energy management method based on reinforcement learning provided in the above embodiment is implemented.
[0073] Furthermore, the vehicle further comprises: The communication interface 503 is used for communication between the memory 501 and the processor 502 .
[0074] The memory 501 is used to store computer programs that can be run on the processor 502 .
[0075] The memory 501 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0076] If the memory 501, processor 502, and communication interface 503 are implemented independently, the communication interface 503, memory 501, and processor 502 can be connected to each other via a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0077] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.
[0078] The processor 502 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.
[0079] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned energy management method based on reinforcement learning.
[0080] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0081] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0082] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0083] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array, a field programmable gate array, etc.
[0084] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0085] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. An energy management method based on reinforcement learning, characterized in that: The following steps are involved: Obtain the current speed, current acceleration, current power battery status, and target power battery status of the target vehicle; Inputting the current speed, the current acceleration, the current power battery state, and the target power battery state into a preset lightweight energy management model to obtain the engine required power, wherein the preset lightweight energy management model is trained by a deep learning training framework; The engine is controlled to operate based on the engine power demand.
2. The method according to claim 1, characterized in that Before inputting the current speed, the current acceleration, the current power battery state, and the target power battery state into the preset lightweight energy management model, the method includes: Build simulation environment and initial energy management model; The initial energy management model is trained based on the simulation environment to obtain a target energy management model. Based on a preset loss function and according to the target energy management model, a preset small-scale neural network is trained to generate a preset lightweight energy management model.
3. The method according to claim 2, characterized in that The training of the initial energy management model based on the simulation environment includes: Construct a deep learning training framework, wherein the training framework includes: state space, action space and reward function, The state space is: ; in, is the state space, is the vehicle acceleration, is the vehicle speed, ; The action space is: engine required power ; The reward function is: in, is the first reward weight, is the second reward weight, is the third reward weight, For fuel prices, For fuel consumption, For electricity price, For power consumption, The real-time state of charge. is the target state of charge, It is the penalty item for low SOC.
4. The method according to claim 2, characterized in that The preset loss function is: ; in, is the preset loss function, Networking for Teachers in State Next, according to the parameters The output, Networking for Students in State Next, according to the parameters Output.
5. The method according to claim 4, characterized in that After generating the preset lightweight energy management model, the method includes: Deploy the lightweight energy management model to a test device, perform a functional test on the test device to obtain a test result, and determine whether the test result meets a preset application condition; If the test result meets the preset application condition, the lightweight energy management model is deployed to the target vehicle.
6. An energy management device based on reinforcement learning, characterized in that: include: An acquisition module is used to obtain the current speed, current acceleration, current power battery status and target power battery status of the target vehicle; a generation module, configured to input the current speed, the current acceleration, the current power battery state, and the target power battery state into a preset lightweight energy management model to obtain a required engine power, wherein the preset lightweight energy management model is trained using a deep learning training framework; A control module is configured to control the engine to operate based on the engine power demand.
7. The device according to claim 6, characterized in that Before inputting the current speed, the current acceleration, the current power battery state, and the target power battery state into the preset lightweight energy management model, the generating module includes: A construction unit, used for constructing a simulation environment and an initial energy management model; A training unit is used to train the initial energy management model based on the simulation environment to obtain a target energy management model, and based on a preset loss function, train a preset small-scale neural network according to the target energy management model to generate a preset lightweight energy management model.
8. The device according to claim 7, characterized in that The training unit is specifically used to: Construct a deep learning training framework, wherein the training framework includes: state space, action space and reward function, The state space is: ; in, is the state space, is the vehicle acceleration, is the vehicle speed, ; The action space is: engine required power ; The reward function is: in, is the first reward weight, is the second reward weight, is the third reward weight, For fuel prices, For fuel consumption, For electricity price, For power consumption, The real-time state of charge. is the target state of charge, It is the penalty item for low SOC.
9. A vehicle, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the energy management method based on reinforcement learning according to any one of claims 1 to 5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the energy management method based on reinforcement learning as described in any one of claims 1 to 5.