Methods and apparatus for training energy management systems in airborne power grid simulation
By training the energy management system in an airborne power grid simulation and utilizing reflection-enhanced reinforcement learning and neural networks, the complex and diverse energy management problems of airborne power grids were solved, achieving efficient and safe energy management strategy adaptability.
Patent Information
- Application Number
- CN202080077322.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-11
- Filing Date
- 2020-10-23
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2040-10-23
AI Technical Summary
The increasing complexity and diversity of airborne power grids make it difficult for existing rule-based operating strategies to meet the demands for efficient, safe, and reliable energy management, especially in complex motor vehicle systems, requiring a more intelligent energy management system training method.
By recording state variables, calculating regenerative power, generating neural network input vectors, designing reward functions, and using reflection-enhanced reinforcement learning to train the energy management system in an airborne power grid simulation, we can achieve prediction and optimization decision-making for unknown system states.
It provides initial training before vehicle delivery, enabling the energy management system to adapt to different equipment variations, implement efficient and safe energy management strategies, and meet the energy needs of complex vehicles.
Smart Images

Figure CN114667520B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and apparatus for training an energy management system in an airborne power grid simulation. Background Technology
[0002] The onboard power grid in motor vehicles has become significantly more complex due to its ever-expanding functional range and increasing number of electronic components and subsystems. This has not only significantly increased demands on vehicle comfort and safety but also placed much higher requirements on energy efficiency and weather resistance, which can only be achieved through complex electronic closed-loop and open-loop control systems, such as those in engine control units and exhaust gas treatment systems. Furthermore, new driver assistance systems are being developed for the most diverse driving conditions, ranging from electronic emergency braking assist systems to automatic parking systems and even fully automated driving.
[0003] These systems, along with their associated controllers, are also associated with high efficiency and reliability requirements for the onboard power grid. Furthermore, there are various forms of multi-voltage onboard grids, high-voltage systems for electric drives, redundant supply architectures for autonomous driving, and numerous possible equipment variations in advanced vehicles requiring complex architectures and customized onboard grid designs. The synergy between subsystems and the onboard grid becomes a complex coordination task. Consequently, the use of simple, rule-based operating strategies for power management is increasingly approaching its limits.
[0004] Machine learning is a crucial method for understanding complexity and variability because it eliminates the need for explicit descriptions of all system states and their associated rules. Instead, it generates basic models based on training data and the learning process, enabling predictions for previously unknown system states. One such method is Reflex-Augmented Reinforcement Learning, which learns operating strategies for energy management in vehicles and leverages artificial intelligence to understand complex and previously unknown system states. In this concept, decisions concerning energy management in a vehicle are made by a so-called agent based on the learned operating strategies. The so-called reflection protects and stabilizes the system in such a way that a decision concerning energy management proposed by the agent is only implemented when it is accepted by reflection. Simultaneously, the agent receives feedback in the form of a reward function, the value of which relates to the impact of the proposed decision and, if necessary, to the intervention of reflection. The reward function is used during the learning process to align the operating strategy with the desired optimization objective. The extension of reflection makes it possible to use reinforcement learning in safety-related systems.
[0005] The concept of reflection-enhanced reinforcement learning can be found in the following document:
[0006] “Reflex-augmented reinforcement learning for electrical energy management in vehicles”, by A. Heimrath, J. Froeschl, and U. Baumgarten, published in Proceedings of the 2018 International Conference on Artificial Intelligence, HRArabnia; D. de LaFuente, E.B. Kozerenko, J.A. Olivas, and F. G. Ttinetti, Eds. CSREA Press, pp. 429–430.
[0007] "Reflex-augmented reinforcement learning for operating strategies in automotive electrical energy management" by A Heimrath, J Froeschl, R Rezaei, M Lamprecht, and U.B. Aumgarten, published in Proceedings of the 2019 International Conference on Computing, Electronics & Communications Engineering (iCCECE), IEEE, pp. 62-67.
[0008] "Künstliche Intelligenz für das elektrische Energymanagement: Zukunft kybernetischer Managementsysteme" by A. Heimrath; J. Froeschl; K. Barbehoen and U. Baumgarten, published in Elektronik Automotiv, pp. 42-46, 2019.
[0009] Reference DE102017214384A1 provides information on how to determine the operating strategy configuration file for vehicle operation by transmitting travel data, and how to determine a global, georeferenced operating strategy configuration file for travel data using a central database device.
[0010] The classifier is known from document DE102016200854A1 as being constructed to assign the value of a feature vector to a category from at least two distinct categories based on the determination of random sample values and the resulting synthetic values. Summary of the Invention
[0011] The objective of this invention is to provide a method and apparatus for training an energy management system in an airborne power grid simulation.
[0012] The task is accomplished by the method and apparatus according to the invention.
[0013] The first aspect of the invention relates to a method for training an energy management system in an onboard power grid simulation, particularly in an onboard power grid simulation of a motor vehicle, the method comprising: (a) simulating a driving cycle with defined regeneration; (b) recording the state variables of the onboard power grid; and (c) using a regeneration current I... reku and battery voltage U bat The regeneration power P is calculated using the following formula. reku :P reku =U bat ·I reku (d) Generate the input vector S of the neural network N; (e) Generate the reward function; and (f) Train the neural network.
[0014] The advantage of this invention is that the energy management system can receive an initial operating strategy for standard equipment variations before vehicle delivery through initial training in an onboard power grid simulation. Starting from this functional state, the operating strategy can be adjusted according to optimization criteria to accommodate additional consumers.
[0015] It is preferable to use a WLTP driving cycle with defined regeneration for the initial training of the energy management system.
[0016] In a preferred embodiment, the regeneration current I is determined using the following steps. reku The working steps include: (a) extracting the battery current change curve I bat (a) All control points attributable to the energy management system's decisions and without external influence on the airborne power grid; (b) Smoothing the battery current variation curve I between the remaining control points. bat (c) By approximating the battery current variation curve I between the remaining control points approx To approximate the battery current variation curve I bat ; and (d) by battery current I bat and approximately the battery current I approx Calculate the regeneration current I using the following formula. reku :I reku =Ibat -I approx .
[0017] The calculations relating regenerative current to the current system behavior of the airborne power grid have an impact on the learning behavior of neural networks.
[0018] Another preferred embodiment can be implemented more simply, in which the regenerative current I... reku Directly equal to battery current I bat .
[0019] In another preferred embodiment, the input vector S of the neural network N is generated using the following steps, the steps including: (a) generating the state input vector S of the neural network N. normal (b) Using the state vector S erweitert The state input vector S of the extended neural network N normal .
[0020]
[0021] In another preferred embodiment, the state vector S erweitert The generation includes: (a) by adjusting the regeneration power P over time t reku The integral of (t) extends from the current time point t0 within the driving cycle up to the time point t0+x·t. vs To calculate the regeneration energy value E reku,x Where x is used to consider the regeneration power P with limited look-ahead. reku (t) look-ahead time t vs (a) percentage share; and (b) generating the state vector S erweitert The state vector includes at least the regeneration energy value E. reku,25% E reku,50% E reku,75% and E reku,100% .
[0022]
[0023] In another preferred embodiment, the state vector S erweitert The generation includes: (a) in the look-ahead time t vs The centroid t of the internal power distribution calculation sp and the predicted regenerative energy value E reku,100% The center of gravity is located at the look-ahead time t. vs (a) The point where the integral of the regenerative power accounts for half of the total regenerative energy; and (b) the point where the state vector S is generated. erweitert The state vector includes the predicted regenerative energy value E. reku,100% and the centroid t of the power distributionsp .
[0024]
[0025] In another preferred embodiment, the state vector S erweitert The generation includes: (a) by adjusting the regeneration power P over time t reku The integral of (t) extends from the current time point t0 within the driving cycle to the end of the driving cycle t. ende To calculate the weighted regenerative energy value E reku,gewichtet Among them, the regeneration power P reku (a) Weighting over time with a weighting factor α(t); and (b) generating the state vector S. erweitert The state vector includes the weighted regenerative energy value E. reku,gewichtet .
[0026]
[0027] A preferred implementation of the extended state vector enables different weighting of the predicted regenerative power within the driving cycle. The last implementation has the advantage that, by choosing a decreasing weighting factor α(t), such regenerative powers in the more distant future can be weighted less, as their occurrence is associated with greater uncertainty, especially with the use of an exponentially decreasing weighting factor α(t).
[0028] In another preferred embodiment, the reward function takes a positive value if (a) the battery state of charge (B) is improved and does not exceed the permissible range, and (C) the predicted regenerative energy can be stored, also within the permissible range of the battery state of charge, and (D) there is no interference from reflection. The reinforcement learning decision is thus implemented only within the range of the state space determined to be safe by reflection. Furthermore, the battery state of charge remains within a reliable range.
[0029] In another preferred embodiment, the neural network is trained according to the Q-learning algorithm. The Q-learning algorithm has proven particularly suitable for the current task.
[0030] A second aspect of the invention relates to an apparatus for performing the method according to the first aspect of the invention.
[0031] Where it is technically meaningful, the features and advantages described with respect to the first aspect of the invention and its advantageous designs also apply to the second aspect of the invention and its advantageous designs. Attached Figure Description
[0032] Other features, advantages, and application possibilities of the invention will become apparent from the following description taken in conjunction with the accompanying drawings.
[0033] At least partially schematic in the accompanying drawings:
[0034] Figure 1 An embodiment of a method for calculating regenerative power in an airborne power grid simulation is shown;
[0035] Figure 2 An embodiment of a method for integrating regenerative forecasting into an energy management system is shown;
[0036] Figure 3 An example of a reflection-enhanced reinforcement learning method in an airborne power grid simulation is shown. Detailed Implementation
[0037] Figure 1 This illustrates the method for calculating regenerative power P in an airborne power grid simulation. reku An embodiment of method 100.
[0038] The input variable is the generator state S. gen Battery current I bat and battery voltage U bat In method step 110, control points affecting the battery current variation curve through the energy management system's operating strategy are identified and extracted. In method step 120, other control point peaks are removed to smooth the battery current variation curve. Subsequently, in method step 130, the remaining control points are used to approximate the battery current variation curve. The approximate battery current variation curve I is then used. approx According to I reku =I bat –I approx To calculate the regeneration current I reku And according to P reku =U bat ·I reku To calculate the regeneration power P reku .
[0039] Figure 2 An embodiment of a method 200 for integrating regenerative prediction in an energy management system is shown.
[0040] The regeneration prediction 300 can be determined from sensor data 240 of the airborne power grid 400 and travel data from the travel database and transmitted to the energy management system 250. The energy management system can make policy decisions based on system state data 220 and the regeneration prediction 230, for example, through reinforcement learning.
[0041] Figure 3 An embodiment of a method 500 for reflection enhancement reinforcement learning in airborne power grid simulation is shown.
[0042] Reflection 600 stabilizes and protects the energy management system by examining all actions 550 proposed by the learning agent 510 and modifying the actions if necessary. Only actions 650 accepted and, if necessary, modified by reflection 600 have a direct impact on the state of the on-board power supply grid 700. The learning agent 510 then receives feedback in the form of a reward 610 according to the reward function on how the actions 550 proposed by the agent affect the on-board power supply grid. Thereby, during the learning process, the operating strategy related to the system state 710 is kept consistent with the desired optimization goal. The intervention of reflection 600 is taken into account in the reward function.
[0043] The following algorithm shows an embodiment for designing a suitable reward function to train the energy management system.
[0044]
[0045]
[0046] Here, the constant Delta represents the deviation of the state of charge SOC from the target value pursued. The deviation can be, for example, 2%. SOC describes the current state of charge, and SOC_ziel describes the optimized state of charge pursued. The optimized state of charge can be, for example, 80% of the maximum state of charge.
[0047] The constant E_Schwellwert can be described as follows:
[0048] SOC + SOC_durch_reku = SOC_ziel + Delta
[0049] SOC_durch_reku = SOC_ziel - SOC + Delta
[0050] SOC: The current SOC value
[0051] SOC_durch_reku: The increase in SOC caused by regeneration
[0052] SOC_ziel: The target SOC, for example 80%
[0053] Delta: Delta indicates how far the SOC should deviate from the target SOC
[0054] This means that only if the required SOC range (SOC_ziel - Delta < SOC < SOC_ziel + Delta) is exceeded without discharging, the battery in the case of expected regenerative energy should discharge.
[0055] E_Schwellwert=SOC_durch_reku*Q_batterie*U_batt_durchschn itt
[0056] E_Schwellwert: Energy threshold
[0057] Q_batterie: Nominal capacity of the battery
[0058] U_batt_durchschnitt: Average battery voltage over the entire cycle.
Claims
1. A method for training an energy management system (500) in an airborne power grid simulation, wherein, The method includes: a. The simulation features a limited regeneration driving cycle; b. Record the state variables of the airborne power grid (700); c. From the regeneration current I reku and battery voltage U bat The regeneration power P is calculated using the following formula. reku : P reku =U bat ·AND reku ; d. Generate the input vector for the neural network (510); e. Generate the reward function (610); f. Train the neural network (510), The generation of the input vector of the neural network (510) includes: The state input vector S of the neural network (510) is generated. normal The state input vector has the following form: With state vector S erweitert The state input vector S of the extended neural network (510) normal Therefore, the total vector S has the following form: Wherein, the state vector S erweitert The generation process includes the following steps: By adjusting the regeneration power P over time t reku The integral of (t) extends from the current time point t0 within the driving cycle up to the time point t0+x·t. vs The regenerative energy value E is calculated using the following integral. reku,x : Where x is used to consider the regeneration power P with limited foresight. reku (t) look-ahead time t vs Percentage share: Generate state vector S erweitert The state vector includes at least the regeneration energy value E. reku,25% E reku,50% E reku,75% and E reku,100% And it has the following form: Among them, if the battery charging status Improvements have been made and the improvement does not exceed the permissible range, and The predicted regenerative energy can be stored, without exceeding the allowable range of the battery's state of charge, and Reflection (600) did not interfere. Then the reward function (610) takes a positive value.
2. The method according to claim 1, wherein, The method is set up to train the energy management system (500) in a simulation of an onboard power grid (700) of a motor vehicle.
3. The method according to claim 1 or 2, wherein, The regenerative current I reku The determination (100) includes: Extracting the battery current change curve I bat The energy can be attributed to the decisions of the energy management system and there are no external influences on all control points of the airborne power grid (110); Smooth the battery current variation curve between the remaining control points. bat (120); Approximate battery current variation curve I between the remaining control points approx To approximate the battery current variation curve I bat (130); From battery current I bat and approximately the battery current I approx Calculate the regeneration current I using the following formula. reku : I reku =I bat -I approx 。 4. The method according to claim 3, wherein, The regenerative current I reku Equal to the battery current I bat .
5. The method according to claim 1 or 2, wherein, The state vector S erweitert The generation process includes the following steps: In the forward time t vs The centroid t of the power distribution is calculated using the following formula. sp and the predicted regenerative energy value E reku,100% : The centroid is located at the look-ahead time t. vs The point where the integral of the regenerative power accounts for half of the total regenerative energy; Generate state vector S erweitert The state vector includes the predicted regenerative energy value E. reku,100% and the centroid t of the power distribution sp And it has the following form:
6. The method according to claim 1 or 2, wherein, The state vector S erweitert The generation process includes the following steps: By adjusting the regeneration power P over time t reku The integral of (t) extends from the current time point t0 within the driving cycle to the end of the driving cycle t. ende The weighted regenerative energy value E is calculated using the following integral. reku,gewichtet : Among them, the regeneration power P reku (t) is weighted over time by a weighting factor α(t); Generate state vector S erweitert The state vector includes the weighted regenerative energy value E. reku,gewichtet And it has the following form: S erweitert =[E reku,gewichtet ]。 7. The method according to claim 1 or 2, wherein, The neural network is trained using the Q-learning algorithm (510).
8. An apparatus for training an energy management system (500) in an airborne power grid simulation, the apparatus being used to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and computing unit for sizing a classifier
DE102016200854A1
Methods and devices for determining an operating strategy profile
DE102017214384A1
Method for calculating energy recovery rate of braking energy recovery system
CN107310397A
Calibrating method of whole-vehicle controller of hybrid power vehicle
CN108515962A