Method and system for controlling an electrified vehicle
Hierarchical reinforcement learning optimizes hybrid vehicle operation by dividing control objectives into sub-objectives, using multiple agents to adapt to individual driving conditions, enhancing fuel and electrical energy efficiency.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-08-03
- Publication Date
- 2026-03-26
AI Technical Summary
Existing methods for controlling hybrid electric vehicles are not adaptable to individual driving conditions and driver styles, leading to inefficiencies in fuel and electrical energy use, and do not proactively consider parameters such as acceleration, speed profiles, weather, or traffic conditions.
A method and system using hierarchical reinforcement learning to divide control objectives into sub-objectives, employing multiple learning agents to optimize the hybrid powertrain's operating mode based on parameters like torque, fuel consumption, battery state, and environmental factors, utilizing cloud computing for interaction and strategy refinement.
Enhances the efficiency and effectiveness of fuel and electrical energy use in hybrid vehicles by optimizing operating modes based on real-time conditions, improving environmental performance.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method and system for controlling an electrified vehicle with or without a hybrid powertrain, wherein the control is based on an operating strategy to achieve a control goal.
[0002] A hybrid electric vehicle is a vehicle with a hybrid drive system powered by both one or more electric motors and an internal combustion engine. It draws its energy from an electrical storage device as well as from fuel carried on board. Hybrid drives are divided into two basic types and various mixed and intermediate forms. A system architecture in which both the electric motor and the internal combustion engine act on the hybrid drivetrain is called a parallel drive. The vehicle can be powered either purely electrically, by only one internal combustion engine, or by both engines in a proportionate manner. In a series hybrid drive, the hybrid drivetrain consists of a series arrangement with an internal combustion engine, a generator, and an electric motor.The combustion engine drives the generator, and the generated electricity is either used entirely to power the electric motor or partially stored in a battery. Various hybrid systems are possible between these two extremes.
[0003] The operating strategy of a hybrid vehicle represents the logical and temporal sequence of all operating states or modes of a hybrid drive system. A hybrid strategy control unit is used to manage the two motors. This unit selects the appropriate operating mode for the hybrid powertrain according to the defined strategy. The following operating states must be configurable by the control unit: 1. a purely electric drive, 2. a purely combustion engine-powered drive, 3. a hybrid drive in which both the electric motor and the combustion engine are used for propulsion, 4. a recuperation of the electrical energy of a regenerative brake, 5. a load point shift, 6. a thrust drive support; 7. Sailing without propulsion.
[0004] In pure combustion engine mode, the electric motor is in a neutral idle state, and the combustion engine must provide all the drive power on its own. With load point shifting, the electric motor's additive generator torque, at a fixed engine speed, shifts the combustion engine's load point along the torque axis. As soon as the electric motor ceases providing all drive power, the vehicle enters pure electric mode. Adding electric motor torque allows for relieving the combustion engine of some of its workload or even increasing the overall drive torque. This is known as drive assistance or boost mode and enhances driving performance. In coasting mode, for example, when rolling to a stop, the vehicle travels emission-free without any propulsion.In this operating mode, the combustion engine is typically switched off, the electric motor is idling, and the wheels are decoupled from the hybrid drivetrain, for example, via an open clutch. In recuperation mode, for instance, during braking or downhill driving where the required drive torque is negative, the electric motor can operate as a generator, thus generating electrical energy without using fuel and simultaneously reducing wear on the brakes.
[0005] For other electrified vehicle types, further operating states can be defined and implemented.
[0006] Choosing the appropriate operating mode ensures good driving performance and optimal energy consumption.
[0007] The choice of operating strategy is determined automatically by a control unit in the vehicle without active input from the driver. When the driver presses the accelerator pedal, signaling a torque demand, a control strategy calculates a target efficiency for the combustion engine and the electric motor. Based on this calculated target efficiency, the control unit (ECU = electronic control unit or ECM = electronic control module) determines which operating mode is most suitable for efficient fuel consumption and the desired driving performance.
[0008] The target efficiency depends on a number of parameters, including the power output of the combustion engine and / or electric motor, the temperature, the vehicle speed, and the battery's state of charge (SoC). Typically, the vehicle control unit (EUC) incorporates algorithms, characteristic curves, and graphical representations to calculate the target efficiency and select the most efficient operating mode.
[0009] Two main approaches are known for determining and predicting the operating mode of a hybrid vehicle. The first approach is based on mathematical models such as regression analyses and uses mathematical methods like artificial neural networks or the method of least squares to calculate the optimal operating mode for achieving a desired torque. Regression methods are statistical analysis techniques that aim to model relationships between a dependent variable and one or more independent variables. The regression models receive input such as gas charge, engine speed, or intake air temperature and use mathematical methods to calculate an estimated value for the torque.
[0010] Another approach uses tables and diagrams, preferably stored in the vehicle's electronic control unit (ECU). However, these tables and diagrams, or the mathematical procedures used, are based on expert knowledge, such as expected fuel consumption or the state of the energy storage system, and are not modified after implementation. Furthermore, they are not adapted to an individual engine but are fixed for a specific model series. They are subject to significant inaccuracies, and individual adjustment to driving conditions and a driver's style is not possible. In particular, it is not possible to proactively consider other parameters such as acceleration and speed profiles, weather conditions, or traffic conditions along a specific route.
[0011] CN 1 05 909 406 B describes a control procedure for an intelligent electronic control unit of a hybrid electric vehicle motor. Initialization is performed upon power-up, and information is collected. The expected torque and speed are determined using a control algorithm, thus optimizing the motor's performance.
[0012] CN 1 03 661 355 B describes an intelligent control system for a hybrid electric vehicle. Control units collect information about the states of the combustion engine, the electric motor, the storage battery, and the transmission, and select an operating mode that is adapted to the vehicle's condition and the current road conditions.
[0013] German patent DE 10 2018 215 017 A1 describes a method for determining an operating strategy for a vehicle, whereby vehicle data is recorded and grouped to train a classifier. The vehicle's operating strategy is then created using the classifier.
[0014] The publication “Marco Wiering, Martijn van Otterlo: “Reinforcement Learning - State of the Art”, Vol. 12, 1st edition, Springer Verlag Berlin, 2012, pp. 1-33, ISBN 978-3-642-27644-6” describes the fundamentals of reinforcement learning and related algorithms such as Markov decision processes, Bellman equations and Monte Carlo methods.
[0015] The publication “X. Lin et al: “Machine Learning-Based Energy Management in a Hybrid Electric Vehicle to Minimize Total Operating Cost”, in: IEEE / ACM International Conference on Computer-Aided Design (ICCAD), Austin, 2015, pp. 627-634, DOI: 10.1109 / ICCAD.2015.7372628) describes an approach to optimizing the energy management of hybrid vehicles using machine learning algorithms.
[0016] The publication “Andreas Heimrath et al.: “Reflex-augmented reinforcement learning for electrical energy management in vehicles”, in: ICAI '18: proceedings of the 2018 International Conference on Artificial Intelligence; July 30 - August 02, 2018, Las Vegas, Nevada, USA, ISBN 1-60132-480-4” describes the application of reflex-augmented reinforcement learning (RARL) for the energy management of electric vehicles.
[0017] The object underlying the invention is to create a method and a system for controlling an electrified vehicle that is characterized by more efficient use of fuel and / or electrical energy and thus contributes to an improvement in the environmental balance.
[0018] This problem is solved according to the invention with respect to a method by the features of claim 1, and with respect to a system by the features of claim 10. The further claims relate to preferred embodiments of the invention.
[0019] According to a first aspect, the invention provides a method for controlling an electrified vehicle along a route. The control is based on an operating strategy for achieving a control goal, and the calibration of the operating strategy comprises the following process steps: - Splitting the control objective G from a learning reinforcement main agent into one or more sub-objectives g i and passing on the sub-goals g i to one or more learning reinforcement sub-agents, each of which uses a reinforcement learning algorithm; - Determining a state s i by a state module, where a state s i by parameter p i such as data and / or measured values of at least one property e i a hybrid powertrain and / or the vehicle and / or the route and / or at least one influencing factor E i is defined, and transmitting the state s i to the main life insurance agent and / or one or more of the sub-agents, - Selecting a calculation function f i and / or an action i based on a guideline for the condition s i for a modification of at least one parameter pi which at least one property e i and / or at least one influencing factor E i from the main LV agent and / or at least one of the LV sub-agents; - Calculating a modeled value for the property e i and / or influencing factor E i using the modified parameter p i from an action module; - Calculating a new state s i+1 from an environment module based on the modeled value for the property e i and / or influencing factor E i ; - Comparing the new state s i+1 with a target state s t and assigning a deviation Δ for a comparison result in the state module; - Determining a reward r i from a reward module for the comparison result; - Adjusting the respective policy of the LV main agent and / or at least one LV sub-agent based on the reward r i , whereby, in the event of a convergence of the respective guidelines, the optimal action a i for the calculated state s i returned, and if the calculated respective policy does not converge, a further calculation function f is used. j and / or another action a j for another condition s j with a further modification of at least one further parameter p j selected by the main LV agent and / or the LV sub-agents until the target state s t of the tax objective and / or a sub-objective g iThe LV main agent, the LV sub-agents, the action module, the environment module, the state module, and the reward module each have one or more technical interfaces and protocols for accessing a cloud computing environment. Multiple LV sub-agents interact with each other via the cloud computing environment.
[0020] In a further development, the control objective represents a selection of an optimal operating mode for the electrified vehicle to increase efficiency, and the operating mode is either a purely electric drive, a purely combustion engine drive, a hybrid drive in which both an electric motor and the combustion engine are used for propulsion, a recuperation of electrical energy from a regenerative brake, a load point shift, a push-pull drive support, or driving without propulsion.
[0021] In particular, it is intended that a positive action A(+) changes the value for the parameter p. i The value of the parameter p is increased for a neutral action A(0). i remains the same, and a negative action A(-) changes the value of the parameter p i reduced.
[0022] In one embodiment, the reinforcement learning algorithm is designed as a Markov decision process (MDP or POMDP), as Temporal Difference Learning (TD-Learning), as Q-Learning, as SARSA, as Monte Carlo simulation, or as Actor-Critic.
[0023] Advantageously, at least the parameter p represents i represents a dimension, a material, a shape, or a measurement.
[0024] In particular, the property e ia target torque of the hybrid powertrain and / or a fuel consumption and / or a state of charge of the battery and / or a vehicle speed and / or a maximum power of the vehicle.
[0025] In particular, one influencing factor E i Weather conditions such as ambient temperature or wind speed, and / or the topography of the chosen route, such as a gradient or incline on a particular stretch, and / or the type of road surface, for example on a motorway or on a country road, and / or the traffic conditions.
[0026] In a further development, it is envisaged that a guideline will assign states s i to actions a i represents, where a positive reward r i for the calculated state s i a probability of election for the previous action a i for this condition s iis increased when there is a negative reward r i for the calculated state s i the probability of voting for the previous action a i for this condition s i is reduced, and in the event of a convergence of the directive, the optimal action a i for this calculated state s i is returned.
[0027] In a further training course, it is stipulated that the calculation results be presented in the form of states s i , actions a i , rewards r i and strategies are stored in the cloud computing environment and are available via the Internet.
[0028] According to a second aspect, the invention provides a system for controlling an electrified vehicle along a route. The control is based on an operating strategy for achieving a control goal. The system comprises a learning reinforcement main agent with a reinforcement learning algorithm, at least one learning reinforcement subagent with a reinforcement learning algorithm, an action module, an environment module, a state module, and a reward module. The learning reinforcement main agent is configured to divide the control goal into one or more subgoals. i to divide and each a sub-goal g i to pass it on to an LV sub-agent. The state module is configured to represent a state s i to determine, where a state s i by parameter p i such as data and / or measured values of at least one property e ia hybrid powertrain and / or the vehicle and / or the route and / or at least one influencing factor E i is defined. The state module transmits the state s i to the main life insurance agent and / or one or more of the sub-agents. The main life insurance agent and / or the sub-agents are trained to perform a calculation function f. i and / or an action a i based on a guideline for the condition s i for modifying at least one parameter p i which at least one property e i and / or influencing factor E i to select. The action module is configured to select a modeled value for the property e. i and / or influencing factor E i using the modified parameter p i to calculate. The environment module is configured to create a new state s i+1 due to the modeled value for the property e i and / or influencing factor Ei to calculate. The state module is formed, the new state s i+1 with a target state s t to compare and assign a deviation Δ to a comparison result. The reward module is formed, a reward r i to determine the comparison result and the reward r i to pass on the comparison result to the LV main agent and / or at least one LV sub-agent who are trained to implement the respective policy based on this reward r i to adapt. In the event of a convergence of the respective guidelines, the optimal action a i for the calculated state s i returned, and in the event of non-convergence of the respective policy, a further calculation function f is used. j and / or another action a j for another condition s j with a further modification of at least one further parameter p jwhich at least one property e j and / or influencing factor E j selected by the main LV agent and / or the LV sub-agents until the target state s t of the tax objective and / or a sub-objective g i The goal is achieved... The LV main agent, the LV sub-agents, the action module, the environment module, the state module, and the reward module each have one or more technical interfaces and protocols for accessing a cloud computing environment. Multiple LV sub-agents interact with each other via the cloud computing environment.
[0029] In a further development, the control objective represents the selection of an optimal operating mode for the hybrid vehicle to increase efficiency, and the operating mode is either a purely electric drive, a purely combustion engine drive, a hybrid drive in which both the electric motor and the combustion engine are used for propulsion, recuperation of electrical energy from a regenerative brake, load point shifting, thrust drive support, or driving without propulsion.
[0030] In particular, one property represents e i a target torque of the hybrid powertrain and / or the fuel consumption and / or the state of charge of the battery and / or the vehicle speed and / or the maximum power of the vehicle, and at least one influencing factor E irepresents the weather conditions such as the ambient temperature or wind strength, and / or the topography of the chosen route, such as the gradient or incline on a particular stretch, and / or the type of road surface, for example on a motorway or on a country road, and / or the traffic conditions.
[0031] In a further training course, it is stipulated that the calculation results be presented in the form of states s i , actions a i , rewards r i and strategies are stored in the cloud computing environment and are available via the Internet.
[0032] According to a third aspect, the invention provides a computer program product comprising executable program code that performs the method according to the first aspect.
[0033] The invention will now be explained in more detail with reference to exemplary embodiments shown in the drawing.
[0034] This shows: Fig. 1. A block diagram to illustrate a system according to the state of the art; Fig. 2 a block diagram to illustrate an embodiment of the system according to the invention; Fig. 3 a block diagram to illustrate a further detail of the system according to the invention Fig. 2; Fig. 4 a flowchart to explain the individual process steps of a process according to the invention; Fig. Figure 5 schematically shows a computer program product according to an embodiment of the third aspect of the invention.
[0035] Additional features, aspects and advantages of the invention or its embodiments become apparent from the detailed description in conjunction with the claims.
[0036] The operating strategy of a hybrid vehicle represents the logical and temporal sequence of all operating states or modes of a hybrid drive system with an internal combustion engine and an electric motor. In particular, the two motors are controlled according to a selected hybrid strategy. The hybrid strategy control unit selects the correct operating mode of the hybrid powertrain based on the chosen strategy. The following operating states must be adjustable by the control unit: 1. purely electric drive, 2. Pure propulsion by the internal combustion engine, 3. Hybrid drive, in which both the electric motor(s) and the combustion engine are used for propulsion, 4. Recuperation of the electrical energy of a regenerative brake, 5. Load point shift, 6. Thrust drive support; 7. Sailing without propulsion.
[0037] In pure combustion engine propulsion mode, the electric motor is in a neutral idle mode, and the combustion engine must provide all the drive power on its own. With load point shifting, the electric motor's additive generator torque, at a fixed engine speed, shifts the combustion engine's load point along the torque axis. As soon as the electric motor ceases providing all drive power, the vehicle enters purely electric mode. Adding electric motor torque allows for relieving the combustion engine of some of its load or even increasing the overall drive torque. This is known as drive assistance or boost mode and serves to enhance driving performance. In coasting mode, for example, when rolling to a stop, the vehicle travels emission-free without any propulsion.In this operating mode, the combustion engine is typically switched off, the electric motor is idling, and the wheels are decoupled from the hybrid drivetrain, for example, via an open clutch. In recuperation mode, for instance, during braking or downhill driving where the required drive torque is negative, the electric motor can be operated as a generator, thus generating electrical energy without using fuel and simultaneously reducing wear on the brakes.
[0038] Fig. Figure 1 shows the state of the art for a system 100 with a control logic for controlling a hybrid vehicle based on reinforcement learning. Only a single reinforcement learning agent 10 is provided, which learns only one global guideline to achieve the control objective, i.e., the selection of the appropriate operating mode.
[0039] The LV agent 10 selects at least one action 20 for a given state 50. The choice of action 20 is based on a policy. For the selected action 20, the agent 10 receives a reward 40 from an environment module 30. The policy is then adjusted based on the received rewards 40. An environment module 30 also calculates the states 50 based on the selected action 20.
[0040] The LV agent 10 only learns to calculate the efficiency of the hybrid powertrain.
[0041] In Fig. Figure 2 shows a system 200 according to the invention, comprising a control logic for controlling a hybrid vehicle. The control logic uses principles of hierarchical reinforcement learning (HRL) to determine a control target G. The control target G is defined in a main LV agent 210 and represents the operating strategy for a specific driving situation in order to improve the efficiency of the hybrid powertrain and thus of the hybrid vehicle. The control target G therefore determines the appropriate operating mode of the hybrid vehicle for a specific driving situation, such as purely electric driving.
[0042] According to the invention, the control target G is divided by the LV main agent 210 into various sub-targets g i divided. These sub-goals g iThey can themselves each represent an operating mode, such as a load point shift, a purely electric mode, a purely combustion engine mode, or they can refer to properties e i of the hybrid powertrain and the vehicle, such as fuel consumption and / or battery charge level and / or vehicle speed and / or maximum combustion engine power, or other influencing factors E i , such as weather conditions, for example ambient temperature or wind speed, and / or topography of the chosen route, for example gradient or incline on a particular stretch, and / or type of road surface, for example on a motorway or on a country road, and / or traffic conditions, etc.
[0043] According to the invention, each sub-goal g irepresented by an LV sub-agent 220, 230 and achieved using values, functions, and policies. The number of LV sub-agents 220, 230 varies according to the number of sub-goals Σg. i Each individual LV sub-agent 220, 230 thus has the task of achieving a sub-target g i to achieve. If all or at least a large part of the sub-goals are achieved. i Once the target G has been achieved, the overall control objective has been reached. Possible architectures for hierarchical reinforcement learning include FeUdal HRL, hierarchical abstract machines, MAXQ, HIRO, HAC, MLSH, H DQN, STRAW, H DRLN, AMDP, IMHOP, and HSP.
[0044] System 200, with a control logic based on hierarchical reinforcement learning, comprises the main LV agent 210 and at least one or more LV sub-agents 220, 230, each of which has a reinforcement learning algorithm. Furthermore, at least one action module 300 with actions a i ∈ A, at least one environment module 400, at least one state module 500 with states s i ∈ S, and at least one reward module 600 with rewards r i ∈ ℝ provided.
[0045] The LV main agent 210, the LV sub-agents 220, 230, the action module 300, the environment module 400, the condition module 500 and the reward module 600 can each be equipped with a processor and / or a memory unit.
[0046] In the context of the invention, a "processor" can be understood to mean, for example, a machine or an electronic circuit. In particular, a processor can be a central processing unit (CPU), a microprocessor, or a microcontroller, such as an application-specific integrated circuit or a digital signal processor, possibly in combination with a memory unit for storing program instructions, etc. A processor can also be understood to mean a virtualized processor, a virtual machine, or a soft CPU.It may, for example, also be a programmable processor that is equipped with configuration steps for executing the said method according to the invention or is configured with configuration steps such that the programmable processor realizes the features of the method, the component, the modules, or other aspects and / or partial aspects of the invention according to the invention.
[0047] In the context of the invention, a "storage unit" or "storage module" and the like can refer, for example, to volatile memory in the form of random-access memory (RAM), persistent storage such as a hard drive or data carrier, or, for example, a removable storage module. The storage module can also be a cloud-based storage solution.
[0048] In the context of the invention, a "module" can be understood to mean, for example, a processor and / or a memory unit for storing program instructions. For example, the processor is specifically configured to execute the program instructions in such a way that the processor and / or the control unit performs functions to implement or realize the method according to the invention or a step thereof.
[0049] In the context of the invention, "measured values" refers to both raw data and processed data, for example, from sensor measurements. These sensors may include accelerometers, velocity sensors, pressure sensors, cameras, charge measuring devices, temperature sensors, chemical sensors, etc.
[0050] According to the invention, the main LV agent 210 has the task of achieving the control objective G. The control objective G represents the calibration of a suitable operating strategy for selecting an optimal operating mode of the hybrid vehicle's powertrain in order to improve the efficiency and effectiveness of the fuel and electrical energy used. The main LV agent 210 divides the control objective G into various sub-objectives g. i divided up, each being allocated to a LV sub-agent 220, 230.
[0051] The LV main agent 210 and the LV sub-agents 220, 230 each select for a specific state s i ∈ S, at least one action a from a set of available states i ∈ A from a set of available actions. The choice of the selected action a i based on their chosen policy. For the selected action a iLV agents 210, 220, and 230 each receive a reward. i ∈ ℝ of the reward module 600. The states s i The LV agents 210, 220, 230 receive ∈ S from the state module 500, which the LV agents 210, 220, 230 can access. A state s i ∈ S is determined by the selection of certain parameter values p i for properties e i and / or influencing factors E i defined and thus by measured and / or calculated values of selected properties e i and / or influencing factors E i marked.
[0052] The policy of the respective LV agent 210, 220, 230 is based on the rewards received. i adapted. The policy specifies which action a i ∈ A from the set of available actions for a given state s i∈ S is to be selected from the set of available states. This creates a new state s. i+1 generated, for which the respective LV agent 210, 220, 230 receives a reward r i a guideline thus defines the assignment between a state s. i and an action a i fixed, so that the directive determines the choice of action to be performed a i for a state s i indicates. The goal of LV's main agent 210 is to collect the rewards obtained. i ,r i+1 ,...,r i+n to maximize in order to achieve his tax target G. The goal of the respective LV sub-agents 220 and 230 is to maximize the rewards earned. i ,r i+1 ,...,r i+n to maximize in order to achieve the respective sub-goal g i to achieve this. A guideline is therefore an assignment of a state s i to one or more actions a i . For each condition s iThe directive specifies the actions to be performed. i The relationship between a next action a i and guideline R can therefore be defined for the LV sub-agents 220, 230 as follows: ai=R(si,gi)
[0053] An example of an algorithm for improving policy are greedy algorithms. These are characterized by their step-by-step selection of the next state that promises the greatest gain or the best result at the time of selection. Evaluation methods such as gradient descent are typically used for this purpose.
[0054] The sub-goal g i A goal is considered achieved when a certain period of time has elapsed and the target state of the sub-goal g has been reached. i has been achieved.
[0055] In action module 300, the actions selected by the respective LV agent 210, 220, 230 are a iexecuted through an action a i For example, an adjustment of a parameter value p i a property e i and / or an influencing factor E i carried out. Regarding the properties e i This could involve the target torque of the hybrid powertrain and / or the fuel consumption and / or the battery's state of charge and / or the vehicle's speed and / or its maximum power output. The influencing factors E i This could involve weather conditions, such as ambient temperature or wind speed, and / or the topography of the chosen route, such as the gradient or incline on a particular stretch, and / or the type of road surface, for example, on a motorway or a country road, and / or traffic conditions, etc. The measured or estimated parameter values p iThe values may have been determined by sensors not described in detail here. Preferably, the parameter values are stored in a table of values or the like. Preferably, action a i to one of the actions A(+), A(0) and A(-). A positive action A(+) is an action that changes the value for a parameter p. i increased, in a neutral action A(0) is an action in which the value of the parameter p i remains the same, while in the case of a negative action A(-) the value of the parameter p changes. i reduced.
[0056] The Environment Module 400 incorporates the data and measurements from the vehicle's sensors and / or other data sources, such as cartographic information, and calculates the actual fuel efficiency. It also determines a modeled fuel efficiency. For this purpose, the Environment Module 400 calculates a based on the selected action. iand taking into account the respective sub-goal defined g i the conditions s i ∈ S.
[0057] In state module 500, a sub-goal g is defined. i a deviation Δ between a target state s t and the calculated state s i calculated. A state s i This can, for example, refer to the efficiency of the hybrid powertrain. The deviation Δ then provides feedback on the difference between the actual efficiency of the hybrid powertrain and the calculated efficiency of the hybrid powertrain. The sub-goal g i is achieved when the calculated states s i equal to or greater than the target states s t are.
[0058] In the reward module 600, the degree of deviation Δ between the calculated value for the state s is taken into account. i and the target value of the state s t a reward r iassigned. Since the degree of deviation Δ depends on the selection of the respective action A(+), A(0), A(-), the reward r is preferably assigned in a matrix or a database for the respective selected action A(+), A(0), A(-). i assigned. A reward r i preferably has the values +1 and -1, with a small or positive deviation Δ between the calculated state s i and the target state s t A significant deviation is rewarded with +1 and thus reinforced, while a substantial negative deviation Δ is rewarded with -1 and thus negatively assessed. However, it is also conceivable that values > 1 and values < 1 could be used.
[0059] Rewards r i This is therefore feedback to the LV agents 210, 220, 230 regarding the successfully or less successfully modified parameter values p. i The feedback is based on the selected actions. ifor the respective states s i , i.e., the deviation between the calibrated and the real value for a parameter p i .
[0060] Preferably, a Markov decision process (MDP or POMDP) is used as the algorithm for the LV agents 210, 220, 230. However, it is also possible to use a Temporal Difference Learning (TD-Learning) algorithm. An LV agent 210, 220, 230 with a TD-Learning algorithm does not adjust the actions A(+), A(0), A(-) only when it receives the reward, but after each action a. i based on an estimated expected reward. Furthermore, algorithms such as Q-Learning and SARSA, or Actor-Critic or Monte Carlo simulations, are also conceivable. The algorithm allows for dynamic programming and strategy adaptation through iterative processes.
[0061] Furthermore, the LV agents 210, 220, 230 and / or the action module 300 and / or the environment module 400 and / or the state module 500 and / or the reward module 600 contain calculation methods and algorithms f i for mathematical regression methods or physical model calculations that establish a correlation between selected parameters p i ∈ P from a set of parameters and the target states s t describe. In the case of mathematical functions f tThese can be statistical methods such as means, minimum and maximum values, lookup tables, models of expected values, linear regression methods or Gaussian processes, fast Fourier transforms, integral and differential calculus, Markov methods, probability methods such as Monte Carlo methods, temporal difference learning, but also extended Kalman filters, radial basis functions, data fields, convergent neural networks, deep neural networks, artificial neural networks and / or feedback neural networks. Based on the actions a i and the rewards r i The LV agents 210, 220, 230 and / or the action module 300 and / or the environment module 400 and / or the state module 500 and / or the bypass module 400 select a state s i one or more of these calculation functions f i out of.
[0062] Now a second cycle begins to achieve the respective sub-goals. i In this case, LV agents 210, 220, 230 can perform a different action. i+1 and / or another calculation function f i+1 and / or another parameter p i+1 Select according to the defined guideline. The result is then fed back into the state module 500, and the result of the comparison is evaluated in the reward module 600. LV agents 210, 220, and 230 repeat the calculations for the respective sub-goal g. i for all planned actions a i ,ai +1> ...,ai +n , Calculation functions f i ,f i+1 ,...,f i+n and parameter p i ,p i+1 ,...,p i+n until the greatest possible agreement is reached between a calculated state s i ,s i+1 ,...,s i+n and a target state s tThe final state of the control target G is reached when the deviation Δ between the actual efficiency of the hybrid powertrain and the modeled efficiency of the hybrid powertrain is in the range of + / -5%.
[0063] LV agents 210, 220, and 230 optimize their behavior and thus the guideline according to which an action a i The selection continues until the policy converges. The LV agents 210, 220, and 230 thus learn which action(s) a i ,a i+1 ,...,a i+n for which condition s i ,s i+1 ,...,s i+n are the best. If they are in a certain condition. i ,s i+1 ,...,s i+n visit very often and each time a different chain of actions a i ,a i+1 ...,a i+n with selected actions a i ,a i+1, ...,a i+nBy trying out conditions that can be both very different and very similar, they gain experience with their respective policy and thus the calibration methodology. When they test the conditions, i ,s i+1 ,...,s i+n have visited often enough and enough actions have taken place i ,a i+1, ...,a i+n If the respective LV agent 210, 220, 230 has been tested, then the policy can converge to the optimal policy. This means that the optimal action(s) a i ,a i+1, ...,a i+n for a specific condition s i ,s i+1 ,...,s i+n be returned to the target state s t to come.
[0064] As in Fig. As shown in Figure 3, it can be specifically provided that the calculation results, in the form of states, actions, rewards, and policies of the LV agents 210, 220, and 230, are stored in a cloud computing environment 700 and are accessible via the internet. The LV agents 210, 220, and 230, the action module 300, the environment module 400, the state module 500, and the reward module 600 have the necessary technical interfaces and protocols for accessing the cloud computing environment 700. This increases computing efficiency because access to previously calculated states, actions, rewards, and strategies is simplified and speeds are improved.
[0065] Furthermore, it is possible to couple the various LV agents 210, 220, and 230, allowing them to dynamically interact with each other via the cloud computing environment 700 and store their results there. This can improve the quality and stability of the operational strategy calculation, as the LV agents 210, 220, and 230 can learn from each other and compensate for potential weaknesses in other LV agents 210, 220, and 230. Overall, this can enhance the convergence behavior of the system 200.
[0066] It can also be provided that the entire software application (computer program product) according to the invention is stored in the cloud computing environment 700. This allows the know-how of the calculation algorithms to be better protected and secured, since these algorithms do not need to be passed on to the environment outside the cloud computing environment 700.
[0067] The reward function R is usually defined as a linear combination of different attributes (features) A i and weights w i shown: R=w1∗A1+w2∗A2+⋯+wn+An with R = Evaluation function w i = Weights f i = Attributes (deviation between modeled and real values)
[0068] Regarding attributes A i Within the scope of this invention, it is particularly concerned with the deviation Δ between a target state s t and a calculated state s i The attributes A i However, they can also represent other categories. Furthermore, other formulas for the reward function R are also possible.
[0069] To develop an optimal reward function R, the individual weights w are i especially adapted by an expert such as an engineer, so that the reward r iThe reward function is maximized. Since this is not an autonomous process of reinforcement learning, such an approach can be called inverse reinforcement learning. In inverse reinforcement learning, the optimal reward function R is thus determined based on the observed behavior of experts, such as engineers and scientists.
[0070] The actions taken by an expert a i based on his experience. It is possible to generalize the expert's behavior using a characteristic curve that can be used to calculate rewards. In the next step, the reward function is then optimized to calculate the rewards.
[0071] Furthermore, optimization algorithms such as yield optimization or entropy optimization, statistical algorithms such as classification and regression algorithms or Gaussian processes, and algorithms from imitative learning can be used to optimize the reward function R.
[0072] Furthermore, algorithms such as Generative Adversarial Networks (GANs) can be used. Generative Adversarial Networks consist of two artificial neural networks that perform a zero-sum game. The first network is called the generator, which creates candidates. The second network is the discriminator, which evaluates the candidates. This allows deviations to be added to the environment. The LV agents 210, 220, and 230 now operate in an environment that has changed slightly due to the added deviation. These changes to the environment are generated by the GAN.
[0073] In Fig. Figure 4 shows the process steps for controlling a hybrid powertrain of a hybrid vehicle, wherein the control is based on an operating strategy to achieve a control target G and calibration of the operating strategy comprises the following process steps: In step S10, the control target G is divided into several sub-targets g by a main LV agent 210. i divided up, each of which is passed on to LV sub-agents 220 and 230.
[0074] In step S20, at least one state s is determined by a state module 500. i of the hybrid powertrain and / or the vehicle and / or the route is determined and transmitted to the LV main agent 210 and / or at least one of the LV sub-agents 220, 230, wherein a state s i by parameter p i such as data and measurements of at least one property e i of the hybrid powertrain and / or at least one influencing factor Ei of the vehicle and / or the route.
[0075] In step S30, the LV main agent 210 and / or at least one of the LV sub-agents 220, 230 select the respective transmitted state s. i at least one calculation function f i and / or an action a i based on a guideline for a condition s i for modifying at least one parameter p i which at least one property e i and / or at least one influencing factor E i .
[0076] In step S40, an action module 300 calculates a modeled value for the property e i and / or the influencing factor E i using the modified parameter p i .
[0077] In step S50, an environment module 400 calculates a new state s i+1 due to the modeled value for the property e iand / or the influencing factor E i .
[0078] In step S60, the state module 500 compares the new state s i+1 with a target state s t of the control objective G and / or a sub-objective g i and assigns it a deviation Δ.
[0079] In step S70, a reward module 600 determines a reward r i for the comparison result.
[0080] In step S80, the policy of LV main agent 210 and LV sub-agents 220, 230 is adjusted based on the reward r. i , where, in the event of a convergence of the directive, the optimal action is a i for the calculated state s i is returned, and in the event of non-convergence of the policy, a further calculation function f j and / or another action a j for another condition s j with a further modification of at least one further parameter pj which at least one property e j and / or influencing factor E j selected by the LV main agent 210 and / or the LV sub-agents 220, 230, until the target state s t of the control objective G and / or a sub-objective g i has been achieved.
[0081] Fig. Figure 5 schematically represents a computer program product 900 comprising an executable program code 950 that performs the method according to the first aspect of the present invention.
[0082] According to the method and system of the present invention, hierarchical reinforcement learning is used to calibrate an operating strategy for a vehicle. For this purpose, a control target G is defined, which is divided into sub-targets g by a main LV agent 210. i is divided. These sub-goals g iEach is processed by an assigned LV sub-agent 220, 230, in order to achieve the control objective G. The LV main agent 210 and the LV sub-agents 220, 230 independently select the parameter p. i actions a i The LV subagents 220 and 230 can interact with each other, so that nonlinear relationships between these parameters p are also possible. i can be recorded. It is an autonomous calibration and calculation method, since the LV main agent 210 and the LV sub-agents 220, 230 perform the respective actions a i choose yourself and receive a reward for each one. i This allows for individual control of the hybrid powertrain in a hybrid vehicle, enabling more efficient use of fuel and electrical energy and thus improving the overall environmental performance. In addition to the properties of e iThe hybrid powertrain also incorporates other aspects such as the topography of the respective traffic route, the traffic conditions, the weather conditions, and the vehicle's speed as further influencing factors. i The LV sub-agents 220 and 230 are taken into account to calculate an optimal operating strategy for controlling the hybrid powertrain of the hybrid vehicle. By using hierarchically arranged LV agents 210, 220, and 230, each performing its calculations using reinforcement learning algorithms, it is possible to implement the operating strategy for controlling a hybrid vehicle autonomously and in a self-optimizing manner. In particular, the system 200 according to the invention makes it possible to handle even complex control tasks.
[0083] One advantage of this solution lies in its suitability for models and control systems in a wide variety of applications. These include combustion engines, electric motors and batteries, as well as their diagnostic functions, aircraft, fuel cells, hybrid drives, trains, boats, robots, motorcycles, electric cars, and other road vehicles. Examples include vehicle dynamics controllers (aerodynamics, suspension, chassis, steering, etc.), safety systems (alarm, central locking, parking brake, etc.), driver assistance systems (blind spot monitoring, adaptive cruise control, lane keeping assist, crosswind stabilization, etc.), passenger compartment equipment (seats, air conditioning, panoramic sunroof, cameras, etc.), and other auxiliary systems (brake booster, front axle lift, gradient control, tire pressure monitoring, windshield wipers, headlights, etc.). Reference sign 10 LV agent 20 Action 30 Environment module 40 Reward 50 condition 100 System 200 system according to the invention for control 210 LV Main Agent 220 LV sub-agent 240 LV sub-agent 300 action module 400 Environment Module 500 State module 600 Reward Module 700 Cloud Computing Environment 900 computer program product 950 program code
Claims
[1] A method for controlling an electrified vehicle on a journey route, wherein the control is based on an operating strategy to achieve a control target (G) and a calibration of the operating strategy comprises the following process steps: - Splitting the control objective (G) (S10) from a learning reinforcement main agent (210) into one or more sub-objectives (g i ) and passing on the sub-goals (g i ) to one or more learning reinforcement sub-agents (220, 230), wherein the LV agents (210, 220, 230) each use a reinforcement learning algorithm; - Determining (S20) a state (s i ) by a state module (500), where a state (s i ) by parameters (p i ) such as data and / or measured values of at least one property (e i ) of a hybrid powertrain and / or the vehicle and / or the route and / or at least one influencing factor (Ei ) is defined, and transmitting the state (s i ) to the main LV agent (210) and / or one or more of the sub-LV agents (220, 230), - Selecting (S30) a calculation function (f i ) and / or an action (a i ) based on a policy for the condition (s i ) for a modification of at least one parameter (p i ) of at least one property (e i ) and / or at least one influencing factor (E i ) from the main LV agent (210) and / or at least one of the sub-LV agents (220, 230); - Calculating (S40) a modeled value for the property (e i ) and / or influencing factor (E i ) using the modified parameter (p i ) from an action module (300); - Calculating (S50) a new state (s i+1 ) from an environment module (400) based on the modeled value for the property (e i) and / or influencing factor (E i ); - Comparing (S60) the new state (s i+1 ) with a target state (s t ) and assigning a deviation (Δ) for a comparison result in the state module (500); - Determining (S70) a reward (r i ) from a reward module (600) for the comparison result; - Adjusting (S80) the respective policy of the LV main agent (210) and / or at least one LV sub-agent (220, 230) based on the reward (r i ), where If the respective guidelines converge, the optimal action (a) i ) for the calculated state (s i ) is returned, and if the calculated policy does not converge, another calculation function (f) is used. j ) and / or another action (a j ) for another state (s j ) with a further modification of at least one further parameter (pj ) selected by the LV main agent (210) and / or the LV sub-agents (220, 230) until the target state (s t ) of the tax objective (G) and / or a sub-objective (g) i ) is reached, wherein the LV main agent (210), the LV sub-agents (220, 230), the action module (300), the environment module (400), the state module (500) and the reward module (600) have one or more technical interfaces and protocols for accessing a cloud computing environment (700), and where several LV sub-agents (220, 230) interact with each other via the cloud computing environment (700). [2] Method according to claim 1, wherein the control objective (G) is a selection of an optimal operating mode for the electrified vehicle to increase efficiency and the operating mode is either a purely electric drive or a purely internal combustion engine drive or a hybrid drive in which at least one electric motor and the internal combustion engine are used for propulsion, or recuperation of electrical energy from a regenerative brake, or load point shifting, or thrust drive support, or driving without propulsion. [3] Method according to claim 1 or 2, wherein a positive action (A(+)) changes the value for the parameter (p i ) increases, in the case of a neutral action (A(0)) the value of the parameter (p i ) remains the same, and a negative action (A(-)) changes the value of the parameter (p i ) reduced. [4] Method according to any of the preceding claims, wherein the reinforcement learning algorithm is designed as a Markov decision process or as Temporal Difference Learning (TD-Learning) or as Q-Learning or as SARSA or as Monte Carlo simulation or as Actor-Critic. [5] Method according to any one of the preceding claims, wherein at least the parameter (p i ) represents a dimension, a material, a shape, or a measurement. [6] Method according to one or more of the preceding claims, wherein at least the property (e i ) represents a target torque of the hybrid powertrain and / or a fuel consumption and / or a state of charge of a battery and / or a vehicle speed and / or a maximum power of the vehicle. [7] Method according to one or more of the preceding claims, wherein at least one influencing factor (E i) Weather conditions such as ambient temperature or wind speed, and / or the topography of the chosen route, such as a gradient or incline on a particular stretch, and / or the type of road surface, for example on a motorway or on a country road, and / or the traffic conditions. [8] Method according to one or more of the preceding claims, wherein the guideline provides for an assignment of states (s i ) to actions (a i ) represents, where a positive reward (r) i ) for the calculated state (s i ) a probability of choosing the previous action (a i ) for this condition (s i ) is increased in the case of a negative reward (r i ) for the calculated state (s i ) the probability of choosing the previous action (a i ) for this condition (s i) is reduced, and at the convergence of the policy, the optimal action (a i ) for this calculated state (s i ) is returned. [9] Method according to any of the preceding claims, wherein the calculation results are in the form of states (s i ), actions (a i ), rewards (r i ) and strategies are stored in the cloud computing environment (700) and are available via the Internet. [10] A system (200) for controlling an electrified vehicle on a route, wherein the control is based on an operating strategy to achieve a control goal (G), comprising a learning reinforcement main agent (210) with a reinforcement learning algorithm, at least one learning reinforcement sub-agent (220, 230) with a reinforcement learning algorithm, an action module (300), an environment module (400), a state module (500) and a reward module (600), wherein the learning reinforcement main agent (210) is configured to divide the control goal (G) into one or more sub-goals (g i ) to divide and each have a sub-goal (g i ) to pass to an LV sub-agent (220, 230), wherein the state module (500) is configured to represent a state (s i ) to determine, whereby the state (s i ) by parameters (p i ) such as data and / or measured values of at least one property (e i) of a hybrid powertrain and / or the vehicle and / or the route and / or at least one influencing factor (E i ) is defined, and the state (s i ) to transmit to the LV main agent (210) and / or one or more of the LV sub-agents (220, 230); wherein the LV main agent (210) and / or the LV sub-agents (220, 230) are configured to perform a calculation function (f i ) and / or an action (a i ) based on a policy for the condition (s i ) for a modification of at least one parameter (p i ) of at least one property (e i ) and / or influencing factor (E i ) to select; where the action module (300) is configured to select a modeled value for the property (e i ) and / or influencing factor (E i ) using the modified parameter (p i ) to calculate; wherein the environment module (400) is configured to generate a new state (s i+1) due to the modeled value for the property (e i ) and / or influencing factor (E i ) to calculate; wherein the state module (500) is configured to determine the new state (s i+1 ) with a target state (s t ) to compare and assign a deviation (Δ) to a comparison result; wherein the reward module (600) is configured, a reward (r) i ) to determine the comparison result and the reward (r i ) to pass on the comparison result to the LV main agent (210) and / or at least one LV sub-agent (220, 230) who are trained to apply the respective policy based on this reward (r) i ) to adapt, whereby in the event of a convergence of the respective policies, the optimal action (a i ) for the calculated state (s i ) is returned, and in the event of non-convergence of the respective policy, a further calculation function (f j) and / or another action (a j ) for another state (s j ) with a further modification of at least one further parameter (p j ) of at least one property (e j ) and / or influencing factor (E j ) is selected by the LV main agent (210) and / or the LV sub-agents (220, 230) until the target state (s t ) of the tax objective (G) and / or a sub-objective (g) i ) is achieved, wherein the LV main agent (210), the LV sub-agents (220, 230), the action module (300), the environment module (400), the state module (500) and the reward module (600) have one or more technical interfaces and protocols for accessing a cloud computing environment (700), and wherein multiple LV sub-agents (220, 230) interact with each other via the cloud computing environment (700). [11] System (200) according to claim 10, wherein the control objective (G) is the selection of an optimal operating mode for the electrified vehicle to increase efficiency and the operating mode is either a purely electric drive or a purely internal combustion engine drive or a hybrid drive in which at least one electric motor and the internal combustion engine are used for propulsion, or recuperation of electrical energy from a regenerative brake, or load point shifting, or thrust drive support, or driving without propulsion. [12] System (200) according to claim 10 or 11, wherein at least one feature (e i ) represents a target torque of the hybrid powertrain and / or an energy consumption and / or a state of charge of a battery and / or a vehicle speed and / or a maximum power of the vehicle, and at least one influencing factor (E i) Weather conditions such as ambient temperature or wind speed, and / or the topography of a chosen route, such as a gradient or incline on a particular stretch, and / or the type of road surface, for example on a motorway or on a country road, and / or the traffic conditions. [13] System (200) according to one of claims 10 to 12, wherein the calculation results are in the form of states (s i ), actions (a i ), rewards (r i ) and strategies are stored in the cloud computing environment (700) and are available via the Internet. [14] Computer program product (900) comprising an executable program code (950) that performs the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
A method for intelligent control of hybrid electric vehicle powertrain
CN103587522B
Energy management method for hybrid electric vehicles based on multi-agent technology
CN103640569B
A hybrid vehicle powertrain intelligent control system
CN103661355B
A control method for a hybrid electric vehicle engine intelligent electronic control unit
CN105909406B
Method for determining an operating strategy for vehicle, system, control unit for a vehicle and a vehicle
DE102018215017A1