Method and device for monitoring an onboard network of a vehicle
Through a machine learning-based energy management system and the use of reference units and reward units trained through reinforcement learning, accurate monitoring and failure prediction of vehicle network components are achieved, solving the difficult problem of failure prediction of vehicle network components and improving vehicle safety and comfort.
Patent Information
- Application Number
- CN202180014666.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-13
- Filing Date
- 2021-02-15
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2041-02-15
AI Technical Summary
Existing technologies have difficulty effectively and accurately predicting the failure of in-vehicle network components, which may lead to safety-critical problems, especially in automated vehicles.
A machine learning-based energy management system is adopted, and reinforcement learning is used to train the energy management system. The vehicle network status is monitored through reference units and reward units. The comparison between reference rewards and actual rewards is used to predict component failure, including the design of reference units and reward units.
It achieves accurate monitoring of vehicle network components and early prediction of failures, improves vehicle safety and comfort, and ensures safe energy supply for automated driving.
Smart Images

Figure CN115135526B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and a corresponding device with which components of an onboard electrical system can be monitored reliably and efficiently, in particular in order to be able to predict failures of onboard electrical system components already in advance. Background Art
[0002] A vehicle includes an onboard (energy) electrical system, via which electrical energy can be supplied to a plurality of electrical consumers in the vehicle. The onboard electrical system typically includes an electrical energy accumulator for storing electrical energy and for supporting the voltage of the onboard electrical system. Furthermore, the onboard electrical system typically includes a generator (driven, for example, by the vehicle's internal combustion engine) configured to generate electrical energy for the onboard electrical system. Furthermore, the onboard electrical system of a hybrid or electric vehicle typically includes a DC voltage converter (powered, for example, by the vehicle's traction battery) configured to generate or provide electrical energy for the onboard electrical system.
[0003] The onboard electrical system can be operated with the aid of an energy management system. The energy management system can be configured to protect the energy supply of electrical consumers via the onboard electrical system. For this purpose, the energy management system can include one or more controllers configured to adjust one or more controlled variables of the onboard electrical system (e.g., the onboard electrical system voltage) to corresponding setpoint values.
[0004] The energy management system may include or may be an energy management system based on machine learning. In particular, the energy management system may include one or more regulators that have been trained based on machine learning.
[0005] Failure of an onboard power supply component of the onboard power supply (eg, an energy store or a generator or a DC voltage converter or an electrical load) can lead to impairment of vehicle operation, which can lead to safety-critical situations, particularly in autonomously driven vehicles. Summary of the Invention
[0006] The technical task addressed here is to predict future and / or impending failures of components of an onboard electrical system in an efficient and precise manner.
[0007] The object is achieved by the technical solution of the present invention. The present invention relates to a device for monitoring an onboard energy system, the onboard energy system comprising different onboard energy system components and being operated by means of a machine-learning energy management system; wherein the energy management system has been trained by means of reinforcement learning for a reference onboard energy system, the reference onboard energy system being able to correspond to an onboard energy system having fault-free and / or undamaged onboard energy system components; wherein the device comprises: a reference unit configured to determine a reference reward for a state of the onboard energy system and for an action caused by the energy management system based on the state, the reference reward being generated when the reference onboard energy system is operated; a reward unit configured to determine an actual reward for the state and for the action, the actual reward being generated when the onboard energy system is operated; and a monitoring unit configured to monitor the onboard energy system based on the actual reward and the reference reward.
[0008] According to one aspect, a device (also referred to in this document as a diagnostic module) for monitoring an onboard energy system, in particular an onboard energy system of a (motor) vehicle, is described. The onboard energy system includes various onboard power system components, such as one or more electrical energy storage devices, one or more electrical consumers, and / or one or more generators or DC voltage converters (the onboard power system components are configured to provide electrical energy externally in the onboard power system).
[0009] The onboard energy system is operated using a machine learning-based energy management system, which has been trained using reinforcement learning. Within the scope of reinforcement learning, rewards can be assigned to actions initiated by the energy management system based on a specific state of the onboard energy system. The rewards can be calculated based on a specific reward function, wherein the reward function depends on one or more measurable variables, in particular state variables, of the onboard energy system. The energy management system can be trained within the scope of reinforcement learning such that, based on the state of the onboard energy system, the energy management system can initiate actions that, as a result of these actions, maximize the cumulative sum of future (possibly discounted) rewards.
[0010] The state of the onboard energy system can be described by one or more (measurable) state variables. Exemplary state variables are:
[0011] Current and / or voltage in the vehicle electrical system and / or on vehicle electrical system components;
[0012] the charge status of the accumulator; and / or
[0013] Loads on the generator and / or DC voltage converter and / or electrical consumers.
[0014] The energy management system can be designed to determine measured values for one or more state variables at specific times. The measured values can then be used to determine the state of the onboard energy system at the specific time. Based on the state at the specific time, actions can then be determined and initiated. Examples of actions include:
[0015] Changing (in particular increasing or decreasing) the current and / or voltage in the vehicle electrical system and / or on vehicle electrical system components; and / or
[0016] Changing (in particular increasing or reducing) the load on the generator and / or the DC voltage converter and / or the load on the electrical consumer.
[0017] The energy management system may include a neural network trained (via reinforcement learning) that receives measured values of the one or more state variables as input values and provides actions to be initiated as output values. Alternatively or additionally, the neural network may be configured to provide a Q value (ascertained via Q-learning) for a pair consisting of the measured values of the one or more state variables and an action. Based on the Q values for a plurality of different possible actions, the action that produces the best (e.g., the largest) Q value may then be selected.
[0018] The above process can be repeated at a series of time points in order to continuously control and / or regulate the onboard energy system. In this case, the corresponding current state can be measured at each time point and an action (such as the action resulting in the corresponding optimal Q value) can be determined based on this.
[0019] The energy management system may have been trained for a reference onboard power system, wherein the reference onboard power system may correspond to an energy onboard power system having fault-free and / or undamaged onboard power system components.
[0020] An energy management system based on machine learning can include at least one controller configured to regulate a measurable (state) variable of the onboard energy system to a target value. The reward or reward function used in training the energy management system (particularly its neural network) can depend on the deviation of the (measured) actual value of the measurable variable from the target value during operation of a reference onboard energy system or during operation of the onboard energy system. The smaller the deviation of the actual value from the target value, the greater the reward. This allows for precise regulation of one or more state variables of the onboard energy system.
[0021] The device includes a reference unit configured to determine a reference reward for a state of the onboard electrical system and for an action initiated by the energy management system based on the state, the reference reward being generated during operation of the reference onboard electrical system. The reference unit may have been trained within the context of a training process for the machine-learning energy management system, in particular based on rewards determined for different combinations of states and actions within the context of the training process for the machine-learning energy management system.
[0022] Thus, a reference unit can be provided that displays, for a state-action pair, the rewards that would result when operating a reference onboard power system (with intact onboard power system components). The reference unit can include at least one neural network (trained within the context of a training process of the energy management system).
[0023] The device also includes a reward unit configured to determine an actual reward for the state and for the action (i.e., for a state-action pair), the actual reward resulting (actually) during operation of the onboard energy system. The actual reward and the reference reward can be determined based on the same reward function.
[0024] As already described above, the reward or reward function, in particular the actual reward and the reference reward, can depend on one or more measurable (state) variables of the onboard electrical system. In particular, the reward or the reward function, and thus the actual reward and the reference reward, can include one or more reward components for the corresponding one or more measurable (state) variables of the onboard electrical system.
[0025] The reward unit can be configured to determine measured values for one or more measurable (state) variables, which result from actions initiated during the operation of the energy onboard power system. An actual reward can then be determined based on the measured values for the one or more measurable variables. Similarly, rewards can also be determined during the training of the energy management system and the reference unit (during the operation of the reference onboard power system).
[0026] The device also includes a monitoring unit configured to monitor the energy vehicle network based on an actual reward and a reference reward, in particular based on a comparison of the actual reward with the reference reward. The monitoring unit can be configured to determine, based on the actual reward and the reference reward, whether an onboard network component of the energy vehicle network is damaged. Furthermore, the device can be configured to output an indication (e.g., a fault notification) regarding the onboard network component if damage is determined to be present.
[0027] The apparatus described herein enables comparison of the (actual) rewards generated for a state-action pair during operation of an energy onboard power system with the (reference) rewards generated for the state-action pair during operation of a corresponding, fault-free reference onboard power system. This enables efficient and precise monitoring of the energy onboard power system. In particular, premature and / or impending failures of onboard power system components can be accurately predicted.
[0028] As described above, the reference reward and / or the actual reward can each include one or more reward components. Exemplary reward components include: a reward component related to the current and / or voltage within the vehicle electrical system and / or on a vehicle electrical system component; a reward component related to the load and / or loading of the vehicle electrical system component; and / or a reward component related to the charge level of an energy storage device in the vehicle electrical system.
[0029] By taking into account different reward components for different onboard power supply components, the accuracy of monitoring the energy onboard power supply system can be further increased. In particular, specific onboard power supply components at risk of failure can be precisely identified.
[0030] The monitoring unit can be configured to determine a deviation between an actual reward and a reference reward, in particular a deviation between a reward component of the actual reward and a corresponding reward component of the reference reward. Based on this deviation, in particular by comparing it with a deviation threshold, it can then be determined precisely whether an onboard power supply component is defective. The deviation threshold can be determined in advance through simulations and / or tests, in particular for a plurality of different onboard power supply components and / or a corresponding plurality of different reward components.
[0031] The actual reward and the reference reward (i.e., in particular, a common reward function) can each include a reward component for a specific onboard power supply component. The monitoring unit can be configured to determine whether the specific onboard power supply component is defective based on a deviation of the reward component of the actual reward for the specific onboard power supply component from the reference reward for the specific onboard power supply component. Thus, by comparing the individual reward components, defective onboard power supply components of the energy onboard power supply (which are at risk of failure) can be identified in a particularly precise manner.
[0032] According to a further aspect, a (road) motor vehicle, in particular a passenger car or a truck or a bus or a motorcycle, is described, which comprises a device as described herein.
[0033] According to another aspect, a method for monitoring an onboard energy system is described. The onboard energy system includes various onboard energy system components and is operated using a machine-learning energy management system. The energy management system has been trained using reinforcement learning on a reference onboard energy system. The method includes determining a reference reward for the state of the onboard energy system and for actions initiated by the energy management system based on this state, the reference reward being generated during the operation of the reference onboard energy system. Furthermore, the method includes determining an actual reward for the state and for the actions, the actual reward being generated during the operation of the onboard energy system. The method also includes monitoring the onboard energy system based on the actual reward and the reference reward, in particular based on a comparison of the actual reward with the reference reward.
[0034] According to another aspect, a software (SW) program is described which can be configured to be executed on a processor (for example, on a control unit of a vehicle) and thereby to execute the method described herein.
[0035] According to another aspect, a storage medium is described. The storage medium may include a SW program that is configured to be executed on a processor and thereby to implement the method described herein.
[0036] It should be noted that the methods, apparatuses, and systems described herein can be used not only individually but also in combination with other methods, apparatuses, and systems described herein. Furthermore, each aspect of the methods, apparatuses, and systems described herein can be combined with one another in a variety of ways. In particular, the features of the present invention can be combined with one another in a variety of ways. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The present invention will be described in more detail below with reference to an embodiment. In the accompanying drawings:
[0038] Figure 1a An exemplary vehicle network is shown;
[0039] Figure 1b An exemplary control loop is shown;
[0040] Figure 2a An exemplary neural network is shown;
[0041] Figure 2b An exemplary neuron is shown;
[0042] Figure 3 An exemplary apparatus for training a regulator is shown;
[0043] Figure 4 An exemplary device for determining the state of an onboard electrical system of a vehicle is shown; and
[0044] Figure 5 A flow chart shows an exemplary method for ascertaining the state of an onboard electrical system of a vehicle. DETAILED DESCRIPTION
[0045] As explained in the introduction, this document relates to reliable and precise prediction of the status of components of a vehicle's onboard electrical system. FIG1 shows a block diagram of an onboard electrical system 100, which includes an electrical energy storage device 105 (e.g., a lithium-ion battery), one or more electrical consumers 106, and / or a generator 107. Furthermore, onboard electrical system 100 includes an energy management system 101, which is configured to regulate one or more (state) variables of onboard electrical system 100, in particular to corresponding setpoint values. An exemplary (state) variable of onboard electrical system 100 is onboard electrical system voltage 111, which can be regulated, in particular to a specific target value, for example.
[0046] A controller may be used to adjust a manipulated variable (eg, onboard electrical system voltage 111 ) to a (time-varying) desired value. Figure 1b An exemplary control loop 150 is shown, in which a manipulated variable 156 is adjusted to a setpoint value 151 by means of a controller 153. Controller 153 is configured to determine an actuating variable 154 based on a control error 152 (i.e., a difference) formed between manipulated variable 156 and the (individually current) setpoint value 151. One or more actuators of onboard electrical system 100 (e.g., generator 107 and / or one or more electrical consumers 106) can be operated based on actuating variable 154. An exemplary actuating variable 154 is the speed at which generator 107 is operated (e.g., by the vehicle's internal combustion engine). Within a control path 155 dependent on the characteristics of onboard electrical system 100, a manipulated variable 156 (e.g., a value of a state variable of onboard electrical system 100) is derived from actuating variable 154.
[0047] One possibility for efficient and flexible regulation and / or adaptation of the controller 153 is training of the controller 153 or of the control function by means of one or more neural networks. Figure 2a and Figure 2b An exemplary component of a neural network 200, in particular a feedforward network, is shown. In the example shown, the network 200 includes two input neurons or input nodes 202, which each record the current value of an input variable at a specific time t as input value 201. The input node or nodes 202 are part of an input layer 211.
[0048] The neural network 200 also includes neurons 220 in one or more hidden layers 212 of the neural network 200. Each neuron in the neurons 220 can have individual output values of neurons in the previous layers 212, 211 (or at least a portion thereof) as input values. Processing is performed in each neuron 220 to determine an output value of the neuron 220 based on the input values. The output values of the neurons 220 in the last hidden layer 212 can be processed in the output neurons or output nodes 220 of the output layer 213 to determine one or more output values 203 of the neural network 200.
[0049] Figure 2b An exemplary signal processing within a neuron 220, in particular within a neuron 202 of one or more hidden layers 212 and / or output layer 213, is shown. Input values 221 of a neuron 220 are weighted by individual weights 222, so that a weighted sum 224 of the input values 221 is determined in a summation unit 223 (this determination is performed, if necessary, taking into account a bias or offset 227). Weighted sum 224 can be mapped to an output value 226 of the neuron 220 via an activation function 225. The activation function 225 can, for example, limit the value range. For example, a sigmoid function, a hyperbolic tangent (tanh) function, or a rectified linear unit (ReLU), such as f(x)=max(0,x), can be used as activation function 225 for a neuron 220. The value of the weighted sum 224 can optionally be offset by an offset 227.
[0050] Thus, neurons 220 have weights 222 and / or possibly biases 227 as neuron parameters. Neuron parameters of neurons 220 of neural network 200 may be trained during a training phase so that neural network 200 approximates a specific function and / or models a specific behavior.
[0051] The training of the neural network 200 can be performed, for example, using a backpropagation algorithm. For this purpose, in the first phase of the qth epoch of the learning algorithm, an output value 203 corresponding to the input value 201 at the input node 202 of the neural network 200 can be determined at the output of the one or more output neurons 220. Based on the output value 203, the value of an optimization function or error function can be determined. In the present case, a temporal difference (TD) error can be used as the optimization function or error function, as explained further below.
[0052] In the second phase of the qth epoch of the learning algorithm, the error or error value is back-propagated from the output to the input of the neural network in order to change the neuron parameters of neurons 220 layer by layer. The error function determined at the output can be derived in part from each individual neuron parameter of neural network 200 in order to determine the magnitude and / or direction for adapting the individual neuron parameters. The learning algorithm can be iteratively repeated for a plurality of epochs until a predefined convergence criterion and / or abort criterion is reached.
[0053] To train the controller 153 or the control function provided for determining the control variable 154 based on the control error 152, so-called (Actor-Critic) reinforcement learning can be used, for example. As another example, Q-learning can be used. Within the scope of Q-learning, a Q-function can be trained (and approximated, for example, by a neural network 200), wherein the Q-function can be used to select the optimal action for a specific state.
[0054] Figure 3 An exemplary device 300 is shown for training a control function 303 for a controller 153, in particular for an energy management system 101. The control function 303 can be approximated, for example, by a neural network 200. Alternatively or additionally, the control function 303 can be described by an analytical function with one or more control parameters. An exemplary control function is:
[0055] u t =π(x t )=kx t ,
[0056] Here, k is a vector with one or more control parameters, and x is the value of a state vector at time t, wherein the values are the values of one or more state variables 306 of the state of onboard electrical system 100. Exemplary state variables 306 are the charge level of energy storage 105, onboard electrical system voltage 111, the load of generator 107, the load of electrical load 106, etc.
[0057] The value of the one or more state variables 306 can indicate the deviation of the respective state variable from the corresponding setpoint value 301. In this case, the value x t A value representing one or more adjustment errors.
[0058] The control function 303 is called "Actor" in the context of (Acer-Critic) reinforcement learning. The control function 303 can be used to determine the current value u of one or more execution parameters or action parameters or actions 304 based on the current value of the one or more state variables 306. tExemplary executed variables or action variables or actions 304 are the required load of the generator 107 and / or the load caused by the electrical load 106 .
[0059] The current value u of the one or more execution parameters or action parameters 304 t It can be used to operate the system to be regulated or the control path 305. In particular, one or more components 106, 107 of the vehicle electrical system 100 can be used to operate the system to be regulated or the control path 305 according to the current value u of the one or more execution variables or action variables 304. t This results in the value x of the one or more measurable state variables 306 at the subsequent time t+1. t+1 .
[0060] Based on the current value x of the one or more measurable state variables 306 t and based on the current value u of the one or more execution parameters or action parameters 304 t The value function can be evaluated. The value function can correspond to the discounted sum of (future) rewards. At each time point t, a reward r(x t ,u t ). For example, the reward can depend on
[0061] How much to adjust the state of charge of the energy storage device 105 to a determined target state of charge; and / or
[0062] • How close the load drawn by the generator 107 is to the determined target load.
[0063] Reward r(x t ,u t ) can have different reward terms or reward components for different control variables and / or state variables 306. The individual reward components can be summarized as a reward vector. The reward r(x t ,u t The current value 302 of ) (ie, the reward function) may be calculated by unit 307 .
[0064] The regulation function 303 can be trained so that the sum of the discounted rewards over time is increased, in particular maximized. Since, due to the unknown regulation path 305, it is not known how the action or execution variable 304 influences the value x of the one or more (measurable) state variables 306, t (ie, the value of the adjustment error), so the state action value function 308 can be trained as a "Critic", which is the state action value function for the state x of the system 305 to be adjusted (ie, the vehicle network 100) t and action u tEach combination of 304 indicates the value Q of the sum of the rewards discounted over time π (x t ,u t )310.
[0065] On the other hand, a state value function can be defined that is t denotes the discounted reward r(x i ,u i )
[0066]
[0067] where the discount factor γ∈[0, 1]. Here, we can assume that V π (x t+1 )=Q π (x t+1 ,u t+1 ),
[0068] Among them, u t+1 =π(x t+1 ), the trained regulation function is π()303.
[0069] The value function can be trained iteratively over time, where after convergence, the value that should be applied is Q π (x t ,u t )=r(x t ,u t )+γV π ( t+1 ),
[0070] As long as convergence has not yet been achieved, the so-called timing difference (TD) error δ 311 can be calculated based on the above formula (e.g., in unit 309) as
[0071] δ=r(x t ,u t )+γV π (x t+1 )-Q π (x t ,u t ),
[0072] Among them, the TD error δ311 is based on the assumption
[0073] V π (x t+1 )=Q π (x t+1 ,u t+1 )
[0074] The reward value r(x t ,u t ) 302 and the value Qπ(x t ,u t ), Qπ(x t+1 ,u t+1 ) 310 is calculated. The reward value 302 can be provided for this purpose in unit 309 (not shown). The TD error δ 311 can be used to iteratively train the state-action-value function 308 and, if necessary, the regulation function 303. In particular, the TD error δ 311 can be used to train the state-action-value function 308. The trained state-action-value function 308 can then be used to train the regulation function 303.
[0075] The state-action-value function 308 can be approximated and / or modeled by the neural network 200 and adapted based on the TD error δ 311. After the state-action-value function 308 is adapted, the control function 303 can be adapted. The apparatus 300 can be configured to iteratively adapt the control function 303 and / or the state-action-value function 308 for a plurality of time points t until a convergence criterion is reached. Thus, the control function 303 for the controller 153 can be determined in an efficient and precise manner.
[0076] In a corresponding manner, multiple controllers 153 can be trained for multiple manipulated variables or state variables 106 of a machine learning-based energy management system 101 .
[0077] As described above, combined Figure 3 The described method is only one example for training the energy management system 101 by means of reinforcement learning. Another example is Q-learning. In this case, a Q-function or state-action-value function 308 can be trained, which is configured to provide a Q-value for a state-action pair (wherein the Q-value represents, for example, the sum of discounted future rewards). Starting from the current state, corresponding Q-values can be determined for a plurality of different possible actions according to the trained Q-function. The action can then be selected for which the best (e.g., maximum) Q-value results.
[0078] Figure 4 The device 450 is shown, which enables monitoring of the states of the various components 105, 106, 107 of the onboard electrical system 100, in particular to enable early prediction of failures of the components 105, 106, 107. The device 450 comprises a reference unit 400, which is configured to, based on the current state x of the onboard electrical system 100, t 306 and based on the action u caused by the energy management system 101 t 304 to seek (reference) reward r(xt ,u t ) 302, in particular a (reference) reward vector, which is generated when the onboard electrical system 100 is functioning properly. The reference unit 400 can be trained during a training process of (the controller 153 of) the energy management system 101. For this purpose, the reference unit 400 can include the neural network 200.
[0079] The device 450 further includes a reward unit 410, which is configured to reward the user based on the current state x of the vehicle network 100. t 306 and based on the action u caused by the energy management system 101 t 304 to obtain the actual reward r(x t ,u t ) 402, in particular, an actual reward vector, the actual reward is actually generated when the vehicle network 100 is running (when the reward function is used).
[0080] Reference reward 302 and actual reward 402 can be compared with one another in a comparison and / or checking unit 420. In particular, individual vector magnitudes of the reference reward vector can be compared with corresponding vector magnitudes of the actual reward vector. Based on this comparison, a state 405 of onboard electrical system 100 can then be determined. In particular, based on this comparison, it can be predicted whether a component 105, 106, 107 of onboard electrical system 100 will fail within a specific upcoming time interval. If necessary, the component 105, 106, 107 that is about to fail can also be identified.
[0081] Thus, herein, diagnosis and / or failure prediction of components 105, 106, 107 in a vehicle with machine learning-based energy management is described. The energy management system 101 can be trained, for example, using reflection-enhanced reinforcement learning (RARL) or reinforcement learning. RARL is described, for example, in Heimrath et al., “Reflection-enhanced reinforcement learning for automotive power management operational strategies,” published in the 2019 IEEE International Conference on Computing, Electronics, and Communications Engineering, pp. 62-67. The contents of this document are hereby incorporated by reference in their entirety.
[0082] Here, an agent, such as deep neural network 200, learns which actions 304 to perform in a certain normal state 306 of vehicle electrical system 100. Action 304 can include, for example, increasing electrical system voltage 111. After performing action 304, state 306 of electrical system 100 changes, and the agent receives feedback in the form of a reward (i.e., reward 302) regarding the quality of its decision (i.e., the resulting action 304). This reward 302 flows through the learning process of energy management system 101. The state of electrical system 100 can include multiple state variables 306, such as the utilization rate of generator 107, the (normalized) current flowing into or out of energy storage 105, the state of charge of energy storage 105 (particularly the state of charge SOC), the temperature of energy storage 105, etc.
[0083] Herein, a diagnostic module 450 is described that is at least partially integrated into the learning process of the energy management system 101. In particular, the reference unit 400 can be trained during the learning process of the energy management system 101. This allows the diagnostic module 450 to predict and / or quantify fault behaviors and / or failures of the components 105, 106, 107 of the onboard electrical system 100. The diagnostic module 450 can be configured (using the reference unit 400) to predict the expected effect of performing the action 304 in the onboard electrical system 100. Furthermore, the diagnostic module 450 can be configured (within the comparison or checking unit 420) to compare the expected behavior of the onboard electrical system 100 with the actual behavior of the onboard electrical system 100.
[0084] During the training of energy management system 101 in (reference) onboard electrical system 100 having functional components 105, 106, 107, reference unit 400 is trained in parallel and independently (using separate neural network 200). During the training of energy management system 101, the agent, starting from a certain state 306, performs an action 304 and receives a reward 302 for doing so. This information is used to train reference unit 400 so that it can predict the expected effect of performing action 304 in functional (reference) onboard electrical system 100 as reward 302.
[0085] After the training of energy management system 101 and / or during use during vehicle operation, the training of reference unit 400 is also completed. During the operation of energy management system 101, energy management system 101 selects an action 304 based on a specific state 306 and performs this action 304. Based on the actually measured variables of onboard electrical system 100, an actually measured reward 402 can then be determined (in reward unit 420).
[0086] The reward 302 predicted by reference unit 400 (for a fault-free (reference) onboard power system 100) can then be compared with the actually measured reward 402 (for onboard power system 100). The respective rewards 302, 402 can have different reward components, which can be compared in pairs. The comparison of the actually measured reward 402 and the predicted reward 302 indicates whether one or more components 105, 106, 107 of onboard power system 100 have an actual behavior that deviates from the target behavior.
[0087] For the differences in the values of rewards 302, 402 and / or individual reward components, tolerances and / or thresholds can be determined, from which a fault is detected. The tolerances and / or thresholds can be ascertained within the scope of simulations and / or based on tests on a vehicle (e.g., within the scope of development).
[0088] If a fault is detected based on the comparison of rewards 302, 402, a notification can be issued to the driver of the vehicle and / or a repair organization for repairing the vehicle. If necessary, the driver of the vehicle can be prompted to manually take over driving the vehicle if it is recognized that automated driving operation is no longer possible due to the detected fault (e.g., a fault in energy storage device 105).
[0089] For example, the reward 302, 402 can be a function of the battery current entering or leaving the energy accumulator 105 and / or the utilization rate of the generator 107. If, for example, the reward 302, 402 is a function of the battery current and the functional availability status of the generator 107 is known (e.g., based on sensor data from a dedicated sensor), the deviation of the calculated predicted reward 302 from the actual reward 402 can be used as a quantitative indicator of the functional availability of the energy accumulator 105. On the other hand, the functional availability of the generator 107 can be inferred from the deviation of the rewards 402, 402 with respect to the utilization rate of the generator 107.
[0090] If reward 302, 402 is a function of battery current and generator utilization, the weighting of these two influencing variables within reward 302, 402 can be used to explain the deviation of predicted reward 302 from actual reward 402. In particular, the weighting of the individual reward components can be used to determine which onboard power supply component 105, 106, 107 is defective.
[0091] Figure 5A flow chart shows an exemplary (possibly computer-implemented) method 500 for monitoring an onboard power system 100 (of a motor vehicle), which includes various onboard power system components 105 , 106 , 107 (e.g., an electrical energy storage device 105 , one or more electrical consumers 106 , and / or a generator 107 ) and is operated using a machine-learning energy management system 101. Energy management system 101 can be trained using reinforcement learning using a reference onboard power system. In this case, the reference onboard power system 100 can correspond to the energy onboard power system 100 in the case where the energy onboard power system 100 has only fault-free and / or undamaged onboard power system components 105 , 106 , 107 .
[0092] Method 500 includes determining 501 a reference reward 302 for the state 306 of the onboard electrical system 100 and for an action 304 initiated by the energy management system 101 based on the state 306 , which would be generated during operation of the reference onboard electrical system. A reference unit 400 can be used for this purpose, which is designed to provide a reference reward for each state-action pair. Reference unit 400 can already be trained during the training of the energy management system 101 using reinforcement learning.
[0093] Furthermore, method 500 includes ascertaining 502 an actual reward 402 for state 306 and for action 304 , which actual reward is generated during operation of onboard energy system 100 . Actual reward 402 may be ascertained based on measured values of one or more measured variables for onboard energy system 100 .
[0094] Method 500 further includes monitoring 503 onboard energy system 100 based on actual reward 402 and reference reward 302, in particular based on a comparison of actual reward 402 and reference reward 302. In this case, in particular, defective onboard energy system components 105, 106, 107 can be identified based on actual reward 402 and reference reward 302.
[0095] The measures described herein can improve the comfort and safety of an onboard energy system 100 (for a vehicle). This allows for a prediction of whether the onboard energy system 100 or its components 105 , 106 , 107 no longer meet the requirements placed on the onboard energy system 100 . Predicted deteriorated components 105 , 106 , 107 can thus be repaired or replaced in advance, before damage to the onboard energy system 100 becomes noticeable. Furthermore, the measures described can provide a secure energy supply for autonomous vehicles. Furthermore, quality fluctuations in the components 105 , 106 , 107 of the onboard energy system 100 can be detected.
[0096] The present invention is not limited to the embodiments shown. In particular, it should be noted that the description and drawings are intended only to illustrate the principles of the proposed methods, devices, and systems.
Claims
1. A device (450) for monitoring an onboard energy network (100), the onboard energy network comprising different onboard network components (105, 106, 107) and operated with the aid of a machine-learning energy management system (101); wherein: The energy management system (101) has been trained by means of reinforcement learning for a reference onboard network, which can correspond to an energy onboard network with fault-free and / or undamaged onboard network components; wherein the device (450) comprises: a reference unit (400) configured to determine a reference reward (302) for a state (306) of the energy onboard network (100) and for an action (304) caused by the energy management system (101) based on the state (306), the reference reward being generated when the reference onboard network is operated; a reward unit (410) configured to determine an actual reward (402) for the state (306) and for the action (304), the actual reward being generated during the operation of the energy onboard network (100); and A monitoring unit (420) is configured to monitor the energy onboard network (100) based on the actual reward (402) and based on the reference reward (302).
2. The device (450) according to claim 1, wherein The monitoring unit is configured to monitor the energy onboard network (100) based on a comparison of the actual reward (402) with the reference reward (302).
3. The device (450) according to claim 1, wherein The reference unit (400) has been trained within the scope of a training process of the machine learning energy management system (101).
4. The device (450) according to claim 3, wherein The reference unit (400) is trained based on rewards (302) that have been generated for different combinations of states (306) and actions (304) within the scope of a training process of the machine learning energy management system (101).
5. The device (450) according to any one of claims 1 to 4, wherein - the reference reward (302) and / or the actual reward (402) respectively include one or more reward components; and - the one or more reward components include: - a bonus component that is dependent on the current and / or voltage in the energy onboard power system (100) and / or on the onboard power system components (105, 106, 107); - a bonus component that is related to the load and / or the load of the onboard network component (105, 106, 107); and / or - a bonus component that is dependent on the charge level of the energy storage device (105) of the onboard energy system (100).
6. The device (450) according to claim 5, wherein The device (450) is configured to: - calculating the deviation of the actual reward (402) from the reference reward (302); and - Based on the deviation, it is determined whether the onboard electrical system component (105, 106, 107) is damaged.
7. The device (450) according to claim 6, wherein The monitoring unit (420) is configured to: - calculating the deviation of the actual reward (402) from the reference reward (302); and - Based on the deviation, it is determined whether the onboard electrical system component (105, 106, 107) is damaged.
8. The apparatus (450) of claim 6, wherein: The device (450) is configured to: - determining the deviation of the reward component of the actual reward (402) from the corresponding reward component of the reference reward (302); and / or - determining whether an onboard electrical system component (105, 106, 107) is defective by comparison with a deviation threshold value.
9. The device (450) according to claim 8, wherein - the actual reward (402) and the reference reward (302) each include a reward component for a specific onboard network component (105, 106, 107); and The device (450) is configured to determine whether the determined onboard power supply component (105, 106, 107) is damaged based on a deviation of a reward component of an actual reward (402) for the determined onboard power supply component (105, 106, 107) from a reward component of a reference reward (302) for the determined onboard power supply component.
10. The device (450) according to claim 9, wherein The monitoring unit (420) is configured to determine whether the determined onboard power supply component (105, 106, 107) is damaged based on a deviation of a reward component of an actual reward (402) for the determined onboard power supply component (105, 106, 107) from a reward component of a reference reward (302) for the determined onboard power supply component.
11. The device (450) according to claim 9 or 10, wherein The deviation threshold value is determined in advance by simulation and / or by testing.
12. The device (450) according to claim 11, wherein The deviation threshold value is determined in advance, specific to a plurality of different onboard power supply components (105, 106, 107) of the onboard power supply system (100) and / or specific to a corresponding plurality of different reward components.
13. The apparatus (450) of claim 5, wherein: - the actual reward (402) and the reference reward (302) depend on one or more measurable parameters of the energy onboard network (100); and - the reward unit (410) is configured to: - determining measured values for the one or more measurable variables, the measured values being generated as a result of actions (304) caused during the operation of the onboard power system (100); and - Determining the actual reward based on the measured values for the one or more measurable variables (402).
14. The apparatus (450) of claim 13, wherein: The actual reward (402) and the reference reward (302) include one or more reward components for one or more corresponding measurable variables of the onboard energy system (100).
15. The device (450) according to any one of claims 1 to 4, wherein The machine-learning energy management system (101) comprises at least one controller (150) which is designed to regulate a measurable variable of the onboard energy system (100) to a setpoint value; and The actual reward (402) and the reference reward (302) depend on a deviation of the actual value of the measurable variable from the setpoint value during operation of the reference onboard power system or during operation of the energy onboard power system (100).
16. The device (450) according to any one of claims 1 to 4, wherein The reference unit (400) includes at least one neural network (200).
17. The device (450) according to any one of claims 1 to 4, wherein The device (450) is configured to: - determining, based on the actual reward (402) and based on the reference reward (302), whether the onboard network component (105, 106, 107) is damaged; and - when it is determined that the onboard network component (105, 106, 107) is damaged, outputting an indication about the onboard network component (105, 106, 107).
18. The apparatus (450) of claim 17, wherein: The monitoring unit (420) is configured to: - determining, based on the actual reward (402) and based on the reference reward (302), whether the onboard network component (105, 106, 107) is damaged; and - when it is determined that the onboard network component (105, 106, 107) is damaged, outputting an indication about the onboard network component (105, 106, 107).
19. A method (500) for monitoring an onboard energy network (100), said onboard energy network comprising different onboard network components (105, 106, 107) and said onboard energy network being operated with the aid of a machine learning energy management system (101); The energy management system (101) has been trained by means of reinforcement learning for a reference onboard network, which can correspond to an energy onboard network with fault-free and / or undamaged onboard network components; wherein the method (500) comprises: - determining (501) a reference reward (302) for the state (306) of the energy onboard network (100) and for an action (304) caused by the energy management system (101) based on the state (306), the reference reward being generated when the reference onboard network is operated; - determining (502) an actual reward (402) for the state (306) and for the action (304), the actual reward being generated during the operation of the energy onboard network (100); and - monitoring (503) the energy on-board network (100) based on the actual reward (402) and based on the reference reward (302).
20. The method according to claim 19, wherein The method (500) includes monitoring (503) the energy onboard network (100) based on a comparison of the actual reward (402) with the reference reward (302).