Method and device for providing a data-based prediction model taking into account dynamic properties
The method improves prediction accuracy in data-based models by extending the loss function to include temporal property terms, addressing the limitations of conventional models in capturing long-term system behavior without increasing model complexity.
Patent Information
- Application Number
- DE102023211149
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-15
AI Technical Summary
Conventional data-based prediction models, such as NARX models, struggle to accurately predict system behavior over longer time horizons due to limited consideration of temporal properties and increased dimensionality of input tensors.
A method for training data-based prediction models that extends the loss function to include additional terms accounting for temporal properties, such as signed differences and variance of model outputs within a training time window, to improve prediction accuracy without increasing the model's complexity.
The proposed method enhances prediction accuracy for future time steps by effectively capturing temporal properties beyond the immediate dynamic behavior of the system, while maintaining or reducing the complexity of the prediction model.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical field
[0001] The invention relates to a method for providing prediction values using a data-based prediction model for mapping a physical system, e.g., to implement a virtual sensor or perform anomaly detection. The invention further relates to a suitable training method for the data-based prediction model, in particular to consider dynamic properties of the physical system in the prediction variable to be modeled. Technical background
[0002] A data-driven prediction model can, for example, be designed as a NARX model based on a neural network. Such NARX models are approaches that use data-driven models for dynamic applications in which the prediction model predicts on time series data.
[0003] For evaluation, a NARX model is fed input values for a number of input variables for a current time step and input values for at least one input variable for one or more previous time steps. Furthermore, if necessary, the output value of the prediction model, which models a prediction value for a subsequent time step, can be traced back to an input of the prediction model. Typically, such a data-based sensor model is trained in a supervised manner, i.e., on a series of training data sets containing short time series of input variables up to a specific time step and labeled with an actual prediction value of the variable to be modeled for a subsequent time step.
[0004] By using the NARX model, it is possible to consider temporal dependencies of various physical effects. However, conventional supervised training of the data-based sensor model, e.g., using a loss value based on an average error, such as the mean square error, log likelihood, or mean absolute error, fails to take certain system properties into account.
[0005] To predict system behavior, an extended approach is therefore necessary that allows for the modeling of certain time-related properties. However, increasing the temporal horizon of the input variables considered on the input side of the prediction model is generally limited, as this would significantly increase the dimensionality of the prediction model's input tensor.
[0006] It is therefore an object of the present invention to provide an improved prediction model for predicting one or more prediction variables of one or more future time steps, wherein the complexity of the prediction model can be maintained or even reduced while improving the prediction accuracy. Disclosure of the invention
[0007] This object is achieved by the method for providing a data-based prediction model for at least one prediction variable according to claim 1 and by a corresponding device according to the independent claim.
[0008] Further embodiments are specified in the dependent claims.
[0009] According to a first aspect, a method is provided for providing a data-based prediction model for modeling a technical system, wherein the prediction model is designed to provide a prediction value for a time step following a time step to be evaluated as a function of an input tensor (such as an input vector), wherein the input tensor comprises, in addition to the value of the at least one input variable for the time step to be evaluated, at least one value of the at least one input variable for one or more time steps prior to the time step to be evaluated within a predetermined time horizon, wherein the prediction model is trained in training steps based on a respective total loss, in particular with a gradient-based training method, wherein the total loss for each training step is determined with the following steps: - Providing training data sets, each formed from time series of state variables of the technical system, wherein a training data set comprises an input tensor and a label value, wherein the input tensor has at least one value of at least one of the state variables at a specific reference time and at least one value of at least one of the state variables at a time step prior to the specific reference time, wherein the label value depends on a state variable at a time step following the reference time; - Determining the total loss based on a plurality of training data sets of a training time window with successively recorded values of the state variables, wherein a main term is determined by an average error, wherein a further term is provided which depicts a temporal property of the time series of the input variables within the training time window.
[0010] Predictive models can be used in many areas. For example, models for virtual sensors used in control systems or models for anomaly detection rely on determining a corresponding state value of a state variable, which indicates the state of a physical system, for one or more subsequent time steps in advance. This enables improved control or earlier anomaly detection.
[0011] For example, recurrent models or NARX models can be used to capture temporal dynamics. However, these approaches are very complex in terms of training and computational effort, and they yield high variability in the prediction value.
[0012] According to the above method for providing a prediction model, it is preferably provided to use a NARX model that takes into account values of one or more input variables with a number of previous time steps and also uses the output state value of the at least one state variable for the next time step for the evaluation in the next time step of the input variable. The training of the data-based prediction model is carried out in a conventional manner based on a loss value, which can be determined depending on labeled training data. The loss value can be used for gradient-based training, such as backpropagation.
[0013] Training is based on training data sets consisting of time series of input variables whose temporal progression ends at a specific reference time (reference time step) and which are assigned a label. The label or label value in the training data sets specifies the target value of the model output and can correspond to one of the input variables or a different state variable, and is related to a time step following the reference time.
[0014] Traditionally, the resulting loss value for each training dataset is derived from an average error, such as a mean-squared approach, which specifies the difference between the label value and at least one prediction value of the prediction model for the consecutive time steps. The training datasets are preferably selected randomly from the time series of the training data, so that during consecutive training, the training datasets are chosen as much as possible to avoid temporally consecutive training datasets. However, this destroys the sequential structure of the training data.
[0015] To take into account temporal properties that depend on temporal developments that consider periods that are larger than can be covered by the temporal horizon of the previous time steps, at least one additional term can be taken into account in addition to the main term above to determine the loss value during the training of the data-based prediction model.
[0016] It can be provided that a first additional term is determined depending on the training data set that determines the value of the main term, wherein the first additional term is determined from the training data sets of the training time window depending on a sum of the signed differences between the model output and the label value of the training data sets of the training time window. In particular, the result of the summation can be subjected to a first normalization function, which can in particular comprise absolute value formation (absolute value formation) or squaring.
[0017] A first additional term can be determined depending on the training data set, which determines the value of the main term. The additional term is formed from training data sets of the same time series as the training data set, which are selected from a defined training time window that ends with the reference time of the training data set for determining the main term. The length of the training time window is preferably chosen such that it covers a longer period than that represented by the time steps of the input variables used on the input side of the data-based prediction model. The first additional term of the loss value is determined after a model evaluation of the number of training data sets in the training time window.
[0018] Furthermore, the first additional term can correspond to a sum of the differences between the model output and the label value of the considered training dataset for all training datasets of the training time window. The result of the summation can be subjected to a first normalization function, which can correspond to an absolute value (absolute value) or a squaring. Summing the differences allows for the positive and negative deviations in the prediction to be compensated.
[0019] It can be provided that a second further term is determined from the training data sets of the training time window, wherein the second further term takes into account or corresponds to a variance of the model outputs of the evaluation of the prediction model for the training data sets of the training time window.
[0020] A second, additional term can be determined analogously to the first, depending on the training data set that determines the value of the main term. The second, additional term is formed from training data sets of the same time series as the training data set, which are selected from a defined training time window that ends with the reference time of the training data set for determining the main term. The length of the training time window is preferably selected such that it covers a longer period than is represented by the time steps of the input variables used on the input side of the data-based prediction model. The second, additional term of the loss value is determined after a model evaluation of the number of training data sets in the training time window.
[0021] The second additional term can include an unsupervised term that takes into account or corresponds to the variance of the model outputs from the prediction model's evaluation for the training data sets of the training time window. The variance of the model outputs from the prediction model's evaluation can also be fed to a second normalization function, so that a high variance of the prediction values within a training time window is penalized during training, i.e., increases the overall loss.
[0022] Furthermore, the total loss can be determined by a weighted sum of the value of the main term and a value of at least the first additional term and the second additional term.
[0023] According to one embodiment, the training time window may comprise a larger number of time steps than the number of time steps of the time horizon.
[0024] It can be provided that the training time windows for the training steps are randomly selected from available acquisition sequences of state variables. The training time windows are thus randomly selected from the time series of the training data sets, with each of the training time windows taking the available training data sets into account for determining the loss value.
[0025] According to a further aspect, a use of a prediction model provided by the above method is provided as a virtual sensor for a temperature estimation of a component of a technical system, in particular an electric motor, depending on time series of state variables of the technical system, in particular an electric current, a movement speed and / or a drive torque or a drive force, in order to model a temperature value inside the component, in particular a rotor of the electric motor.
[0026] According to a further aspect, an apparatus for carrying out the above method is provided. Brief description of the drawings
[0027] Embodiments are explained in more detail below with reference to the attached drawings. They show: Fig. 1 a schematic representation of a NARX model for predicting a prediction value for a first time step; and Fig. 2 a flowchart illustrating a procedure for training the data-based prediction model. Description of embodiments
[0028] Fig. Figure 1 shows a schematic representation of a data-based prediction model, designed as a NARX model 1, for evaluating time series of a (perhaps arbitrary) number of input variables E1, E2, E3. The architecture of the NARX model is based on a deep neural network. Such time series of input variables can be provided, for example, by technical systems and represent measured variables and / or state variables of a technical system.
[0029] The input variables are selected from the state variables of the technical system. To evaluate temporal dynamics, values of the input variables E1, E2, E3 for the current time step t are fed to the data-based prediction model. Furthermore, for at least one of the input variables, one or more input variable values from previous time steps t-1 ... tn can be applied to the prediction model, where n can be specified as the time horizon. Thus, for one or more of the input variables, not only the value of the current time step but also values for one or more previous time steps t-1 ..., tn can be considered in the input tensor of the data-based prediction model. tn ... t represents the temporal horizon of the temporal dynamics considered in the input tensor in the prediction model.
[0030] The prediction model 1 can be designed to determine, for a prediction variable P(t+1) to be modeled, the value of the state variable to be predicted for the next time step following the reference time or a difference value between the value of a state variable to be predicted for the current time step and a value of the state variable to be predicted for the next time step t+1 following the reference time, which indicates a development of the prediction variable P.
[0031] Typically, such a NARX model is designed to consider a time horizon of n time steps for the one or more input variables E1, E2, E3. This time horizon is sufficient to represent the immediate dynamic behavior of the system, but keeps the dimensionality of the input tensor within limits. While a simple temporal dynamics of a system can be represented by a relatively small number of consecutive values of the input variables E1, E2, E3, there are longer-term temporal properties of the system that cannot be described by the purely dynamic behavior. To represent these temporal properties, and to avoid increasing the time horizon and thus the dimensionality of the input tensor, the training procedure for the data-based prediction model 1 provides for expanding the loss function for training the data-based prediction model.To train the data-based prediction model, a procedure is carried out as shown in the flow chart of the . Fig. 2 is described in more detail. The method is carried out using a conventional data processing device.
[0032] In step S1, time series of input variables E1, E2, E3 are acquired, which may include measured variables and / or state variables and correspond to physical variables or model variables. Furthermore, a synchronous time series of a prediction variable to be modeled by the prediction model, which may correspond to an input-side state variable or another state variable, is provided.
[0033] In step S2, training data sets are first prepared from the time series provided in step S1. Each of the training data sets refers to a reference time point (reference time step). According to the time horizon for which the input side of the prediction model is designed, an input tensor is formed that includes not only the current values of the input variables for the reference time point but also the historical values of the input variables considered in prediction model 1.The input tensor is assigned as a label value the value of the input variable to be predicted or the state variable to be predicted for the next time step t+1 following the reference time t or a difference value of a value between the value of the input variable to be predicted or the state variable to be predicted for the current time step and a value of the input variable to be predicted or the state variable to be predicted for the next time step t+1 to be determined, modeled state variable or output variable for the time step following the reference time.
[0034] The label can also be extended to several subsequent time steps.
[0035] These training data sets correspond to the input variables E for the above example 1 (tn ... t), E2 (tn ... t) E 3(tn ... t), where the label value corresponds to the quantity to be modeled P(t+1) and n specifies a time horizon with which dynamic effects in the technical system are taken into account by the prediction model 1.
[0036] According to the acquired time series and the respective reference time t, temporal sequences of training data sets can be created over a common acquisition sequence (a series of measurements of a certain time duration that is longer than the time horizon n) of input variables / state variables and label values.
[0037] First, in step S3, training data sets are randomly selected from the time series with reference times for which historical values of the input variables / state variables and label values exist that extend back (into the past) for a specified training time window (with a specified number of time steps) that is larger than the time horizon considered by the prediction model with the input tensor. Preferably, the duration of the acquisition sequence is an integer multiple of the duration of the training time window.
[0038] Randomization can be achieved by randomly selecting training time windows from the acquisition sequences. Another loss term can, for example, correspond to the variance of the model outputs within the training time window.
[0039] For training the data-based prediction model, a total loss value is determined in step S4, which is used in a backpropagation process to train the data-based prediction model.
[0040] The total loss value can include an average error determined in a conventional manner as the main term. The average error can be specified as a mean squared error (Euclidean norm), which indicates the average squared error between the model output and the corresponding label value, or as a sum norm. The average error is calculated over the successive time steps of the training time window.
[0041] Furthermore, a first additional loss term can be considered additively, resulting from the summed signed differences between the model output and the label value for the training data sets of the considered training time window. Furthermore, the resulting sum can be subjected to a first normalization function.
[0042] A second additional loss term can be provided that takes into account or corresponds to the variance of the model outputs from the evaluation of the prediction model for the training data sets of the training time window. The variance of the model outputs from the evaluation of the prediction model can also be fed into a second normalization function to obtain the value of the second additional term. With the variance, it should be provided that a high variance of the model outputs in a training time window should lead to a higher overall loss. In this way, a high variance of the prediction values within a training time window can be penalized during training, as it increases the overall loss. The second additional loss term is unsupervised and can be used to detect sequences of state variables that must be assumed to be anomalies due to their high variance.
[0043] By choosing the training time window with a number of time steps that is larger than the time steps considered on the input side of the data-based prediction model, temporal properties or dependencies can be taken into account that go beyond the pure dynamics of the technical system.
[0044] The additional loss terms can be added to the main term (e.g. additively), in particular with a respective weighting, in order to obtain a total loss.
[0045] In step S5, a training step based on the total loss now takes place.
[0046] The procedure for performing further training steps is then continued with step S4 until a known termination criterion is met.
[0047] The above prediction model can be used in many technical fields. For example, it can be used as a virtual sensor for temperature estimation, e.g., of components of an electric motor, depending on time series of motor state variables, such as motor current, speed, torque, and the like, to model a temperature value inside a motor component, such as the rotor. This allows the motor temperature to be monitored so that appropriate throttling or shutdown measures can be taken in the event of overheating, i.e., when a temperature limit is reached.
[0048] The prediction model is particularly suitable in control systems as an observer model for the variable to be controlled.
Claims
[1] Computer-implemented method for providing a data-based prediction model (1) for modeling a technical system, wherein the prediction model (1) is designed to provide a prediction value for a time step following a time step to be evaluated as a function of an input tensor, wherein the input tensor comprises, in addition to the value of the at least one input variable (E1, E2, E3) for the time step to be evaluated, at least one value of the at least one input variable (E1, E2, E3) for one or more time steps prior to the time step to be evaluated within a predetermined time horizon, wherein the prediction model is trained (S5) in training steps based on a respective total loss, in particular with a gradient-based training method, wherein the total loss for each training step is determined with the following steps: - Providing (S1, S2) training data sets, each formed from time series of state variables of the technical system, wherein a training data set comprises an input tensor and a label value, wherein the input tensor has at least one value of at least one of the state variables at a specific reference time (t) and at least one value of at least one of the state variables at a time step (tn) prior to the specific reference time (t), wherein the label value depends on a state variable at a time step (t+1) following the reference time (t); - Determining (S4) the total loss based on a plurality of training data sets of a training time window with successively recorded values of the state variables, wherein a main term is determined with an average error, wherein at least one further term is provided which maps a temporal property of the time series of the input variables (E1, E2, E3) within the training time window. [2] Method according to claim 1, wherein a first further term is determined as a function of the training data set which determines the value of the main term, wherein the first further term is determined from the training data sets of the training time window as a function of a sum of the signed differences between the model output and the label value of the training data sets of the training time window, wherein in particular the result of the summation is subjected to a first normalization function which can in particular comprise an absolute value formation (absolute value formation) or a squaring. [3] Method according to one of claims 1 to 2, wherein a second further term is determined from the training data sets of the training time window, wherein the second further term takes into account or corresponds to a variance of the model outputs of the evaluation of the prediction model for the training data sets of the training time window, wherein in particular the variance of the model outputs of the evaluation of the prediction model is additionally subjected to a second normalization function. [4] The method of claim 3, wherein the total loss is determined by a weighted sum of the value of the main term and a value of at least the first further term and the second further term. [5] Method according to one of claims 1 to 4, wherein the training time window comprises a larger number of time steps than the number of time steps of the time horizon. [6] Method according to claim 5, wherein the training time windows for the training steps are randomly selected from available acquisition sequences of state variables (S3). [7] Method according to claims 1 to 6, wherein the total loss is determined by a weighted sum of the value of the main term and a value of at least the first further term and the second further term. [8] The method of claim 1 to 7, wherein the average error is expressed as a mean squared error (Euclidean norm), which indicates the average squared error between the model output and the corresponding label value, as a log likelihood, or as a mean absolute error. [9] Use of a prediction model (1) provided with a method according to one of claims 1 to 8 as a virtual sensor for a temperature estimation of a component of a technical system, in particular an electric motor, depending on time series of state variables of the technical system, in particular an electric current, a movement speed and / or a drive torque or a drive force, in order to model a temperature value of the component, in particular a rotor of the electric motor. [10] Apparatus for carrying out one of the methods according to one of claims 1 to 8. [11] Computer program product comprising instructions which, when the program is executed by at least one data processing device, cause the latter to carry out the steps of the method according to one of claims 1 to 8. [12] Machine-readable storage medium comprising instructions which, when executed by at least one data processing device, cause the device to carry out the steps of the method according to one of claims 1 to 8.