Uncertainty quantification of an artificial intelligence model

FR3154523B1Active Publication Date: 2025-09-05VITESCO TECHNOLOGIES GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2023011527
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-10-24
Publication Date
2025-09-05
Estimated Expiration
2043-10-24

AI Technical Summary

Technical Problem

Existing prediction processes in artificial intelligence systems lack quantification of uncertainty, leading to unreliable decisions in real-time scenarios, while physical theoretical models are costly and complex, or imprecise and tedious to calibrate.

Method used

A process combining an automatic learning model with a physical theoretical model to predict a technical system's state, using confidence intervals to determine the reliability of predictions, switching between models based on confidence interval thresholds and critical values.

Benefits of technology

Provides reliable predictions with reduced treatment time by ensuring precision through model switching, leveraging confidence intervals and physical models to enhance decision-making reliability.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to a method for predicting a state of a technical system embedded in a motor vehicle from real data, said method comprising the steps: (a) receiving (EA) said real data from said sensor of the technical system; (c) predicting (EB) a first value representative of the state of the technical system and a confidence interval associated with said predicted value, using a first model, said first model corresponding to a machine learning model; (d) determining (ED) whether the length of the confidence interval is greater than a threshold value or whether said interval includes a critical value; and (e) if not, defining (EE) a final predicted value as being said first predicted value, or, if so, predicting (EE') a second value representative of the state of the technical system using a second model, said final predicted value being defined as being said second predicted value.Abstract Figure: Figure 2.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Quantification of uncertainty in an artificial intelligence model

[0001] This document relates to a method for predicting the state of a technical system from real data from at least one sensor of said technical system. State of the art

[0002] In automotive technology, artificial intelligence (AI) enables vehicles to achieve significant levels of efficiency, safety and automation.

[0003] Real-time embedded artificial intelligence systems are becoming increasingly common in various fields, including autonomous vehicles, healthcare, and industrial automation. This recent emergence is due to their ability to process large amounts of data and make complex decisions in real time, combined with the advent of the Internet of Things and advances in the computing capabilities of embedded systems. This capability is driving the development of autonomous vehicles, advanced driver assistance systems (ADAS), and predictive maintenance technologies.

[0004] As artificial intelligence systems make increasingly autonomous decisions, the reliability of these decisions becomes a significant issue. Trust in these systems depends largely on their ability to accurately predict outcomes and react appropriately in real time.

[0005] Quantifying the uncertainty of predictions made by artificial intelligence system models is an essential aspect of reliability. The term "uncertainty" of a model refers to the level of confidence in the predictions made by that model. It is a measure of what the model "knows" or "does not know".

[0006] In scenarios where decisions must be made in real time, understanding the uncertainty of a prediction made by an artificial intelligence model can make the difference between a correct and timely decision and a catastrophic failure.

[0007] Prior art methods for predicting the state of a technical system based on a model of an artificial intelligence system are known. These methods have the advantage of being able to be carried out efficiently with a low prediction cost.

[0008] However, these processes are carried out without quantifying the uncertainty of the predictions and therefore without verifying the reliability of said predictions.

[0009] It is also known from the prior art of methods for predicting the state of a system The technique is based on a theoretical physical model. These processes offer very good reliability, but are costly in terms of processing time.

[0010] Furthermore, methods for predicting the state of a technical system based on a simplified theoretical physical model are also known in the prior art, since the theoretical physical model is either not easily embedded or too complex in terms of processing time and / or computing power. However, this simplified theoretical physical model lacks precision and is tedious to calibrate.

[0011] The objective of this document is to couple in the same process predictions according to a model of an artificial intelligence system and a theoretical physical model following the quantification of the uncertainty of the predictions. Description of the invention

[0012] To this end, the present document relates to a method for predicting the state of a technical system embedded in a motor vehicle from real data obtained from at least one sensor of said technical system, said method comprising the following steps:

[0013] (a) receive said actual data from said sensor of the technical system,

[0014] (c) predict a first representative value of the state of the technical system and an in confidence interval associated with said predicted value, using a first model, said first model corresponding to a machine learning model,

[0015] (d) determine if the length of the confidence interval is greater than a value threshold, or if said interval includes a critical value, and

[0016] (e) in the negative, define a final predicted value as being said first predicted value, or, if so, predict a second value representing the state of the technical system using a second model, said final predicted value being defined as said second predicted value.

[0017] Step (a) may include a step of putting the technical system into operation and measuring said real data using the sensor(s).

[0018] The first model may be a model of an artificial intelligence system.

[0019] The first model can be a statistical model or more generally a model machine learning.

[0020] The life cycle of such a model is classically divided into three main phases: the training phase, the testing phase and the inference phase.

[0021] The training phase aims to adapt a model to a training dataset so that it can learn to perform a specific task (e.g., classification, regression). To this end, the model is exposed to a training dataset and adjusts its internal parameters to minimize error on this data. The training dataset is generally a subset of the available data. It is important that it be representative of the complete dataset.

[0022] The testing phase aims to evaluate the performance of the trained model on test data, not seen during training, to estimate how the model will behave in real-world situations. To this end, after training the model, it is evaluated on a separate test dataset, not used during training. Performance metrics are calculated, and these metrics may depend on the nature of the model. The test dataset is a separate subset of the available data, distinct from the training dataset.

[0023] The inference phase (also called the prediction phase) aims to use the trained model to make predictions or inferences on new data.

[0024] During inference, the model parameters are fixed and not modified. In other words, the model does not learn any new information during this phase.

[0025] The second model can be a physical representation model.

[0026] The second model can be a theoretical model describing the behavior of the technical system. For example, the second model can describe the temperature of a battery or the current of a battery.

[0027] The second model can be based on: - Expert rules: Expert rules are instructions or guidelines derived from experience and human expertise. They can be used to guide decision-making or the evaluation of the behavior of a technical system. - Nomograms: Nomograms are graphs or tables that allow for the rapid identification of solutions to specific problems based on certain variables. They are often used to solve engineering problems without the need for complex calculations, and - Simple regression models: Simple linear or logistic regression models can be used to establish relationships between variables and predict the behavior of a technical system based on these variables.

[0028] Actual data refers to an observation that has actually been collected. It is authentic data originating from a concrete source and reflecting real events, behaviors, or phenomena. This actual data can include a variety of data types, including time series or other types of data from sensors, etc.

[0029] A technical system is an interconnected set of components, devices or processes, whether mechanical, electronic, electrical, computer-related, or of a different nature, working together to perform a specific function. This technical system may be equipped with sensors. This technical system is capable of providing information or a signal that indicates the state of this technical system.

[0030] Below is defined, by way of example, a non-exhaustive list of physical quantities or properties that can be measured by sensors in technical systems, as well as output information that can be provided by different technical systems (information represented by the signal from the sensor): - mechanical systems: position (encoders, position sensors), speed (speed sensors, revolution counters), acceleration (accelerometers), force (force sensors, dynamometers), pressure (pressure gauges, pressure sensors), deformation (deformation sensors), vibration (vibration sensors), temperature (thermometers), rotation (rotation sensors), level (level sensors) - Electronic or electrical systems: voltage (voltmeters), current (ammeters), resistance (ohmmeters), frequency (frequency meters), power (wattmeters), power factor (power factor indicators), frequency response curve (spectrum analyzers), ripple (oscilloscopes), threshold voltage (threshold detectors), luminous intensity (photometers) - Computer systems or software: data (data streams), status (system state, errors), execution time (timestamp), CPU usage (performance monitors), memory used (memory monitors), latency (latency measurements), network throughput (traffic analyzers), events (event logs) - Environmental systems: ambient temperature (thermometers), humidity (hygrometers), atmospheric pressure (barometers), noise level (sound level meters), air quality (air quality sensors), solar radiation (pyranometers), wind speed (anemometers), precipitation (rain gauges) - Control and automation systems: commands (control signals), states (system states), alarms (anomaly alerts), errors (error messages), feedback signals (system feedback), switching states (states of switches, relays), activation states (states of motors, actuators).

[0031] The technical system can be, for example, a battery.

[0032] A confidence interval (CI) is a range of values ​​within which one estimates that an unknown parameter of a probability distribution lies with a certain predefined probability. It is used to express the uncertainty associated with a statistical estimate (random uncertainty) and / or a model (epistemic uncertainty).

[0033] More specifically, a confidence interval is constructed, for example, around a point estimate (such as the mean, median, variance, etc.) to indicate the margin of error or uncertainty associated with that estimate. It is expressed as an interval; for example, a 95% confidence interval (CI) indicates that the prediction is estimated to be within that interval with 95% confidence. Confidence intervals can thus be useful for assessing the reliability of an estimate or prediction. The narrower the confidence interval, the more precise the estimate.

[0034] Step (d) refers to certain conditions or criteria evaluated in statistics or data analysis, specifically with regard to confidence intervals. In particular, the length of a confidence interval is the difference between the upper and lower bounds of that interval. This provides a measure of the precision or certainty of the estimate. The threshold value is a pre-specified value that serves as a criterion for evaluating whether the length of the confidence interval is acceptable. If the length of the interval is greater than this threshold value, it may indicate that the estimate is not precise enough. A critical value is an unacceptable value. This may indicate a prediction error in the model. For example, in the case of a battery, a critical value may be an excessively high current through the battery.

[0035] Step (d) determines whether the first or second model should be used to provide the predicted final value. In particular, if the first model fails, i.e., if the length of the confidence interval exceeds a threshold value or if the interval includes a critical value, then the second model will be used instead of the first model to predict the predicted final value.

[0036] Such a method makes it possible to determine the final predicted value with satisfactory reliability while offering a low processing time.

[0037] Said method may further include a step (b) carried out between steps (a) and (c), said step (b) comprising substep (bl) consisting of verifying that the real data are consistent with training data of the first model.

[0038] Consistency between real data and training data refers to the similarity or compatibility of distributions, characteristics, and trends between these two datasets. Consistency between these datasets is useful to ensure that the model trained on the training data can effectively generalize to new real data.

[0039] For example, consistent data can be understood as included data in the interval [ / / — 3*cr ; / / + 3*cr] where is the mean and ° is the standard deviation.

[0040] Step (b) may further include a substep (b2) carried out following substep (b1), said substep (b2) consisting of: - if the actual data is consistent, preprocess the actual data; and - if at least some of the real data are inconsistent, predicting a second representative value of the state of the technical system using a second model, a final predicted value being defined as said second predicted value, and stop the process.

[0041] Preprocessing can be normalization. Data normalization is a technique used to change the values ​​of the numeric columns in a dataset to a common scale, without distorting the differences in value ranges or the relationships between features. Normalization is useful to ensure that each feature contributes equally during model training. Several methods can be used to achieve such normalization.

[0042] The preprocessing can be a calculation of metrics from the real data.

[0043] The normalization of substep (b2) can, for example, be a normalization according to the following methods: - Min-Max normalization: the data is transformed by resizing all numerical values ​​to a scale between 0 and 1 (or -1 to 1). The formula for this normalization is as follows: X_normalized = (X - X_min) / (X_max - X_min) - Z-score normalization (or standardization): the data is transformed by subtracting the mean of each characteristic, then dividing by the standard deviation. The formula for this normalization is as follows: X_normalized = (X- [i)i &

[0044] The removal of inconsistent data (also called outliers) can be carried out using the variance (and therefore the standard deviation). This often involves using the distance in terms of the number of standard deviations from the mean.

[0045] For example, assuming that the data follow a normal distribution, the removal of outliers may include the successive steps of: - calculate the mean (^) and variance (<7 2 ) of the dataset. The variance is the mean of the squared deviations from the mean; - calculate the standard deviation (SD) which is the square root of the variance; - Identify outliers using a threshold based on the standard deviation. A common threshold is two or three standard deviations above and below the mean. In other words, any value outside the interval / / ± k*a (where k (where x is the number of standard deviations) can be considered an outlier. In other words, the value is an outlier if x < / 7 - k* <? ou si x < pi + k*<r ; - remove from the dataset or replace values ​​identified as outliers with limit values ​​such as / / ± k *cr ; and - Optionally, it is possible to recalculate the mean, variance, and standard deviation to check the consistency of the dataset.

[0046] The threshold value can be between 5% and 15% of the first value representing the state of the technical system.

[0047] The first model can be a Gaussian process model.

[0048] A Gaussian process model is a probabilistic model used in the field of machine learning, particularly for regression and classification. It is a non-parametric model based on the theory of stochastic processes.

[0049] In particular, a Gaussian process is a collection of random variables, where every finite subset of these variables has a joint (normal) Gaussian distribution. Formally, a Gaussian process is defined by a mean function and a covariance function.

[0050] The mean function gives the average of the Gaussian distribution at each point, while the covariance function (or kernel) gives the covariance between the values ​​of two points. The kernel describes how the correlation between the data decreases with distance.

[0051] In a Gaussian process model, the inference phase is carried out by examining the posterior distribution of the functions, given the observed data. This allows predictions to be made with measures of uncertainty.

[0052] Gaussian process models provide a measure of uncertainty with their predictions. They are also very flexible and can model a wide range of functions simply by choosing an appropriate kernel. Finally, they can adapt to the complexity of the data by learning the covariance function.

[0053] The first model can be a Bayes neural network.

[0054] A Bayesian neural network, or Bayesian neural network, is a fusion of the concepts of neural networks and Bayesian inference. Bayesian neural networks are distinguished by their ability to operate efficiently with less data while improving generalization. This characteristic is crucial, particularly in situations where data is limited or expensive to obtain. They allow the construction of complex architectures capable of learning hierarchical representations from data. Unlike traditional neural networks that use point parameters, Bayesian neural networks employ probability distributions to represent uncertainties in the parameters. network parameters. In other words, instead of having a fixed value for each weight in the network, a Bayesian neural network has a probability distribution for each weight, thus reflecting the uncertainty about the value of these weights.

[0055] One of the main advantages of Bayesian neural networks lies in their ability to provide measures of uncertainty, which can be very useful in many fields, where understanding uncertainty is crucial.

[0056] The document "Hands-On Bayesian Neural Networks - A Tutorial for Deep Learning Users" (Jospin et al.) describes Bayesian neural networks.

[0057] The first model can be a Monte Carlo dropout type model.

[0058] It is recalled that regularization is a technique used in machine learning to prevent overfitting, that is to say to help the model to generalize well from training data to unseen data.

[0059] It is also recalled that dropout is a common regularization technique which consists of randomly "switching off" certain neurons during training, which forces the network to learn more robust representations.

[0060] Typically, Monte Carlo Dropout is a technique that extends the regular dropout method used during neural network training to the testing phase. The basic concept of Monte Carlo Dropout is an implementation where the use of regular dropout can be interpreted as a Bayesian approximation of a well-known probabilistic model, the Gaussian process.

[0061] The main idea is to generate random predictions and interpret them as samples of a probabilistic distribution, thus allowing a Bayesian interpretation of the predictions.

[0062] During the test phase, dropout is applied, which means that some neurons are switched off (or "dropped out") randomly during each forward pass through the network.

[0063] Each dropout configuration corresponds to a different sample of the posterior distribution over the network weights, and the process is repeated several times to obtain a prediction distribution.

[0064] In other words, when dropout is applied during the testing phase, it generates different predictions on each pass, depending on which neurons have been turned off. This makes it possible to obtain a distribution of predictions, rather than a single prediction, from which statistical measures such as the mean and standard deviation can be calculated.

[0065] This prediction distribution can be used to quantify the uncertainty associated with the model predictions.

[0066] This technique aims to represent the uncertainty of the model in deep learning by allowing a Bayesian approximation, which is crucial for obtaining uncertainty estimates on the model's predictions. One of the advantages of Monte Carlo Dropout is that it provides a simple and computationally efficient way to obtain uncertainty estimates.

[0067] Furthermore, Monte Carlo dropout requires fewer computing resources than Bayesian neural networks. However, it is only capable of capturing epistemic uncertainty.

[0068] The document "Dropout as a Bayesian Approximation: Representing Model Un-certainty in Deep Leaming" (Gai et al.) describes a Monte Carlo Dropout technique.

[0069] The first model can be formed by a set of learning sub-models.

[0070] Such a model thus uses a set method, or ensemble method, which is an approach which consists of combining the predictions of several sub-models (called "learners") to improve the overall performance of the model.

[0071] The fundamental idea behind set methods is that combining several models can compensate for the individual weaknesses of each model, thus resulting in better generalization and a reduction in prediction error.

[0072] In the present document, this combination allows the confidence interval to be calculated.

[0073] For example, the most commonly used set-theoretic methods are bagging, boosting or stacking.

[0074] A wide variety of basic learning submodels can be used. In particular, it is possible to use learning submodels of the following types: decision trees, neural networks, linear or polynomial regressions, support vector machines (SVMs), K-nearest neighbors (K-NNs), Bayes networks, time series models (ARIMA or recurrent neural network models, RNNs).

[0075] The first model can be a Bayesian Physics-Informed Neural Network (also called Bayesian Physics-Informed Neural Networks (B-PINNs)).

[0076] Physics-informed neural networks are neural networks that use physical data to improve their predictions and reliability. These neural networks are designed to incorporate known physical laws or equations into their architecture. This allows them to provide data-driven solutions to physical problems, while ensuring that these solutions remain consistent with known physics.

[0077] B-PINNs combine elements of Bayesian deep learning and Physics Informed Neural Networks (PINNs). Training is performed with physical data, and the inference phase is Bayesian; Bayes' methods are used to determine the uncertainty of the predictions.

[0078] The invention also relates to a computer program comprising instructions for implementing the method according to the aforementioned type, when this program is executed by a processor.

[0079] The invention also relates to a non-transient computer-readable recording medium on which a program is recorded for the implementation of the aforementioned method, when this program is executed by a processor.

[0080] The invention also relates to a computer system comprising

[0081] - an input interface for receiving data representative of a system technical,

[0082] - a memory for storing at least the instructions of a computer program of the aforementioned type,

[0083] - a processor accessing memory to read said instructions and execute then the process of the aforementioned type,

[0084] - an output interface to provide the final predicted value of step (e).

[0085] The invention may also relate to a motor vehicle comprising a computer system of the aforementioned type.

[0086] Of course, this same process can be used in other technical fields. Brief description of the drawings

[0087] Other features, details and advantages will become apparent upon reading the detailed description below, and upon analysis of the attached drawing showing various figures, in which: - [Fig. 1] schematically represents an example of a computer system according to this document. - [Fig.2] illustrates the different stages of the process according to one embodiment of this document, and - [Fig.3] is a graph of the error as a function of the number of iterations training based on a method of predicting the state of a technical system. Detailed description

[0088] Figure 1 schematically represents a computer system 1 comprising

[0089] - an input interface 2 for receiving data from at least one sensor of a technical system and representative of a technical system,

[0090] - a memory 3 for storing at least the instructions of a program computer,

[0091] - a processor 4 accessing memory 3 to read said instructions and execute then the process illustrated in [Fig.2],

[0092] - an output interface 5.

[0093] Reference is now made to [Fig.2] which illustrates the different stages of the process according to an embodiment of the present document.

[0094] This method aims to predict a state of a technical system, for example a battery of a motor vehicle, from real data from at least one sensor of said technical system, using a first model and a second model.

[0095] This method includes an EA step in which real data are measured using at least one sensor of the technical system and received by the input interface 2. The sensor provides, for example, a time series of voltage, current or temperature values, in the case of a battery.

[0096] Then, during an EB step, it is checked (EB1 step) whether the real data are consistent with training data.

[0097] If the actual data are consistent (result Y), then this actual data is preprocessed (step EB2). Conversely, if at least some of the actual data are not consistent (result N), a second value representing the state of the technical system is predicted using a second model and the process is stopped (step EB2').

[0098] Then, during an EC step, a first representative value of the state of the technical system is predicted and is associated with a confidence interval, using the first model.

[0099] Then, during an ED step, it is determined whether the length of the confidence interval is greater than a threshold value or whether said interval includes a critical value.

[0100] If the length of the confidence interval is not greater than a threshold value and if said interval does not include a critical value (result N), then the final predicted value is said first predicted value (value predicted by the first model - step EE).

[0101] Conversely, if the length of the confidence interval is greater than a threshold value or if said interval includes a critical value, a second value representing the state of the technical system is predicted using the second model (result Y), said final predicted value is defined as this second predicted value (step EE').

[0102] As mentioned previously, the first model can be a Gaussian process model, a Bayes neural network, or a set of learning submodels. The first model can also use a Monte Carlo dropout regularization technique.

[0103] As previously also indicated, the second model can be a theoretical model describing the behavior of the technical system based, for example, on expert rules, nomograms or simple regression models.

[0104] Fig. 3 represents the evolution of the error as the training step progresses, in three scenarios.

[0105] The graph in [Fig. 3] has the number of iterations as its x-axis and the y-axis as its y-axis. number of prediction errors.

[0106] In the first case (Clin curve) only the first model is used, the use of the second model being disabled.

[0107] In the second case (curve C2), the transition to the second model is carried out with a high threshold value of the length of the confidence interval.

[0108] Finally, in the third case (curve C3), the transition to the second model is carried out with an average threshold value of the length of the confidence interval.

[0109] It is observed that the use of the second model makes it possible to significantly reduce the prediction error in the case of the process according to the invention. It is also observed that the choice of the threshold has an influence on such an error.

Claims

Claims

1. Method for predicting a state of a technical system embedded in a motor vehicle from real data from at least one sensor of said technical system, said method comprising the steps: (a) receiving (EA) said real data from said sensor of the technical system, (c) predicting (EC) a first value representative of the state of the technical system and a confidence interval associated with said predicted value, using a first model, said first model corresponding to a machine learning model, (d) determining (ED) whether the length of the confidence interval is greater than a threshold value or whether said interval includes a critical value, and (e) if not, defining (EE) a final predicted value as being said first predicted value, or, if so, predicting (EE') a second value representative of the state of the technical system using a second model,said final predicted value being defined as said second predicted value.,

2. Method according to the preceding claim which further comprises a step (b) (EB) carried out between step (a) and (c), said step (b) comprising the sub-steps consisting of: (bl) (EB1) checking that the real data are consistent with training data of the first model, and (b2) (EB2) if the real data are consistent, preprocessing the real data; and if at least some real data are not consistent, predicting a second value representative of the state of the technical system using a second model, a final predicted value being defined as being said second predicted value, and stopping the method.

3. Method according to one of the preceding claims, in which the second model is a physical representation model.

4. Method according to one of the preceding claims, in which the threshold value is between 5% and 15% of the first value representative of the state of the technical system.

5. Method according to one of claims 1 to 4, in which the first model is a Gaussian process model.

6. Method according to one of claims 1 to 4, in which the first model is a Bayesian neural network.

7. Method according to one of claims 1 to 4, in which the first model is a Monte Carlo dropout type model.

8. Method according to one of claims 1 to 4, in which the first model is formed by a set of learning sub-models.

9. Method according to one of claims 1 to 4, in which the first model is a Bayesian neural network of the Physics Informed type.

10. Computer program comprising instructions for implementing the method according to one of the preceding claims, when this program is executed by a processor.

11. A non-transitory computer-readable recording medium on which is recorded a program for implementing the method according to one of claims 1 to 9, when this program is executed by a processor.

12. Computer system (1) comprising - an input interface (2) for receiving data representative of a technical system, - a memory (3) for storing at least the instructions of a computer program according to claim 10, - a processor (4) accessing the memory (3) to read said instructions and then execute the method according to one of claims 1 to 9, - an output interface (5) for providing the final predicted value of step (e).