Method for validating and enabling a prediction program that evaluates vehicle data, in particular an on-board prediction program
The method addresses the challenge of validating on-board prediction programs for new-generation vehicles by using a combination of real and simulated data sets, enabling reliable validation and approval even without extensive fleet data, thus ensuring the safety and effectiveness of vehicle assistance functions.
Patent Information
- Application Number
- PCT/EP2024/074466
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-20
- Filing Date
- 2024-09-02
- Publication Date
- 2025-05-30
AI Technical Summary
New-generation vehicles lack comprehensive fleet data for validating and approving on-board prediction programs, which are essential for providing advanced assistance functions while ensuring safety and reliability.
A method is developed to validate and release vehicle data-evaluating prediction programs by using three data sets: real vehicle data from new-generation vehicles, real vehicle data from older-generation vehicles, and simulated vehicle data. The prediction program creates predictions based on these data sets, and if the prediction from the new-generation vehicle data falls within predefined limits of the other two data sets, the program is validated and released.
This method allows for the reliable validation and approval of prediction programs for new-generation vehicles, even when comprehensive fleet data is not yet available, ensuring the safety and effectiveness of assistance functions.
Smart Images

Figure EP2024074466_30052025_PF_FP_ABST
Abstract
Description
[0001] Method for validating and approving a vehicle data-evaluating prediction program, particularly an on-board prediction program
[0002] TECHNICAL FIELD OF THE INVENTION
[0003] The invention relates to a method for validating and releasing a vehicle data evaluating prediction program, in particular an on-board prediction program, for a vehicle of a new generation.
[0004] BACKGROUND OF THE INVENTION
[0005] Today's vehicles have powerful on-board computers that can perform a variety of very different assistance functions, on the one hand to make driving not only more pleasant and comfortable, but above all also significantly safer, and on the other hand to make maintenance processes more efficient and environmentally friendly, for example by not simply carrying out maintenance after a certain mileage, but only when this seems sensible due to the actual condition of the vehicle, which results from the individual driving behavior of a user, among other things.
[0006] Corresponding assistance functions often use so-called "machine learning models", hereinafter referred to as "prediction programs" due to their function, which are able to evaluate vehicle data recorded by the vehicle itself using corresponding sensors, typically supplemented by further data, under certain conditions, particularly on board a vehicle, but also elsewhere, e.g. cloud-based, and thus make predictions about certain events, e.g. the degree of wear of a certain component, which can be relevant on very different levels, e.g. for driving or maintaining the vehicle.
[0007] Prediction programs of the type in question here are generated using so-called "machine learning algorithms" and corresponding training data. They generally perform better—that is, their predictions are more accurate—the more real vehicle data is available for training. Since modern vehicles typically collect a large amount of vehicle data and transmit it to the vehicle manufacturer wirelessly or wired (e.g., during a workshop visit), depending on customer approval, the manufacturer can collect a large amount of data (so-called fleet data) about the respective vehicle type of a particular generation and its behavior, which allows for the creation of highly precise prediction programs.
[0008] Before such predictive programs can be deployed, they must be validated—that is, tested for the reliability of their predictions—and approved for use. Depending on their type, they not only affect safety-relevant aspects but can also contribute significantly to customer satisfaction and brand loyalty. Customers are generally impressed by the reliability of their vehicle if, for example, a precisely functioning program prompts them to visit a workshop, which in retrospect actually proves to be worthwhile.
[0009] DISCLOSURE OF THE INVENTION
[0010] When new-generation vehicles are to be brought to market, the problem arises that, on the one hand, fleet data—that is, large data sets on the real-life everyday behavior of the new vehicles—is lacking because only a few new-generation vehicles are currently in test operation. On the other hand, customers should already be provided with the most comprehensive assistance functions possible. While appropriate approval procedures are established for assistance functions that use software that rarely changes (e.g., radio control), they are not for prediction programs, which frequently improve and are retrained based on new data.
[0011] Based on this, the invention is based on the object of specifying a method for validating and releasing a vehicle data-evaluating prediction program, in particular an on-board prediction program, for a new-generation vehicle. This method enables the validity of a prediction program to be checked and, if necessary, released in a standardized, i.e., comprehensible and uniform manner, even in cases where comprehensive fleet data is not yet available. This object is achieved by a method having the features of claim 1. Advantageous embodiments and further developments are the subject of the dependent claims. The subordinate claim 13 relates to a computer program product for executing certain steps of the method according to the invention.
[0012] In particular, the object is achieved by a method for validating and releasing a vehicle data evaluating prediction program, in particular an on-board prediction program, for a vehicle of a first group, in particular of a new generation, wherein the method comprises the following steps:
[0013] Creating the prediction program to create a prediction based on vehicle data,
[0014] Providing an initial dataset of real vehicle data collected by one or more new generation vehicles,
[0015] Providing a second data set with real vehicle data collected by a plurality of vehicles of a second group, in particular of an older generation,
[0016] Providing a third dataset with simulated vehicle data generated by a computer simulation of the vehicles in the first group,
[0017] Creating a prediction regarding a specific event using the prediction program based on the first data set,
[0018] Creating a prediction regarding the specific event using the prediction program based on the second data set,
[0019] Creating a prediction regarding the specific event using the prediction program based on the third data set,
[0020] Comparing the prediction based on the first data set with the predictions based on the second and third data sets,
[0021] Confirming the validity and releasing the prediction program if the prediction based on the first data set is within predefined limits around the predictions based on the second and third data sets.
[0022] The invention has the advantage of enabling the reliable validation and approval of vehicle data-evaluating prediction programs designed for vehicles whose everyday behavior is still unknown due to their novelty. It should be noted here that the term "prediction program" refers to all types of machine learning models on board a vehicle that allow so-called inferences to be performed—that is, to make predictions about certain events based on data collected by the vehicle, typically supplemented by additional data previously stored in the vehicle.
[0023] The predictions and events can be of a wide variety of types and serve both convenience and safety. For example, predictions can be made about the expected wear and tear of a specific vehicle part, such as a starter battery, and the respective customers can be encouraged to visit a workshop within a certain time or mileage period. Depending on the customer's approval, the vehicle can then also communicate directly with the workshop and / or the manufacturer, for example, to announce a workshop visit or to reserve certain parts or check their availability. The invention can be advantageously used for completely different machine learning models.
[0024] In a method according to the invention for validating and releasing a vehicle data-evaluating prediction program, in particular an on-board prediction program, for a new-generation vehicle, the prediction program is first created to generate a prediction based on vehicle data, typically using a learning algorithm and a training data set. This procedure is well known to those skilled in the field of artificial intelligence and self-learning programs. The learning algorithm used is irrelevant to the invention.
[0025] Three data sets are then provided: a first data set with real vehicle data from a test fleet, collected by one or more new-generation vehicles; a second data set with real vehicle data collected by a large number of older-generation vehicles (e.g., diagnostic data collected during workshop visits); and a third data set with simulated vehicle data generated by a computer simulation of the new-generation vehicles. The terms "new-generation vehicles" and "older-generation vehicles" are to be understood in a use-case-specific manner and not in the traditional sense, according to which a new generation is an improved or at least visually modified version of an existing vehicle type. Rather, "new-generation vehicles" encompass all vehicles in whatever form, e.g.,Vehicles that have been modified through retrofitting or conversion, or that have been created entirely new, for which no large fleet is yet in operation, and for whose behavior no extensive data sets are yet available. Vehicles of an older generation may then, under certain circumstances, also be vehicles of a completely different type, for which fleet data is already available and which share certain features with the new generation vehicles, so that they are likely to behave similarly in certain respects.
[0026] The prediction program then creates predictions regarding a specific event, based on the first data set, the second data set, and the third data set. The term "event" is defined specifically for each prediction program. For example, if maintenance issues are involved, an event could be the question of the condition of a specific component at a specific mileage, driving behavior, (past) environmental conditions, material aging, etc.
[0027] The prediction based on the first data set is then compared with the predictions based on the second and third data sets. A beneficial approach is to calculate a key performance indicator for each prediction and compare the key performance indicators directly or indirectly. An indirect comparison can be performed, for example, by first determining a comparison value from the key performance indicators of the predictions based on the second and third data sets, and then comparing the key performance indicator of the prediction based on the first data set with this comparison value.
[0028] If the prediction based on the first dataset falls within predefined limits around the predictions based on the second and third datasets, the prediction program is released. If it is not released, an error message may be generated indicating that the prediction program needs further improvement, e.g., training with new or supplemented training data.
[0029] The release process can also be such that the prediction program is initially released only partially, namely only for a predefined subset of the new generation of vehicles, which will then be delivered and can collect additional data during everyday operation. This real vehicle data can then advantageously supplement the initial dataset and be used to validate and release further versions of the prediction program.
[0030] Further details and advantages of the invention will become apparent from the following purely exemplary and non-limiting description of an embodiment in conjunction with the drawing.
[0031] SHORT DESCRIPTION OF THE DRAWING
[0032] Fig. 1 shows a highly schematic flow chart of a method according to the invention.
[0033] DESCRIPTION OF PREFERRED EMBODIMENTS
[0034] Figure 1 schematically illustrates the process for validating and releasing a vehicle data-evaluating prediction program, in this case an on-board prediction program for a new-generation vehicle. First, a machine learning model, referred to here as the prediction program 12, is created using training data and a machine learning algorithm 10.
[0035] The prediction program 12 is provided with three data sets 14, 16, and 18, wherein the first data set 14 contains real vehicle data collected by one or more new-generation vehicles (the so-called test fleet), the second data set 16 contains real vehicle data collected by a plurality of older-generation vehicles, and the third data set contains simulated vehicle data generated by a computer simulation of the new-generation vehicles. Such data sets may each contain a variety of data about vehicles and their respective driving histories, which a prediction program may make use of in whole or in part. In order to make a prediction about a specific event, some data may be irrelevant. In practice, the amount of data present in the second data set relating to the specific event may be increased by a factor of between 10 1 and 10 5be greater than the amount of data present in the first data set relating to the specific event, and the amount of data present in the third data set relating to the specific event may be greater by a factor of between 10 1 and 10 5 be larger than the amount of data present in the second data set relating to the particular event.
[0036] The prediction program 12 then creates three predictions regarding a specific event, namely a prediction 20 based on the first data set, a prediction 22 based on the second data set, and a prediction 24 based on the third data set. In programming terms, the predictions are inferences, i.e., conclusions that a machine learning model draws from given data sets. In the example shown in Fig. 1, the predictions based on the second and third data sets are created off-board the vehicle, while the prediction based on the first data set is created on-board a new generation vehicle. The prediction created on-board the vehicle can be sent wirelessly to an off-board data center, where it can be automatically compared with the predictions based on the second and third data sets.
[0037] In the illustrated embodiment, key performance indicators 26, 28, and 30 are calculated for each prediction 20, 22, and 24 in the form of so-called KPIs ("Key Performance Indicators"). These KPIs serve to evaluate the quality of the respective prediction 20, 22, and 24. The basic calculation of such key performance indicators 26, 28, and 30 is known to those skilled in the art. Which key performance indicators are applied in a specific case depends on the type of event. For example, if traffic signs are to be recognized in an image, one key performance indicator could be the percentage of correctly recognized signs. The invention advantageously allows those skilled in the art to select the optimal key performance indicators for the respective application.
[0038] In the illustrated embodiment, in the next step, the performance indicators 28 and 30 of the predictions 22 and 24 based on the second and third data sets 16 and 18 are compared with each other, and an average value, also referred to as a "benchmark," hereinafter referred to as comparison value 32, is formed. In step 34, a check is carried out to determine whether the performance indicators 28 and 30 deviate from each other or, depending on the implementation, from the formed comparison value 32 by more than a predetermined, application-dependent, and user-defined amount, e.g., 10%. If this is the case, the prediction program 12 is fed to the machine learning algorithm 10 for further improvement ("retraining"), as indicated by arrow 36. If the deviation lies within predetermined limits, the comparison value 32 formed from the performance indicators 28 and 30 is compared in step 38 with the performance indicator 26 based on the first data set 14.Depending on the comparison, in step 40, the prediction program 12 is either evaluated as valid and released (arrow 42) or passed on to the machine learning algorithm 10 for further training, as indicated by arrow 44. The criteria for release also depend on the specific application. In the example of traffic sign recognition, it may be specified that release is only granted if there is only a small difference between the stated mean value and the performance indicator 26, while in other cases, higher tolerance limits can be set. Release can be granted automatically, by a higher-level program instance, or manually, by a human decision-maker.
[0039] The released prediction program 12 can then, as indicated by step 46, initially be loaded only onto a predefined subset of new-generation vehicles that can collect additional data during everyday operation. This real vehicle data can, as indicated by arrow 48, advantageously supplement the initial data set and be used to validate and release further versions of the prediction program.
[0040] A specific application example is the monitoring of a wear part in a vehicle. Based on data points from driving behavior / use, external conditions, and expected material aging, a prediction program 12 can estimate the level of individual wear. For example, if the wear is on the starter battery, a message indicating that the battery should be replaced is generated at a certain wear level and displayed, for example, via a user app on a mobile phone and / or in the vehicle itself.
[0041] If a test fleet of vehicles on which the prediction program 12 is running is available, the prediction program 12 can estimate the condition of the starter battery based on various vehicle data (prediction 20). The same prediction program 12 is applied to test and diagnostic data, which was obtained, for example, during workshop visits (prediction 22), and to simulation data from the laboratory (prediction 24). The data type and question are identical for all three inferences. The key performance indicators 26, 28, and 30 calculated from the predictions are compared as described above. If the key performance indicators 28 and 30 deviate from each other by more than a predetermined first limit, e.g., 10%, the prediction program 12 is retrained. If the key performance indicators 28 and 30 deviate by exactly or less than the first limit, e.g.,10% from each other, the mean of the two key performance indicators 28 and 30 is compared with the key performance indicator 26. If the deviation is greater than a second predefined limit, which can correspond to the first limit, e.g. 10%, the prediction program 12 is retrained. If the deviation is equal to or less than the second limit, the prediction program is released, preferably initially only for a larger test fleet. The process is then repeated iteratively until the prediction program 12 has been rolled out to the entire vehicle fleet. With each iteration, the first and second limit values for the deviations can be redefined, typically lower (i.e. 5% instead of 10% in the second iteration).
[0042] The invention was described using the example of the validation and approval of an on-board prediction program that evaluates vehicle data. However, the invention can also be used to validate and approve vehicle data-evaluating prediction programs that do not run on board a vehicle, but rather, for example, are cloud-based.
Claims
PATENT CLAIMS 1. A method for validating and releasing a vehicle data evaluating prediction program, in particular an on-board prediction program, for a vehicle of a new generation, comprising the steps of Creating the prediction program (12) for creating a prediction based on vehicle data, Providing a first data set (14) with real vehicle data collected by one or more new generation vehicles, Providing a second data set (16) with real vehicle data collected by a plurality of vehicles of an older generation, Providing a third data set (18) with simulated vehicle data generated by a computer simulation of the new generation vehicles, creating a prediction (20) regarding a specific event by means of the prediction program based on the first data set, Creating a prediction (22) regarding the specific event by means of the prediction program based on the second data set, Creating a prediction (24) regarding the specific event by means of the prediction program based on the third data set, Comparing the prediction (20) based on the first data set with the predictions (22, 24) based on the second and third data sets, confirming the validity and releasing (40) the prediction program (12) if the prediction (20) based on the first data set is within predefined limits around the predictions (22, 24) based on the second and third data sets.
2. Method according to claim 1, characterized in that in order to compare the predictions (20, 22, 24) with each other, a performance indicator (26, 28, 30) is calculated for each prediction and the performance indicators are compared directly or indirectly with each other.
3. Method according to claim 2, characterized in that for comparing the predictions (20, 22, 24) with one another, a comparison value (32) is determined from the performance indicators (28, 30) of the predictions (22, 24) based on the second and the third data set (16, 18), and the performance indicator (26) of the prediction based on the first data set is compared with this comparison value (32).
4. Method according to one of claims 1 to 3, characterized in that the amount of data present in the second data set (16) with regard to the specific event is reduced by a factor of between 10 1 and 10 5 is greater than the amount of data present in the first data set (14) relating to the specific event.
5. Method according to one of claims 1 to 4, characterized in that the amount of data present in the third data set (18) with regard to the specific event is reduced by a factor of between 10 1 and 105 is greater than the amount of data present in the second data set (16) relating to the specific event.
6. Method according to one of claims 1 to 5, characterized in that the predictions (22, 24) based on the second and third data sets are created externally of the vehicle.
7. Method according to one of claims 1 to 6, characterized in that the prediction (20) based on the first data set is created on board a new generation vehicle.
8. The method according to claim 7, characterized in that the prediction (20) created on board the vehicle is sent wirelessly to a data center external to the vehicle and is automatically compared there with the predictions (22, 24) based on the second and third data sets.
9. The method according to any one of claims 1 to 8, further comprising a step of generating an error message if the prediction (20) based on the first data set (14) does not fall within predefined limits around the predictions (22, 24) based on the second and third data sets (16, 18).
10. The method according to any one of claims 1 to 9, wherein the creation of the prediction program (12) is carried out using a learning algorithm and a training data set.
11. The method according to claim 10, wherein the prediction program (12), if its validity is not confirmed and it is not released, is fed back to the learning algorithm.
12. The method according to any one of claims 1 to 11, wherein the releasing (42) of the prediction program comprises a step of partially releasing the prediction program (12) only for a predefined part of the new generation vehicles.
13. The method according to claim 12, wherein after the partial release of the prediction program (12), the first data set (14) is supplemented by real vehicle data from new generation vehicles that use the partially released prediction program (12).
14. A computer program product for carrying out at least the steps of comparing the prediction (20) based on the first data set (14) with the predictions (22, 24) based on the second and third data sets (16, 18) and confirming the validity of the prediction program (12) in a method according to one of claims 1 to 13.
Citation Information
Patent Citations
Anomaly detection in cyber-physical systems
US20220067535A1
Training configuration-agnostic machine learning models using synthetic data for autonomous machine applications
US20230110713A1
Adaptive model pruning to improve performance of federated learning
US20230177404A1