Method and system for generating a classifier for indirectly measuring physical states of a test object

The method generates a classifier using a classification algorithm trained on operational data to predict the physical states and defects in vehicles and batteries, addressing the challenge of identifying malfunctions in complex products by reducing irrelevant features and improving prediction accuracy.

WO2025245555A1PCT designated stage Publication Date: 2025-12-04AVL LIST GMBH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/AT2025/060218
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-29
Filing Date
2025-05-28
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing methods struggle to accurately determine the physical states of complex products like high-voltage batteries and vehicles, particularly in identifying underlying malfunctions based on symptoms, making it difficult to predict and prevent defects.

Method used

A method and system for generating a classifier using a classification algorithm trained on operational data from multiple devices, involving feature selection methods like Random Ranking, Backward Ranking, Cluster-Based Selection, and Subset Refinement to identify relevant features, which are then used to predict the physical states and potential defects.

Benefits of technology

Enables accurate prediction of physical states and potential defects in vehicles and components like batteries, reducing overfitting and computational demands, allowing for faster training and more effective root cause analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AT2025060218_04122025_PF_FP_ABST
    Figure AT2025060218_04122025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for generating a classifier for indirectly measuring physical states of a test object of a type of battery or vehicle by training a classification algorithm, having the following steps: detecting operating data of a plurality of technical devices of the type, said operating data comprising value curves of time-resolved measurement parameters and characterizing the operating behavior and the environment of a respective technical device from a field operation, and state data of the respective technical device during the field operation; generating features by processing at least some of the operating data by means of mathematical operations and / or by selecting data regions from the operating data; selecting relevant features from the generated features by chaining feature selection methods; and training the classification algorithm using the relevant features and the respective associated state data, the classifier thus being generated. The state data characterizes the physical states of the technical devices, and at least two, preferably three or particularly preferably all of the following feature selection methods are selected: - random ranking (103A); - backward ranking (103B); - cluster-based selection (103C); and - subset refinement (103D), the respective selected feature selection methods being carried out according to the order defined by the aforementioned enumeration.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method and system for generating a classifier for the indirect measurement of physical states of a test object

[0002] The invention relates to a system and a computer-implemented method for generating a classifier for the indirect measurement of physical states of a test object of a type of technical equipment, in particular a physical system (e.g., HV battery) or a vehicle, by training a classification algorithm, wherein operating data of a plurality of technical equipment of the type, which include value profiles of time-resolved measurement parameters and characterize an operating behavior and environment of a respective technical equipment from field operation, and state data of the respective technical equipment during field operation are acquired, wherein the classification algorithm is trained on the basis of these data, thereby creating the classifier.Furthermore, the invention relates to a corresponding computer-implemented classifier and a system and computer-implemented method for applying such a classifier.

[0003] Quality is one of the most important factors influencing a customer's choice between competing products in a given product segment. Therefore, understanding the causal relationship between individual components of a product within the overall product suite and potential defects throughout its entire lifecycle is crucial for the business success of a product and, consequently, of a company.

[0004] An example of the requirements for a product, such as a vehicle, is the lifespan of individual components, such as the battery. This lifespan has a direct impact on the total cost of ownership of a product, which is very important for the end user. In this respect, the product should not exceed a predetermined value.

[0005] To ensure the functionality of a product even after its manufacture and delivery to the customer, it is therefore of great interest to the manufacturer to be able to determine the properties of the product, and in the case of units with multiple technical components, also the properties of individual components within the unit. Such properties are characterized by the physical states of the product.

[0006] The more individual technical components a product or unit comprises, the more important it generally becomes to understand its properties. High-voltage batteries for electric vehicles, or the electric vehicles themselves, are examples of such complex products. With these types of products, it is often difficult to deduce, or even identify, the underlying physical parameter or technical component causing a malfunction, based on a symptom that suggests one or more malfunctions.

[0007] The object of the invention is to provide a system and a method for determining the physical states of a technical product under investigation. In particular, it aims to determine those physical states that are difficult or impossible to measure using previously known methods.

[0008] These tasks are solved through the doctrine of independent claims. Advantageous configurations are claimed in dependent claims.

[0009] A first aspect of the invention relates to a computer-implemented method for generating a classifier for the indirect measurement of physical states of a test specimen of a class of technical devices, which are preferably batteries or vehicles, by training a classification algorithm, comprising the following steps:

[0010] • Acquisition of operational data from a large number of technical devices of the type, which include value profiles of time-resolved measurement parameters and characterize the operational behavior and environment of a respective technical device from field operation, and of state data of the respective technical device during field operation; • Generation of features by processing at least a part of the operational data by means of mathematical operations and / or by selecting data ranges from the operational data;

[0011] • Selection of relevant features from the generated features using a chain of feature selection methods; and

[0012] • Training the classification algorithm using the relevant features and their respective state data, thereby generating the classifier; wherein the state data characterize the physical states of the technical equipment, and wherein at least two, preferably three, or particularly preferably all of the following feature selection methods are selected:

[0013] • Random Ranking (103A);

[0014] • Backward ranking (103B);

[0015] • Cluster-Based Selection (103C);

[0016] • Subset Refinement (103D); and wherein the selected feature selection methods are executed in the order defined by the enumeration above.

[0017] For the purposes of this disclosure, data acquisition preferably involves reading measured operating data via a data interface. Alternatively or additionally, data acquisition includes determining a measurement signal using a sensor and / or post-processing a measurement signal to generate the operating data.

[0018] Measurement parameters within the meaning of the present disclosure preferably include an ambient temperature, battery temperatures, cell temperatures, dwell times in different speed ranges, a number of cold starts, a number of journeys with a certain minimum duration, a number of regenerations, states of charge, charging / discharging processes, battery control parameters, a difference between maximum and minimum voltage per cell, a balancing of the frequency numbers of all cell IDs, a 2D histogram of the thermal management in "idle" mode and a battery contactor temperature above 25 degrees Celsius, particularly during charging, a 2D dwell time histogram of the cell temperature and the SOC based on the HV bus voltage, particularly during charging, a differential coolant temperature during charging processes, and / or a differential cell voltage over the battery's discharge processes.

[0019] A mathematical operation within the meaning of the present disclosure is statistics, in particular mean, median, minimum, maximum or standard deviation, or aggregation, addition, subtraction, multiplication, division, differentiation, gradient calculation and / or integration over certain durations and / or multivariate integrals, in which measurement parameters are multiplied and then integrated.

[0020] A class of technical devices within the meaning of this disclosure is preferably a group of technical devices that are identical in their essential features and are therefore preferably of identical construction. Preferably, the essential components of the technical devices of a class are identical in construction. A specific technical device is thus preferably an implementation of the class of technical devices. In particular, technical devices of a class differ by tolerances, especially manufacturing tolerances and / or aging or wear effects.

[0021] An indirect measurement within the meaning of the present disclosure preferably involves determining a desired measured value of a parameter using other available information from a system to be measured. More preferably, the desired measured value of a parameter is determined from at least one other physical parameter.

[0022] A vehicle within the meaning of the present disclosure may preferably be or comprise a passenger car, a truck, a motorcycle, a construction machine, a ship, an agricultural vehicle, a train or another means of transport.

[0023] The invention enables the creation of a predictive model for the condition of a vehicle from measurement data, particularly historical measurement data. This predictive model can then be used to determine the vehicle's physical condition. Furthermore, the method can also be applied to individual vehicle components, such as a battery, a fuel cell, an engine, a transmission, a braking system, or others.

[0024] The physical states can specify properties of the vehicle component, which are defined as physical quantities or as probabilities for a property.

[0025] The predictive model is designed as a classifier and is trained as follows: Time-resolved, labeled measurement data is imported. Based on this time-resolved measurement data, features are generated ("feature generation" or "feature engineering"). This process primarily involves statistical calculations based on the time-resolved measurement data. From the generated features, the most relevant features—that is, those features that are causally related to the physical state indicated by the labels—are selected ("feature selection"). Using these features and their corresponding labels, a classification algorithm is trained (modeling). The trained algorithm is the classifier.In other words, a classification model is trained using an algorithm such as linear regression (lasso), decision trees, random forest (RF), or XGBoost (XGB) with a small number of estimators, or a generalized linear model (GLM), taking the relevant features into account. This model is preferably evaluated using standard measures.

[0026] The trained classifier is then applied to further time-resolved measurement data from vehicles or vehicle components. This determines the physical states of the vehicles or vehicle components. The time-resolved measurement data is available as aggregated data. If the time-resolved data is available in a different format, it is first aggregated to enable the application of the classifier.

[0027] According to the invention, for an accurate prediction of the physical state of a vehicle or vehicle component, a chain of at least two, three, or four feature selection methods is executed sequentially. This reduces the number of channels with measurement data and the number of features derived or generated from this measurement data, so that only a comparatively small number of relevant features remain for training the classification algorithm. Removing irrelevant or redundant features improves model performance. This reduces overfitting, as the model is not influenced by irrelevant data and can focus more effectively on the important patterns.

[0028] Fewer features mean less data to process. This leads to faster training times and lower demands on storage and computing power. This is particularly important for large datasets and complex models. Furthermore, the remaining features can be used to perform root cause analysis related to a physical state of the device under test. This allows the features, or the underlying time-resolved measurement data, to be examined for anomalies that may have triggered a physical state.

[0029] In a further advantageous embodiment of the procedure, the classification algorithm is a probabilistic classification algorithm, and the procedure further comprises the following steps:

[0030] • Deriving a probability of failure for the respective technical equipment based on the relevant characteristics and the condition data of the multitude of technical equipment, taking into account the probability of failure of the respective technical equipment when training the classification algorithm.

[0031] If the classification algorithm is trained with a sufficiently large number of time-resolved measurement data from various vehicles or vehicle components, a failure probability for the vehicle or vehicle components can be derived and taken into account when training the probabilistic classification algorithm. This allows failure probabilities related to a physical condition of the vehicle or vehicle component to be specified when the trained classifier is applied.

[0032] In a further advantageous embodiment, the method includes the following step: evaluating the computer-implemented classifier using at least one standard measure, in particular from the following group of standard measures: Area under the ROC Curve, Recall, Precision, True Positive Rate.

[0033] An evaluation of the trained classifier can verify its accuracy in predicting physical states.

[0034] According to the invention, at least two, preferably three or all of the following feature selection methods are selected:

[0035] • Random Ranking;

[0036] • Backward Ranking or Backward Elimination;

[0037] • Cluster-Based Selection;

[0038] • Subset Refinement; wherein the selected feature selection methods are executed in the order defined by the enumeration above.

[0039] The selection of features has a significantly greater impact on the quality of the trained classifier than the selection of the classification algorithm. As explained above, the feature selection methods mentioned are executed sequentially in a chain. This allows the most relevant features to be reliably identified and, moreover, significantly reduces the number of features that need to be processed when training the classification algorithm.

[0040] In a further advantageous embodiment of the procedure, the Random Ranking method comprises the following steps:

[0041] • Selecting a random subset of features;

[0042] • Performing multiple, in particular five-fold (k=5), cross-validation for a preliminary machine model with the selected features, determining the relevance of the selected features for each trained preliminary machine model;

[0043] • Store the relevance of the selected features for each preliminary machine learning model trained during cross-validation; • Repeat the preceding steps in several iterations until all features have been selected at least once;

[0044] • Ordering the features according to their average relevance; and

[0045] • Output a defined number of features with the highest relevance to the subsequent feature selection method, preferably 15% of all generated features.

[0046] The defined number of features is preferably a relative number of the total number of features, but can also be an absolute number of features.

[0047] A preliminary machine learning model within the meaning of the present disclosure is preferably based on one of the following classification algorithms: Linear Regression (lasso), Decision Trees, Random Forest (RF) or XGBoost (XGB) with a small number of estimators, or Generalized Linear Model (GLM).

[0048] In a further advantageous embodiment of the procedure, the cross-validation comprises the following steps:

[0049] • Training the preliminary machine model using the selected features via k-1 / k of the operating data and the state data;

[0050] • Evaluating the trained preliminary machine learning model using 1 / k of the operational and state data; and

[0051] • Determining the relevance of the features based on the evaluation of the trained preliminary machine model.

[0052] For further details of cross-validation as defined in this disclosure, reference is made to the publication by Ian H. Witten, Eibe Frank and MarkA. Hall: Data Mining: “Practical Tools and Techniques of Machine Learning”, 3rd edition. Morgan Kaufmann, Burlington, MA 2011, ISBN 978-0-12-374856-0 (waikato.ac.nz).

[0053] In a further advantageous embodiment of the procedure, the backward ranking method comprises the following steps:

[0054] • Training a preliminary machine learning model with the features;

[0055] • Remove the least relevant feature; • Repeat the preceding steps in several iterations until all features have been removed; and

[0056] • Ordering the features according to their relevance, with the feature removed first receiving the lowest rank and the feature removed last receiving the highest rank; and

[0057] • Output a defined number of features with the highest relevance to the subsequent feature selection method, preferably 15% of all generated features.

[0058] For further details of a backward ranking as defined in the present disclosure, reference is made to the publication Zhou, Ji-Yuan, et al.: "Prediction of hepatic inflammation in chronic hepatitis B patients with a random forest-backward feature elimination algorithm.", World Journal of Gastroenterology 27.21 (2021): 2910).

[0059] In a further advantageous embodiment of the procedure, the cluster-based selection method comprises the following steps:

[0060] • Calculating a pairwise correlation between all features submitted from the previous feature selection method and the associated state data using Pearson correlation coefficients;

[0061] • Clustering the features based on their correlation using agglomerative clustering techniques;

[0062] • Selection of features from each cluster at random, so that subgroups of features are created whose information is neither redundant nor incomplete;

[0063] • Training and validating preliminary machine nyem models using the selected subset on a single-digit number of folds, particularly five, and cross-validation to determine which subsets lead to the best model performance; and

[0064] • Store all evaluation metrics for each trained preliminary machine learning model, as well as the meaning of the features.

[0065] For further details of feature clustering as described in this disclosure, please refer to Chormunge, Smita, and Sudarson Jena: "Correlation-based feature selection with clustering for high-dimensional data", Journal of Electrical Systems and Information Technology 5.3 (2018): 542-549.

[0066] In a further advantageous embodiment of the procedure, the Subset Refinement method comprises the following steps, which are applied to at least some subgroups of the features to improve the relevance of the respective subgroup:

[0067] • If the relevance of a feature changes when retraining a preliminary machine learning model, remove the feature;

[0068] • If two features have a high Pearson correlation, remove the feature with the lower relevance; and

[0069] • If a feature has a high Pearson correlation with another feature that is ranked higher in backward ranking, replace the feature with the other feature; and

[0070] • Training the preliminary machine learning model with each of the improved subgroups; selecting the subgroup that most accurately determines the state data based on the features.

[0071] In a further advantageous embodiment of the method, the classification algorithm is a probabilistic classification algorithm.

[0072] By using a probabilistic classification algorithm, a probability of occurrence can be output for the occurrence of the respective physical states.

[0073] In a further advantageous embodiment, the method also comprises the following steps:

[0074] • Comparing the status data of the respective technical equipment with a limit value; and

[0075] • Marking the technical equipment as intact or defective based on comparison. Marking the technical equipment as defective or intact allows a user to better understand the prediction.

[0076] In a further advantageous embodiment, the method includes the following step:

[0077] • Determining the threshold for distinguishing between intact and defective technical equipment based on a cost-benefit analysis between a reactive and a proactive replacement of the technical equipment.

[0078] By taking into account a cost-benefit analysis when marking the technical equipment, in particular the vehicle or vehicle component, a decision can be made that is particularly advantageous for ensuring the function of the technical equipment.

[0079] In a further advantageous embodiment of the method, the condition data divides the technical equipment into a first group in which a defect occurs and a second group in which this defect does not occur, with the equipment of the first group being marked as defective and the equipment of the second group as intact.

[0080] The advantages and features described above with regard to the first aspect of the invention apply accordingly to the other aspects of the invention and vice versa.

[0081] A second aspect of the invention relates to a computer-implemented classifier for the indirect measurement of physical states of a technical device of a class of technical devices, which are preferably batteries or vehicles, wherein the classifier is generated by training a classification algorithm, wherein the classification algorithm was configured by the following steps, which are performed for each training input of a plurality of training inputs:

[0082] • Acquisition of operational data of a large number of technical facilities, wherein the operational data include value profiles of time-resolved measurement parameters and characterize an operational behavior and environment of a respective technical facility from a field operation, and of condition data of the respective technical facility;

[0083] • Generating features by processing at least part of the operational data using mathematical operations and / or by selecting data ranges from the operational data;

[0084] • Selection of relevant features from the generated features using a chain of feature selection methods; and

[0085] • Training the classification algorithm using the relevant features and their respective state data, thereby generating the classifier; where the state data characterize the physical states of the technical equipment.

[0086] A third aspect of the invention relates to a computer-implemented method for the indirect measurement of physical states of a technical device of a certain type to be analyzed, in particular a battery, by means of a classifier, comprising the following steps:

[0087] • Recording operational data of the technical equipment to be analyzed, wherein the operational data includes value profiles of time-resolved measurement parameters and characterizes the operational behavior of the technical equipment and its environment from field operation; and

[0088] • Determining output data by applying the classifier to input data based on operational data, where the output data characterizes the physical states of the technical equipment to be analyzed.

[0089] Using the method for the indirect measurement of physical states, configurations of a technical device can be determined from its operating data without having to determine the physical states of the technical device itself. Preferably, the cause of a defect can also be determined based on the determined state data by analyzing the values / value trends of the measurement parameters and / or the generated characteristics. In a further advantageous embodiment, the method also includes the following step:

[0090] • Generating features by processing at least part of the operating data by means of a mathematical operation and / or by selecting data ranges from the measurement parameters and the processed measurement parameters, wherein the input data includes the features.

[0091] In a further advantageous embodiment of the method, the classification algorithm is a probabilistic classification algorithm, wherein the output data indicates a probability of failure of the technical equipment.

[0092] In a further advantageous embodiment of the method, the output data of the technical equipment are divided into a first group in which a defect occurs and a second group in which this defect does not occur, wherein the equipment of the first group is marked as defective and the equipment of the second group as intact.

[0093] In a further advantageous embodiment of the procedure, the classification algorithm is selected from the following group:

[0094] Naive Bayes, logistic regression, gradient boosting, random forest.

[0095] A fourth aspect of the invention relates to a system for generating a classifier for indirectly measuring the states of a test specimen of a type of technical equipment, in particular a battery or a vehicle, by training a classification algorithm, comprising:

[0096] • Means for recording operational data of a large number of technical devices of the type, wherein the operational data comprise value profiles of time-resolved measurement parameters and characterize an operational behavior and environment of a respective technical device from field operation, and of state data of the respective technical device during field operation;

[0097] • Means for generating features by processing at least a part of the operational data by means of mathematical operations and / or by selecting data ranges from the operational data; means for selecting relevant features from the generated features by means of a chain of feature selection methods, wherein the chain comprises at least two, preferably three or particularly preferably all of the following feature selection methods:

[0098] Random Ranking (103A);

[0099] Backward ranking (103B);

[0100] Cluster-Based Selection (103C);

[0101] Subset Refinement (103D); and wherein the selected feature selection methods are executed in the order defined by the enumeration above; and

[0102] • Means for training the classification algorithm using the relevant features and the respective associated state data, whereby the classifier is generated; wherein the state data characterize the physical states of the technical equipment.

[0103] A fifth aspect of the invention relates to a system for the indirect measurement of physical states of a technical device to be analyzed, comprising the following steps:

[0104] • Means for recording operational data of the technical equipment, wherein the operational data comprise value profiles of time-resolved measurement parameters and characterize an operational behavior of the technical equipment and an environment of the technical equipment from a field operation; and

[0105] • Means of determining output data by applying the classifier to input data based on operational data, wherein the output data characterizes the physical states of the technical equipment.

[0106] A means according to the invention can be designed using hardware and / or software and, in particular, comprises a processing unit, preferably a microprocessor (CPU), preferably connected to a storage and / or bus system via data or signals, and / or one or more programs or program modules. The CPU can be configured to execute instructions implemented as a program stored in a storage system, to acquire input signals from a data bus, and / or to output signals to a data bus. A storage system can comprise one or more, in particular different, storage media, especially optical, magnetic, solid-state, and / or other non-volatile media. The program can be designed such that it embodies the methods described herein.is capable of executing such procedures, so that the CPU can perform the steps of such procedures and thus, in particular, analyze at least one technical device or train a classification algorithm.

[0107] The term "means" as used herein encompasses all structures, materials, or actions set forth herein, as well as all equivalents thereof. Furthermore, the structures, materials, or actions and their equivalents include everything described in the abstract, the brief description of the figures, the detailed description, the summary, and the claims themselves. A system and / or its means may preferably take the form of a pure hardware variant, a pure software variant (including firmware, resident software, microcode, etc.), or a combination of software and hardware aspects, generally referred to as a "circuit," "module," or "system." Any combination of one or more computer-readable media may be used. The computer-readable medium may be a computer-readable signaling medium or a computer-readable storage medium.

[0108] The systems and methods according to the present disclosure can preferably be implemented in conjunction with a suitably configured computer, a programmed microprocessor or microcontroller and one or more peripheral integrated circuit elements, an ASIC or other integrated circuit, a digital signal processor, a hard-wired electronic or logic circuit, such as a circuit with discrete elements, a programmable logic device or gate arrangement, such as a programmable logic device (PLD), a programmable logic array (PLA), a field-programmable gate arrangement (FPGA), a programmable logic arrangement (PAL), or a comparable means.In general, any device or means capable of implementing the methodology presented herein may be used to implement the various aspects of this disclosure. Exemplary hardware includes computers, handheld devices, telephones (e.g., cellular, internet-enabled, digital, analog, hybrid, and others), and other hardware known in engineering. Some of these devices include processors (e.g., a single or multiple microprocessors), memory, non-volatile memory, input devices, and output devices. Furthermore, alternative software implementations, including but not limited to distributed processing or distributed processing of components / objects, parallel processing, or processing by virtual machines, may be developed to implement the procedures described herein.

[0109] Further features and advantages will become apparent from the following description of exemplary embodiments with reference to the figures. These show, at least partially schematically:

[0110] Figure 1 shows a functional block diagram of an embodiment of a system for generating a classifier for the indirect measurement of the states of a test object;

[0111] Figure 2 shows a flowchart of an embodiment of a method for generating a classifier for the indirect measurement of the states of a test specimen;

[0112] Figure 3 shows a functional block diagram of an embodiment of a system for the indirect measurement of the physical states of a test specimen; and

[0113] Figure 4 shows a flowchart of an embodiment of a method for the indirect measurement of physical states of a test specimen.

[0114] In the following, an embodiment of a training of a classification algorithm is explained with reference to Figure 1 and Figure 2, where Figure 1 shows a functional block diagram of a system 10 for generating a classifier and Figure 2 shows a flowchart of a procedure 100 executable by means of the system 10 for generating the classifier.

[0115] The description refers here to a battery 2 of an electric passenger car 3 as a technical device. However, it is obvious to a person skilled in the art that the described method 100 and system 10 can also be used with regard to other technical devices, in particular the passenger cars 3 themselves.

[0116] For example, the classifier can be used to predict a thermal runaway in a vehicle battery in a given application. If a high probability of thermal runaway is detected, the battery must be replaced, as otherwise a fire is likely.

[0117] In a first step 101 of the method 100, operating data with value profiles of time-resolved measurement parameters, which characterize operating behavior and an environment, are recorded for a plurality of batteries 2 of the same type, in particular of the same design. Preferably, this is accomplished by means of a plurality of sensors 11, which are arranged on the passenger car 3 or in the vicinity of the passenger car 3. More preferably, the value profiles can also be historical value profiles of time-resolved measurement parameters. More preferably, these operating data are so-called telemetry data obtained from field operation.

[0118] The value profiles of time-resolved measurement parameters, or time series data, describe the use of the passenger car. 3. Suitable measurement parameters are those available via a vehicle control bus system, particularly a CAN network. These parameters can be selected from the following group, with the specific selection depending on the application: ambient temperature, engine temperature, durations spent in certain speed ranges, number of cold starts, number of journeys of a specific minimum duration, number of regenerations, state of charge, number of charging and / or discharging cycles, battery temperatures, cell temperatures, battery control parameters, difference between maximum and minimum voltage per cell, balancing of the frequency counts of all cell IDs, and a 2D histogram of the thermal management in "idle" mode. 1and a battery contactor temperature above 25 degrees Celsius, especially during charging, 2D heat map of cell temperature and SOC based on HV bus voltage, especially during charging, differential coolant temperature during charging, and / or differential cell voltage over battery discharge processes.

[0119] If the category of technical equipment is vehicle batteries, the measurement parameters preferably include at least a number of charging / discharging cycles, battery temperatures, cell temperatures, battery control parameters, a difference between maximum and minimum voltage per cell, a balancing of the frequency numbers of all cell IDs, a 2D histogram of the thermal management in "idle" mode and a battery contactor temperature above 25 degrees Celsius, particularly during charging, a 2D heat map of the cell temperature and SOC based on the HV bus voltage, particularly during charging, a differential coolant temperature during charging, and / or a differential cell voltage over the battery's discharge cycles.

[0120] If the category of technical equipment is vehicles, the measurement parameters preferably include at least an outside temperature, an engine temperature, dwell times in certain speed ranges, number of cold starts and / or a number of journeys with a certain minimum duration.

[0121] If the category of technical equipment is "battery-powered vehicles", the measurement parameters preferably include at least an outside temperature, a motor temperature, dwell times in certain speed ranges, number of cold starts and / or a number of journeys with a certain minimum duration, a number of charging / discharging cycles, battery temperatures, cell temperatures, battery control parameters, a difference between maximum and minimum voltage per cell, a balancing of the frequency numbers of all cell IDs, a 2D histogram of the thermal management in "idle" mode and a battery contactor temperature above 25 degrees Celsius, especially during charging, a 2D heat map of the cell temperature and SOC based on the HV bus voltage, especially during charging, a differential coolant temperature during charging, and / or a differential cell voltage over the battery's discharge cycles.

[0122] Furthermore, in the first step, 101 status data of battery 2 and / or vehicle 3 are recorded during field operation. This status data characterizes the physical states of battery 2. Such a physical state can, in particular, be a component specification, a dimension, a tolerance, or a change in age. Preferably, a physical state is a property of battery 3, especially one of the states "intact" or "defective." More preferably, the status data can also be specified as the probability of occurrence, in particular the probability of failure of battery 3, of a property.

[0123] Preferably, the operating data and the state data of battery 2 are acquired by means of sensors 11, which are arranged in the area of ​​battery 2 and its surroundings or in the area of ​​vehicle 3 and its surroundings and are configured to determine measurement parameters. Furthermore, preferably, the operating data and / or the state data are acquired via an interface 11, in particular a data interface.

[0124] In the aforementioned use case of identifying an impending thermal runaway of a battery 3, the value profiles of time-resolved measurement parameters constitute a time series with a defined sampling rate, in particular 30 seconds, which were recorded on thousands of battery-powered vehicles 3 as a type of technical equipment. Sixty of the vehicles 3 are identified as having a defective physical condition and are marked or labeled accordingly; the remaining vehicles 3 are identified as having an intact physical condition and are marked or labeled accordingly. Fifty of the vehicles 3 marked as "defective" have defective batteries 2 that therefore need to be replaced; ten more have battery defects that pose a safety risk.

[0125] In a second step, features are generated that are suitable for processing by a classification algorithm (Feature Generation / Feature Engineering).

[0126] The time-resolved measurement parameter profiles contained in the operational data consist of individual measurements per journal, for example, one temperature value per minute. These measurements are preferably taken and stored using a physical measurement system, for example, a sensor 11. However, the time-resolved measurement parameters are not typically used to train the classification rhythm; instead, so-called features are calculated from the time-resolved operational data. These features can be numerical, for example, size, weight; categorical, for example, color, type; temporal, for example, time, date; or geographical, for example, location coordinates.

[0127] Preferably, the features are generated using mathematical operations and / or by selecting data ranges from the operating data. The features are preferably aggregated values ​​and an alternative description of the time series data. Examples of features include statistics, such as mean, minimum, and maximum values, of parameters like ambient temperature, engine temperature, etc.; duration spent in specific speed ranges; number of cold starts; number of journeys with a certain minimum duration; number of regenerations, etc. Further features include, for example, an elevation profile of a road network, road conditions, ambient temperature, or humidity. Accordingly, the features can be categorized into features that describe the technical equipment or its operation—in this embodiment, the battery 2 or the vehicle 3—and features that describe the environment of the technical equipment.

[0128] The generation of the features is preferably carried out using means 12 for generating features. In particular, these means are configured to select data ranges from the operational data and to process them using mathematical operations.

[0129] In the aforementioned use case of identifying thermal runaway, more than 4800 features are generated for each of the vehicles of type 3. Some of the features contain one-dimensional or two-dimensional histograms of the recorded signals. Examples of such signals or time-resolved measurement parameters are cell temperature, voltage, current, ambient temperature, etc. Other features represent summaries of signals and recordings over a year or specific events. Such specific events can be, for example, fast charging of battery 2, normal charging, battery discharges, self-discharge of battery 2, etc. In a third step, relevant features are selected from the features generated in the second step using a chain of so-called feature-select methods.In this process step, the relevant characteristics are selected to be relevant to a specific defect. These characteristics are generated by processing at least a portion of the operational data using mathematical operations and / or by selecting data ranges from the operational data, and thus exhibit technical references.

[0130] In this third step, 103, the aim is to identify those characteristics that contribute most to characterizing the physical states of the vehicle 3 or the battery 2. For example, when selecting characteristics, it can be examined which characteristics contribute most to distinguishing "healthy" from "bad" operating data of the vehicle 3 or the battery 2. "Healthy" operating data characterizes an intact technical device, in particular an intact vehicle 3 or an intact battery 2, insofar as a measurement has shown that the specific defect has not occurred. However, defects other than the specific defect may have occurred in intact technical devices, intact vehicles, or intact batteries."Sick" operating data characterizes a defective technical device, in particular a defective vehicle 3 or a defective battery 2, insofar as a measurement has revealed that the specific defect has occurred. The selection is preferably made exclusively on the basis of the recorded operating data or value trends of time-resolved measurement parameters. The most important characteristics identified can also provide an indication of an underlying malfunction of the "sick" operating data. Therefore, this data can be used for a root cause analysis after the detection of a defective vehicle 3 or a defective battery 2.

[0131] Preferably, two, preferably three, or all of the following feature selection methods are used: Random Ranking 103A, Backward Ranking or Backward Elimination 103B, Cluster-Based Selection 103C, Subset Refinement 103D. In particular, two, preferably three, or especially preferably all four feature selection methods are applied in the order listed above. Accordingly, the following combinations can be used: Random Ranking and Backward Ranking, Random Ranking and Cluster-Based Selection, Random Ranking and Subset Refinement, Random Ranking and Backward Ranking and Cluster-Based Selection, Random Ranking and Backward Ranking and Cluster-Based Selection and Subset Refinement, Backward Ranking and Cluster-Based Selection, Backward Ranking and Subset Refinement, Backward Ranking and Cluster-Based Selection and Subset Refinement, Cluster-Based Selection and Subset Refinement.In this context, backward ranking and backward elimination are considered two variants of the same feature selection method.

[0132] The random ranking method 103A preferably comprises the following steps, as shown in Fig. 2:

[0133] Selections 103A-1 of a random subset of features;

[0134] Performing 103A-2 a multiple, in particular five-fold (k=5), cross-validation for a preliminary machine model with the selected features, whereby the relevance of the selected features for each trained preliminary machine learning model is determined;

[0135] Store 103A-3 a relevance of the selected features for each preliminary machine learning model trained during cross-validation;

[0136] Repeat steps 103A-4 of the preceding steps in several iterations until all features have been selected at least once;

[0137] Order 103A-5 of the features according to their average relevance; and

[0138] Output 103A-6 of a defined number of features with the highest relevance to the subsequent feature selection method, preferably 15% of all generated features.

[0139] Cross-validation preferably comprises the following steps:

[0140] Training 103A-2-1 of the preliminary machine intelligence model using the selected features via k-1 / k of the operational data and the state data;

[0141] Evaluate 103A-2-2 of the trained preliminary machine model using 1 / k of the operational data and the state data; and

[0142] Determine the relevance of features based on the evaluation of the trained preliminary machine model. For further details of cross-validation as described in this disclosure, reference is made to the publication by Ian H. Witten, Eibe Frank, and Mark A. Hall: Data Mining: Practical Tools and Techniques of Machine Learning, 3rd edition. Morgan Kaufmann, Burlington, MA 2011, ISBN 978-0-12-374856-0 (waikato.ac.nz).

[0143] The Backward Ranking Method 103B preferably includes the following steps:

[0144] Training 103B-1 of a preliminary machine learning model with the features;

[0145] Remove 103B-2, the least relevant feature;

[0146] Repeat step 103B-3 of the preceding steps in several iterations until all features have been removed; and

[0147] Order 103B-4 the features according to their relevance, with the feature removed first receiving the lowest rank and the feature removed last receiving the highest rank; and

[0148] Output 103B-5 of a defined number of features with the highest relevance to the subsequent feature selection method, preferably 15% of all generated features.

[0149] For further details of a backward ranking as defined in the present disclosure, reference is made to the publication Zhou, Ji-Yuan, et al.: "Prediction of hepatic inflammation in chronic hepatitis B patients with a random forest-backward feature elimination algorithm.", World Journal of Gastroenterology 27.21 (2021): 2910).

[0150] The Cluster-Based Selection method 103C preferably includes the following steps:

[0151] Calculate 103C-1 a pairwise correlation between all features transmitted from the previous feature selection method and the associated state data using Pearson correlation coefficients;

[0152] Clustering 103C-2 of the features based on their correlation using agglomerative clustering techniques;

[0153] Selecting 103C-3 features from each cluster at random, such that subsets of features are formed whose information is neither redundant nor incomplete; training and validating 103C-4 preliminary machine learning models using the selected subset on a single-digit number of folds, in particular five, a cross-validation to determine which subsets lead to the best model performance; and

[0154] Store 103C-5 all evaluation metrics for each trained preliminary machine learning model as well as the meaning of the features.

[0155] The subset refinement method 103D preferably comprises the following steps, which are applied to at least some subsets of features to improve the relevance of the subset: if the relevance of a feature changes upon retraining a preliminary machine learning model, remove 103D-1 of the feature; if two features have a high Pearson correlation, remove 103D-2 of the feature with the lower relevance; and if a feature has a high Pearson correlation with another feature that is ranked higher in backward ranking, replace 103D-3 of the feature with the other feature; and

[0156] Train 103D-4 of the preliminary machine learning model with each of the improved subgroups; selecting the subgroup that most accurately determines the state data.

[0157] The following algorithms can be used as preliminary machine learning models: Linear Regression (lasso), Decision Trees, Random Forest (RF) or XGBoost (XGB) with a small number of estimators, or Generalized Linear Model (GLM).

[0158] For each feature selection method 103A, 103B, 103C, 103D, the same or a different machine learning model can be used.

[0159] In the aforementioned use case of identifying thermal runaway, operating data from vehicles whose batteries have burned out have been flagged for the defect "thermal runaway." This defect is immediately and unambiguously recognizable to a technician who sees the defective vehicle. Furthermore, as shown in Fig. 1, within the chain of feature selection methods 103, the number of more than 4800 features is reduced in a first step to approximately 890 features by random ranking 103A. In a second step, this number is reduced to approximately 260 features by backward ranking 103B. In a third step, this number is further reduced to approximately 130 features by cluster-based selection 103C. Finally, in a fourth step, this number can be reduced to six relevant features by subset refinement 103D. These remaining relevant features are as follows:

[0160] EVENTS_CHARGING_MEAN_DELTA_CELL_VOLTAGE_STD_LAST_YEAR: For each charging cycle, the mean difference between the maximum and minimum voltage per cell is calculated. The standard deviation of the mean values ​​for the past year's data yields the characteristic "Standard deviation of the delta cell voltage across charging events" (aggregated as the average standard deviation of all charging events). This characteristic analyzes the dispersion of the cell voltage difference during charging (delta decrease and increase). It takes into account both large voltage differences and voltage fluctuations during charging. Since the standard deviation is larger for defective vehicles or vehicle batteries than for intact ones, the relevance of this characteristic indicates that the individual cells are being charged with a highly unbalanced voltage.The cause of this relevant signal can be micro-short circuits at the electrode level, which affect a cell via significantly increased self-discharge. If the voltage difference is a strongly curved line, this indicates a large state-of-charge (SoC) difference during the charging process, which persists even after equalization has occurred. The standard deviation of this signal represents the time that the different SoCs persist. In critical vehicles, this difference can persist for an extended period.

[0161] AGG_LAST_YEAR_STD_FREQUENCY_COUNTS_BALANCING_CELL_ID: The signals of interest, which are part of the operational data, are the battery cell balancing status (plural) along with the battery cell IDs of those battery cells that triggered battery cell balancing. This counts how often each battery cell has triggered battery cell balancing, including zeros, i.e., battery cells that have never triggered battery cell balancing. Subsequently, a standard deviation of the number of triggered battery cell balancings is calculated for all cell IDs. Only data from the last year is considered. In vehicles with a high risk of thermal runaway, the critical cell IDs are grouped around a few battery cells. This feature takes into account both the balancing duration (with the standard deviation) and the "centralization" of balancing efforts onto a few problematic cells.

[0162] HEAT_BatteryThermalManagementModeHvBatteryContactorTemp_charge_ldle_Higher HVCondatorTemp_PERCENT: This parameter aggregates the vehicle's dwell time in "idle" thermal management mode when the HV battery contactor reaches warm / hot temperatures. The duration the HV battery contactor remains above a specific threshold, here T>25 °C, during the charging process is measured. The purpose of this feature is to identify high battery temperatures in thermal mode during "idle" operation. The HV contactor temperature can be affected by various component failures in the battery terminal box. If one of these components is defective, the temperature rises. It appears that this condition is particularly pronounced in faulty vehicles during low-current operation (thermal system idle). The most likely triggers for a thermal event are the cells located beneath the EE unit.

[0163] HEAT_CellTempAvgSocHvBus_charge_HigherCellTempAvg_SOCHVBus_PE CENT:

[0164] The average cell temperature, calculated from the minimum and maximum cell temperatures, and the state of charge (SOC) based on the HV bus voltage are used as operating data. From these, the vehicle's dwell times at warm / hot battery cell temperatures above a threshold and low SOC during charging are aggregated. This feature correlates high cell temperatures with low bus voltages. This feature is intended to detect whether one or more cells in a battery pack are faulty (especially with improved self-discharge). We consider the SoC based on the total voltage of the battery pack and find that faulty vehicles usually have a lower SoC, even when charged at the same cutoff voltage, and simultaneously exhibit a higher average temperature.The higher the cell temperature and the lower the state of charge of a vehicle, the more likely it is that more than one cell will be affected.

[0165] EVENTS_CHARGING_STD_DELTA_COOLANT_TEMP_AVG_LAST_YEAR: The mean delta cooling temperature is calculated across all charging events, aggregated as the standard deviation of all charging events of the past year. The delta cooling temperature is the battery cell coolant temperature at the inlet minus the battery cell coolant temperature at the outlet. The standard deviation of the average delta cooling temperature was calculated for all charging events of the past year. This allows us to determine the variance of the average coolant delta during the charging process (decreases and increases in the delta). Furthermore, it allows us to detect changes between heating and cooling during the charging process. This measure takes into account both the distribution of the cell temperature within the vehicle battery and its dynamics, which are caused by the control of the thermal system.At first glance, it might seem that the cooling effort during low-power charging unnecessarily increases the temperature distribution within the battery pack. The cause of this issue has been identified as internal cell defects, which often lead to greater self-heating of an affected cell compared to others. If a cell is not well connected to the cooling system, its temperature will not respond well to the cooling system's operation. Calculating the standard deviation of the coolant inlet temperature versus the coolant outlet temperature can provide a signal that illustrates these two effects, which are amplified by the thermal system's control strategy. The standard deviation is high when the cooling system frequently switches to keep the highest cell below its threshold without overcooling the lowest cell.

[0166] E VE NTS_BATTE RY_DI SCH ARG E_ST D_DE LTA_CE LL_VO LTAG E_AVG_G RADTOTAL: The delta cell voltage signal (maximum minus minimum cell voltage) is used to calculate the average gradient of this signal for each discharge event. The standard deviation of the average gradient is calculated for all discharge events over the entire lifespan of the vehicle battery. The cell voltage fluctuation is an important indicator of internal short circuits at the electrode level (burrs and dendrites caused by electrode poisoning). This is a high-quality measure of voltage noise despite the low data sampling rate. If a battery cell has a faulty condition that leads to an increase in cell resistance, this is not easily detected in the low resolution of data points with 30s intervals.If a delta exists between the minimum and maximum cell voltage during discharge, a resistance difference may be present, even if it cannot be quantified. However, the gradient of this signal indicates a developing difference during operation.

[0167] The third step 103 is preferably performed by means 13 to select relevant features. These means 13 execute several feature selection methods sequentially, in particular in a kind of chain.

[0168] In a fourth step 104, a threshold value for distinguishing between intact and defective technical equipment of the respective type is preferably determined based on the condition data that characterize the physical state of the vehicle 3 and / or the battery 2. Preferably, a cost-benefit analysis is performed for this purpose, comparing reactive and proactive replacement of the respective technical equipment 3, 4 of the type. Furthermore, preferably, the threshold value is set such that if the probability of failure of the vehicle 3 and / or the battery 2 is comparatively high, the respective technical equipment 2, 3 is marked as defective.

[0169] The fourth step 104 is preferably carried out by means 14 to determine a limit value. Preferably, the limit value is determined on the basis of a cost-benefit analysis between a reactive and a proactive replacement of the technical equipment 2, 3.

[0170] In a fifth step 105, the status data of a respective technical device 2, 3 of the type are preferably compared with the limit value. If the limit value is exceeded or not reached by the respective technical device 2, 3, the respective technical device 2, 3 is preferably marked as defective in a sixth step 106.

[0171] The fifth work step 105 and the sixth work step 106 are each preferably carried out by means for comparing the state data 15 and by means 16 for marking a technical device 2, 3.

[0172] In a seventh step 107, the classification algorithm is trained using the relevant features selected in the third step and the respective associated state data, which are preferably classified or labelled according to the fourth to sixth steps 104 to 106.

[0173] For training the classification algorithm, which is a type of modeling, a technique called supervised learning is used. Depending on the type of condition data, this can involve training with binary condition data classified as "defective" or "intact," or with multi-class condition data containing further classifications, each relating to a specific defect. The result is either a binary classifier or a multi-class classifier adapted to classify operational data according to the probability of occurrence of that specific defect. The classifier uses the correlation between the relevant features selected based on the condition data and the operational data being analyzed.

[0174] An example of a classification algorithm that can be used within the scope of the disclosure is logistic regression. This involves applying a logistic function to linear regression, in particular Lasso, Ridge, or Elastic. When trained with binary state data, a binary response is obtained, in the examples given, "intact" or "defective".

[0175] Other possible classification algorithms are generalized linear models, Random Forest, XGBoost, support vector machines, and stacked classification algorithms. Preferably, interpretable classification algorithms that can also handle outliers well are used. The seventh step 107 is preferably performed by means 17 to train a classification algorithm. In particular, such means 17 has an interface, especially a data interface, to output the generated and selected features as well as the associated state data, which characterize the physical states of the vehicle 3 or the battery 2, to the classification algorithm to be trained.

[0176] In an eighth step, the trained classifier is preferably evaluated. Standard metrics such as Area Under the ROC Curve, Recall, Precision, True Positive Rate, etc., are preferably used for this evaluation. In particular, the quality of the trained classifier is determined. Typically, the classification algorithm is trained on a first subset of the operational data, and the trained classifier is then evaluated on a second subset of the available operational data. This avoids overfitting.

[0177] The eighth step 108 is preferably performed by means 18 to determine an evaluation of the trained classifier.

[0178] The trained classifier is adapted to be applied to unknown operating data of technical equipment 2, 3 of the same type in order to determine the physical states of this equipment 2, 3. Furthermore, the trained classifier can also be applied to those technical equipment 2, 3 on which the underlying classification algorithm was trained. This allows the physical states of these equipment 2, 3 to be re-determined. Given the large number of technical equipment 2, 3 on which the classifier was trained, the trained classifier can be used to determine the probability of occurrence of a specific physical state, in particular a probability of failure, for the technical equipment 2, 3 under investigation.The physical state can be specified here both by physical quantities and by the probability of occurrence of a particular physical state. The application of a trained classifier is explained below with reference to exemplary embodiments of a computer-implemented method 200 and a system 20 for the indirect measurement of physical states of a technical device 2, 3, using Figure 3 and Figure 4 as examples. Figure 3 shows a functional block diagram of the system 20, and Figure 4 shows a flowchart of the method 200.

[0179] In a first step 201, operating data of the technical equipment 2, 3 to be analyzed—in the described embodiment, a battery-powered vehicle 3 or a battery 2 of such a vehicle 3—are recorded. This operating data includes time-resolved measurement parameter value profiles. These operating data also characterize the operating behavior of the technical equipment to be analyzed and its environment from a field operation. However, unlike the training phase, the operating data do not need to be labeled. Therefore, the physical states of the vehicle 3 or battery 2 to be analyzed are normally unknown.

[0180] Preferably, the system 20 also includes corresponding means 21 for acquiring operating data. Preferably, these means 21 are configured as an interface, in particular a data interface. More preferably, these means 21 are configured as sensors that can directly acquire the measurement parameters.

[0181] In a second step 202, features are preferably generated by processing at least a part of the operating data by means of a mathematical operation and / or by evaluating data ranges from the measurement parameters and revised measurement parameters.

[0182] Preferably, at least those measurement parameters necessary for generating the relevant features are recorded. Furthermore, preferably, the measurement parameters and / or the respective relevant features are the same as those already mentioned in relation to method 100 for training a classification algorithm.

[0183] Preferably, the system 20 comprises means 22 for generating features. In a third step 203, output data is determined by applying the classifier to the input data. The input data is based on the operating data and is, in particular, the generated features and preferably the features identified as relevant. The output data characterizes the physical states of the vehicle 3 or battery 2 under investigation.

[0184] Preferably, the system 20 includes means 23 for determining output data. In the embodiment shown in Fig. 3, physical states of the battery 2 are determined, as indicated by the battery 2 shown in the box 24 in Fig. 3.

[0185] Preferably, the output data is output via a further interface 24, in particular a user interface or a data interface, as shown in Fig. 3.

[0186] In a fifth step 205, a limit value can preferably be introduced which defines whether the vehicle 3 or the battery 2 are "defective" or "intact" as technical equipment. This limit value can also be determined in a cost-benefit analysis, preferably comparing the costs of a reactive replacement with the costs of a proactive replacement.

[0187] In a sixth step 206, the condition data of the technical equipment to be analyzed, determined using method 200, can be compared with the limit value. In a seventh step 207, the technical equipment to be analyzed, in particular the test specimen, can then be marked according to its physical condition with respect to the limit value, preferably as "intact" or "defective". Accordingly, system 20 further preferably comprises means 25 for determining a limit value, means 26 for comparison, and means 27 for marking. The marking can also be output via interface 24.

[0188] This method 200 can be used for various purposes, such as predictive maintenance, risk assessment and continuous monitoring of a vehicle fleet, vehicle components and other technical equipment.

[0189] In the described use case of identifying an impending thermal runaway of battery 3, the root cause was determined to be a manufacturing defect. Analysis of the relevant parameters showed that calendar aging had only a minimal impact (comparison of Delta_cell_temp_avg (during battery discharge), delta_cell_avg (during normal charging), min_cell_voltage_avg (during battery discharge), and delta_soc_bus_per_kWh (during normal charging)). This suggested that a healthy and a defective battery population existed from the outset, and that the underlying problem stemmed from a manufacturing defect. Analysis of low_cell_voltage and high_cell_voltage deltas, as well as high_temperatures and high_temperature deltas, further supported this conclusion.

[0190] Another application is the investigation of oil dilution: In this application, based on the measurement parameters and the selected characteristics (evaluation of diesel particulate filter (DPF) regenerations, number and duration of successful and failed DPF or NOx regenerations), it could be concluded that the characteristics of defective vehicles differ significantly from those of intact vehicles. In the case of failed regeneration, increasing amounts of fuel enter the oil, diluting it and leading to engine failures.

[0191] Another application is the investigation of the efficiency of the catalytic converter system: The evaluation focuses on determining which driving behaviors lead to catalytic converter failure. The main differences between intact and defective vehicles were observed with regard to idling events (more idling events, longer idling duration, more idling events at very cold and hot ambient temperatures, etc.), engine operation (higher percentage of operation under very high engine load), and related P-codes.

[0192] It should be noted that the exemplary embodiments are merely examples and are not intended to restrict the scope of protection, applications, or structure in any way. Rather, the preceding description provides the person skilled in the art with a guideline for implementing at least one exemplary embodiment, whereby various modifications, particularly with regard to the function and arrangement of the described components, can be made without departing from the scope of protection as defined by the claims and these equivalent combinations of features.

Claims

Claims 1. Computer-implemented method (100) for generating a classifier (1) for the indirect measurement of physical states of a test specimen of a class of technical equipment, which are preferably vehicle batteries (2) or vehicles (3), by training a classification algorithm, comprising the following steps: • Recording (101) operational data of a large number of technical equipment of the type, wherein the operational data include value profiles of time-resolved measurement parameters and characterize an operational behavior and environment of a respective technical equipment from field operation, and state data of the respective technical equipment during field operation; • Generating (102) features by processing at least part of the operational data using mathematical operations and / or by selecting data ranges from the operational data; • Selecting (103) relevant features from the generated features using a chain of feature selection methods; and • Training (107) the classification algorithm using the relevant features and their respective state data, generating the classifier (1); wherein the state data characterize the physical states of the technical equipment (2, 3), and wherein at least two, preferably three, or particularly preferably all of the following feature selection methods are selected: • Random Ranking (103A); • Backward ranking (103B); • Cluster-Based Selection (103C); • Subset Refinement (103D); and wherein the selected feature selection methods are executed in the order defined by the enumeration above.

2. The method according to claim 1, further comprising the following step: evaluating (108) the computer-implemented classifier using at least one standard measure, in particular from the following group of standard measures: Area under the ROC Curve, Recall, Precision, True Positive Rate.

3. A method according to any of the preceding claims, wherein the random ranking method (103A) comprises the following steps: • Selections (103A-1 ) of a random subset of features; • Performing (103A-2) multiple, in particular fivefold (k=5), cross-validation for a preliminary machine model with the selected features, determining the relevance of the selected features for each trained preliminary machine learning model; • Store (103A-3) a relevance of the selected features for each preliminary machine learning model trained during cross-validation; • Repeat (103A-4) the preceding steps in several iterations until all features have been selected at least once; • Order (103A-5) the features according to their average relevance; and • Output (103A-6) a defined number of features with the highest relevance to the subsequent feature selection method, preferably 15% of all generated features.

4. Method according to any of the preceding claims, wherein the cross-validation (103A-2) comprises the following steps: • Training (103A-2-1) of the preliminary machine intelligence model using the selected features via k-1 / k of the operational data and the state data; • Evaluate (103A-2-2) the trained preliminary machine learning model using 1 / k of the operational data and the state data; and • Determine (103A-2-3) the relevance of the features based on the evaluation of the trained preliminary machine model.

5. A method according to any of the preceding claims, wherein the backward ranking method (103B) comprises the following steps: • Training (103B-1) a preliminary machine learning model with the features; • Removal (103B-2) of the least relevant feature; • Repeat (103B-3) the preceding steps in several iterations until all features have been removed; and • Ordering (103B-4) the features according to their relevance, with the feature removed first receiving the lowest rank and the feature removed last receiving the highest rank; and • Output (103B-5) a defined number of features with the highest relevance to the subsequent feature selection method, preferably 15% of all generated features.

6. Method (100) according to any of the preceding claims, wherein the cluster-based selection method (103C) comprises the following steps: • Calculate (103C-1) a pairwise correlation between all features submitted from the previous feature selection method and the associated state data using Pearson correlation coefficients; • Clustering (103C-2) the features based on their correlation using agglomerative clustering techniques; • Select (103C-3) features from each cluster at random, so that subgroups of features are created whose information is neither redundant nor incomplete; • Training and validating (103C-4) preliminary machine learning models using the selected subset on a single-digit number of folds, in particular five, a cross-validation to determine which subsets lead to the best model performance; and • Store (103C-5) all evaluation metrics for each trained preliminary machine learning model as well as the meaning of the features.

7. Method (100) according to any of the preceding claims, wherein the subset refinement method (103D) comprises the following steps, which are applied to at least some subgroups of the features to improve the relevance of the subgroup: • if the relevance of a feature changes when retraining a preliminary machine learning model, remove (103D-1) the feature; • if two features have a high Pearson correlation, remove (103D-2) the feature with the lower relevance; and • if a feature has a high Pearson correlation with another feature that is ranked higher in backward ranking, replace (103D-3) the feature with the other feature; and • Training (103D-4) the preliminary machine learning model with each of the improved subgroups; selecting the subgroup that most accurately determines the state data.

8. A computer-implemented classifier (1) for the indirect measurement of physical states of a test specimen of a class of technical equipment, which are preferably vehicle batteries (2) or vehicles (3), wherein the classifier (1) is generated by training a classification algorithm, wherein the classification algorithm was configured by the following steps, which are performed for each training input of a plurality of training inputs: • Acquisition of operational data of a large number of technical devices of the type, wherein the operational data include value profiles of time-resolved measurement parameters and characterize an operational behavior and environment of a respective technical device from a field operation, and of condition data of the respective technical device; • Generating features by processing at least part of the operational data using mathematical operations and / or by selecting data ranges from the operational data; • Selection of relevant features from the generated features using a chain of feature selection methods; and • Training the classification algorithm using the relevant features and their respective state data, generating the classifier (1); wherein the state data characterize the physical states of the technical equipment (2, 3); wherein at least two, preferably three, or particularly preferably all of the following feature selection methods are selected: • Random Ranking (103A); • Backward ranking (103B); • Cluster-Based Selection (103C); • Subset Refinement (103D); and wherein the selected feature selection methods are executed in the order defined by the enumeration above.

9. Computer-implemented method (200) for indirectly measuring the physical states of a test specimen of a class of technical equipment, which are preferably vehicle batteries (2) or vehicles (3), by means of a classifier (1), in particular according to claim 9 and / or which is generated by means of a method (100) according to one of claims 1 to 9, comprising the following steps: • Acquisition (201) of operational data of the test object, wherein the operational data include value profiles of time-resolved measurement parameters and characterize an operational behavior of the test object and an environment of the test object from a field operation; • Determining (203) output data by applying the classifier (1) to input data based on the operating data, wherein the output data characterize the physical states of the device under test; and Output (204) the output data.

10. The method according to claim 9, further comprising the following step: • Generating (202) features by processing at least part of the operating data by means of a mathematical operation and / or by selecting data ranges from the measurement parameters and the processed measurement parameters, wherein the input data comprise the features.

11. Method (200) according to claim 9 or 10, wherein the classification algorithm is a probabilistic classification algorithm, wherein the output data indicates a failure probability of the test specimen.

12. Method (200) according to any one of claims 9 to 11, wherein the state data and / or the output data divide the technical equipment of the genus into a first group in which a defect occurs and a second group in which this defect does not occur, wherein the technical equipment of the first group is marked as defective and the technical equipment of the second group as intact.

13. Method (200) according to any one of claims 9 to 12, further comprising the following steps: Comparing (105; 205) the status data or output data of the respective technical equipment (2, 3), in particular the device under test, with a limit value; and Marking (106; 206) the respective technical equipment (2, 3) as intact or as defective on the basis of the comparison.

14. Method according to any one of claims 9 to 13, further comprising the following step: Determining (104; 204) the limit value for distinguishing between intact and defective technical equipment on the basis of a cost-benefit analysis between a reactive and a proactive exchange of the technical equipment.

15. Method (100; 200) according to any one of claims 9 to 14, wherein the classification algorithm is selected from the following group: naive Bayes, logistic regression, gradient boosting, random forest.

16. System (10) for generating a classifier for indirectly measuring the states of a test specimen of a certain type of technical equipment, which are preferably vehicle batteries (2) or vehicles (3), by training a classification algorithm, comprising: • Means for recording (11) operational data of a plurality of technical equipment of the genus, wherein the operational data comprise value profiles of time-resolved measurement parameters and characterize an operational behavior and environment of a respective technical equipment from field operation, and of state data of the respective technical equipment during field operation; • Means for generating (12) features by processing at least a part of the operational data by means of mathematical operations and / or by selecting data ranges from the operational data; Means for selecting (13A, 13B, 13C, 13D) relevant features from the generated features by means of a chaining of feature selection methods, wherein the chaining comprises at least two, preferably three or particularly preferably all of the following feature selection methods: Random Ranking (103A); Backward ranking (103B); Cluster-Based Selection (103C); Subset Refinement (103D); and wherein the selected feature selection methods are executed in the order defined by the enumeration above; and • Means for training (17) the classification algorithm using the relevant features and the respective associated state data, whereby the classifier is generated; where the state data characterize the physical states of the technical equipment (2, 3).

17. System (20) for the indirect measurement of physical states of a test specimen of a class of technical equipment, which are preferably vehicle batteries (2) or vehicles (3), by means of a classifier (1), according to claim 8, comprising: • Means for recording (21) operational data of the technical equipment, wherein the operational data comprise value profiles of time-resolved measurement parameters and characterize an operational behavior of the technical equipment and an environment of the technical equipment from a field operation; and • Means for determining (23) output data by applying the classifier (1) to input data based on the operating data, wherein the output data characterize the physical states of the technical equipment.

Citation Information

Patent Citations

  • Analysis of vehicle data to predict component failure

    US20160035150A1