Method and system for indirectly measuring a probability of the occurrence of a defect in a test object
A method using time-resolved measurement data and probabilistic classification algorithms addresses the challenge of defect prediction in complex products, enhancing reliability through automated defect identification and proactive maintenance.
Patent Information
- Application Number
- PCT/AT2025/060217
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-29
- Filing Date
- 2025-05-28
- Publication Date
- 2025-12-04
AI Technical Summary
Existing methods struggle to accurately determine the probability of defects in complex products like high-voltage batteries and vehicles, particularly those that are difficult to measure directly, which affects product reliability and customer satisfaction.
A computer-implemented method using time-resolved measurement data to generate features, apply a probabilistic classification algorithm, and select relevant features through a chain of methods to predict defect probabilities, enabling automated defect identification and proactive maintenance.
Enables predictive modeling for vehicle condition assessment, reducing overfitting and computational demands, allowing for timely corrective actions and improved product reliability.
Smart Images

Figure AT2025060217_04122025_PF_FP_ABST
Abstract
Description
[0001] Method and system for indirectly measuring the probability of a defect occurring in a test object
[0002] The invention relates to a system and a computer-implemented method for indirectly measuring the probability of a defect occurring in a test specimen of a type of technical equipment, in particular a high-voltage battery or a vehicle, using operating data from a plurality of technical equipment of the type and the test specimen. This data includes time-resolved measurement parameter profiles and characterizes the operating behavior and environment of a respective technical equipment from field operation. A classification algorithm is trained on this data, resulting in a classifier, which is then applied to input data based on the operating data of the test specimen. Furthermore, the invention relates to a corresponding system for indirectly measuring the probability of a defect occurring.
[0003] Quality is one of the most important factors influencing a customer's choice between competing products in a given product segment. Therefore, understanding the causal relationship between individual components of a product within the overall product suite and potential defects throughout its entire lifecycle is crucial for the business success of a product and, consequently, of a company.
[0004] An example of the requirements for a product, such as a vehicle, is the lifespan of individual components, such as the battery. This lifespan has a direct impact on the total cost of ownership of a product, which is very important for the end user. In this respect, the product should not exceed a predetermined value.
[0005] To ensure the functionality of a product even after its manufacture and delivery to the customer, it is therefore of great interest to the manufacturer to be able to determine the properties of the product, and in the case of units with multiple technical components, also the properties of individual components within the unit. Such properties are characterized by the physical states of the product.
[0006] The more individual technical components a product or unit comprises, the more important it generally becomes to understand its properties. High-voltage batteries for electric vehicles, or the electric vehicles themselves, are examples of such complex products. With these types of products, it is often difficult to deduce, or even identify, the underlying physical parameter or technical component causing a malfunction, based on a symptom that suggests one or more malfunctions.
[0007] It is an object of the invention to provide a system and a method for determining the probability of a specific defect occurring. In particular, it is intended to determine probabilities for the occurrence of defects that are difficult or impossible to measure using previously known methods.
[0008] These tasks are solved through the doctrine of independent claims. Advantageous configurations are claimed in dependent claims.
[0009] A first aspect of the invention relates to a computer-implemented method for indirectly measuring the probability of a defect occurring in a test specimen of a class of technical equipment, preferably batteries or vehicles, comprising the following steps: a) measuring state data in a plurality of technical equipment of the class, wherein the state data includes whether the defect has occurred; b) measuring operating data of the plurality of technical equipment of the class and of the test specimen, wherein the operating data includes value profiles of time-resolved measurement parameters and characterizes the operating behavior and environment of a respective technical equipment from field operation, wherein the operating data of those technical equipment in which the defect has occurred are marked as defective according to the result of the measurement in step a).c) Generating features by processing at least a portion of the operational data using mathematical operations and / or by selecting data ranges from the operational data; d) Selecting features relevant to the defect from the generated features using a chain of feature selection methods; e) Training a probabilistic classification algorithm using the selected relevant features and their respective state data, thereby generating a classifier; and f) Generating output data, comprising a probability for the occurrence of the defect in the device under test, by applying the classifier to input data based on the operational data of the device under test.
[0010] A measurement within the meaning of this disclosure preferably involves determining a measurement signal using a sensor and / or post-processing a measurement signal to generate operating data and to determine whether the defect has occurred. A measurement may also include a technical inspection by a qualified person.
[0011] Measurement parameters within the meaning of the present disclosure preferably include an outside temperature, an engine temperature, dwell times in certain speed ranges, a number of cold starts, a number of journeys with a certain minimum duration, a number of regenerations, states of charge, charging / discharging processes, battery temperatures, cell temperatures, battery control parameters, a difference of maximum to minimum voltage per cell, a balancing of the frequency numbers of all cell IDs, a 2d histogram of the thermal management in "idle" mode and a battery contactor temperature above 25 degrees Celsius, particularly during charging, a 2d dwell time histogram of the cell temperature and state of charge (SOC) based on the HV bus voltage, particularly during charging, a differential coolant temperature during charging processes, and / or a differential cell voltage over the battery's discharge processes.
[0012] A mathematical operation within the meaning of the present disclosure is statistics, in particular mean, median, minimum, maximum or standard deviation, or aggregation, addition, subtraction, multiplication, division, differentiation, gradient calculation and / or integration over certain durations and / or multivariate integrals, in which measurement parameters are multiplied and then integrated.
[0013] A class of technical devices within the meaning of this disclosure is preferably a group of technical devices that are identical in their essential features and are therefore preferably of identical construction. Preferably, the essential components of the technical devices of a class are identical in construction. A specific technical device is thus preferably an implementation of the class of technical devices. In particular, technical devices of a class differ by tolerances, especially manufacturing tolerances and / or aging or wear effects.
[0014] An indirect measurement within the meaning of this disclosure is the determination of a desired value, in particular the parameter probability of the occurrence of the defect, using other available information from a system to be measured. The desired measured value of this parameter is determined from other physical parameters.
[0015] For the purposes of the invention, a defect is a specific defect that occurs in different technical devices in the same or similar way, i.e., with the same or similar fault pattern. The defect could, for example, be thermal runaway of a battery in a battery-powered vehicle, increasing dilution of oil with fuel in internal combustion engine vehicles, a reduction in the efficiency of a catalytic converter system, a failure of a specific mechanical element in a transmission, or another type of defect.
[0016] A vehicle within the meaning of the present disclosure may preferably be or comprise a passenger car, a truck, a motorcycle, a construction machine, a ship or an agricultural vehicle.
[0017] The invention enables the creation of a predictive model for the condition of a vehicle from measurement data, particularly historical measurement data, provided that time-resolved operating data from multiple technical devices of the same type or category can be measured or are available, and that the defect has already occurred multiple times. In particular, the method is therefore suitable for fault identification in vehicles within a fleet. The physical condition of the vehicle can then be determined using this predictive model. Furthermore, the method can also be applied to individual vehicle components, such as a battery or fuel cell, an engine, a braking system, a transmission, or others. The condition data contains at least the information on whether one of the states is "defective" or "intact," i.e., whether the defect has occurred or not.The acquisition of condition data can be automated, particularly through a defect monitoring system and / or in a workshop. Additionally, the condition data can include other physical states such as component specifications, dimensions, and tolerances of the technical equipment.
[0018] The probability of the defect occurring indicates the physical state of the operational capability of the technical equipment with respect to this defect.
[0019] The method for indirectly measuring the probability of a defect occurring is based on generating features from time-resolved measurement data ("feature generation" or "feature engineering"). This process primarily involves statistical calculations based on the time-resolved measurement data. From the generated features, the most relevant features—that is, those features that are causally related to the physical state indicated by the labels—are selected ("feature selection"). A classification algorithm is then trained (modeling) using these features and their corresponding labels. This model can be evaluated, particularly using standard metrics.
[0020] The trained classifier is then applied to further time-resolved measurement data of the test object, in particular the vehicle or vehicle component. This determines the probabilities of the defect occurring in the test object. The time-resolved measurement data is available as aggregated data. If the time-resolved data is available in a different format, it is first aggregated to enable the application of the classifier. The test object, like the technical equipment, belongs to the category of technical equipment. Unlike the technical equipment that forms the data basis of the method, the state of the test object with regard to whether the defect has occurred is unknown, or it is known insofar as the defect has not occurred in the test object. In a particular embodiment of the invention, the test object can therefore also be a technical piece of equipment belonging to the category of technical equipment.
[0021] According to the invention, to predict the probability of a defect occurring, a chain of at least two, three, or four feature selection methods is executed sequentially. This reduces the number of channels with measurement data and the number of features derived or generated from this measurement data, so that only a comparatively small number of relevant features remain for training the classification algorithm. Removing irrelevant or redundant features improves model performance. This reduces overfitting, as the model is not influenced by irrelevant data and can better focus on the important patterns.
[0022] Fewer features mean less data to process. This leads to faster training times and lower demands on storage and computing power. This is particularly important for large datasets and complex models. Furthermore, the remaining data can be used to perform root cause analysis related to a physical state of the device under test. This allows the features, or the underlying time-resolved measurement data, to be examined for anomalies that may have triggered a physical state.
[0023] Further advantages are achieved if the procedure also includes the step of outputting the probability of the defect occurring in the test specimen.
[0024] Preferably, the process can be automated. Automated execution includes, in particular, the automatic generation of features, which is a time-consuming process manually, and the application of chaining feature selection methods to select the relevant features.
[0025] In a further advantageous embodiment of the procedure, the procedure further comprises the following step:
[0026] Implementing a defect rectification measure, especially if the probability of the defect occurring exceeds a threshold.
[0027] Depending on the type of defect, various corrective measures can be implemented. These measures may include, among other things, a change in the manufacturing process of the technical equipment, particularly the battery or the vehicle; a modification of operating parameters, such as charging parameters; a recall; a repair; maintenance; battery replacement; a software update; vehicle decommissioning; and / or battery monitoring.
[0028] In a further advantageous embodiment, the method includes the following step: evaluating the computer-implemented classifier using at least one standard measure, in particular from the following group of standard measures: Area under the ROC Curve, Recall, Precision, True Positive Rate.
[0029] An evaluation of the trained classifier can verify its accuracy in predicting physical states.
[0030] In a further advantageous embodiment of the method, at least two, preferably three or all of the following feature selection methods are selected:
[0031] • Random Ranking;
[0032] • Backward Ranking or Backward Elimination;
[0033] • Cluster-Based Selection;
[0034] • Subset Refinement; wherein the selected feature selection methods are executed in the order defined by the enumeration above.
[0035] The choice of feature selection methods and their order has a significantly greater impact on the quality of the trained classifier than the choice of the classification algorithm itself. As explained above, the aforementioned feature selection methods are executed sequentially in a chain. This selection and order reliably identify the most relevant features and, moreover, substantially reduce the number of features that need to be processed when training the classification algorithm.
[0036] In a further advantageous embodiment of the procedure, the Random Ranking method comprises the following steps:
[0037] • Selecting a random subset of features;
[0038] • Performing multiple, in particular five-fold (k=5), cross-validation for a preliminary machine model with the selected features, determining the relevance of the selected features for each trained preliminary machine model;
[0039] • Storing a relevance of the selected features for each preliminary machine learning model trained during cross-validation;
[0040] • Repeat the preceding steps in several iterations until all features have been selected at least once; • Order the features according to their average relevance; and
[0041] • Outputting a defined number of features with the highest relevance to the subsequent feature selection method, preferably 20% or less, particularly preferably 15% or less of all generated features.
[0042] The defined number of features is preferably a relative number of the total number of features, but can also be an absolute number of features.
[0043] The defined number of features is preferably a relative number of the total number of features, but can also be an absolute number of features.
[0044] A preliminary machine learning model within the meaning of the present disclosure is preferably based on one of the following classification algorithms: Linear Regression (lasso), Decision Trees, Random Forest (RF) or XGBoost (XGB) with a small number of estimators, or Generalized Linear Model (GLM).
[0045] A preliminary machine learning model within the meaning of the present disclosure is preferably based on one of the following classification algorithms: Linear Regression (lasso), Decision Trees, Random Forest (RF) or XGBoost (XGB) with a small number of estimators, or Generalized Linear Model (GLM).
[0046] In a further advantageous embodiment of the procedure, the cross-validation comprises the following steps:
[0047] • Training the preliminary machine model using the selected features via k-1 / k of the operating data and the state data;
[0048] • Evaluating the trained preliminary machine learning model using 1 / k of the operational and state data; and
[0049] • Determining the relevance of the features based on the evaluation of the trained preliminary machine model.
[0050] The value k corresponds to a proportion of all operational and state data and can take positive values of the real numbers, excluding 0. For further details of cross-validation as described in this disclosure, reference is made to the publication by Ian H. Wtten, Eibe Frank, and Mark A. Hall: Data Mining: Practical Tools and Techniques of Machine Learning, 3rd edition. Morgan Kaufmann, Burlington, MA 2011, ISBN 978-0-12-374856-0 (waikato.ac.nz). In another advantageous embodiment of the procedure, the backward ranking method comprises the following steps:
[0051] • Training a preliminary machine learning model with the features;
[0052] • Removing the least relevant feature;
[0053] • Repeat the preceding steps in several iterations until all features have been removed; and
[0054] • Ordering the features according to their relevance, with the feature removed first receiving the lowest rank and the feature removed last receiving the highest rank; and
[0055] • Outputting a defined number of features with the highest relevance to the subsequent feature selection method, preferably 20% or less, particularly preferably 15% or less of all generated features.
[0056] For further details of a backward ranking as defined in the present disclosure, reference is made to the publication Zhou, Ji-Yuan, et al.: "Prediction of hepatic inflammation in chronic hepatitis B patients with a random forest-backward feature elimination algorithm.", World Journal of Gastroenterology 27.21 (2021): 2910).
[0057] In a further advantageous embodiment of the procedure, the cluster-based selection method comprises the following steps:
[0058] • Calculating a pairwise correlation between all features submitted from the previous feature selection method and the associated state data using Pearson correlation coefficients;
[0059] • Clustering the features based on their correlation using agglomerative clustering techniques;
[0060] • Selecting features from each cluster at random, so that subgroups of features are created whose information is neither redundant nor incomplete;
[0061] • Training and validating preliminary machine nyem models using the selected subset on a single-digit number of folds, particularly five, and cross-validation to determine which subsets lead to the best model performance; and
[0062] • Storing all evaluation metrics for each trained preliminary machine learning model, as well as the significance of the features. For further details on feature clustering as described in this disclosure, see Chormunge, Smita, and Sudarson Jena: "Correlation-based feature selection with clustering for high-dimensional data", Journal of Electrical Systems and Information Technology 5.3 (2018): 542-549.
[0063] In a further advantageous embodiment of the procedure, the Subset Refinement method comprises the following steps, which are applied to at least some subgroups of the features to improve the relevance of the respective subgroup:
[0064] • If the relevance of a feature changes when retraining a preliminary machine learning model, remove the feature;
[0065] • If two features have a high Pearson correlation, remove the feature with the lower relevance; and
[0066] • If a feature has a high Pearson correlation with another feature that is ranked higher in backward ranking, replace the feature with the other feature; and
[0067] • Training the preliminary machine learning model with each of the improved subgroups; selecting the subgroup that most accurately determines the state data based on the features.
[0068] By using a probabilistic classification algorithm, a probability for the occurrence of the defect can be determined from the input data.
[0069] In a further advantageous embodiment, the method also comprises the following steps:
[0070] • Comparing the probability of the defect occurring in the test specimen with a limit value; and
[0071] • Marking the test item as intact or defective based on the comparison.
[0072] By marking the technical equipment as defective or intact, a user can better understand the prediction. In particular, corrective actions can also be carried out based on this comparison. In a further advantageous embodiment, the method includes the following step:
[0073] • Determining the limit value for distinguishing between intact and defective technical equipment based on a cost-benefit analysis between a reactive and a proactive replacement of the test object.
[0074] By taking into account a cost-benefit analysis when marking the test object, in particular the vehicle or vehicle component, a decision can be made that is particularly advantageous for ensuring the function of the test object.
[0075] Further advantages are achieved if the condition data includes a lifetime after which the defect occurred, and in step f), the output data includes a probable lifetime at which the defect will occur in the device under test. A lifetime can be a point in time or a period of time after commissioning, an operating duration, or a combination of both. The condition data includes the lifetime if the defect has actually occurred. For those technical devices where the defect has not occurred, no lifetime value or a very high lifetime value is specified. The term lifetime is also used for defects where continued operation of the technical device or the device under test is possible.
[0076] A second aspect of the invention relates to a system for indirectly measuring the probability of a defect occurring in a test specimen of a class of technical equipment, which are preferably vehicle batteries or vehicles, comprising:
[0077] • Means for measuring condition data in a variety of technical equipment of the type technical equipment, wherein the condition data includes whether the defect has occurred
[0078] • Means for measuring operating data of the multitude of technical equipment of the type and the test specimen, wherein the operating data comprise value profiles of time-resolved measurement parameters and characterize the operating behavior and environment of a respective technical equipment from field operation, such that the operating data of those technical equipment in which the defect occurred are defect-marked; • Means for generating features by processing at least a part of the operating data by means of mathematical operations and / or by selecting data ranges from the operating data;
[0079] • Means of selecting features relevant to the defect from the generated features by means of a chain of feature selection methods; and
[0080] • Means for training a probabilistic classification algorithm using the selected relevant features and their respective state data, thereby generating a classifier; and
[0081] • Means of generating output data, comprising a probability of the occurrence of the defect in the device under test, by applying the classifier to input data based on the operating data of the device under test.
[0082] Features and details described in connection with the method according to the invention are of course also described in connection with the system according to the invention, and vice versa, so that with regard to the disclosure of the individual aspects of the invention, mutual reference is always made or can be made.
[0083] Further advantages are achieved if the system also includes: means of outputting the probability of the defect occurring in the test object.
[0084] A device according to the invention can be configured as hardware and / or software and, in particular, comprises a processing unit, preferably a microprocessor (CPU), preferably connected to a storage and / or bus system via data or signals, and / or one or more programs or program modules. The CPU can be configured to execute instructions implemented as a program stored in a storage system, to acquire input signals from a data bus, and / or to output signals to a data bus. A storage system can comprise one or more, in particular different, storage media, especially optical, magnetic, solid-state, and / or other non-volatile media. The program can be configured to embody the methods described herein.is capable of executing such procedures, enabling the CPU to perform the steps of such methods and, in particular, to analyze at least one technical device or train a classification algorithm. The term "means" as used herein encompasses all structures, materials, or actions set forth herein, as well as all equivalents thereof. Furthermore, the structures, materials, or actions and their equivalents include everything described in the summary, the brief description of the figures, the detailed description, the abstract, and the claims themselves. A system and / or its means may preferably take the form of a pure hardware variant, a pure software variant (including firmware, resident software, microcode, etc.), or a combination of software and hardware aspects, which are generally referred to as a "circuit," "module," or "system."Any combination of one or more computer-readable media can be used. The computer-readable medium can be a computer-readable signaling medium or a computer-readable storage medium.
[0085] The systems and methods according to the present disclosure can preferably be implemented in conjunction with a suitably configured computer, a programmed microprocessor or microcontroller and one or more peripheral integrated circuit elements, an ASIC or other integrated circuit, a digital signal processor, a hard-wired electronic or logic circuit, such as a circuit with discrete elements, a programmable logic device or gate arrangement, such as a programmable logic device (PLD), a programmable logic array (PLA), a field-programmable gate array (FPGA), a programmable logic arrangement (PAL), or a comparable means.In general, any device or means capable of implementing the methodology presented herein can be used to implement the various aspects of this disclosure. Exemplary hardware includes computers, handheld devices, telephones (e.g., cellular, internet-enabled, digital, analog, hybrid, and others), and other hardware known in the art. Some of these devices include processors (e.g., a single or multiple microprocessors), memory, non-volatile memory, input devices, and output devices. Furthermore, alternative software implementations, including but not limited to distributed processing or distributed processing of components / objects, parallel processing, or processing by virtual machines, can be developed to implement the methods described herein. In another particular embodiment of the invention, the term "comprise" can also mean "be."
[0086] Further features and advantages will become apparent from the following description of exemplary embodiments with reference to the figures. These show, at least partially schematically:
[0087] Figure 1 shows a functional block diagram of an embodiment of a system for indirectly measuring the probability of a defect occurring in a test specimen; and
[0088] Figure 2 shows a flowchart of an embodiment of a method for indirectly measuring the probability of a defect occurring in a test specimen.
[0089] In the following, an embodiment of a training of a classification algorithm is explained with reference to Figure 1 and Figure 2, where Figure 1 shows a functional block diagram of a system 10 for the indirect measurement of a probability of the occurrence of a defect in a test object and Figure 2 shows a flow diagram of a procedure 100 executable by means of the system 10 for the indirect measurement of a probability.
[0090] The description refers here to a battery 2 of an electric passenger car 3 as a technical device. However, it is obvious to a person skilled in the art that the described method 100 and system 10 can also be used with regard to other technical devices, in particular the passenger cars 3 themselves.
[0091] For example, the method can be used to determine the probability of a vehicle battery experiencing thermal runaway in a given application. If a high probability of thermal runaway is detected, the battery must be replaced, as otherwise a fire is imminent.
[0092] In a preliminary step 090 of procedure 100, condition data is measured in a large number of technical devices of the type "technical equipment." This condition data includes whether a specific defect has occurred in the respective technical devices. It is sufficient to detect if the specific defect has occurred in some of the technical devices. Preliminary step 090 also records in which technical devices, in this case batteries, the defect "thermal runaway" has occurred. The detection of the defect "thermal runaway" can be carried out after a measurement and a damage report from a workshop or a police report. The damage pattern is clearly and easily measurable in the case of a fire originating from the battery in the vehicle. Other damage patterns can be detected, in particular, using vehicle sensors or in a workshop.This preliminary step 090 is preferably carried out, as in the example shown, using means 11 for measuring condition data in technical facilities.
[0093] In parallel and / or subsequently, in a first step 101 of the method 100, operating data with value profiles of time-resolved measurement parameters, which characterize operating behavior and an environment, are recorded for a plurality of batteries 2 of the same type, in particular of the same design. Preferably, this is accomplished by means of a plurality of sensors 11, which are arranged on the passenger car 3 or in the vicinity of the passenger car 3. Furthermore, preferably, the value profiles can also be historical value profiles of time-resolved measurement parameters. A further advantage is that these operating data are obtained from field operation, so-called telemetry data.
[0094] Based on the results of the measurements in the preliminary work step 090, in the first work step 101, those operating data of the technical equipment, in this case batteries, are marked where the defect "thermal runaway" has occurred. Accordingly, the operating data of those technical equipment where the defect occurred are marked as defective. The operating data of those technical equipment where the defect did not occur remain unmarked or can be marked as intact. This allows for a distinction between "defective" and "intact" data. The recording of the defect "thermal runaway" can be carried out after a measurement and a damage report from a workshop or even a police report. The damage pattern is clearly and easily measurable due to a fire originating from the battery in the vehicle.Other types of damage can also be detected by measurement, particularly using vehicle sensors or in a workshop. The trend lines of time-resolved measurement parameters or time series data describe the use of the passenger car. Suitable measurement parameters include those available via a vehicle control bus system, especially a CAN network.The measurement parameters can be selected from the following group of measurement parameters, with the selection depending on the specific application: ambient temperature, engine temperature, dwell times in specific speed ranges, number of cold starts, number of journeys with a specific minimum duration, number of regenerations, state of charge, number of charging and / or discharging cycles, battery temperatures, cell temperatures, battery control parameters, difference between maximum and minimum voltage per cell, balancing of the frequency numbers of all cell IDs, 2D histogram of the thermal management in "idle" mode. 1and a battery contactor temperature above 25 degrees Celsius, especially during charging; a 2D heat transfer map of the cell temperature and SOC based on the HV bus voltage, especially during charging; differential coolant temperature during charging; and / or differential cell voltage over the battery's discharge cycles. Depending on the application, other measurement parameters can also be used.
[0095] If the category of technical equipment is vehicle batteries, the measurement parameters preferably include at least a number of charging / discharging cycles, battery temperatures, cell temperatures, battery control parameters, a difference between maximum and minimum voltage per cell, a balancing of the frequency numbers of all cell IDs, and a 2D histogram of the thermal management in "idle" mode. 1and a battery contactor temperature above 25 degrees Celsius, especially during charging, a 2D heat map of the cell temperature and SOC based on the HV bus voltage, especially during charging, a differential coolant temperature during charging, and / or a differential cell voltage over the battery's discharge processes.
[0096] If the category of technical equipment is vehicles, the measurement parameters preferably include at least an outside temperature, an engine temperature, dwell times in certain speed ranges, number of cold starts and / or a number of journeys with a certain minimum duration.If the category of technical equipment is "battery-powered vehicles", the measurement parameters preferably include at least an outside temperature, a motor temperature, dwell times in certain speed ranges, number of cold starts and / or a number of journeys with a certain minimum duration, a number of charging / discharging cycles, battery temperatures, cell temperatures, battery control parameters, a difference between maximum and minimum voltage per cell, a balancing of the frequency numbers of all cell IDs, a 2D histogram of the thermal management in "idle" mode and a battery contactor temperature above 25 degrees Celsius, especially during charging, a 2D dwell heat map of the cell temperature and SOC based on the HV bus voltage, especially during charging, a differential coolant temperature during charging, and / or a differential cell voltage over the battery's discharge cycles.
[0097] Furthermore, in the first step, 101 status data of battery 2 and / or vehicle 3 are recorded during field operation. This status data characterizes the physical states of battery 2. Such a physical state can be, in particular, a component specification, a dimension, a tolerance, or a change due to age.
[0098] Preferably, the operating data and the state data of battery 2 are acquired by means of sensors 11, which are arranged in the area of battery 2 and its surroundings or in the area of vehicle 3 and its surroundings and are configured to determine measurement parameters. Furthermore, preferably, the operating data and / or the state data are acquired via an interface 11, in particular a data interface.
[0099] In the aforementioned use case of identifying the probability of a battery 3 thermal runaway, the value profiles of time-resolved measurement parameters are time series with a defined sampling rate, specifically 30 seconds, which were recorded on thousands of battery-powered vehicles 3 as a type of technical equipment. Sixty of the vehicles 3 are identified as having a "defective" physical state, and their operating data are marked or labeled accordingly. The remaining vehicles 3 are identified as having an "intact" physical state and are marked or labeled accordingly. The vehicles 3 associated with the operating data marked as "defective" have defective batteries 2 in that they are no longer operational due to thermal runaway.
[0100] In a second step, features are generated that are suitable for processing by a classification algorithm (Feature Generation / Feature Engineering).
[0101] The time-resolved measurement parameter profiles contained in the operational data consist of individual measurements per journal, for example, one temperature value per minute. These measurements are preferably acquired using operational data acquisition equipment. This equipment can comprise a physical measurement system with one or more sensors, configured to measure and store the time-resolved profiles of the measurement parameters. However, the time-resolved measurement parameters are not typically used for training the classification algorithm. Instead, so-called features are calculated from the time-resolved operational data. These features can be numerical (e.g., size, weight), categorical (e.g., color, type), temporal (e.g., time, date), or geographical (e.g., location coordinates).
[0102] Preferably, the features are generated using mathematical operations and / or by selecting data ranges from the operating data. The features are preferably aggregated values and an alternative description of the time series data. Examples of features include statistics, such as mean, minimum, and maximum values, of parameters like ambient temperature, engine temperature, etc.; duration spent in specific speed ranges; number of cold starts; number of journeys with a certain minimum duration; number of regenerations, etc. Further features include, for example, an elevation profile of a road network, road conditions, ambient temperature, or humidity. Accordingly, the features can be categorized into features that describe the technical equipment or its operation—in this embodiment, the battery 2 or the vehicle 3—and features that describe the environment of the technical equipment.
[0103] The generation of the features is preferably carried out using feature generation means 12. In particular, these means are configured to select data ranges from the operational data and process them using mathematical operations. The means 12 can be configured to generate the features automatically.
[0104] In the aforementioned use case of identifying thermal runaway, more than 4800 features are generated for each of the three vehicles of type 3. Some of the features contain one-dimensional or two-dimensional histograms of the recorded signals. Examples of such signals or time-resolved measurement parameters are cell temperature, voltage, current, ambient temperature, etc. Other features represent summaries of signals and recordings over a year or specific events. Such specific events might include, for example, fast charging of battery 2, normal charging, battery discharges, self-discharge of battery 2, etc. Due to their connection to physical measurement data and thus to physical devices, the features are correlated with possible causes of the defect.
[0105] In a third step 103, features relevant to the defect are selected from the features generated in the second step 102 using a chain of so-called feature selection methods. In the exemplary embodiment, this is achieved by means 13 of selecting features relevant to the defect from the generated features using a chain of feature selection methods. These are explicitly means of selection using random ranking 13A, means of selection using backward ranking 13B, means of selection using cluster-based selection 13C, and means of selection using subset refinement 13D. The selected feature selection methods are executed according to the order defined by the enumeration above.
[0106] In this third step, 103, the aim is to identify those features that contribute most to characterizing the physical states of the vehicle 3 or the battery 2. For example, when selecting features, it can be examined which features contribute most to distinguishing "healthy" from "bad" operating data of the vehicle 3 or the battery 2. "Healthy" operating data characterizes an intact technical device, in particular an intact vehicle 3 or an intact battery 2, insofar as a measurement has shown that the specific defect has not occurred.In intact technical equipment, vehicles, or batteries, defects other than the specific defect may have occurred. In contrast, "sick" operating data characterizes a technical device, particularly a vehicle 3 or a battery 2, insofar as a measurement has revealed the occurrence of the defect whose probability of occurrence is to be measured. The selection is preferably made exclusively on the basis of the recorded operating data or value profiles of time-resolved measurement parameters. The most important characteristics identified can also provide an indication of an underlying malfunction of the "sick" operating data. Therefore, in addition to the indirect measurement of the failure probability in a test object, this data can also be used for a root cause analysis after the detection of a defective vehicle 3 or battery 2 (a so-called "root cause analysis").
[0107] Preferably, two, preferably three, or all of the following feature selection methods are used: Random Ranking 103A, Backward Ranking or Elimination Algorithm 103B, Cluster-Based Selection 103C, Subset Refinement 103D. Preferably, the two, preferably three, or all four feature selection methods are applied in the order listed above. Accordingly, the following combinations can be used: Random Ranking and Backward Ranking, Random Ranking and Cluster-Based Selection, Random Ranking and Subset Refinement, Random Ranking and Backward Ranking and Cluster-Based Selection, Random Ranking and Backward Ranking and Cluster-Based Selection and Subset Refinement, Backward Ranking and Cluster-Based Selection, Backward Ranking and Subset Refinement, Backward Ranking and Cluster-Based Selection and Subset Refinement, Cluster-Based Selection and Subset Refinement.In this context, backward ranking and backward elimination are considered two variants of the same feature selection method.
[0108] The random ranking method 103A has the following steps, as shown in Fig. 2:
[0109] Selections 103A-1 of a random subset of features;
[0110] Performing 103A-2 a multiple, in particular five-fold (k=5), cross-validation for a preliminary machine model with the selected features, whereby the relevance of the selected features for each trained preliminary machine learning model is determined;
[0111] Store 103A-3 a relevance of the selected features for each preliminary machine learning model trained during cross-validation; Repeat 103A-4 the preceding steps in several iterations until all features have been selected at least once;
[0112] Order 103A-5 of the features according to their average relevance; and
[0113] Output 103A-6 of a defined number of features with the highest relevance to the subsequent feature selection method, preferably 20% or less, particularly preferably 15% or less of all generated features.
[0114] The cross-validation process involves the following steps:
[0115] Training 103A-2-1 of the preliminary machine intelligence model using the selected features via k-1 / k of the operational data and the state data;
[0116] Evaluate 103A-2-2 of the trained preliminary machine model using 1 / k of the operational data and the state data; and
[0117] Determine 103A-2-3 the relevance of the features based on the evaluation of the trained preliminary machine model.
[0118] For further details of cross-validation as defined in this disclosure, reference is made to the publication by Ian H. Witten, Eibe Frank and MarkA. Hall: Data Mining: “Practical Tools and Techniques of Machine Learning”, 3rd edition. Morgan Kaufmann, Burlington, MA 2011, ISBN 978-0-12-374856-0 (waikato.ac.nz).
[0119] The Backward Ranking Method 103B preferably includes the following steps:
[0120] Training 103B-1 of a preliminary machine learning model with the features;
[0121] Remove 103B-2, the least relevant feature;
[0122] Repeat step 103B-3 of the preceding steps in several iterations until all features have been removed; and
[0123] Order 103B-4 the features according to their relevance, with the feature removed first receiving the lowest rank and the feature removed last receiving the highest rank; and
[0124] Output 103B-5 of a defined number of features with the highest relevance to the subsequent feature selection method, preferably 20% or less, particularly preferably 15% or less of all generated features.
[0125] For further details of a backward ranking as defined in the present disclosure, reference is made to the publication Zhou, Ji-Yuan, et al.: "Prediction of hepatic inflammation in chronic hepatitis B patients with a random forest-backward feature elimination algorithm.", World Journal of Gastroenterology 27.21 (2021): 2910).
[0126] The Cluster-Based Selection method 103C preferably includes the following steps:
[0127] Calculate 103C-1 a pairwise correlation between all features transmitted from the previous feature selection method and the associated state data using Pearson correlation coefficients;
[0128] Clustering 103C-2 of the features based on their correlation using agglomerative clustering techniques;
[0129] Selections 103C-3 of features from each cluster are made at random, such that subgroups of features are created whose information is neither redundant nor incomplete;
[0130] Training and validating preliminary machine learning models using the selected subset on a single-digit number of folds, in particular five, a cross-validation to determine which subsets lead to the best model performance; and
[0131] Store 103C-5 all evaluation metrics for each trained preliminary machine learning model as well as the meaning of the features.
[0132] The subset refinement method 103D preferably comprises the following steps, which are applied to at least some subsets of features to improve the relevance of the subset: if the relevance of a feature changes upon retraining a preliminary machine learning model, remove 103D-1 of the feature; if two features have a high Pearson correlation, remove 103D-2 of the feature with the lower relevance; and if a feature has a high Pearson correlation with another feature that is ranked higher in backward ranking, replace 103D-3 of the feature with the other feature; and
[0133] Train 103D-4 of the preliminary machine learning model with each of the improved subgroups, selecting the subgroup that most accurately determines the state data. The following algorithms can be used as preliminary machine learning models: linear regression (lasso), decision trees, random forest (RF), or XGBoost (XGB) with a small number of estimators, or a generalized linear model (GLM).
[0134] For each feature selection method 103A, 103B, 103C, 103D, the same or a different machine learning model can be used.
[0135] In the aforementioned use case of indirectly measuring the probability of thermal runaway, operating data from vehicles whose batteries have burned out have been flagged for the defect "thermal runaway." This defect is immediately and unambiguously recognizable to a technician who sees the defective vehicle. Furthermore, as shown in Fig. 1, within the chain of feature selection methods 103, the number of more than 4800 features is reduced in a first step to approximately 890 features by random ranking 103A. In a second step, this number is reduced to approximately 260 features by backward ranking 103B. In a third step, this number is further reduced to approximately 130 features by cluster-based selection 103C. Finally, in a fourth step, this number can be reduced to six relevant features by subset refinement 103D. These remaining features are as follows:
[0136] EVENTS_CHARGING_MEAN_DELTA_CELL_VOLTAGE_STD_LAST_YEAR:
[0137] For each charging cycle, the mean difference between the maximum and minimum voltage per cell is calculated. The standard deviation of the mean values for the past year's data yields the characteristic "Standard deviation of the delta cell voltage across charging events" (aggregated as the average standard deviation of all charging events). This characteristic analyzes the dispersion of the cell voltage difference during charging (delta decrease and increase). It takes into account both large voltage differences and voltage fluctuations during charging. Since the standard deviation is larger for defective vehicles or vehicle batteries than for intact ones, the relevance of this characteristic indicates that the individual cells are being charged with significant unbalance. The cause of this relevant signal can be micro-short circuits at the electrode level, which affect a cell through greatly increased self-discharge.If the voltage difference is a strongly curved line, this indicates a significant state-of-charge (SoC) difference during charging, which persists even after equalization has occurred. The standard deviation of this signal represents the duration of the SoC difference. In critical vehicles, this difference can persist for an extended period.
[0138] AGG_LAST_YEAR_STD_FREQUENCY_COUNTS_BALANCING_CELL_ID:
[0139] The signals of interest, which are part of the operational data, are battery cell balancing status (plural) along with the battery cell IDs of those cells that triggered a balancing cycle. The system counts how often each cell has triggered a balancing cycle, including cells that have never triggered one. Subsequently, a standard deviation of the number of triggered balancing cycles for all cell IDs is calculated. Only data from the past year is considered. In vehicles with a high risk of thermal runaway, the critical cell IDs are grouped around a few cells. This feature takes into account both the balancing duration (with the standard deviation) and the "centralization" of balancing efforts onto a few problematic cells.
[0140] HEAT_BatteryThermalManagementModeHvBatteryContactorTemp_charge_ldle_Higher HVCondatorTemp-PERCENT:
[0141] This feature aggregates the vehicle's dwell time in "idle" thermal management mode when the high-voltage (HV) battery contactor reaches warm / hot contact temperatures. The duration the HV battery contactor remains above a specific threshold, here T>25 °C, during charging is measured. The purpose of this feature is to identify high battery temperatures in thermal mode during "idle" operation. The HV contactor temperature can be affected by various component failures in the battery terminal box. If one of these components is defective, the temperature rises. It appears that this condition is particularly pronounced in faulty vehicles during low-current operation (thermal system idle). The most likely triggers for a thermal event are the cells located beneath the electrical / electronic (EE) unit.
[0142] HEAT_CellTempAvgSocHvBus_charge_HigherCellTempAvg_SOCHVBus_PERCENT: The average cell temperature, calculated from the minimum and maximum cell temperatures, and the state of charge (SOC) based on the HV bus voltage are used as operating data. From these, the vehicle's dwell times at warm / hot battery cell temperatures above a threshold and low SOC during charging are aggregated. This feature correlates high cell temperatures with low bus voltages. This feature is intended to detect whether one or more cells in a battery pack are faulty (especially with improved self-discharge). We consider the SoC based on the total voltage of the battery pack and find that faulty vehicles usually have a lower SoC, even when charged at the same cutoff voltage, and simultaneously exhibit a higher average temperature.The higher the cell temperature and the lower the state of charge of a vehicle, the more likely it is that more than one cell will be affected.
[0143] EVENTS_CHARGING_STD_DELTA_COOLANT_TEMP_AVG_LAST_YEAR:
[0144] The mean delta cooling temperature is calculated across all charging events, aggregated as the standard deviation of all charging events from the past year. The delta cooling temperature is the battery cell coolant temperature at the inlet minus the battery cell coolant temperature at the outlet. The standard deviation of the average delta cooling temperature was calculated for all charging events of the past year. This allows us to determine the variance of the average coolant delta during the charging process (the decrease and increase of the delta). Furthermore, it enables the detection of switching between heating and cooling during the charging process. This measure takes into account both the distribution of the cell temperature within the vehicle battery and its dynamics, which are influenced by the thermal system's control.At first glance, it might seem that the cooling effort during low-power charging unnecessarily increases the temperature distribution within the battery pack. The cause of this issue has been identified as internal cell defects, which often lead to greater self-heating of an affected cell compared to others. If a cell is not well connected to the cooling system, its temperature will not respond well to the cooling system's operation. Calculating the standard deviation of the coolant inlet temperature versus the coolant outlet temperature can provide a signal that illustrates these two effects, which are amplified by the thermal system's control strategy. The standard deviation is high when the cooling system frequently switches to keep the highest cell below its threshold without overcooling the lowest cell.
[0145] E VE NTS_BATTE RY_DI S CH ARG E_ST D_DE LTA_CE LL_VO LTAG E_AVG_G RADTOTAL:
[0146] The delta cell voltage signal (maximum minus minimum cell voltage) is used to calculate the average gradient of this signal for each discharge event. The standard deviation of the average gradient is then calculated for all discharge events over the entire lifespan of the vehicle battery. Cell voltage fluctuations are an important indicator of internal short circuits at the electrode level (burrs and dendrites caused by electrode poisoning). This provides a high-quality measure of voltage noise despite the low data sampling rate. If a battery cell has a faulty condition leading to an increase in cell resistance, this is not easily detected with the low resolution of data points spaced 30 seconds apart.If a delta exists between the minimum and maximum cell voltage during discharge, a resistance difference may be present, even if it cannot be quantified. However, the gradient of this signal indicates a developing difference during operation.
[0147] The third step 103 is preferably performed by means 13 to select relevant features. These means 13 execute several feature selection methods sequentially, particularly in a chained manner. An exemplary chain is shown in Figure 2.
[0148] In a fourth step 104, a classifier 1 is generated by training a probabilistic classification algorithm using the selected relevant features and the respective associated state data.
[0149] The fourth step 104 is carried out by means 14 to train a probabilistic classification algorithm.
[0150] In a fifth step, the output data is determined by applying classifier 1 to input data based on the operating data of the device under test. The output data includes a probability for the occurrence of the defect in the device under test.
[0151] Optionally, the probability of the defect occurring can be compared to a threshold value. If the threshold value is exceeded or not reached by the respective technical device 2, 3, the respective technical device 2, 3 is preferably marked as defective in a sixth step 106. A threshold value is determined to distinguish between intact and defective technical devices of the respective type. Preferably, a cost-benefit analysis is performed for this purpose, comparing a reactive versus a proactive replacement of the test specimen 3, 4 of the respective type. Furthermore, preferably, the threshold value is set such that the respective technical device 2, 3 is marked as defective if there is a comparatively high probability of failure of the vehicle 3 and / or the battery 2.
[0152] Preferably, the limit value is determined on the basis of a cost-benefit analysis between a reactive and a proactive exchange of the test specimen 2, 3.
[0153] The fifth work step 105 and sixth work step 106 are each preferably carried out by means 15 for generating output data and by means (not shown) for marking a technical device 2, 3.
[0154] In an optional seventh step (107), a corrective action is performed on the test specimen if the probability of the defect occurring exceeds a threshold. For the defective batteries, a threshold of 2% reliably distinguished between defective and intact batteries. As a corrective action, a battery replacement was performed. Depending on the number of defective technical equipment, the defect itself, or its cause, an effective corrective action may also include a change in the production of the technical equipment, a change in operating parameters, a recall, repair, maintenance, replacement, a software update, decommissioning, and / or monitoring of the technical equipment.
[0155] The classification algorithm, which is a type of modeling, is trained using supervised learning. The state data consists of binary data classified as "defective" or "intact" with respect to a specific defect. The result is a multi-class classifier suitable for classifying data according to the probability of that defect occurring.
[0156] An example of a classification algorithm that can be used within the scope of the disclosure is logistic regression. This involves applying a logistic function to linear regression, in particular Lasso, Ridge, or Elastic. When trained with binary state data, a binary response is obtained, in the examples given, "intact" or "defective".
[0157] Other possible classification algorithms include generalized linear models, Random Forest, XGBoost, support vector machines, and stacked classifiers. Preferably, interpretable classification algorithms that can also handle outliers well are used.
[0158] The seventh step 107 is preferably carried out by means 17 to perform a defect rectification measure. In a particular embodiment of the invention, such means 17 has an interface, in particular a data interface, for adjusting the operating parameters of a battery management system.
[0159] In a further, optional eighth step, the trained classifier is preferably evaluated. Standard metrics such as Area Under the ROC Curve, Recall, Precision, True Positive Rate, etc., are preferably used here. In particular, the quality of the trained classifier is determined. Typically, the classification algorithm is trained on a first subset of the operational data, and the trained classifier is then evaluated on a second subset of the available operational data. This avoids overfitting.
[0160] The eighth step is preferably performed by means 18 to determine an evaluation of the trained classifier.
[0161] This method 100 can be used for various purposes, such as predictive maintenance, risk assessment, and the continuous monitoring of a vehicle fleet, vehicle components, and other technical equipment. In the described use case of identifying an impending thermal runaway of battery 3, in addition to determining the probability of failure in a majority of electric vehicles, it was found that the root cause was a manufacturing defect. The evaluation of the relevant characteristics showed that calendar aging has only a minimal influence (comparison of Delta_cell_temp_avg (during battery discharge), delta_cell_avg (during normal charging), min_cell_voltage_avg (during battery discharge), delta_soc_bus_per_kWh (during normal charging)).This suggested that there had been a healthy and a defective population from the very beginning, and that the underlying problem stemmed from a production error. The analyses of low_cell_voltage and high_cell_voltage_deltas, as well as high_temperatures and high_temperature_deltas, confirmed this finding. In addition to repairing the defective vehicles by replacing their batteries, this method also allowed the production error to be identified and eliminated by adjusting the production process.
[0162] A wide variety of applications are possible with this method. Two further examples are given below.
[0163] Another application is the investigation of undesirable oil dilution in a type of vehicle, which led to its failure: In this application, based on the selected characteristics (evaluation of diesel particulate filter (DPF) regenerations, number and duration of successful and failed DPF or NOx regenerations), it could be concluded that these characteristics differed significantly between defective and functioning vehicles. Specifically, with more failed regenerations, a slightly higher amount of fuel entered the oil, diluting it and leading to the failures.
[0164] Another exemplary use case is the investigation of inefficient catalytic converter systems in a different type of vehicle: Here, the relevant characteristics indicated that a specific driving behavior led to catalytic converter failure. The main differences between intact and defective vehicles were identified with regard to idling events (more idling events, longer idling duration, more idling events at very cold and hot ambient temperatures, etc.), engine operation (higher percentage of operation at very high engine load), and related P-codes.
[0165] It should be noted that the exemplary embodiments are merely examples and are not intended to restrict the scope of protection, applications, or structure in any way. Rather, the preceding description provides the person skilled in the art with a guideline for implementing at least one exemplary embodiment, whereby various modifications, particularly with regard to the function and arrangement of the described components, can be made without departing from the scope of protection as defined by the claims and these equivalent combinations of features.
Claims
Claims 1. A computer-implemented method (100) for indirectly measuring the probability of a defect occurring in a test specimen of a class of technical equipment, preferably batteries (2) or vehicles (3), comprising the following steps: a) measuring state data in a plurality of technical equipment of the class of technical equipment, wherein the state data include whether the defect has occurred; b) measuring (101) operating data of the plurality of technical equipment of the class and of the test specimen, wherein the operating data include value profiles of time-resolved measurement parameters and characterize an operating behavior and environment of a respective technical equipment from field operation, wherein the operating data of those technical equipment in which the defect has occurred are defect-marked according to the result of the measurement in step a);c) Generating (102) features by processing at least a part of the operational data using mathematical operations and / or by selecting data ranges from the operational data; d) Selecting (103) features relevant to the defect from the generated features using a chain of feature selection methods; e) Training a probabilistic classification algorithm using the selected relevant features and their respective state data, generating a classifier (1); and f) Generating (203) output data, comprising a probability for the occurrence of the defect in the device under test, by applying the classifier (1) to input data based on the operational data of the device under test.
2. The method according to claim 1, wherein the method is carried out automatically.
3. Method according to one of the preceding claims, further comprising the step of: carrying out a defect rectification measure, in particular if the probability of the occurrence of the defect exceeds a limit value.
4. Method according to claim 3, wherein the defect rectification measure comprises calibration of the test specimen, repair of the test specimen, replacement of the test specimen or part of the test specimen and / or disposal of the test specimen.
5. Method (100) according to any of the preceding claims, wherein at least two, preferably three or particularly preferably all, of the following feature selection methods are selected: • Random Ranking (103A); • Backward ranking (103B); • Cluster-Based Selection (103C); • Subset Refinement (103D); wherein the selected feature selection methods are executed in the order defined by the enumeration above.
6. The method of claim 5, wherein the random ranking method (103A) comprises the following steps: • Selecting (103A-1) a random subset of features; • Performing (103A-2) multiple, in particular fivefold (k=5), cross-validation for a preliminary machine model with the selected features, determining the relevance of the selected features for each trained preliminary machine learning model; • Store (103A-3) a relevance of the selected features for each preliminary machine learning model trained during cross-validation; • Repeat (103A-4) the preceding steps in several iterations until all features have been selected at least once; • Order (103A-5) the features according to their average relevance; and • Output (103A-6) a defined number of features with the highest relevance to the subsequent feature selection method, preferably 20% or less, particularly preferably 15% or less of all generated features.
7. The method of claim 6, wherein the cross-validation (103A-2) comprises the following steps: • Training (103A-2-1 ) of the preliminary machine learning model using the selected features by means of k-1 / k of the operational data and the state data; • Evaluate (103A-2-2) the trained preliminary machine learning model using 1 / k of the operational data and the state data; and • Determine (103A-2-3) the relevance of the features based on the evaluation of the trained preliminary machine model.
8. A method according to any one of claims 5 to 7, wherein the backward ranking method (103B) comprises the following steps: • Training (103B-1) a preliminary machine learning model with the features; • Removal (103B-2) of the least relevant feature; • Repeat (103B-3) the preceding steps in several iterations until all features have been removed; and • Ordering (103B-4) the features according to their relevance, with the feature removed first receiving the lowest rank and the feature removed last receiving the highest rank; and • Output (103B-5) a defined number of features with the highest relevance to the subsequent feature selection method, preferably 15% of all generated features.
9. Method (100) according to any one of claims 5 to 8, wherein the cluster-based selection method (103C) comprises the following steps: • Calculate (103C-1) a pairwise correlation between all features submitted from the previous feature selection method and the associated state data using Pearson correlation coefficients; • Clustering (103C-2) the features based on their correlation using agglomerative clustering techniques; • Select (103C-3) features from each cluster at random, so that subgroups of features are created whose information is neither redundant nor incomplete; • Training and validating (103C-4) preliminary machine learning models using the selected subset on a single-digit number of folds, in particular five, a cross-validation to determine which subsets lead to the best model performance; and • Store (103C-5) all evaluation metrics for each trained preliminary machine learning model as well as the meaning of the features.
10. Method (100) according to any one of claims 5 to 9, wherein the Subset Refinement Method (103D) comprises the following steps which are applied to at least some subgroups of the features to improve the relevance of the subgroup: • if the relevance of a feature changes when retraining a preliminary machine learning model, remove (103D-1) the feature; • if two features have a high Pearson correlation, remove (103D-2) the feature with the lower relevance; and • if a feature has a high Pearson correlation with another feature that is ranked higher in backward ranking, replace (103D-3) the feature with the other feature; and • Training (103D-4) the preliminary machine learning model with each of the improved subgroups; selecting the subgroup that most accurately determines the state data.
11. Method (200) according to any of the preceding claims, further comprising the following work steps: (105; 205) compare the probability of the defect occurring in the test specimen with a limit value; and Marking (106; 206) the test item as intact or as defective based on the comparison.
12. The method according to claim 11, further comprising the following step: Determine (104; 204) the limit value for distinguishing between intact and defective technical equipment on the basis of a cost-benefit analysis between a reactive and a proactive replacement of the test item.
13. Method (100; 200) according to one of the preceding claims, wherein the classification algorithm is selected from the following group: naive Bayes, logistic regression, gradient boosting, random forest.
14. Method (100; 200) according to one of the preceding claims, wherein the state data comprise a lifetime after which the defect occurred, and in step f) the output data comprise a probable lifetime at which the defect occurs in the test specimen.
15. System (10) for indirectly measuring the states of a test specimen of a certain type of technical equipment, which are preferably batteries (2) or vehicles (3), by training a classification algorithm, comprising: • Means for recording (11) operational data of a plurality of technical equipment of the genus, wherein the operational data comprise value profiles of time-resolved measurement parameters and characterize an operational behavior and environment of a respective technical equipment from field operation, and of state data of the respective technical equipment during field operation; • Means for generating (12) features by processing at least a part of the operational data by means of mathematical operations and / or by selecting data ranges from the operational data; • Means for selecting (13A, 13B, 13C, 13D) relevant features from the generated features by means of a chain of feature selection methods; and • Means for training (14) the classification algorithm using the relevant features and the respective associated state data, whereby the classifier is generated; wherein the state data characterize the physical states of the technical equipment (2, 3).
6. System (20) for indirectly measuring the probability of a defect occurring in a test specimen of a class of technical equipment, which are preferably batteries (2) or vehicles (3), in particular according to An, comprising: • Means for measuring condition data in a variety of technical equipment of the type technical equipment, wherein the condition data includes whether the defect has occurred; • Means for measuring operating data of the multitude of technical equipment of the type and the test object, wherein the operating data comprise value profiles of time-resolved measurement parameters and characterize an operating behavior and environment of a respective technical equipment from a field operation, such that the operating data of those technical equipment in which the defect occurred are defect-marked; • Means for generating (102) features by processing at least a part of the operational data by means of mathematical operations and / or by selecting data ranges from the operational data; • Means for selecting (103) features relevant to the defect from the generated features by means of a chain of feature selection methods • Means for training a probabilistic classification algorithm using the selected relevant features and their respective state data, thereby generating a classifier; and • Means of generating output data, comprising a probability of the occurrence of the defect in the device under test, by applying the classifier to input data based on the operating data of the device under test.
Citation Information
Patent Citations
Disk fault prediction method and system
CN113722130A
Fault early warning method and device, equipment and storage medium
CN113742163A
Energy storage fault prediction method and device
CN115034471A
Fault diagnosis method and device for oil pumping unit based on multi-source data fusion
CN117909881A
Methods and systems for computation of probabilistic loss of function from failure mode
US20100088538A1