Method and system for identifying the cause of a defect
A method for identifying defect causes in vehicle batteries using operational data and feature selection methods addresses the challenge of symptom-based malfunction identification, enhancing defect detection and maintenance efficiency.
Patent Information
- Application Number
- PCT/AT2025/060216
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-29
- Filing Date
- 2025-05-28
- Publication Date
- 2025-12-04
AI Technical Summary
Existing methods struggle to identify the underlying physical cause of recurring defects in complex products like vehicle batteries or vehicles, particularly high-voltage batteries for electric vehicles, due to the difficulty in deducing or identifying malfunctions based on symptoms.
A computer-implemented method that records operational data, generates features through mathematical operations, selects relevant features using a chain of feature selection methods, and identifies the cause of defects as an intersection of correlated causes and effects, utilizing a classification algorithm trained on these features.
Enables precise identification of defect causes, reduces data processing requirements, and allows for proactive maintenance decisions, improving product reliability and reducing operational costs by identifying the root cause of defects.
Smart Images

Figure AT2025060216_04122025_PF_FP_ABST
Abstract
Description
[0001] Method and system for identifying the cause of a defect
[0002] The invention relates to a computer-implemented method and a system for identifying the cause of a recurring defect in technical equipment of a type, preferably vehicle batteries or vehicles, wherein operating data of a plurality of technical equipment of the type, comprising value profiles of time-resolved measurement parameters and characterizing the operating behavior and environment of a respective technical equipment from field operation, and state data of the respective technical equipment during field operation are acquired, wherein features are generated on the basis of these data by means of mathematical operations which are correlated with different possible causes of the defect and subsequently relevant features are identified by means of a chain of feature selection methods, so that a cause results as an intersection of the possible causes.
[0003] Quality is one of the most important factors influencing a customer's choice between competing products in a given product segment. Therefore, understanding the causal relationship between individual components of a product within the overall product suite and potential defects throughout its entire lifecycle is crucial for the business success of a product and, consequently, of a company.
[0004] An example of the requirements for a product, such as a vehicle, is the lifespan of individual components, such as the battery. This lifespan has a direct impact on the total cost of ownership of a product, which is very important for the end user. In this respect, the product should not exceed a predetermined value.
[0005] To ensure the functionality of a product even after its manufacture and delivery to the customer, it is therefore of great interest to the manufacturer to be able to determine the properties of the product, and in the case of units with multiple technical components, also the properties of individual components within the unit. Such properties are characterized by the physical states of the product.
[0006] The more individual technical components a product or unit comprises, the more important it generally becomes to understand its properties. High-voltage batteries for electric vehicles, or the electric vehicles themselves, are examples of such complex products. With these types of products, it is often difficult to deduce, or even identify, the underlying physical parameter or technical component causing a malfunction, based on a symptom that suggests one or more malfunctions.
[0007] The object of the invention is to provide a system and a method for identifying the cause of a recurring defect in a technical product under investigation. In particular, it aims to determine physical states that are difficult or impossible to measure using previously known methods.
[0008] These tasks are solved through the doctrine of independent claims. Advantageous configurations are claimed in dependent claims.
[0009] A first aspect of the invention relates to a computer-implemented method for identifying the cause of a recurring defect in technical equipment of a type, which are preferably vehicle batteries or vehicles, comprising the following steps:
[0010] • Recording operational data of a large number of technical devices of the type, wherein the operational data include value profiles of time-resolved measurement parameters and characterize an operational behavior and environment of a respective technical device from a field operation and state data of the respective technical device, wherein the operational data of those technical devices in which the defect occurred are defect-marked;
[0011] • Generating features by processing at least part of the operational data using mathematical operations and / or by selecting data ranges from the operational data, wherein the features are each correlated with different possible causes and / or effects of the defect;
[0012] • Selecting features relevant to the defect from the generated features using a chain of feature selection methods, and
[0013] • Identifying the cause of the defect as an intersection or intersections of the different possible causes and / or effects of the defect with which the relevant features are correlated, wherein the procedure is adapted to carry out a defect rectification measure on at least one technical device of the genus.
[0014] For the purposes of this disclosure, data acquisition preferably involves reading measured operating data via a data interface. Alternatively or additionally, data acquisition includes determining a measurement signal using a sensor and / or post-processing a measurement signal to generate the operating data.
[0015] Measurement parameters within the meaning of the present disclosure preferably include an outside temperature, an engine temperature, dwell times in different speed ranges, a number of cold starts, a number of journeys with a certain minimum duration, a number of regenerations, states of charge, charging / discharging processes, battery temperatures, cell temperatures, battery control parameters, a difference of maximum to minimum voltage per cell, a balancing of the frequency numbers of all cell IDs, a 2d histogram of the thermal management in "idle" mode and a temperature of the battery contactor of above 25 degrees Celsius, particularly during the charging process, a 2d dwell time histogram of the cell temperature and the SOC based on the HV bus voltage, particularly during the charging process, a differential coolant temperature during charging processes, and / or a differential cell voltage over the discharge processes of the battery.
[0016] A mathematical operation within the meaning of the present disclosure is statistics, in particular mean, median, minimum, maximum or standard deviation, or aggregation, addition, subtraction, multiplication, division, differentiation, gradient calculation and / or integration over certain durations and / or multivariate integrals, in which measurement parameters are multiplied and then integrated.
[0017] A class of technical devices within the meaning of this disclosure is preferably a group of technical devices that are identical in their essential features and are therefore preferably of identical construction. Preferably, the essential components of the technical devices of a class are identical in construction. A specific technical device is thus preferably an implementation of the class of technical devices. In particular, technical devices of a class differ by tolerances, especially manufacturing tolerances and / or aging or wear effects.Identifying the cause of a recurring defect, as described in the present disclosure, is preferably achieved through indirect measurement by determining a cause-and-effect relationship between the specific defect and a physical-technical trigger, and / or between the defect and an effect that follows the defect, i.e., is triggered by the defect. The invention is based on the premise that every defect has a cause, which can be identified if sufficient data is available, and that defects manifest themselves in effects that are also correlated with the defect. Provided that a sufficient amount of data is available to generate features that can encompass any number of defect causes but at least include the actual defect cause, each cause of the defect can be identified. Deviations from operating data can indicate specific defect causes.If a defect occurs repeatedly, the cause of the defect is generally unknown initially. Since the cause is unknown, correlations between data or data features and the defect are also generally unknown at first. The invention is based on the surprising finding that the cause of the defect can be identified by analyzing the value profiles of time-resolved measurement parameters and correlating features generated from these profiles with different defect causes. The relevant features selected by the method are preferably adapted to identify the cause of the defect.
[0018] An intersection of the relevant features, as defined by the invention, generally refers to a technical overview or causal relationship. In particular, an intersection can be an overlap of possible causes and / or effects, after which only one possible cause of the defect remains. Alternative intersections between the possible causes of the defect can also arise in connection with basic technical knowledge about the technical device, whereby, from the relevant features in relation to each other, only one possible cause of the defect remains, which explains the defect in conjunction with all relevant features. The term "intersection" is to be interpreted broadly. Some defects have more than one cause or manifest themselves through more than one effect.This manifests itself in the fact that several intersections of the different possible causes and / or effects of the defect, with which the relevant features are correlated, are identified side by side. Such behavior can occur, for example, if two causes only lead to a defect if they happen to occur together, but are only weakly or not at all correlated, such as a material defect and operation at a particularly low ambient temperature. Even in this case, identification of the cause of the defect is possible with the method according to the invention, whereby in this case the identification takes place as multiple intersections, i.e., as a combination of several intersections.
[0019] For the purposes of this disclosure, a vehicle is preferably a passenger car, truck, motorcycle, construction machinery, ship, agricultural vehicle, train or other means of transport.
[0020] The invention enables the creation of a predictive model for the condition of a vehicle from measurement data, particularly historical measurement data. This predictive model can then be used to determine the vehicle's physical condition. Furthermore, the method can also be applied to individual vehicle components, such as a battery or fuel cell, an engine, a braking system, a transmission, or others. The condition data includes at least the information on whether a condition is "defective" or "intact," i.e., whether the defect has occurred or not. The acquisition of the condition data can be automated, particularly by a defect monitoring system, and / or performed in a workshop. Additionally, the condition data can include other physical states, such as component specifications, dimensions, and tolerances of the technical equipment.
[0021] The physical states can specify properties of the vehicle component, which are defined as physical quantities or as probabilities for a property.
[0022] The identification method is based on generating features from time-resolved measurement data ("feature generation" or "feature engineering"). This involves, in particular, statistical calculations based on the time-resolved measurement data. From the generated features, the most relevant features—that is, those features that are causally related to the physical state indicated by the labels—are selected ("feature selection"). According to the invention, for the precise identification of the cause of a recurring defect, and thus of a physical state of a technical device, a chain of at least two, three, or four feature selection methods is performed sequentially. This allows for the number of channels with measurement data, or the number of features derived from this measurement data, to be determined.The number of generated features is reduced so that only a comparatively small number of relevant features remain to identify the cause of the defect. Removing irrelevant or redundant features improves model performance. This reduces overfitting, as the model is not influenced by irrelevant data and can focus more effectively on the important patterns. Fewer features mean less data to process, resulting in faster training times and lower storage and processing power requirements. This is especially important for large datasets and complex models.
[0023] Further advantages are achieved if the procedure also includes the step of training a classification algorithm using the relevant features and their respective state data, thereby generating a classifier.
[0024] The classification algorithm is trained (modeled) using previously selected features relevant to the defect and their corresponding labels. This model is preferably evaluated using standard metrics. The trained classifier can then be applied to further time-resolved measurement data from technical equipment of the same type. This process determines the physical states of the vehicles or vehicle components.
[0025] It is preferably possible that the classification algorithm is a probabilistic classification algorithm, the procedure further comprising the following steps:
[0026] • Deriving a probability of failure for the respective technical equipment based on the relevant characteristics and the condition data of the multitude of technical equipment, taking into account the probability of failure of the respective technical equipment when training the classification algorithm.
[0027] If the classification algorithm is trained with a sufficiently large number of time-resolved measurement data from various vehicles or vehicle components, a probability of the defect occurring in the vehicle or vehicle components can be derived.
[0028] In a further advantageous embodiment, the method includes the following step:
[0029] • Evaluating the computer-implemented classifier using at least one standard measure, in particular from the following group of standard measures: Area under the ROC Curve, Recall, Precision, True Positive Rate.
[0030] An evaluation of the trained classifier can verify its accuracy in predicting physical states.
[0031] In a further advantageous embodiment, the process is carried out automatically. The automated execution includes, in particular, the automatic generation of the features, which is a time-consuming process manually, as well as the application of chaining feature selection methods to select the relevant features.
[0032] In a further advantageous embodiment of the method, at least two, preferably three or all of the following feature selection methods are selected:
[0033] • Random Ranking;
[0034] • Backward Ranking or Backward Elimination;
[0035] • Cluster-Based Selection;
[0036] • Subset Refinement; wherein the selected feature selection methods are executed in the order defined by the enumeration above.
[0037] The choice of feature selection methods and their order have a significant impact on the speed of the process and / or, depending on the use case, on the quality of the identified features. This selection and order allow the most relevant features to be reliably identified in different use cases and, furthermore, substantially reduce the number of features that need to be checked when identifying the cause of the defect.
[0038] In a further advantageous embodiment of the procedure, the Random Ranking method comprises the following steps: • Selection of a random subset of features;
[0039] • Performing multiple, in particular five-fold (k=5), cross-validation for a preliminary machine model with the selected features, determining the relevance of the selected features for each trained preliminary machine model;
[0040] • Storing a relevance of the selected features for each preliminary machine learning model trained during cross-validation;
[0041] • Repeat the preceding steps in several iterations until all features have been selected at least once;
[0042] • Ordering the features according to their average relevance; and
[0043] • Outputting a defined number of features with the highest relevance to the subsequent feature selection method, preferably 20% or less, particularly preferably 15% or less of all generated features.
[0044] The defined number of features is preferably a relative number of the total number of features, but can also be an absolute number of features. The number of relevant features sufficient to uniquely identify the cause of the defect is often less than 10, and most frequently between 4 and 8.
[0045] A preliminary machine learning model within the meaning of the present disclosure is preferably based on one of the following classification algorithms: Linear Regression (lasso), Decision Trees, Random Forest (RF) or XGBoost (XGB) with a small number of estimators, or Generalized Linear Model (GLM).
[0046] In a further advantageous embodiment of the procedure, the cross-validation comprises the following steps:
[0047] • Training the preliminary machine model using the selected features via k-1 / k of the operating data and the state data;
[0048] • Evaluating the trained preliminary machine learning model using 1 / k of the operational and state data; and
[0049] • Determining the relevance of the features based on the evaluation of the trained preliminary machine model.
[0050] For further details of cross-validation as defined in this disclosure, reference is made to the publication by Ian H. Witten, Eibe Frank and Mark A. Hall: Data Mining: “Practical Tools and Techniques of Machine Learning”, 3rd edition. Morgan Kaufmann, Burlington, MA 2011, ISBN 978-0-12-374856-0 (waikato.ac.nz).
[0051] In a further advantageous embodiment of the procedure, the backward ranking method comprises the following steps:
[0052] • Training a preliminary machine learning model with the features;
[0053] • Removing the least relevant feature;
[0054] • Repeat the preceding steps in several iterations until all features have been removed; and
[0055] • Ordering the features according to their relevance, with the feature removed first receiving the lowest rank and the feature removed last receiving the highest rank; and
[0056] • Outputting a defined number of features with the highest relevance to the subsequent feature selection method, preferably 20% or less, particularly preferably 15% or less of all generated features.
[0057] The defined number of features is preferably a relative number of the total number of features, but can also be an absolute number of features.
[0058] For further details of a backward ranking as defined in the present disclosure, reference is made to the publication Zhou, Ji-Yuan, et al.: “Prediction of hepatic inflammation in chronic hepatitis B patients with a random forest-backward feature elimination algorithm.”, World Journal of Gastroenterology 27.21 (2021): 2910.
[0059] In a further advantageous embodiment of the procedure, the cluster-based selection method comprises the following steps:
[0060] • Calculating a pairwise correlation between all features submitted from the previous feature selection method and the associated state data using Pearson correlation coefficients;
[0061] • Clustering the features based on their correlation using agglomerative clustering techniques;
[0062] • Selecting features from each cluster at random, so that subgroups of features are created whose information is neither redundant nor incomplete;
[0063] • Training and validating preliminary machine nyem models using the selected subset on a single-digit number of folds, particularly five, and cross-validation to determine which subsets lead to the best model performance; and
[0064] • Store all evaluation metrics for each trained preliminary machine learning model, as well as the meaning of the features.
[0065] For further details of feature clustering as described in this disclosure, please refer to Chormunge, Smita, and Sudarson Jena: “Correlation-based feature selection with clustering for high-dimensional data”, Journal of Electrical Systems and Information Technology 5.3 (2018): 542-549.
[0066] In a further advantageous embodiment of the procedure, the Subset Refinement method comprises the following steps, which are applied to at least some subgroups of the features to improve the relevance of the respective subgroup:
[0067] • If the relevance of a feature changes when retraining a preliminary machine learning model, remove the feature;
[0068] • If two features have a high Pearson correlation, remove the feature with the lower relevance; and
[0069] • If a feature has a high Pearson correlation with another feature that is ranked higher in backward ranking, replace the feature with the other feature; and
[0070] • Training the preliminary machine learning model with each of the improved subgroups; selecting the subgroup that most accurately determines the state data based on the features.
[0071] In a further advantageous embodiment of the method, the classification algorithm is a probabilistic classification algorithm.
[0072] By using a probabilistic classification algorithm, a probability of occurrence can be output for the occurrence of the respective physical states.
[0073] In a further advantageous embodiment, the method also comprises the following steps: • Comparing the status data of the respective technical equipment with a limit value; and
[0074] • Marking the technical equipment as intact or defective based on comparison.
[0075] By marking the technical equipment as defective or intact, a user can better classify the prediction.
[0076] In a further advantageous embodiment, the method includes the following step:
[0077] • Determining the threshold for distinguishing between intact and defective technical equipment based on a cost-benefit analysis between a reactive and a proactive replacement of the technical equipment.
[0078] By taking into account a cost-benefit analysis when marking the technical equipment, in particular the vehicle or vehicle component, a decision can be made that is particularly advantageous for ensuring the function of the technical equipment.
[0079] A second aspect of the invention relates to a system for identifying the cause of a recurring defect in technical equipment of a type of technical equipment, in particular a battery or a vehicle, comprising:
[0080] • Means for recording operational data of a large number of technical equipment of the type, wherein the operational data include value profiles of time-resolved measurement parameters and characterize an operational behavior and environment of a respective technical equipment from field operation, and defect-related condition data of the respective technical equipment during field operation, so that the operational data of those technical equipment in which the defect occurred can be defect-marked;
[0081] • Means of generating features by processing at least a portion of the operational data using mathematical operations and / or by selecting data ranges from the operational data, wherein the features are each correlated with different possible causes of the defect; • Means of selecting features relevant to the defect from the generated features using a chain of feature selection methods; and
[0082] • Means of identifying the cause of the defect as an intersection of the different possible causes of the defect with which the relevant features are correlated.
[0083] Features and details described in connection with the method according to the invention are of course also described in connection with the system according to the invention, and vice versa, so that with regard to the disclosure of the individual aspects of the invention, mutual reference is always made or can be made.
[0084] A means according to the invention can be designed using hardware and / or software and, in particular, comprises a processing unit, preferably a microprocessor (CPU), preferably connected to a storage and / or bus system via data or signals, and / or one or more programs or program modules. The CPU can be configured to execute instructions implemented as a program stored in a storage system, to acquire input signals from a data bus, and / or to output signals to a data bus. A storage system can comprise one or more, in particular different, storage media, especially optical, magnetic, solid-state, and / or other non-volatile media. The program can be designed such that it embodies the methods described herein.is capable of executing such procedures, so that the CPU can perform the steps of such procedures and thus, in particular, analyze at least one technical device or train a classification algorithm.
[0085] The term "means" as used herein encompasses all structures, materials, or actions set forth herein, as well as all equivalents thereof. Furthermore, the structures, materials, or actions and their equivalents include everything described in the abstract, the brief description of the figures, the detailed description, the summary, and the claims themselves. A system and / or its means may preferably take the form of a pure hardware variant, a pure software variant (including firmware, resident software, microcode, etc.), or a combination of software and hardware aspects, generally referred to as a "circuit," "module," or "system." Any combination of one or more computer-readable media may be used. The computer-readable medium may be a computer-readable signaling medium or a computer-readable storage medium.
[0086] The systems and methods according to the present disclosure can preferably be implemented in conjunction with a suitably configured computer, a programmed microprocessor or microcontroller and one or more peripheral integrated circuit elements, an ASIC or other integrated circuit, a digital signal processor, a hard-wired electronic or logic circuit, such as a circuit with discrete elements, a programmable logic device or gate arrangement, such as a programmable logic device (PLD), a programmable logic array (PLA), a field-programmable gate arrangement (FPGA), a programmable logic arrangement (PAL), or a comparable means.In general, any device or means capable of implementing the methodology presented herein may be used to implement the various aspects of this disclosure. Exemplary hardware includes computers, handheld devices, telephones (e.g., cellular, internet-enabled, digital, analog, hybrid, and others), and other hardware known in engineering. Some of these devices include processors (e.g., a single or multiple microprocessors), memory, non-volatile memory, input devices, and output devices. Furthermore, alternative software implementations, including but not limited to distributed processing or distributed processing of components / objects, parallel processing, or processing by virtual machines, may be developed to implement the procedures described herein.
[0087] In another particular embodiment of the invention, the term “comprise” can also mean “be”.
[0088] Further features and advantages will become apparent from the following description of exemplary embodiments with reference to the figures. These show, at least partially schematically:
[0089] Figure 1 shows a functional block diagram of an embodiment of a system for identifying the cause of a recurring defect in technical equipment of a type; and Figure 2 shows a flowchart of an embodiment of a method for identifying the cause of a recurring defect in technical equipment of a type;
[0090] In the following, an embodiment of a training of a classification algorithm is explained with reference to Figure 1 and Figure 2, where Figure 1 shows a functional block diagram of a system 10 for identifying a cause of a recurring defect and Figure 2 shows a flowchart of a procedure 100 executable by means of the system 10 for identifying a cause of a recurring defect in technical equipment of a type of technical equipment.
[0091] The description refers here to a battery 2 of an electric passenger car 3 as a technical device. However, it is obvious to a person skilled in the art that the described method 100 and system 10 can also be used with regard to other technical devices, in particular the passenger cars 3 themselves.
[0092] For example, the method can be used to identify the cause of a vehicle battery's thermal runaway in a specific application. If a cause for the thermal runaway is identified, various corrective actions can be implemented on at least one piece of equipment of the type. The corrective action is selected based on the cause of the defect and may include, among other things, a change in the battery or vehicle manufacturing process, a modification of operating parameters (e.g., charging parameters), a recall, repair, maintenance, battery replacement, a software update, vehicle scrapping, and / or battery monitoring.
[0093] In a first step 101 of the method 100, operating data with value profiles of time-resolved measurement parameters, which characterize the operating behavior and environment, are recorded for a plurality of batteries 2 of the same type, in particular of the same design. Preferably, this is accomplished by means of a plurality of sensors (not explicitly shown) which are arranged on or in the vicinity of the passenger car 3. Furthermore, the value profiles can also be historical value profiles of time-resolved measurement parameters. A further advantage is that these operating data are obtained from field operation, so-called telemetry data.
[0094] The first step involves recording which technical components, in this case batteries, have experienced the "thermal runaway" defect. Accordingly, the operating data of those components where the defect occurred are marked as defective. The operating data of those components where the defect did not occur remains unmarked or can be marked as intact. This allows for a distinction between "defective" and "intact" data. The "thermal runaway" defect can be recorded following a measurement and a damage report from a workshop or even a police report. The damage pattern is clearly and easily measurable in the case of a battery fire within the vehicle. Other types of damage can also be detected, particularly using vehicle sensors or through measurements in a workshop.
[0095] The value profiles of time-resolved measurement parameters or time series data describe the use of the passenger car 3. In particular, those parameters that are available via a bus system of the vehicle control, especially a CAN network, are suitable as measurement parameters.The measurement parameters can be selected from the following group of parameters, with the selection depending on the specific application: ambient temperature, engine temperature, dwell times in specific speed ranges, number of cold starts, number of journeys with a specific minimum duration, number of regenerations, states of charge, number of charge / discharge cycles, battery temperatures, cell temperatures, battery control parameters, difference between maximum and minimum voltage per cell, balancing of the frequency numbers of all cell IDs, 2D histogram of the "Thermal Management" in "idle" mode and a battery contactor temperature above a threshold, especially during charging, 2D heat map of the cell temperature and SOC based on the HV bus voltage, especially during charging, differential coolant temperature during charging, and / or differential cell voltage over the battery's discharge cycles.
[0096] If the category of technical equipment is vehicle batteries, the measurement parameters preferably include at least a number of charging / discharging cycles, battery temperatures, cell temperatures, battery control parameters, a difference between maximum and minimum voltage per cell, a balancing of the frequency numbers of all cell IDs, a 2D histogram of the thermal management in "idle" mode and a battery contactor temperature above 25 degrees Celsius, particularly during charging, a 2D heat map of the cell temperature and SOC based on the HV bus voltage, particularly during charging, a differential coolant temperature during charging, and / or a differential cell voltage over the discharge cycles of the vehicle battery.
[0097] If the category of technical equipment is vehicles, the measurement parameters preferably include at least an outside temperature, an engine temperature, dwell times in specific speed ranges, the number of cold starts, and / or a number of journeys of a specific minimum duration. In addition to vehicle batteries and vehicles, the causes of a recurring defect can, of course, also be identified in other technical components, particularly other vehicle components, within a technical device.
[0098] If the category of technical equipment is "battery-powered vehicles", the measurement parameters preferably include at least an outside temperature, a motor temperature, dwell times in certain speed ranges, number of cold starts and / or a number of journeys with a certain minimum duration, a number of charging / discharging cycles, battery temperatures, cell temperatures, battery control parameters, a difference between maximum and minimum voltage per cell, a balancing of the frequency numbers of all cell IDs, a 2D histogram of the thermal management in "idle" mode and a battery contactor temperature above 25 degrees Celsius, especially during charging, a 2D heat map of the cell temperature and SOC based on the HV bus voltage, especially during charging, a differential coolant temperature during charging, and / or a differential cell voltage over the battery's discharge cycles.
[0099] Furthermore, in the first step, defect-related condition data of battery 2 and / or vehicle 3 are recorded during field operation. This condition data characterizes physical states, particularly with regard to the presence of the defect in battery 2. Such a physical state can also include a component specification, a dimension, a tolerance, or a change in age. Preferably, a physical state is a property of battery 3 of one of the states "intact" or "defective".
[0100] Preferably, the operating data and the state data of battery 2 are acquired by means of sensors which are arranged in the area of battery 2 and its surroundings or in the area of vehicle 3 and its surroundings and which are configured to determine measurement parameters. Furthermore, preferably, the operating data and / or the state data are acquired via an interface, in particular a data interface.
[0101] In the aforementioned use case of identifying the cause of a defect, which was also present in batteries that experienced thermal runaway, the value profiles of time-resolved measurement parameters are time series with defined sampling rates, in particular 30 seconds, which were recorded on thousands of battery-powered vehicles 3 as a type of technical equipment. Sixty of the vehicles 3 are identified as having a defective physical state, and their operating data is marked or labeled accordingly. The remaining vehicles 3 are identified as having an intact physical state and are marked or labeled accordingly. The vehicles 3 associated with the operating data marked as "defective" have defective batteries 2 in that they are no longer operational due to thermal runaway.
[0102] In a second step, features are generated, each correlated with different possible causes of the defect. Furthermore, these features can be suitable for processing by a classification algorithm (feature generation / feature engineering). For example, a feature relating to temperature measurement in a specific operating state can be correlated, depending on the operating state, with excessively high external temperatures, internal short circuits, a particular state of aging, mechanical damage, and / or a misconfigured battery management system. These different possible causes can provide clues to the cause if the feature is selected as relevant to the defect using a chain of feature selection methods.If, however, this or other similarly correlated features are not selected, another cause can be identified by the method according to the invention. The value profiles of time-resolved measurement parameters contained in the operating data consist of individual measured values per journal, for example, one temperature value per minute. These measured values are preferably acquired using means for acquiring operating data 11. These can comprise a physical measuring system with one or more sensors, which is configured to measure and store the value profiles of the measurement parameters in a time-resolved manner. However, for training the classification rhythm, the time-resolved measurement parameters are not normally used; instead, so-called features are calculated from the time-resolved operating data.These characteristics can be numerical, for example size, weight; categorical, for example color, type; temporal, for example time of day, date; or geographical, for example location coordinates.
[0103] Preferably, the features are generated using mathematical operations and / or by selecting data ranges from the operating data. The features are preferably aggregated values and an alternative description of the time series data. Examples of features include statistics, such as mean, minimum, and maximum values, of parameters like ambient temperature, engine temperature, etc.; duration spent in specific speed ranges; number of cold starts; number of journeys with a certain minimum duration; number of regenerations, etc. Further features include, for example, an elevation profile of a road network, road conditions, ambient temperature, or humidity. Accordingly, the features can be categorized into features that describe the technical equipment or its operation—in this embodiment, the battery 2 or the vehicle 3—and features that describe the environment of the technical equipment.
[0104] The generation of the features is preferably carried out using feature generation means 12. In particular, these means are configured to select data ranges from the operational data and process them using mathematical operations. The means 12 can be configured to generate the features automatically.
[0105] In the aforementioned use case of identifying thermal runaway, more than 4800 features are generated for each of the three vehicles of type 3. Some of the features contain one-dimensional or two-dimensional histograms of the recorded signals. Examples of such signals or time-resolved measurement parameters are cell temperature, voltage, current, ambient temperature, etc. Other features represent summaries of signals and records over a year or specific events. Such specific events might include, for example, fast charging of battery 2, normal charging, battery discharges, self-discharge of battery 2, etc. Due to their connection to physical measurement data and thus to physical devices, the features are correlated with possible causes of the defect.
[0106] In a third step 103, features relevant to the defect are selected from the features generated in the second step 102 using a chain of so-called feature selection methods. In the exemplary embodiment, this is achieved by means 13 of selecting features relevant to the defect from the generated features using a chain of feature selection methods. These are explicitly means of selection using random ranking 13A, means of selection using backward ranking 13B, means of selection using cluster-based selection 13C, and means of selection using subset refinement 13D. The selected feature selection methods are executed according to the order defined by the enumeration above.
[0107] In this third step, 103, the characteristics that are essential with regard to the cause of the defect in vehicle 3 or battery 2 are to be identified. For this purpose, the selection of characteristics is examined to determine which characteristics allow for the differentiation between "healthy" and "bad" operating data of vehicle 3 or battery 2. "Healthy" operating data characterizes an intact technical device, in particular an intact vehicle 3 or battery 2, insofar as a measurement has shown that the specific defect has not occurred. However, defects other than the specific defect may have occurred in intact technical devices, vehicles, or batteries. In contrast, "bad" operating data characterizes a technical device, in particular a vehicle 3 or battery 2, insofar as a measurement has shown that the specific defect has occurred.The selection is preferably made exclusively on the basis of the recorded operating data or value trends of time-resolved measurement parameters. The most important characteristics identified can also provide an indication of a malfunction underlying the abnormal operating data. Therefore, these can be used for a root cause analysis after the detection of a defective vehicle 3 or a defective battery 2 (a so-called "root cause analysis").
[0108] Preferably, two, preferably three, or all of the following feature selection methods are used: Random Ranking 103A, Backward Ranking or Backward Elimination 103B, Cluster-Based Selection 103C, Subset Refinement 103D. In particular, two, preferably three, or especially preferably all four feature selection methods are applied in the order listed above. Accordingly, the following combinations can be used: Random Ranking and Backward Ranking, Random Ranking and Cluster-Based Selection, Random Ranking and Subset Refinement, Random Ranking and Backward Ranking and Cluster-Based Selection, Random Ranking and Backward Ranking and Cluster-Based Selection and Subset Refinement, Backward Ranking and Cluster-Based Selection, Backward Ranking and Subset Refinement, Backward Ranking and Cluster-Based Selection and Subset Refinement, Cluster-Based Selection and Subset Refinement.In this context, backward ranking and backward elimination are considered two variants of the same feature selection method.
[0109] The random ranking method 103A has the following steps, as shown in Fig. 2:
[0110] Selections 103A-1 of a random subset of features;
[0111] Performing 103A-2 a multiple, in particular five-fold (k=5), cross-validation for a preliminary machine model with the selected features, whereby the relevance of the selected features for each trained preliminary machine learning model is determined;
[0112] Store 103A-3 a relevance of the selected features for each preliminary machine learning model trained during cross-validation;
[0113] Repeat steps 103A-4 of the preceding steps in several iterations until all features have been selected at least once;
[0114] Order 103A-5 of the features according to their average relevance; and
[0115] Output 103A-6 of a defined number of features with the highest relevance to the subsequent feature selection method, preferably 20% or less, particularly preferably 15% or less of all generated features.
[0116] The cross-validation process involves the following steps: Training 103A-2-1 of the preliminary machine model using the selected features via k-1 / k of the operational data and the state data;
[0117] Evaluate 103A-2-2 of the trained preliminary machine model using 1 / k of the operational data and the state data; and
[0118] Determine 103A-2-3 the relevance of the features based on the evaluation of the trained preliminary machine model.
[0119] For further details of cross-validation as defined in this disclosure, reference is made to the publication by Ian H. Witten, Eibe Frank and MarkA. Hall: Data Mining: “Practical Tools and Techniques of Machine Learning”, 3rd edition. Morgan Kaufmann, Burlington, MA 2011, ISBN 978-0-12-374856-0 (waikato.ac.nz).
[0120] The Backward Ranking Method 103B preferably includes the following steps:
[0121] Training 103B-1 of a preliminary machine learning model with the features;
[0122] Remove 103B-2, the least relevant feature;
[0123] Repeat step 103B-3 of the preceding steps in several iterations until all features have been removed; and
[0124] Order 103B-4 the features according to their relevance, with the feature removed first receiving the lowest rank and the feature removed last receiving the highest rank; and
[0125] Output 103B-5 of a defined number of features with the highest relevance to the subsequent feature selection method, preferably 20% or less, particularly preferably 15% or less of all generated features.
[0126] For further details of a backward ranking as defined in the present disclosure, reference is made to the publication Zhou, Ji-Yuan, et al.: “Prediction of hepatic inflammation in chronic hepatitis B patients with a random forest-backward feature elimination algorithm.”, World Journal of Gastroenterology 27.21 (2021): 2910.
[0127] The Cluster-Based Selection method 103C preferably includes the following steps:
[0128] Calculate 103C-1 a pairwise correlation between all features submitted from the previous feature selection method and the associated state data using Pearson correlation coefficients; cluster 103C-2 the features based on their correlation using agglomerative clustering techniques;
[0129] Selections 103C-3 of features from each cluster are made at random, such that subgroups of features are created whose information is neither redundant nor incomplete;
[0130] Training and validating preliminary machine learning models using the selected subset on a single-digit number of folds, in particular five, a cross-validation to determine which subsets lead to the best model performance; and
[0131] Store 103C-5 all evaluation metrics for each trained preliminary machine learning model as well as the meaning of the features.
[0132] The subset refinement method 103D preferably comprises the following steps, which are applied to at least some subsets of features to improve the relevance of the subset: if the relevance of a feature changes upon retraining a preliminary machine learning model, remove 103D-1 of the feature; if two features have a high Pearson correlation, remove 103D-2 of the feature with the lower relevance; and if a feature has a high Pearson correlation with another feature that is ranked higher in backward ranking, replace 103D-3 of the feature with the other feature; and
[0133] Train 103D-4 of the preliminary machine learning model with each of the improved subgroups; selecting the subgroup that most accurately determines the state data.
[0134] The following algorithms can be used as preliminary machine learning models: Linear Regression (lasso), Decision Trees, Random Forest (RF) or XGBoost (XGB) with a small number of estimators, or Generalized Linear Model (GLM).
[0135] For each feature selection method 103A, 103B, 103C, 103D, the same or a different machine learning model can be used. In the aforementioned use case of identifying the cause of the thermal runaway defect, operating data from vehicles whose batteries have burned out have been flagged for the "thermal runaway" defect. This defect is immediately and unambiguously recognizable to a technician who sees the defective vehicle. Furthermore, in a specific class of battery-electric vehicles, as shown in Fig. 1, the number of features (4845) was reduced to 892 in a first step using random ranking (103A) within the chain of feature selection methods 103. In a second step, this number is reduced to 268 features using backward ranking (103B). In a third step, this number is further reduced to 128 features using cluster-based selection (103C).Finally, in a fourth step, this number can be reduced to six relevant features using Subset Refinement 103D. These remaining features are as follows in the example shown here:
[0136] EVENTS_CHARGING_MEAN_DELTA_CELL_VOLTAGE_STD_LAST_YEAR:
[0137] For each charging cycle, the mean difference between the maximum and minimum voltage per cell is calculated. The standard deviation of the mean values for the past year's data yields the characteristic "Standard deviation of the delta cell voltage across charging events" (aggregated as the average standard deviation of all charging events). This characteristic analyzes the dispersion of the cell voltage difference during charging (delta decrease and increase). It takes into account both large voltage differences and voltage fluctuations during charging. Since the standard deviation is larger for defective vehicles or vehicle batteries than for intact ones, the relevance of this characteristic indicates that the individual cells are being charged with significant unbalance. The cause of this relevant signal can be micro-short circuits at the electrode level, which affect a cell through greatly increased self-discharge.If the voltage difference is a strongly curved line, this indicates a significant state-of-charge (SoC) difference during charging, which persists even after equalization has occurred. The standard deviation of this signal represents the duration of the SoC difference. In critical vehicles, this difference can persist for an extended period.
[0138] AGG_LAST_YEAR_STD_FREQUENCY_COUNTS_BALANCING_CELL_ID: The signals of interest, which are part of the operational data, are battery cell balancing status (plural) along with the battery cell IDs of those cells that triggered battery cell balancing. This counts how often each battery cell has triggered battery cell balancing, including zeros (cells that have never triggered balancing). Subsequently, a standard deviation of the number of triggered battery cell balancings is calculated for all cell IDs. Only data from the last year is considered. In vehicles with a high risk of thermal runaway, the critical cell IDs are grouped around a few battery cells. This feature takes into account both the balancing duration (with the standard deviation) and the "centralization" of balancing efforts onto a few problematic cells.
[0139] HEAT_CellTempAvgSocHvBus_charge_HigherCellTempAvg_SOCHVBus_PE CENT: This parameter aggregates the vehicle's dwell time in "idle" thermal management mode when the high-voltage (HV) battery contactor reaches warm / hot contact temperatures. The duration the HV battery contactor remains above a specific threshold, here T>25 °C, during the charging process is measured. The purpose of this feature is to identify high battery temperatures in thermal mode during "idle" operation. The HV contactor temperature can be affected by various component failures in the battery terminal box. If one of these components is defective, the temperature rises. It appears that this condition is particularly pronounced in faulty vehicles during low-current operation (thermal system idle). The most likely triggers for a thermal event are the cells located beneath the electrical / electronic (EE) unit.
[0140] EVENTS_CHARGING_STD_DELTA_COOLANT_TEMP_AVG_LAST_YEAR:
[0141] The mean delta cooling temperature is calculated across all charging events, aggregated as the standard deviation of all charging events from the past year. The delta cooling temperature is the battery cell coolant temperature at the inlet minus the battery cell coolant temperature at the outlet. The standard deviation of the average delta cooling temperature was calculated for all charging processes of the past year. This allows us to determine the variance of the average coolant delta during the charging process (decreases and increases in the delta). Furthermore, it enables the detection of switching between heating and cooling during the charging process. This measure takes into account both the distribution of the cell temperature within the vehicle battery and its dynamics, which are caused by the thermal system's control.At first glance, it might appear that the cooling effort during low-power charging unnecessarily increases the temperature distribution within the battery pack. The cause of this issue has been identified as internal cell defects, which often lead to greater self-heating of an affected cell compared to others. If a cell is not well connected to the cooling system, its temperature will not respond well to the cooling system's operation. Calculating the standard deviation of the coolant inlet temperature versus the coolant outlet temperature can provide a signal that illustrates these two effects, which are amplified by the thermal system's control strategy. The standard deviation is high when the cooling system frequently switches to keep the highest cell below its threshold without overcooling the lowest cell.
[0142] E VE NTS_BATTE RY_DI SCH ARG E_ST D_DE LTA_CE LL_VO LTAG E_AVG_G RADTOTAL: The delta cell voltage signal (maximum minus minimum cell voltage) is used to calculate the average gradient of this signal for each discharge event. The standard deviation of the average gradient is calculated for all discharge events over the entire lifespan of the vehicle battery. The cell voltage fluctuation is an important indicator of internal short circuits at the electrode level (burrs and dendrites caused by electrode poisoning). This is a high-quality measure of voltage noise despite the low data sampling rate. If a battery cell has a faulty condition that leads to an increase in cell resistance, this is not easily detected in the low resolution of data points spaced 30s apart.If a delta exists between the minimum and maximum cell voltage during discharge, a resistance difference may be present, even if it cannot be quantified. However, the gradient of this signal indicates a developing difference during operation.
[0143] The third step 103 is preferably performed by means 13 to select relevant features. These means 13 execute several feature selection methods sequentially, particularly in a chained manner. An exemplary chain is shown in Figure 2.
[0144] In a fourth step, 104, a cause of the defect is identified as Intersection 1 of the different possible causes of the defect with which the relevant characteristics are correlated. While each characteristic spans a space of different possible causes of the defect, Intersection 1 of these, after the procedure has been carried out, is a specific cause of the defect that is identified so precisely that it can be eliminated with a defect rectification measure. In some cases, the single cause of the defect may also be a combination of different causes. These, too, can be identified by the procedure in the same way.
[0145] The fourth step 104 is preferably carried out by means 14 to identify a cause of the defect as an intersection 1 of the different possible causes of the defect with which the relevant features are correlated. In the illustrated embodiment, a manufacturing defect was identified as the cause of the defect.
[0146] In an optional fifth step (105), a modification of the battery manufacturing process was therefore implemented as a defect rectification measure for the identified defect. This eliminated defects in a large number of technical devices of this type, depending on the identified cause of the defect. However, these devices were either already in production or were scheduled for future production. Depending on the defect, the defect rectification measure could also include a change in operating parameters, a recall, repair, maintenance, replacement, a software update, decommissioning, and / or monitoring of the technical device.
[0147] In a further optional sixth step 106, a classification algorithm is trained using the relevant features selected in the third step and the respective associated state data.
[0148] For training the classification algorithm, which is a type of modeling, a technique called supervised learning is used. As described, the condition data consists of binary data classified as "defective" or "intact" with respect to a specific defect. The result can also be a multi-class classifier suitable for classifying data according to the probability of this defect occurring. A multi-class classifier can be adapted to classify operational data according to the probability of this specific defect occurring. The classifier uses the correlation between the relevant features selected based on the defect-marked operational data and the operational data to be tested.
[0149] An example of a classification algorithm that can be used within the scope of the disclosure is logistic regression. This involves applying a logistic function to linear regression, in particular Lasso, Ridge, or Elastic. When trained with binary state data, a binary response is obtained, in the examples given, "intact" or "defective".
[0150] Other possible classification algorithms include generalized linear models, Random Forest, XGBoost, support vector machines, and stacked classifiers. Preferably, interpretable classification algorithms that can also handle outliers well are used.
[0151] The sixth step 106 is preferably carried out by means 17 for training a classification algorithm. In particular, such means 17 has an interface, especially a data interface, to output the generated and selected features as well as the associated state data, which characterize physical states of the vehicle 3 or the battery 2, to the classification algorithm to be trained.
[0152] In a further optional seventh step (107), the trained classifier is evaluated. Here, standard metrics such as Area Under the ROC Curve, Recall, Precision, True Positive Rate, etc., are preferably used. In particular, the quality of the trained classifier is determined. Typically, the classification algorithm is trained on a first subset of the operational data, and the trained classifier is then evaluated on a second subset of the available operational data. This avoids overfitting.
[0153] The eighth step 107 is preferably performed by means of determining an evaluation of the trained classifier.
[0154] A trained classifier is adapted to be applied to unknown operating data of technical equipment 2, 3 of the same type in order to determine the probability of a defect occurring in this equipment 2, 3. In particular, the trained classifier can also be applied to those technical equipment 2, 3 on which the underlying classification algorithm was trained. This allows the probabilities of the defect occurring in these equipment 2, 3 to be determined. In other words, given the large number of technical equipment 2, 3 on which the classifier was trained, the trained classifier can determine the probability of a specific physical state occurring, in particular a failure probability, for the technical equipment 2, 3 under investigation.The probability of a defect occurring in a particular technical device is a physical state to which a specific value can be assigned and which can be measured indirectly using this method.
[0155] This method 100 can be used for various purposes, such as predictive maintenance, risk assessment and continuous monitoring of a vehicle fleet, vehicle components and other technical equipment.
[0156] In the described use case of identifying an impending thermal runaway of battery 3, it was determined that the root cause was a manufacturing defect. Analysis of the relevant characteristics revealed that calendar aging had only a minimal influence (comparison of Delta_cell_temp_avg (during battery discharge), delta_cell_avg (during normal charging), min_cell_voltage_avg (during battery discharge), and delta_soc_bus_per_kWh (during normal charging)). This indicated that a healthy and a defective population had existed from the outset and that the underlying problem stemmed from a manufacturing defect. Analysis of low_cell_voltage and high_cell_voltage deltas, as well as high_temperatures and high_temperature_deltas, confirmed this indication.
[0157] Numerous use cases are possible. Two further examples are given below.
[0158] Another application is the investigation of oil dilution in a specific vehicle category: In this case, based on the measurement parameters and the selected characteristics (evaluation of diesel particulate filter (DPF) regenerations, number and duration of successful and failed DPF and NOx regenerations), it could be concluded that the characteristics of defective vehicles differed significantly from those of functioning vehicles. Specifically, with more failed regenerations, a slightly higher percentage of fuel entered the oil, diluting it and leading to breakdowns.
[0159] Another exemplary use case is the investigation of the inefficiency of the catalytic converter system in a different type of vehicle: Here, the relevant characteristics indicated that a specific driving behavior led to catalytic converter failure. The main differences between intact and defective vehicles were identified with regard to idling events (more idling events, longer idling duration, more idling events at very cold and hot ambient temperatures, etc.), engine operation (higher percentage of operation at very high engine load), and related P-codes.
[0160] It should be noted that the exemplary embodiments are merely examples and are not intended to restrict the scope of protection, applications, or structure in any way. Rather, the preceding description provides the person skilled in the art with a guideline for implementing at least one exemplary embodiment, whereby various modifications, particularly with regard to the function and arrangement of the described components, can be made without departing from the scope of protection as defined by the claims and these equivalent combinations of features.
Claims
Claims 1. Computer-implemented method (100) for identifying a cause of a recurring defect in technical equipment of a type, which are preferably vehicle batteries (2) or vehicles (3), comprising the following steps: • Recording (101) operational data of a plurality of technical equipment of the type, wherein the operational data comprise value profiles of time-resolved measurement parameters and characterize an operational behavior and environment of a respective technical equipment from a field operation, and defect-related condition data of the respective technical equipment, wherein the operational data of those technical equipment in which the defect has occurred are defect-marked; • Generating (102) features by processing at least part of the operational data by means of mathematical operations and / or by selecting data ranges from the operational data, wherein the features are each correlated with different possible causes and / or effects of the defect; • Selections (103) of features relevant to the defect from the generated features by means of a chain of feature selection methods, and • Identifying (104) the cause of the defect as an intersection (1) or intersections of the different possible causes and / or effects of the defect with which the relevant features are correlated, the procedure being adapted to carry out a defect rectification measure on at least one technical device of the genus.
2. The method of claim 1, wherein the method further comprises the step of: performing a defect rectification measure on at least one technical device of the genus, wherein the defect rectification measure preferably involves a change in a manufacturing process, a change in operating parameters, a recall, a repair, a maintenance, a replacement, a software update, a decommissioning and / or a monitoring of the technical device.
3. Method according to one of claims 1 or 2, further comprising the step: training a classification algorithm using the relevant features and the respective associated state data, wherein a classifier is generated.
4. Method (100) according to any one of the preceding claims, wherein at least two, preferably three or particularly preferably all of the following feature selection methods are selected: • Random Ranking (103A); • Backward ranking (103B); • Cluster-Based Selection (103C); • Subset Refinement (103D); wherein the selected feature selection methods are executed in the order defined by the enumeration above.
5. The method of claim 3, wherein the random ranking method (103A) comprises the following steps: • Selecting (103A-1) a random subset of features; • Performing (103A-2) multiple, in particular fivefold (k=5), cross-validation for a preliminary machine model with the selected features, determining the relevance of the selected features for each trained preliminary machine learning model; • Store (103A-3) a relevance of the selected features for each preliminary machine learning model trained during cross-validation; • Repeat (103A-4) the preceding steps in several iterations until all features have been selected at least once; • Order (103A-5) the features according to their average relevance; and • Output (103A-6) a defined number of features with the highest relevance to the subsequent feature selection method, preferably 20% or less, particularly preferably 15% or less of all generated features.
6. The method of claim 5, wherein the cross-validation (103A-2) comprises the following steps: • Training (103A-2-1 ) of the preliminary machine learning model using the selected features by means of k-1 / k of the operational data and the state data; • Evaluate (103A-2-2) the trained preliminary machine learning model using 1 / k of the operational data and the state data; and • Determine (103A-2-3) the relevance of the features based on the evaluation of the trained preliminary machine model.
7. A method according to any one of claims 4 to 6, wherein the backward ranking method (103B) comprises the following steps: • Training (103B-1) a preliminary machine learning model with the features; • Removal (103B-2) of the least relevant feature; • Repeat (103B-3) the preceding steps in several iterations until all features have been removed; and • Ordering (103B-4) the features according to their relevance, with the feature removed first receiving the lowest rank and the feature removed last receiving the highest rank; and • Output (103B-5) a defined number of features with the highest relevance to the subsequent feature selection method, preferably 20% or less, particularly preferably 15% or less of all generated features.
8. Method (100) according to any one of claims 4 to 7, wherein the cluster-based selection method (103C) comprises the following steps: • Calculate (103C-1) a pairwise correlation between all features submitted from the previous feature selection method and the associated state data using Pearson correlation coefficients; • Clustering (103C-2) the features based on their correlation using agglomerative clustering techniques; • Select (103C-3) features from each cluster at random, so that subgroups of features are created whose information is neither redundant nor incomplete; • Training and validating (103C-4) preliminary machine learning models using the selected subset on a single-digit number of folds, in particular five, a cross-validation to determine which subsets lead to the best model performance; and • Store (103C-5) all evaluation metrics for each trained preliminary machine learning model as well as the meaning of the features.
9. Method (100) according to any one of claims 4 to 8, wherein the Subset Refinement Method (103D) comprises the following steps which are applied to at least some subgroups of the features to improve the relevance of the subgroup: • if the relevance of a feature changes when retraining a preliminary machine learning model, remove (103D-1) the feature; • if two features have a high Pearson correlation, remove (103D-2) the feature with the lower relevance; and • if a feature has a high Pearson correlation with another feature that is ranked higher in backward ranking, replace (103D-3) the feature with the other feature; and • Training (103D-4) the preliminary machine learning model with each of the improved subgroups; selecting the features of the subgroup that most accurately determine the state data.
10. Method (100; 200) according to any one of claims 3 to 9, wherein the classification algorithm is selected from the following group: naive Bayes, logistic regression, gradient boosting, random forest.
11. System (10) for identifying a cause of a recurring defect in technical equipment of a type of technical equipment, which are preferably vehicle batteries (2) or vehicles (3), comprising: • Means for recording (11) operational data of a large number of technical devices of the type, wherein the operational data comprise value profiles of time-resolved measurement parameters and operational behavior and characterize the environment of a respective technical device from field operation, and defect-related condition data of the respective technical device so that the operating data of those technical devices in which the defect occurred can be marked as defective; • Means for generating (12) features by processing at least a part of the operational data by means of mathematical operations and / or by selecting data ranges from the operational data, wherein the features are each correlated with different possible causes of the defect; • Means for selecting (13A, 13B, 13C, 13D) features relevant to the defect from the generated features by means of a chain of feature selection methods; and • Means of identifying the cause of the defect as an intersection (1) of the different possible causes of the defect with which the relevant features are correlated.
12. System (20) for the indirect measurement of physical states of a technical device of a class of technical devices, which are preferably vehicle batteries (2) or vehicles (3), comprising: • a system according to claim 11; • Means for training a classification algorithm using the relevant features and the respective associated state data, whereby a classifier adapted to determine a probability of occurrence of the defect is generated. • Means of applying the classifier to operational data of a technical device of the type in which the defect has not occurred.
Citation Information
Patent Citations
Remote fault solution method and system for battery swap station
CN115310552A
Method, device, equipment, system and medium for determining fault cause of power supply module
CN117290151A
Fault positioning method and device, electronic equipment and storage medium
CN117439868A
Fault diagnosis method and device for oil pumping unit based on multi-source data fusion
CN117909881A
Vehicle fault root cause diagnosis
US20200184742A1