Predictive Maintenance for Industrial Machines

JP2024522982A5Pending Publication Date: 2025-06-17PAUL WURTH SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023572536
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-11
Filing Date
2022-06-10
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing predictive maintenance models for industrial machines lack accuracy due to missing data, lack of expert annotations, and erroneous data associations, leading to incorrect maintenance schedules and machine downtime.

Method used

A modular structure of processing modules is employed, including first and second dependent modules and an operating mode classifier, trained through cascade training, to process machine data and predict failures with improved accuracy by determining intermediate status indicators and operating modes.

Benefits of technology

The modular approach enhances predictive accuracy by processing machine data to identify failure times, types, and remaining service life, allowing for more precise maintenance scheduling and reduced downtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The computer implemented failure predictor has a modular structure (373) with first and second subordinate modules (313, 323) subordinate to an output module (363). The first and second subordinate modules process data from the industrial machine to determine first and second intermediate status indicators. A third subordinate module (333) determines an operating mode indicator, and the output module (363) processes the status indicator and the operating mode indicator to predict failure of the industrial machine. The modular structure is trained by cascade training including training the subordinate modules (312, 322, 332), followed by operating the trained subordinate modules, and followed by training the output module.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates generally to industrial machines, and more particularly, the present disclosure relates to a computer system, method and computer program product for predicting failures of industrial machines. [Background technology]

[0002] Industrial machines that operate continuously without interruption are as rare as perpetual motion machines.

[0003] Briefly, there are at least two main reasons for interruptions: A machine operator usually shuts down the machine for maintenance at regular intervals, or the machine may stop due to a malfunction.

[0004] During the past decades, computer models have made great strides in predicting failures. If a failure is predicted, so-called predictive maintenance models allow operators to shut down machines for maintenance. Such an approach can increase the overall time that a machine is in operation and reduce the time that the machine is not in operation.

[0005] A computer model receives sensor data (and other data) from the machine and predicts failure details such as time to failure, type of failure, etc. A computer model needs to know the causal relationships, but in many cases, such relationships are not known, so the computer is trained on training data (usually a combination of historical sensor data and historical failure data). The training approximates those relationships.

[0006] The accuracy of prediction is important. For example, a computer can predict that a failure will occur within a week and the operator is likely to shut down the machine for immediate maintenance. A false prediction is fatal. If the prediction is wrong, the immediate maintenance was not actually needed and the machine could have worked normally without interruption.

[0007] In improving accuracy, experts face many challenges and constraints, such as possible lack of data (e.g., sensor and failure data), possible lack of expert annotation (to identify past failures), possible differences in annotations from different experts, possible incorrect relevance assessment of data, etc. Other challenges are discussed below, but in general there is a need to improve the accuracy of prognosis.

[0008] Stich et al. describe the use of multiple computer models to classify subcomponents of a complex industrial system, a wafer fab (STICH PETER ET AL: “Yield prediction in semiconductor manufacturing using an AI-based cascading classification system”, 2020 IEEE INTERNATIONAL CONFERENCE ON ELECTRO INFORMATION TECHNOLOGY (EIT), IEEE, 31 July 2020 (2020-07-31), pages 609-614).

[0009] US 2013 / 0132001 A1 describes fault detection and prediction in industrial equipment using models. This document discusses detailed examples and also refers to training of the models.

[0010] summary In short, prognosis does not emerge from a single functional module that receives machine data and provides prognostic data, but from a modular structure with output modules and dependent modules. In that sense, the modular structure is implementing a meta-model in which output modules predict failures by processing intermediate indicators from dependent modules (or base models).

[0011] The hierarchical arrangement of modules also has a training effect: subordinate modules are trained before higher level modules.

[0012] More specifically, the modular arrangement includes first and second intermediate modules subordinate to the output module. At least the first and second subordinate modules process the machine data to determine first and second intermediate status indicators, respectively. Such status indicators can be related to operating configurations of the industrial machine.

[0013] In parallel with this, a further subordinate module - an operating mode classifier - also receives the sensor data and determines the operating mode (operating mode indicator) of the industrial machine. An output module processes the intermediate status indicators and the operating mode indicator and predicts faults of the industrial machine. Since the faults are related to different operating modes, the prediction accuracy can be improved compared to the above-mentioned single functional module.

[0014] The figure also illustrates a computer program or computer program product that, when loaded into a computer's memory and executed by at least one processor of the computer, causes the computer to perform the steps of a computer-implemented method. That is, the program provides instructions for the modules. Similarly, a computer system including a number of processing modules performs the steps of the computer-implemented method when they are executed by the computer system.

[0015] The invention relates to a computer-implemented method for predicting failures of an industrial machine as claimed in claim 1. The computer-implemented method for predicting failures of an industrial machine is in this case a method in which a computer uses a structure of processing modules (for the sake of brevity, the attribute "processing" may be omitted from the text). The computer receives machine data from the industrial machine by means of first, second and third subordinate processing modules, which are arranged to provide intermediate data to an output processing module. This structure has been pre-trained by cascade training. The computer processes the machine data by means of the first subordinate module to determine a first intermediate status indicator. The computer processes the machine data by means of the second subordinate module to determine a second intermediate status indicator. The computer processes the machine data by means of a third subordinate module, which is an operation mode classifier module, to determine an operation mode indicator of the industrial machine. The computer processes the first and second intermediate status indicators and the operation mode indicator by means of an output module. This allows the output module to predict failures in industrial machines by providing prognostic data.

[0016] Optionally, the computer uses the trained structure according to the following training sequence: train a third dependent module with the historical machine data; process the historical machine data to execute the trained third dependent module to obtain a historical mode indicator; train the first and second dependent modules with the historical machine data and the historical mode indicator; process the historical machine data to execute the trained first and second dependent modules to obtain first and second intermediate status indicators; and train an output module with the historical mode indicator, the historical machine data, and the historical failure data.

[0017] Optionally, in determining the operation mode index, the computer uses an operation mode classifier that is trained based on historical machine data annotated by experts.

[0018] Optionally, the historical machine data annotated by the experts is sensor data.

[0019] Optionally, the operation mode classifier has been trained based on historical machine data. During training, the operation mode classifier has grouped operation times of the machine into clusters of time series segments.

[0020] Optionally, clusters of time series segments are assigned or selected to operational mode indicators, either automatically or through interaction with an expert.

[0021] Optionally, the operating mode indicator is provided by a number of mode changes over time.

[0022] Optionally, the status indicator is selected from a current indicator indicating a current status, and a predictive indicator indicating a future status.

[0023] Optionally, the output module predicts failure of the industrial machine selected from time to failure, type of failure, remaining useful life, and time between failures.

[0024] Optionally, the operating mode indicator serves as a bias that is further processed by both the first and second subordinate processing modules.

[0025] Optionally, the computer receives the machine data by receiving a sub-set comprising the sensor data, and the computer determines first and second intermediate status indicators by first and second subordinate modules processing the sub-set comprising the sensor data.

[0026] Optionally, the computer receives the machine data, which includes receiving the data through data harmonizers that provide the machine data through virtual sensors or filter the incoming machine data - depending on the contribution of the machine data to the failure prediction.

[0027] Optionally, the computer receives the machine data via a data harmonizer, the act including receiving the machine data from a harmonizer with a module pre-trained by transfer learning.

[0028] Optionally, the computer receives machine data at least partially augmented with data obtained from the simulation.

[0029] From a broader perspective, current methods for predicting failures of industrial machines can be adapted for use in transferring predictive data to a machine controller, which can cause the industrial machine to assume a mode that predicts at the latest a time until failure occurs, and the controller can / enable the industrial machine to assume a mode that predicts at the latest a time until maintenance of the machine occurs.

[0030] Additionally, the industrial machine may be adapted to provide machine data to the computer (i.e., adapted to perform the method). The industrial machine may be further adapted to receive prognostic data from the computer. In such a scenario, the industrial machine is associated with a machine controller that switches the operating mode of the industrial machine according to predefined optimal goals.

[0031] Optionally, select a predefined optimality goal from: It avoids maintenance as much as possible and operates in a mode where failure is predicted to occur at the latest.

[0032] Industrial machines can be selected from chemical reactors, metallurgical furnaces, vessels, pumps, motors and engines.

[0033] There is also a computer-implemented method for training a modular structure having first, second and third dependent modules coupled to an output module to enable the modular structure to provide fault indications with fault prognosis for an industrial machine, the method including applying cascade training to train the dependent modules, followed by operating the trained dependent modules, followed by training the output module.

[0034] Optionally, the cascade training includes training a third sub-ordinate module with the historical machine data: executing the trained third sub-ordinate module to obtain a historical mode indicator by processing the historical machine data; training the first and second sub-ordinate modules with the historical machine data and the historical mode indicator; executing the trained first and second sub-ordinate modules by processing the historical machine data to obtain first and second intermediate status indicators, and training an output module with the historical mode indicator, the historical machine data, and the historical failure data.

[0035] From a further perspective, the computer implemented failure predictor has a modular structure with first and second subordinate modules subordinate to an output module. The first and second subordinate modules process data from the industrial machine to determine first and second intermediate status indicators. A third subordinate module determines an operating mode indicator, and the output module processes the status indicator and the operating mode indicator to predict failure of the industrial machine. The modular structure is trained by cascade training, which includes training a subordinate module, then operating the trained subordinate module, and thereafter training the output module. [Brief description of the drawings]

[0036] Embodiments of the present invention will now be described in detail with reference to the accompanying drawings, in which: [Figure 1A] 1 shows an industrial machine and a modular structure. [Figure 1B] 1 shows an industrial machine and a modular structure. [Diagram 2] 1 shows a modular structure with subordinate modules hierarchically below the output module. [Diagram 3] 1 shows a time diagram for the operation of an industrial machine in combination with failure intervals in failure prediction. [Figure 4] 1 shows a time diagram for the operation of an industrial machine in combination with mode-specific failure intervals. [Diagram 5] 1 shows a block diagram of an industrial machine. [Figure 6] Show multivariate time series with historical data. [Figure 7] 1 shows a simplified time diagram for cascade training. [Figure 8] 1 shows a simplified time diagram for cascade training in a variation. [Figure 9] 1 shows a flowchart of a computer-implemented method for predicting failures in industrial machines. [Figure 10] As an example, a time series with mode indices for two modes is shown, for optional determination of mode change rates. [Figure 11] A state transition diagram involving mode transitions is shown. [Figure 12] 1 illustrates multiple industrial machines and a historical time series with machine data and a historical time series with fault data. [Figure 13] 1 shows different industrial machines in an approach to reconcile machine data (and possible fault data Q). [Figure 14] 1 shows machine data in a time series with data provided by a sensor and data provided by a data processor. [Figure 15] A general purpose computer is shown. Detailed Description

[0037] Overview and Conventions This description uses a top-down approach by showing the industrial machine and module structure in Figures 1A, 1B and 2, discussing accuracy related to operation modes with simplified time diagrams in Figures 3-4, and showing the details of the industrial machine in Figure 5. Figure 6 discusses time series with machine data separated by operation modes. Training is then discussed with reference to Figures 7-8, and prognosis is discussed with the flow chart in Figure 9. Further aspects are mentioned in Figures 10-15 as well.

[0038] The description uses phrases such as "executing a module" or "executing a computer" to describe computer activities, and words containing "operates" to describe machine activities.

[0039] Industrial Machine Modular Construction 1A and 1B provide an overview of the approach in space (FIG. 1A) and time (FIG. 1B).

[0040] 1A shows an industrial machine 113 and a computer with a modular structure 373. The machine 113 provides (current) machine data 153 {{X1...XM}}N (or {{X...}}N for short) to the input of the modular structure 373. The modular structure 373 provides (current) predictive data {Z...} at its output.

[0041] The notation "computer" (singular, no reference) refers to computational functions or to the functions of modules implemented on a computer. These functions may be distributed across different physical computers.

[0042] As used herein, a "module" is a functional unit (or computational unit) that uses one or more internal variables obtained by training.

[0043] Those skilled in the art know such various modules and may refer to them as "machine learning tools" or "ML tools". The description does not use "ML" etc. for the simple reason that "M" represents a computer that performs the calculations. As used herein, an (industrial) machine is associated with machine data X, but the machine itself does not perform the calculations.

[0044] From another perspective, the diagram illustrates modules of a computer system that includes a number of processing modules that, when executed by the computer system, perform steps of a computer-implemented method. Industrial machines are not considered computer modules.

[0045] Modules execute algorithms to solve tasks like regression, classification, and clustering.

[0046] Considering their internal structure, they are: Neural networks (symbol in Fig. 1A with nodes organized in layers with variables as weights), Decision tree structures with a single tree or multiple trees (e.g. random forests) or other modules It could be.

[0047] Those skilled in the art can implement the internal structure using frameworks such as Tensorflow, libraries such as Keras, programming languages ​​such as Python, R, or Julia.

[0048] The diagram symbolizes the potential recipient of the predictive data with an operator 193. The operator (or the person in charge of the industrial machine) can take appropriate measures, such as maintaining the machine in a timely manner, running the machine until a failure is predicted, or changing the operation, such as putting the machine into an operating mode that will delay the occurrence of the failure.

[0049] However, the predictive data {Z...} can be transferred to other computers, which can then trigger (semi-)automatic countermeasures.

[0050] The predictive data {Z...} may take several forms, for example: t_fail_a (the earliest predicted future time for failure), · t_fail_b (the latest predicted future time for failure to occur), failure_type (indicating the type of failure, for example by identifying the machine element that is likely to fail), or · The foresight that the machine will operate without failure for at least a specific time interval in the future.

[0051] Figure 1B shows a matrix with machines, computers and users in the rows and the progression of time (from left to right) in the columns. Figure 1B can be thought of as Figure 1A rotated 90 degrees.

[0052] Very simply, the machine provides the machine data, the computer executes the methods 702, 802 and 203, and the user receives the prognostic data {Z...}.

[0053] step Therefore, for convenience, the diagrams and descriptions distinguish at least the following stages: A preparation phase starting at approximately t1**1 by collecting data over time while the machine is running; A training phase **2 carried out at t2, training the modular structure regardless of whether the machine is running at **2 or not, see method 702 or 802 of FIG. 7-FIG. 8; and · Operational phase **3, showing the collection of data used for machine operation and failure prediction, t3 is the time to perform the prediction in method 203 (subordinate and output modules, see FIG. 2).

[0054] Time Series Data (e.g., machine data) are available in the form of a time series, i.e., a sequence of data values ​​indexed in chronological order at successive time points. Figure 1A introduces the time series by a short notation (a "rounded corner" rectangle 153) and a matrix below the rectangle, while Figure 1B repeats the rectangular notation from the time perspective.

[0055] The notation {X1...XM} denotes a single (i.e., univariate) time series with data elements Xm (or "elements" for short). Elements Xm are available from time 1 to time M: X1, X2, ..., Xm, ...XM (i.e., the "measurement time series"). The index m is the index of the time points. Time point m is usually followed by time points (m+1) at equally spaced intervals Δt. The notation {X...} is a contraction.

[0056] An example is the rotational speed of a machine drive at time M: {1400...1500}. One skilled in the art can pre-process the data values, for example, to normalized values ​​[0,1], or {0.2...1}. The data format is not limited to scalars or vectors, {X1...XM} can also represent a sequence of M image or audio samples taken at time 1 to time M.

[0057] The notation {{X1...XM}}N (or {{X...}}N for short) represents a multivariate time series with a vector of data elements {X_m}N from time point 1 to time point M. The vector has cardinality N (the number of variables, i.e., parameters, whose data are available) which means that at any time point from 1 to M, N data elements are available. A matrix denotes the variable index n as the row index (x_1 to x_N).

[0058] For example, a single time series for rotation may be accompanied by a single time series for temperature, a further single time series for data regarding the chemical composition of the material, and so on.

[0059] Those skilled in the art will understand that this is a simplified explanation. In reality, the number of variables N can reach or exceed two times 1000. The time series is not ideal. Sometimes elements are missing, but those skilled in the art can handle such situations.

[0060] The selection of the time interval Δt and the number of time points M depends on the process or activity performed by the machine. The total duration of the time series Δt*M (i.e., the window size) corresponds to the machine parameter shift that takes the longest time.

[0061] Because the time tm identifies a time for processing by a modular structure (or a component thereof), some of the data may be pre-processed. For example, a temperature sensor may provide data every minute, but for Δt=15 minutes (for example), some of the data may be discarded, averaged over Δt, or otherwise pre-processed.

[0062] The chronological notation {...} is applicable to the following: - explained, machine data {X...}, intermediate data {Y...} generated during the computation of the modules, in particular of the subordinate modules; Failure prediction data {Z...} at the output of the module structure, Failure data {Q...} that actually occur or represent failures as they occur ({Q...} are not predicted).

[0063] The X, Y, Z, and Q data can also be used as multivariate time series.

[0064] However, univariate and multivariate time series are only examples of data formats, and one skilled in the art can process data in other formats.

[0065] Machine Data X As the labels suggest, machine data X is associated with an industrial machine. Data X is processed because predicted failures are associated with the operation of the machine. Not all variates of machine data contribute to the prediction, so there is a rough division according to the relationship of the data source to the machine.

[0066] Machine data can be categorized as follows: Data obtained from sensors associated with the machine ("Sensor Data"); Data obtained from other sources (“further data” or “characteristic data”).

[0067] Further data can describe the objects the machine processes (including characteristics such as object type, object material, loading conditions, etc.) or the tools belonging to the machine (especially as they change over time). Further data can be environmental data during operation (temperature, etc.). A further example is maintenance data.

[0068] Sensor data may potentially be hidden from the machine operator or from other users in the sense that the operator / user does not associate specific sensor data with a specific meaning. As a result, an expert user may not be able to label such data. Further data is potentially more open. For example, a sensor reading representing the vibration of a particular component may not have meaning to an expert, but an expert may have a very good understanding of the effect of environmental temperature on the machine.

[0069] Calendar Time As mentioned, the index m is a time point index, the time series representation is convenient, and one skilled in the art can easily convert the time representation to actual calendar points. The time series are available in sequence (FIG. 1B with the sequenced Ω time series), and the calendar interval can be much longer than Δt*M.

[0070] Training and division of current and historical data Because the module obtains internal variables (such as weights or other machine learning related variables) through training 702 / 802 with data, the description distinguishes between "historical data" and "current data." Historical data is data that can be used to train the module (FIG. 1B with methods 702 and 802 in FIGS. 7-8). Therefore, historical data must be available before training. In other words, the data shown on the left side of methods 702 / 802 is historical data (historical machine data, historical failure data).

[0071] 1B shows training with a single box 702 / 802, with the width of the box representing the execution time between t2 and t2'. Training can be repeated with newly arrived data (i.e., "multilying" the box to the right, as shown in the box at t2''). Over time, as the amount of historical data increases, the module can be retrained (by repeating methods 702, 802) to achieve more accurate predictive outcomes.

[0072] Figure 1B shows a continuous time series with indices (1), (2)...(Ω). It is convenient to process a single total duration Δt*M of historical data at a time (i.e., N*M data values ​​for the N*M inputs of the structure being trained, plus M data values ​​for Q), but one skilled in the art could apply the data to the module otherwise. The Ω numbers (in the time series) are increasing over time.

[0073] In contrast, current data is data that the trained module can process in order to predict possible future failures (method 203 in FIG. 9). FIG. 1B illustrates this by the time series 153 with {X...} that will be processed during the execution of the prediction method 203. In theory, it is possible to process current data that actually overlaps with historical data (see the second box ending with t2'').

[0074] Original data As shown, the module structure receives raw data, i.e. data that has not yet been processed by the module (except for pre-processing to harmonize the data format). During training in method 702 / 802, the module structure receives raw historical data and obtains variables (or "weights"). Once training is complete, in prognosis method 203, the module structure receives current raw data and provides prognosis data {Z...}. During training 702 / 802 and prognosis 203, the structure module provides and processes intermediate data, so the raw data has already been mentioned here. In general, the historical data remains historical data and the current data remains current data.

[0075] Prediction and the division of past and future The run-time of the computer executing the prognosis method 203 can be negligible / small (compared to the M interval in the time series). Therefore, this description takes t3 as the earliest time point at which an operator can be notified about the failure prediction {Z...}. Thus, FIG. 1B also shows the prediction as a time series. As will be explained in more detail below, one element of the failure prediction data {Z...} is the identification of the failure time (t_fail).

[0076] By t3 (but not before), the operator can confirm / know the precognition.

[0077] The future time can also be given relative to the execution time of the computer (see t3 in Figure 3). "Time to failure" refers to the interval or duration from t3 to the earliest failure time.

[0078] Output predictability can be considered as timing accuracy, type accuracy, etc. These aspects are interrelated. For simplicity, the discussion will focus on improving timing accuracy.

[0079] Collecting data for training 1A also shows reference numeral 111 for the industrial machine during historical operation, reference numeral 151 for the historical machine data (and historical failure data) of stage 1. It also shows reference numeral 372 for the structure during training.

[0080] Modular structure 2 shows a modular structure 373 with subordinate modules 313, 323, 333 (in the hierarchy) subordinate to an output module 363 (of relatively higher rank). The subordinate module 333 has the special function of an operating mode classifier.

[0081] In the description, the label "classifier" is used for simplicity, but the label also includes the meaning of "grouping (clustering)". The subordinate module 333 can act as a classifier (assigning machine operating times to classes such as MODE_1 and MODE_2), but it can also act as a cluster tool (separating machine operating times according to the data observed during different operating times).

[0082] The allocation of particular clusters to particular modes is optional.

[0083] For example, the module 333 can process data and group the operation times (i.e., time points m) into a first and a second cluster. The computer can then automatically allocate these clusters to a first and a second operation mode (acting as a class). In other words, "cluster" and "mode" have different meanings. The module monitors the operation of the machine and partitions the operation times into (non-overlapping) clusters. There is an allocation (from the first cluster to the first mode, from the second cluster to the second mode, etc.) and the mode can be set as a classification target. The module can then train to partition the operation times according to the target (classification, not grouping). During further iterations with different data, the module 333 can then determine whether the machine operates in the first or second mode.

[0084] Optionally, an expert can be involved in the allocation of clusters to classes (e.g., the expert simply names the clusters with mode names, or the expert recognizes an association with a fault, etc.). The allocation can be more complex (two clusters may belong to the same mode). But generally, it is not necessary to involve an expert. It may be advantageous not to involve the user. The differences in the operating modes may be "invisible" to the expert (or not at all difficult to detect, see e.g. Figure 5). In other words, the clusters and / or modes may be hidden from the expert. However, the differences may affect the prognosis (and the behavior of the machine, see Figure 4) and the computer can recognize the existence of such differences. Again, the differences may be hidden from the user, but not from the computer.

[0085] Grouping is not required, and experts can annotate historical machine data with operational modes, such as by annotating sensor data.

[0086] Different Modules Different modules perform different tasks (regression, classification / grouping, etc.). The use of dependent modules (specialized for a particular task) in the arrangement can improve prediction accuracy compared to a single module (i.e., a module without dependent modules). Prediction accuracy is described as an example for time accuracy in relation to Figures 3-4.

[0087] Since the module structure 373 has several components that may require specific data as input, the following description further describes the following optional approaches among them: Compensating for data gaps using data from virtual sensors (see Figure 13 for an approach); Compensating for the lack of expert knowledge to classify operating modes by automatically classifying the modes, optionally starting with groupings (see Figures 7-8 for the use of such automatically obtained data); · Cascade training the modular structure in a specific training order (starting with the mode classifier, see Figures 7-8); Compensating for the lack of data by at least partially simulating the behavior of the industrial machine (see FIG. 14) or otherwise foreseeing the behavior of the machine; Augmenting the training data with human-annotated labels (not discussed further), Compensating for data shortages (or data surpluses) by transferring data, for example by harmonizing the availability of data variants, when data from different physical machines must be processed (see FIG. 13 for historical data), or Using a bias indicating the confidence of the inputs (instead of binary classification) to train the output module, allowing different modules to compete for accuracy (e.g. disjunct mode metrics or probabilistic metrics as described below).

[0088] From an overall perspective, the modular structure 373 receives machine data 153 from the industrial machine 113 (see FIG. 1A) and predicts failures of the industrial machine (data {Z...}).

[0089] Viewed from its topology, the module structure 373 includes two or more modules that are subordinate to an output module. The subordinate modules may differ (among peers) in the following ways: Machine data origin can be module specific: for example, dependent modules 313 and 323 can process machine data from different machine components, e.g., module 313 can receive {{X...}}N1 where subset ∈ {{X...}}N, module 323 can receive subset {{X...}}N2, etc. (see Figure 2). The sets of weights (or other machine learning variables) that the dependent modules apply during processing can be different. Intermediate data (like {Y...}) can also be module specific: the figure shows 1{Y...} at the output of module 313 as a first intermediate status indicator, 2{Y...} at the output of module 323 as a second intermediate status indicator, and 3{Y...} at the output of operation mode classifier 333 as an operation mode indicator.

[0090] The topology affects data availability: output modules can process intermediate data as it becomes available (a pipeline structure from left to right in the diagram).

[0091] The topology also influences the training: the dependent modules are trained before the output module can be trained, as described below in connection with Figures 7-8. The same principles apply to further rank hierarchies, and to training in the order of sub-sub-ordinated, dependent, and supra-ordinating modules.

[0092] The topology accommodates individual modules performing different tasks, for example, module 333 provides groupings (or classifications into MODEs) and therefore biases for output modules.

[0093] Mixed Forms In relation to FIG. 1, this description has already introduced modules for performing tasks such as regression, classification, grouping, etc. It is convenient but not necessary to separate the tasks. The predictive failure data {Z...} has a regression form (time from future successive times to the obtained failure) and a classification form (such as the type of the failure). Similarly, module 333 can provide mode indicators that can be separated (e.g., as a result of the classification, either MODE_1 or MODE_2) or can be probability classifiers (more on this later).

[0094] step Unless otherwise indicated, the industrial machine and modular structure are shown in the operational phase 3. Training 2 will be described in relation to Figures 7-8. For convenience, Figure 2 also shows references that can be applied during training: the modular structure 372 is trained by the subordinate modules 312, 322 and 322, as well as the output module 362, which are all trained (see Figures 7-8 for details).

[0095] FIG. 2 also shows an optional index-derived module 374, which is described in relation to FIGS.

[0096] Timing accuracy of fault time prediction FIG. 3 shows a time diagram for the operation of an industrial machine 113 (of FIGS. 1A and 1B) in combination with the fault intervals in fault prediction by a module. The module can be a conventional module (independent) or can be a module structure 373.

[0097] The horizontal line indicates the operation of the industrial machine in a simplified operation scenario. · Scenario 1: The machine operates until it fails at t_fail_1 < t_fail_a. The module did not provide a satisfactory indication. · Scenario 2: The machine operates until it fails within the predicted fault interval [t_fail_a, t_fail_b]. The module did not provide a satisfactory indication, but the operator decided not to perform maintenance on the machine. · Scenario 3: After the fault prediction interval [t_fail_a, t_fail_b], it operates until the machine fails at t_fail_3. · Scenario 4: The machine operates until maintenance break “stop”. Maintenance is started a little before the predicted t_fail_a. The machine resumes operation and eventually fails at t_fail_4. This is an almost ideal situation.

[0098] There is a desire to make the prediction more accurate. The figure shows this by a corrected predicted fault interval [t_fail_a’, t_fail_b’] that is shorter than the original fault interval. The operator can delay maintenance until just before t_fail_a’. Such an improvement is achievable for a module structure (cascaded module, see FIG. 2).

[0099] The module structure operates at a run-time t3 (see FIG. 2), and the calculation duration (the time required for the computer to calculate {Z...}) can be ignored. The interval [t_fail_a, t_fail_b] is the predicted fault interval.

[0100] The figures are simplified and one skilled in the art can derive other indices therein. · Remaining Useful Life (RUL). Failures can be different and not all types of failures will take the machine out of service. For example, if a bearing shows "no oil", the operator has the opportunity to perform maintenance on that bearing and the machine can continue to operate. The operator gets the RUL by collecting further data that indicates a failure beyond a simple lack of oil (such as a motor failure). · Time to failure (TTF) is the interval (from t3) to t_fail_a (shorter TTF) or t_fail_b (longer TTF). · The risk of failure as a measure of severity can be derived from t_type (optionally taking time into account).

[0101] As will be explained, a single module receiving data from substantially all available machine data {{X...}}N may provide inadequate prognostic data {Z...} for an operator to make appropriate decisions.

[0102] FIG. 4 shows a time diagram of the operation of the industrial machine (of FIG. 1A) in combination with the mode-specific failure intervals for prognosis by the mode-specific module.

[0103] The modular structure allows the predictive failure interval to be divided by mode, and the figure shows (t_fail_1, t_fail_2) for MODE_1 and MODE_2 separately.

[0104] A machine operator may understand operating modes that reflect easy-to-detect states such as ON (the machine is running), STAND-BY (the machine is running at low energy but not producing any product), FULL-LOADED, etc. However, modes are associated with predictive failures, and the operator does not need to know that the machine switches modes. Nor does the machine need to implement mode switching. Modes are attributes that describe the behavior of the machine.

[0105] In a simplified example, a machine in MODE_1 will fail sooner than a machine in MODE_2. This information can be important for an operator. As shown below, the operator is informed at t3 (operation time of the module structure) about the predictive failure interval for both modes individually and optionally for a combination of both modes ("MODE_1 or _2").

[0106] Up until t3, the operator could control the machine to operate in MODE_1 or MODE_2, or the machine could assume either mode without being explicitly controlled to be in a particular mode.

[0107] Potentially, the operator could continue in MODE_2 until t4 (just before t_fail_1 in MODE_1). Maintenance could be delayed or the operator could allow the machine to run exclusively in MODE_2 from about t4 onwards.

[0108] The diagram is very simplified, during the machine operation after t3 (represented by the current data from t3 to t4), the computer updates the prognosis. Continuing to operate the machine in MODE_1 (after t3) would probably move t_fail (for MODE_1) towards the left. Therefore, the operator decides to switch to MODE_2 only immediately after t3 (and not t4).

[0109] Note that the operator does not need to know the mode in advance, he can switch the machine to perform a different operation and a mode indicator will inform the operator of the mode.

[0110] A modular structure that partitions the operating modes allows for a more accurate identification of the (overall) failure interval. This description goes into detail on increasing the prediction accuracy in relation to FIG. 5, but digresses briefly to describe an application scenario in which failure prediction data {Z...} is combined with mode identification data to control a machine.

[0111] (Semi)automatic mode applied Figure 4 and its description can be understood as an example for establishing a control rule. The machine controller can process the failure prediction data {Z...} (available at T3) against the actual control commands to control the operation of the machine. This rule can be reinforced by higher level optimality goals. For example, if the optimality goal is "avoid maintenance as much as possible", the controller will operate the machine in any mode up to t4 and will not allow operation in MODE_1 from t4 onwards.

[0112] Expert involvement is kept to a minimum (eg, to define t4 to be earlier than t_fail, within some predefined window).

[0113] The controller sending the control command to the machine may change the mode, but the (trained) modular structure (or at least its mode classifier) ​​can establish the mode (or at least the cluster) so that it can reverse the command if necessary at virtually any time, or the controller checks the command for its possible effect on the mode.

[0114] In other words, the prediction made by the structure (method 203 see FIG. 1B) can be used to send {Z...} to the machine controller to make the machine assume a mode of predicting at the latest the time until failure, a mode of predicting at the latest the time until maintenance, or according to other criteria.

[0115] From another perspective, industrial machines may be associated with a machine controller that switches between operating modes according to predefined optimal goals. The above criteria can also be formulated as goals such as avoiding maintenance (as much as possible), operating the machine in a mode where a failure is predicted to occur at the latest (compared to other modes), etc.

[0116] Machine Example 5 shows a block diagram of an industrial machine 110. The machine is fictitious in the sense that it has symbolic components that represent real components of real machines. Examples of non-fictitious machines include chemical reactors, metallurgical furnaces, vessels, pumps, motors, engines, etc.

[0117] The machine 110 has a drive 120. A vibration sensor 130 is attached to the drive and provides a signal in the form of a time series {X...}. In this simplified example, the machine data need only include sensor data. The machine uses exchangeable tools (or actuators) 140-1 / 140-2. The diagram symbolizes the tools by showing that the machine works alternately with tool 1 or tool 2 ("arrow tool" or "triangle tool"). The machine interacts with an object 150 (here, in the example, via the tool). During the interaction, the object needs to change its shape (the machine is for example a metalworking lathe), its position (a transport machine), its color (a painting robot), etc.

[0118] In the simplified diagram of Figure 5, the choice of tool determines the machine configuration (e.g., first and second configuration). In a more realistic scenario, a machine can have much more components connected in multiple configurations. The more complex the configuration, the more complex the causal relationships mentioned above become, and the more complex the failure prediction becomes. For simplicity, this explanation focuses on vibration as the only possible cause of failure. The occurrence of mechanical vibrations (represented by signals {X...}) during operation is normal. To simplify greatly, industrial machines emit sounds. Depending on the tool / object combination and configuration, the machine emits different sounds (see different frequencies diagram).

[0119] The figure also shows a highly simplified frequency diagram (e.g. obtained by Fast Fourier Transform of the sensor signal as known in the art). Of course, the frequency distribution changes over time for many reasons (e.g. the shape of the object changes), but this diagram gives an approximate view of the prevailing frequencies.

[0120] In general, vibrations do not always lead to failure. However, there are notable exceptions. Natural frequencies (or resonant frequencies, here fR) are those where the amplitude of vibrations becomes relatively large and therefore the risk of failure increases. Again, this is a simplified explanation: in realistic scenarios, different resonant frequencies are found.

[0121] As shown in the diagram, using tool 1 ("arrow") will make the machine vibrate close to the resonant frequency, and using tool 2 ("triangle") will make it vibrate at some other frequency. In this simplified form we cannot exclude the risk that the machine will eventually vibrate at fR, but the risk is higher with tool 1. Small variations (in some properties such as the tool's Young's modulus) can occur, causing the vibration to be at fR.

[0122] Domain experts can study the vibrations and find correlations between the use of different tools and different frequencies (frequency). However, in the realistic scenarios described above, industrial machines are more complex (many different tools, many different objects) and expert knowledge is generally not available.

[0123] As will be explained, a computer can distinguish between modes of operation (or at least group operating times) even when a human expert cannot distinguish between the modes. The explanation is simplified to a first mode of operation and a second mode of operation, and the meaning of the tool is not important to the computer.

[0124] In a simplified example, the two operating modes are separated by different parts of the frequency band: in the first mode, the frequencies dominate in the low band (below fR), and in the second mode, the frequencies dominate in the high band (above fR).

[0125] The resonant frequency can be reached in both modes, although with different probabilities.

[0126] Returning to FIG. 2, the operational mode classifier 333 provides operational mode indices 3{Y...}. Note that although the description uses the term "indices" in the singular, they may change over time. Therefore, they are shown as time series. Examples of 3{Y...} changing over time are shown in FIGS. 10-11.

[0127] In principle, there are several options. The operating mode classifier 333 may operate as a dedicated classifier that outputs a variate corresponding to the operating mode (e.g., mode 1 XOR mode 2). Or, in case of multiple operating modes, the operating mode classifier 333 is a predefined value from a set of values ​​(MODE_1, MODE_2, MODE_3, etc.). Alternatively, the number of modes is not predefined, but is determined as the number of clusters. The operating mode classifier 333 may operate as a probabilistic classifier that outputs a variable with the probability of the operating mode (e.g., 80% for mode 1, 20% for mode 2). The operating mode classifier 333 can be a combination of both: it can be a combination of predefined values ​​and probability ranges. For example, 3{Y...} can be implemented as a vector with two variables, a bivariate time series 3{{Y...}2: the first variable indicates the mode and the second variable the probability. For example, at a certain time tm, the mode is MODE_1 with 80% probability.

[0128] Optionally, splitting historical machine data Assuming that the operating mode classifier 332 / 333 (see FIG. 2) has already been trained, at least by preliminary training, it processes the historical machine data {{X...}}N (multivariate time series, or {{X...}}N3) into two subseries of historical machine data, the details of which are described in connection with FIGs. 6 and 8.

[0129] Figure 6 shows the historical multivariate time series {{X...}}N of Figure 1B. The operating mode classifier may distinguish between the modes (here MODE_1 and MODE_2) of the operating mode indices 3 {Y...}.

[0130] As a result, the X-data can be distributed into two (or more) multivariate time series. In this example, MODE_1 was found for m=1,2,3,... and MODE_2 for m=4,5,8,9.

[0131] Variations can be applied: for example, the mode division can be achieved only with relatively low probability (see discussion above), but certain data can be allocated to both modes.

[0132] For mode-specific timelines, time can be made to appear to advance in successive time slots by ignoring excluded time slots. Those skilled in the art can introduce new time counters, etc.

[0133] In that sense, historical data {{X...}}N turns into modal annotated historical data {{X...@1}}N and {{X...@2}}N. However, no expert supervision is required.

[0134] Although not shown here, the split can also be applied to the failure data. There may be historical failures that occur during operation in mode 1 or mode 2.

[0135] The partitioning of the historical machine data (or failure data) can be used in step 852 of FIG.

[0136] A partition of historical data (machine data or failure data) can be thought of as a grouping. A grouping results in distinct time series segments (e.g. 3{Y...}). It is useful to automatically assign specific clusters to specific modes. In this example we use two clusters assigned to two modes.

[0137] The figure shows, purely by way of example, segm_1 (MODE_1), segm_2 (MODE_2), segm_3 (also MODE_1), segm_4 (also MODE_2), etc. The time series segments can have different durations (e.g. segm_1 at 3*Δt, segm_2 at 2*Δt, etc.). The segments are separated into a first cluster with (segm_1, segm_3, ...) and a second cluster with (segm_2, segm_4, ...).

[0138] In terms of dividing the operating times (of an industrial machine) into different clusters, the grouping is useful because the operating modes are functions of time (3{...} is a time series).

[0139] Original data revisited As mentioned above (FIG. 1B), the modules can be trained and then used to process data. During training of the module structure (see FIG. 2 with two-layer hierarchy), the subordinate modules convert the raw data (machine data {X...}, failure data {Q...}, etc.), all of which are historical data, into intermediate data {Y...}. The output module processes the intermediate and raw data, which are also historical data.

[0140] Once the module structure is trained, it receives the raw data (e.g. {{X...}}N) and provides predictions {Z...} which are the current data. However, at least the output module can receive the raw data and intermediate data, both of which are current data.

[0141] A higher-level module (e.g., an output module) receives the raw data (i.e., data that has not yet been processed) in combination with intermediate data, - The intermediate data has a specific function, and The availability of such intermediate data is cascaded (during training and prediction), It may be advantageous to

[0142] At least one example scenario is shown: since annotating the original data by experts is difficult, intermediate data - such as mode indicators - can serve as de facto annotations. The sequence remains the same: the output module uses the de facto annotations when they are available at the latest.

[0143] This approach describes a two-tier hierarchy (see FIG. 2), although further tiers can be introduced.

[0144] Cascade Training FIG. 7 shows a simplified time diagram of cascade training 702. The bold horizontal lines indicate the availability of data during training. The vertical arrows indicate the use of data during training. There may be cases where multiple vertical lines come from one and the same horizontal line, but this does not mean that the uses require the same data. Sometimes repeated use of data may mean use from different variables (see {{X...}}N possibly not from all N variables but from different subsets of variables). Once data has been used, it will continue to be available: the horizontal lines change from solid to dotted. Reusing data is useful when repeating some training steps.

[0145] Time progresses from left to right, with time t2 indicating the start of phase 2 and time t3 indicating the time of operation phase 3 (see Figure 3, where t3 indicates the execution time of the computer performing the prediction).

[0146] The boxes symbolize method steps 712, 722, 732, but the width of the boxes is not adjusted for time. The boxes may have bold vertical lines 742 and 762 on the right side, which symbolize the trained (subordinated) modules running to provide an output.

[0147] The discussion occasionally switches back from FIG. 1A (providing references to the machine 111, historical machine data 151) to FIG. 2 (topology, where references **2 apply) and FIG. 5 (an example of a machine with two modes).

[0148] In this description, the term "preliminary" is used to indicate an optional repetition of the steps of the method. In other words, individual training steps can be repeated. For convenience, the description refers to the semantics of the data (e.g., failures in frequency or fR), but the computer does not need to take such semantics into account.

[0149] The historical data is available from the beginning (i.e., before t2). The historical data can have, for example, a time-series format. In this figure, the historical data is partitioned into historical failure data {Q...} and historical machine data {{X...}}N (received from the industrial machine 111 or from another machine).

[0150] The failure data are given as univariate time series {Q...}, whereas different types of failures (i.e., failure variables) can be represented by multivariate time series (e.g. {{Q...}}).

[0151] Step 712 / 742 In step 712, the computer (preliminarily) trains a mode classifier (i.e., subordinate module 333 in FIG. 2) using the historical machine data (and optionally failure data, not shown). Once trained, the operating mode classifier 333 can use the historical machine data to calculate the historical mode indices 3{Y...}. No supervision (i.e., processing expert annotations) is required in this step.

[0152] In step 742, the computer calculates the history mode index 3{Y...}. Since the history machine data {{X...}}N is available synchronously with the history mode index 3{Y...}, the time tm is not changed and both data form a data pair (here with a mode index in the sense of the automatically generated annotation).

[0153] For example, 3{Y...} may be a time series indicating alternative operating mode 1 during a first 24 hour interval and mode 2 during a second 24 hour interval.

[0154] It may be advantageous that no identification of reasons (such as the use of tool 1 or 2, or other meanings) is required: the computer uses the available data, but without training with supervision or other forms of expert involvement.

[0155] Step 722 / 762 In step 722, the computer uses the historical machine data {{X...}}N and (optionally) the historical mode index 3{Y...} to train the slave module 313, 323. Once trained, the slave module 313, 323 can provide intermediate status indexes 1{Y...} and 2{Y...}. For example, the intermediate status indexes 1{Y...} and 2{Y...} are values ​​that indicate frequency changes, such as increases or decreases over time.

[0156] Although this step is shown as a single box in the diagram, this step is performed separately (serial or parallel) for both subordinate modules.

[0157] In step 762, the computer again uses the historical machine data {{X...}}N to calculate intermediate status indicators 1{Y...} and 2{Y...}, which are of course historical indicators. For example, both intermediate status indicators may indicate a history of frequency increases (regardless of what that may mean).

[0158] Step 732 Historical failure data Q (actual failure data) is available earlier, but can potentially be used to compare with the intermediate status indicators. Such failure data can be obtained automatically. In a simple implementation, the failures are represented by sensor signals {Q...}, also represented as a time series indicating the time of the (actually occurring) failure.

[0159] In step 732 , the computer uses the historical failure data {Q...}, the intermediate status indicators 1 {Y...} and 2 {Y...}, and the mode indicator 3 {Y...} to train the output module 362 .

[0160] Upon training, output module 362 changes to output module 363 (FIG. 2), and dependent modules change to modules with references **3 as well. To stay with the example semantics, module structure 373 may detect failures in MODE_1 for frequency increases (where frequency just approaches fR) occurring between 10 and 14 hours from mode change, via t_fail_a and t_fail_b. For MODE_2, frequency similarly increases (but moves away from fR), and t_fail may be different.

[0161] In other words, by partitioning the operating modes, the modular structure 373 can provide prognosis with greater timing accuracy.

[0162] Cascade training with split historical data FIG. 8 shows a simplified timing diagram for cascade training 802 in a modification of the training described for FIG.

[0163] The steps correspond to those described for FIG. 7, except that the computer performs the additional step 852 (to split the historical machine data, see FIG. 6), and step 722 (of FIG. 7) is performed as step 822@1 for dependent modules 312 / 313 and as step 822@2 for dependent modules 322 / 323.

[0164] Once training of the mode classifier module is complete (at step 812), the computer calculates the historical mode indices 3{Y...} at step 842. As described in Figure 6, 3{Y...} is then used to split the historical machine data into mode annotated historical data {{X...@1}}N and {{X...@2}}N. (Steps 842 and 852 may be performed in combination.)

[0165] The subordinate networks are then trained separately to provide intermediate status indices 1{Y...} and 2{Y...} (steps 822@1, 822@2).

[0166] It is useful not to split the fault history data {Q...} (environmental faults in MODE_1 can also occur when the machine is running in MODE_2, and vice versa).

[0167] Methodology Overview Figure 9 shows a flow chart of a computer-implemented method 203 for predicting failures in an industrial machine. In performing the method 203, the computer uses a structure of processing modules such as the modular structure 373 of Figure 2, or a structure with further hierarchy. For convenience, the figure illustrates the flow chart with a symbolic copy of Figure 2 with X, Y and Z data.

[0168] In a receiving step 213, the computer receives machine data ({{X...}}N) from the industrial machine 113 by the first, second and third subordinate processing modules 313, 323, 333 arranged to provide intermediate data 1{Y...}, 2{Y...}, 3{Y...} to the output processing module 363. The structure 373 has been previously trained by cascade training, see 702 / 802 in Figs. 7-8.

[0169] The computer processes 223A the machine data using a first dependent module 313 to determine a first intermediate status indicator 1 {Y...}; processes 223B the machine data using a second dependent module 323 to determine a second intermediate status indicator 2 {Y...}; and processes 223C the machine data using a third dependent module 333 - an operating mode classifier module - to determine operating mode indicators 3 {Y...} (for all tree indicators) of the industrial machine 113.

[0170] In processing step 243, the computer processes the first and second intermediate status indicators 1{Y...}, 2{Y...} and the operation mode indicator 3{Y...} by the output module 363. Thereby, the output module 363 predicts a failure of the industrial machine 113 by providing prognostic data {Z...}.

[0171] Example of operation Now, the module structure 373, having received the current machine data 153 (see FIGS. 1-2), identifies - for the actual point in time t3 (see FIG. 3) - the mode (module 333) and the status indicators (modules 313, 323).

[0172] Selecting Machine Data As mentioned before, the machine data {{X...}} can be sensor data and further data.

[0173] Assume that the experts cannot select the subset of machine data that is relevant (for failure prediction), so the selection is made by the modules (during training of the modules): some process machine data with heavier weights, others process sensor data with lighter weights.

[0174] For non-sensor data, experts may have deeper insight to make choices (in which case they can label some data as not relevant).

[0175] In implementations, the subsets {{X...}}N1 and {{X...}}N2 can be further divided by grouping the time series according to a variable, element-of-notation ∈ see FIG. 2.

[0176] Modules - using derived indicators (such as mode indicators) In modern industrial environments, industrial machines are likely to change operating modes frequently, one of the reasons being the trend towards lower production volumes. Mode change rates (number of mode changes per hour) may be associated with failures of some, but not all, machines.

[0177] Mode changes as derivative mode indicators Figure 10 shows the time series for two modes (MODE_1 "black" and MODE_2 "white") with mode index 3 {Y...}. Time windows (equally spaced, with a predefined number of time intervals Δt per window) are associated with the number of mode changes (MODE_1 to MODE_2 or vice versa). This approach can be seen as a derivation of a mode function over time.

[0178] The computer can process the output of the operating mode classifier (see FIG. 2) to determine a mode change rate, which can be a further input to the output module 363. The mode change rate can be calculated on current and historical data. To symbolize this optional operation, FIG. 2 shows a mode index derivation module 374 between the classifier 333 and the output module 363.

[0179] Although FIG. 10 is simplified by showing only two modes, the mode change can be quantified for other scenarios as well.

[0180] In the alternative, the number of time intervals need not be predefined.Grouping to identify clusters according to different window durations and / or different occurrences of mode changes is also possible.

[0181] FIG. 11 shows a state transition diagram (with five modes or states) and a state transition diagram with mode transitions. A diagram can be applied to one time window (of FIG. 10) and can show the occurrence of mode transitions (e.g., A to B, B to C, C to D, and vice versa). The diagram symbolizes the number of occurrences of transitions by line thickness, with D to A being the prominent transition. Of course, the numbers can change for other time windows. Again, the number of transition occurrences for each particular transition can be input to output module 362 / 363.

[0182] The calculations can be performed, for example, by the index derivation module 374 (see FIG. 2).

[0183] In an alternative approach, groupings are also possible here, for example by grouping the transitions and separating the modes by, for example, high or low sub-mode transitions.

[0184] Multiple machines providing historical data Figure 12 shows multiple industrial machines 111α, 111β and 111γ and their historical time series with machine data {{X...}}N and their historical time series with failure data {Q...}. For simplicity, not all available metrics are used in the figure.

[0185] As mentioned above, data may not be available in sufficient quantities. Therefore, the diagram shows several industrial machines providing historical machine data X and historical failure data Q. The diagram symbolizes that under ideal conditions, time series with data are available in the number of time series per machine multiplied by the number of machines (the three machines, α, β, γ, are just for simplicity).

[0186] To train the method 702 / 802, the computer (structure under training 372) processes a time series {{X...]}}N and a time series {Q...} on N+1 input variables at a time. The computer then moves on to the next time series.

[0187] In some cases, the computer processes a continuous time series (1), (2) through (Ω), such as {{X...}}N, as well as {Q...} in the "one-time input" described in Figure 1B. The practitioner can arrange for iterations of α, β, γ, or even have the computer process α{{X...}}N, β{{X...}}N, γ{{X...}}N, α{Q...}, β{Q...}, γ{Q...} all at once. Other processing options are also available.

[0188] Compensating for missing variables by enhancing virtual sensors and transfer learning A multiple machine scenario, such as the one illustrated in FIG. 12, works ideally with machine data (and failure data) from substantially equal sources.

[0189] For example, a univariate time series α{X...}n will be similar to a univariate time series β{X...}n because the sensor for variable n is the same type of sensor on both machine α and machine β. However, not all machines have the same sensors. Here we describe an approach to address this constraint.

[0190] Figure 13 illustrates an approach to reconciling machine data (and potentially fault data Q) for different industrial machines. Harmonization can be applied to historical data (stage **1) and current data (stage **3).

[0191] The figure shows that we repeat industrial machines 111α, 111β and 111γ (from FIG. 12), but with different machine data availability: machine α has the usual N variables, machine β is missing one variable (N-1 variables) and machine γ needs to have one more variable (N+1 variables).

[0192] The figure shows data harmonizers 382β and 382γ. Data harmonizer 382β provides missing data via virtual sensors (here Xn), and data harmonizer 382γ filters the incoming data (i.e., removes redundant data).

[0193] The diagram is simplified, and the omissions and surpluses of data depend on the contribution of specific variables to the prediction. Some machine data (i.e., some variables of that data) are not relevant to the prediction of failure at all.

[0194] Each harmonizer employs modules that have been previously trained by transfer learning. For example, machines α and γ can be masters that teach harmonizer 382β how to virtualize sensors Xn. Or machines α and β are masters that learn to ignore certain data sets.

[0195] As shown in the figure, the harmonizer does not modify the fault data {Q...}.

[0196] A domain-adaptive machine learning model trained by transfer learning processes historical machine data (obtained as multivariate time series from multiple industrial machines of a certain type, but from multiple domains). The historical machine data reflects the state of each machine in multiple domains. Typically, hundreds or thousands of sensors per machine measure operational parameters such as temperature, pressure, chemical content, etc. (see a relatively high number of variables N). Such measured parameters at a specific point in time define the respective state of the machine at that point in time. Due to the presence of multiple characteristics of each machine (e.g., operating mode, size, input materials like material composition, etc.), two machines (source machine and target machine) cannot be directly compared without applying dedicated transformations to the multivariate time series data.

[0197] Various approaches can be used for transfer learning. For example, a domain-adaptive machine learning model can be implemented by a deep learning neural network with convolutional and / or recurrent layers trained to extract domain-invariant features from historical machine data as an initial domain-invariant dataset. Transfer learning can be implemented to extract domain-invariant features from historical machine data. Deep learning features consist in abstract representations of specific machine features extracted from multivariate time series data generated by the operation of a specific machine. Transfer learning can be applied to extract domain-invariant features from multiple real-world machines that are independent of a specific type (i.e., independent of different domains).

[0198] In an alternative approach, a domain-adaptive machine learning model has been trained to learn to perform multiple mappings of corresponding raw data from multiple machines into a reference machine. The reference machine can be a virtual machine or a real machine that represents a kind of average machine. Each mapping represents a transformation from each particular machine to the reference machine. In this approach, the multiple mappings correspond to a first domain-invariant data set. For example, such a domain-adaptive machine learning model can be implemented by a generative deep learning architecture based on the CycleGAN architecture, which is popular in another application field for generating artificial (or "fake") images. CycleGAN is an extension of the GAN architecture, which includes the simultaneous training of two generator (generative) models and two discriminator models. One generator takes data from a first domain as input and outputs data for a second domain, and the other generator takes data from the second domain as input and generates data for the first domain. The discriminator model is then used to determine how plausible the generated data is and update the generator model accordingly. CycleGAN uses an additional extension to the architecture called cycle consistency. The idea behind it is that the data output by the first generator can be used as input to the second generator, and the output of the second generator should match the original data. The reverse is also true: the output from the second generator can be fed as input to the first generator, and the result should match the input to the second generator.

[0199] Cycle consistency is a concept in machine translation where a phrase translated from English to French should be the same as the original phrase translated from French to English. The inverse process should also be true. CycleGAN promotes cycle consistency by adding an additional loss to measure the difference between the generated output of the second generator and the original image, or vice versa. This acts as a regularizer for the generator model, guiding the image generation process in new domains to image translation. To adapt the original CycleGAN architecture from image processing to processing multivariate time series data and obtain the first domain-invariant dataset, the following modifications can be made by using recurrent layers (LSTM as an example) combined with convolutional layers to learn the time dependencies of multivariate time series data, as described in detail in C. Schockaert H. Hoyez, (2020), “MTS-CycleGAN: An Adversarial-based Deep Mapping Learning Network for Multivariate Time Series Domain Adaptation Applied to the Ironmaking Industry”, in arXiv: 2007.07518:

[0200] An overview of transfer learning is available in Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, Qing He: “A Comprehensive Survey on Transfer Learning” arXiv:1911.02685.

[0201] Compensation by simulation FIG. 14 shows machine data in a bivariate time series {{X...}N for N=2 with a first time series provided by sensor 135 (as in the normal situation, see sensor 130 in FIG. 5) and a second time series provided by data processor 165.

[0202] For example, a tool (140 in Figure 5) loses sharpness over time, and there may not be a sensor that can measure it, and setting up a virtual sensor may be difficult (measuring sharpness is difficult, so a master may not be found).

[0203] The data processor 165 may be implemented by a computer using a formula created by an expert. For example, the expert may correlate with existing data to calculate the deterioration of sharpness over time (and therefore the point at which the tool needs to be replaced (or sharpened)). By way of example, such data may include the time the tool has been inserted into the machine, the number of operations, or the number of objects, etc.

[0204] Alternatively, the data processor 165 can be implemented as a computer running a simulation. In that sense, the computer can operate as described above to predict tool failure ("no longer sharp" being a failure condition) but not to predict full machine failure. Setting up the simulator can potentially require minimal interaction with an expert.

[0205] The principles of failure detection described above can also be applied to machine parts. Tools will eventually fail. There are two outcomes: First, the tool failure is of a specific failure type (and can be predicted as such) · Second, tool failures can be simulated and used as input.

[0206] Mode Specific Training FIG. 7 in combination with FIG. 8 shows that the dependent modules can be trained separately for different modes.

[0207] Assuming we have two dependent modules (as in Figure 2), the mode classifier can partition the historical data according to mode, such that the first module is trained on MODE_1 data and the second module is trained on MODE_2 data.

[0208] For the current data, both modules provide intermediate status indicators (such as 1{Y...} and 2{Y...}) and they do not receive a mode indicator, see Fig. 2. Thus, the first module creates "garbage" every time the machine runs in MODE_2 (and vice versa for the second module). However, the operating mode classifier 333 provides a mode indicator (current data) 3{Y...}, so the output network (if trained) will ignore some intermediate data.

[0209] More generally, as the mode classifier module performs grouping, the number of clusters can be more than 2. Depending on the number of mode clusters, subordinate modules (not the mode classifier) ​​can be dynamically added or removed.

[0210] Mode-Specific Bias 2, the operating mode index 3{Y...} goes to output module 363. In implementation, the index may also act as a bias for subordinate modules 313 and 323.

[0211] General-purpose computers FIG. 15 illustrates an example of a general-purpose computing device that may be used with the techniques described herein. The diagram illustrates an example of a general-purpose computing device 900 and a general-purpose mobile computing device 950 that may be used with the techniques described herein. Computing device 900 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 950 is intended to represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, driving assistance systems, or on-board computers of vehicles and other similar computing devices. For example, computing device 950 may be used as a front end by a user (e.g., an operator of an industrial machine) to interact with computing device 900. The components illustrated herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the practice of the inventions described and / or claimed herein.

[0212] The computing device 900 includes a processor 902, a memory 904, a storage device 906, a high-speed interface 908 connecting to the memory 904 and a high-speed expansion port 910, and a low-speed interface 912 connecting to a low-speed bus 914 and the storage device 906. Each of the components 902, 904, 906, 908, 910, and 912 are interconnected using various buses and may be mounted on a common motherboard or otherwise as needed. The processor 902 may process instructions for execution within the computing device 900, including instructions stored in the memory 904 or the storage device 906, to display graphical information for a GUI on an external I / O device, such as a display 916 coupled to the high-speed interface 908. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memories, as needed. Multiple computing devices 900 may also be connected (e.g., as a bank of servers, a group of blade servers, or a multiprocessor system), with each device providing a portion of the required operations.

[0213] The memory 904 stores information within the computing device 900. In one implementation, the memory 904 is a volatile memory unit or units. In another implementation, the memory 904 is a non-volatile memory unit or units. The memory 904 may also be another form of computer-readable medium, such as a magnetic disk or an optical disk.

[0214] The storage device 906 can provide mass storage for the computing device 900. In one implementation, the storage device 906 can be or include a computer-readable medium such as a floppy disk drive, a hard disk drive, an optical disk drive, or a tape drive, a flash memory or other similar solid-state storage device, or an array of devices including devices in a storage area network or other configuration. The computer program product can be embodied in an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer or machine-readable medium such as the memory 904, the storage device 906, or a memory on the processor 902.

[0215] The high-speed controller 908 manages the bandwidth-intensive operations of the computing device 900, while the low-speed controller 912 manages the less bandwidth-intensive operations. This allocation of functionality is merely exemplary. In one implementation, the high-speed controller 908 is coupled to the memory 904, the display 916 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 910, which may accept various expansion cards (not shown). In that implementation, the low-speed controller 912 is coupled to the storage device 906 and the low-speed expansion port 914. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled, for example, via a network adapter, to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a network device, such as a switch or router.

[0216] The computing device 900 may be implemented in many different forms, as shown in the figure. For example, it may be implemented as a standard server 920, or multiple times within a group of such servers. It may also be implemented as part of a rack server system 924. In addition, it may be implemented in a personal computer, such as a laptop computer 922. Alternatively, components from the computing device 900 may be combined with other components in a mobile device (not shown), such as device 950. Each such device may include one or more computing devices 900, 950, and the overall system may consist of multiple computing devices 900, 950 in communication with each other.

[0217] Computing device 950 includes, among other components, a processor 952, memory 964, input / output devices such as a display 954, a communication interface 966, and a transceiver 968. Device 950 may also include a storage device, such as a microdrive or other device, to provide additional storage. Each of the components 950, 952, 964, 954, 966, and 968 are interconnected using various buses, and some of the components may be mounted on a common motherboard or in other manners as desired.

[0218] The processor 952 can execute instructions within the computing device 950, including instructions stored in the memory 964. The processor may be implemented as a chipset of chips including separate and multiple analog and digital processors. The processor may provide for coordination of other components of the device 950, such as control of a user interface, execution of applications by the device 950, and wireless communication by the device 950.

[0219] The processor 952 may communicate with a user via a control interface 958 and a display interface 956 coupled to a display 954. The display 954 may be, for example, a TFT LCD (thin film transistor liquid crystal display) or an OLED (organic light emitting diode) display, or other suitable display technology. The display interface 956 may include appropriate circuitry for driving the display 954 to present graphical and other information to the user. The control interface 958 may receive commands from the user and convert them for providing to the processor 952. Additionally, an external interface 962 may be provided in communication with the processor 952 to enable near field communication of the device 950 with other devices. The external interface 962 may provide, for example, for wired communication in some implementations or for wireless communication in other implementations, and multiple interfaces may also be used.

[0220] The memory 964 stores information within the computing device 950. The memory 964 may be embodied as one or more computer-readable mediums or media, one or more volatile memory units, or one or more non-volatile memory units. An expansion memory 984 may also be provided and connected to the device 950 through an expansion interface 982, which may include, for example, a Single In Line Memory Module (SIMM) card interface. Such expansion memory 984 may provide additional storage space for the device 950 or may also store applications or other information for the device 950. In particular, the expansion memory 984 may include instructions for performing or supplementing the processes described above, and may also include security information. Thus, for example, the expansion memory 984 may act as a security module for the device 950 and may be programmed with instructions that enable secure use of the device 950. In addition, secure applications may be provided via the SIMM card, along with additional information, such as structuring the SIMM card with identification information in an unhackable manner.

[0221] The memory may include, for example, flash memory and / or NVRAM memory, as discussed below. In one implementation, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is, for example, a computer or machine-readable medium, such as memory 964, expansion memory 984, or memory on processor 952, which may be received via transceiver 968 or external interface 962.

[0222] The device 950 may communicate wirelessly via a communication interface 966, which may include digital signal processing circuitry as needed. The communication interface 966 may provide for communication under various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS. Such communication may occur, for example, via a radio frequency transceiver 968. In addition, short-range communication may occur, such as using Bluetooth, WiFi, or other such transceivers (not shown). In addition, a GPS (Global Positioning System) receiver module 980 may provide additional navigation and location related wireless data to the device 950, which may be used as needed by applications executing on the device 950.

[0223] Device 950 can also communicate audibly using audio codec 960, which can also receive voice information from a user and convert it into usable digital information. Audio codec 960 can also generate audible sounds for the user, such as through a speaker in a handset of device 950. Such sounds can include sounds from a voice telephone call, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications running on device 950.

[0224] The computing device 950 may be implemented in many different forms, as shown in the figure, including as a mobile phone 980. It may also be implemented as part of a smartphone 982, personal digital assistant, or other similar mobile device.

[0225] Various implementations of the systems and techniques described herein may be realized in digital electronic circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs executable and / or interpretable on a programmable system that includes at least one programmable processor, which may be special or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0226] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in high level procedural and / or object oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine readable medium" and "computer readable media" refer to any computer program product, apparatus and / or device (e.g., magnetic disks, optical disks, memories, programmable logic devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including machine readable media that receive machine instructions as a machine readable signal. The term "machine readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0227] To provide for user interaction, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide for user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, such as acoustic input, speech input, tactile input, etc.

[0228] The systems and techniques described herein may be implemented in a computing device that includes back-end components (e.g., as a data server), or that includes middleware components (e.g., an application server), or that includes front-end components (e.g., a client computer having a graphical user interface or web browser through which a user can interact with an implementation of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, etc.

[0229] Computing devices may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0230] A number of embodiments have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the invention.

[0231] Additionally, the logic flows depicted in the figures do not require the particular order or sequence shown to achieve desirable results. Additionally, other steps may be provided or removed from the described flows, and other components may be added or removed to the described systems. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A method (203) implemented on a computer using the structure (373) of processing modules (313, 323, 333, 363) to predict the failure of an industrial machine (113), comprising: Receiving (213) machine data ({{X...}}N) from the industrial machine (113) by first, second, and third dependent processing modules (313, 323, 333) arranged to provide intermediate data (1{Y...}, 2{Y...}, 3{Y...}) to an output processing module (363), wherein the training by cascade training (702 / 802) of the structure (373) is:[ By a first dependent processing module (313) that processes the machine data (223A) to determine a first intermediate status indicator (1{Y...}); By a second dependent processing module (323) that processes the machine data (223B) to determine a second intermediate status indicator (2{Y...}); By a third dependent processing module (333), which is an operating mode classification module that processes the machine data (223C) to determine an operating mode indicator (3{Y...}) of the industrial machine (113); which has been completed beforehand, and By an output processing module (363), which processes the first and second intermediate status indicators (1{Y...}, 2{Y...}) and the operating mode indicator (3{Y...}) (243), wherein the output processing module (363) predicts the failure of the industrial machine by providing prediction data ({{Z...}}). A method implemented on a computer, including this.[

2. Training (712, 812) the third dependent processing module (333) with historical machine data ({{X...}}N); Executing (742) the trained third dependent processing module (333) to obtain a historical mode indicator (3{Y...}) by processing the historical machine data ({{X...}}N); Train the first and second dependent processing modules (312, 322) with historical machine data ({{X...}}N) and a historical mode indicator (3{Y...}) (722, 822); Execute the trained first and second dependent processing modules (312, 322) (762, 862) to obtain first and second intermediate status indicators (1{Y...}, 2{Y...}) by processing the historical machine data ({{X...}}N); and Train the output processing module (362) with the historical mode indicator, historical machine data, and historical fault data ({Q...}) (732, 832), The method according to claim 1, wherein the computer uses a structure (373) in which the training in the above order is completed.

3. The determination of the operation mode indicator (3{Y...}) is performed by a trained operation mode classifier (333) based on historical machine data annotated by an expert, according to the method of claim 1 or 2.

4. The method according to claim 3, wherein the historical machine data annotated by an expert is sensor data.

5. The operation mode classifier (333) is trained based on historical machine data, and therefore the operation mode classifier (333) has grouped the operation time (tm) of the machine into clusters of time series segments (segm_1 / 3, segm_2 / 4) during training, according to the method of claim 1 or 2.

6. The clusters of time series segments (segm_1 / 3, segm_2 / 4) are assigned or selected automatically or through interaction with an expert and assigned to operation mode indicators (MODE_1, MODE_2), according to the method of claim 5.

7. The operation mode indicator is provided by the number of mode changes over time, according to the method of claim 1 or 2.

8. The status indicator (1{Y...}, 2{Y...}) is the method according to claim 1 or 2, selected from a current indicator indicating the current status and a predictive indicator indicating the future status.

9. The output processing module (363) is the method according to claim 1 or 2, predicting the failure of an industrial machine selected from time to failure, type of failure, remaining useful life, failure interval.

10. The operation mode indicator (3{Y...}) acts as a bias that is further processed by both the first and second subordinate processing modules (313, 323), the method according to claim 1 or 2.

11. The reception of machine data is performed by receiving a subset comprising sensor data, and the determination of the first and second intermediate status indicators is performed by the first and second subordinate processing modules that process the subset comprising sensor data, the method according to claim 1 or 2.

12. The reception of machine data (213) includes receiving data via a data harmonizer (382β, 382γ) that provides virtual machine data by a virtual sensor or filters incoming machine data - depending on the contribution of the machine data to failure prediction, the method according to claim 1 or 2.

13. The reception of machine data (213) via a data harmonizer (382β, 382γ) includes receiving machine data from a harmonizer that includes a processing module pre-trained by transfer learning, the method according to claim 12.

14. The reception of machine data (213) includes receiving machine data that is at least partially reinforced by data obtained from a simulation, the method according to claim 1 or 2.

15. Use of a method for predicting a failure of an industrial machine (113) according to claim 1 or 2, wherein prediction data ({Z...}) is transferred to a machine controller for controlling the machine.

16. Use of a method for predicting a failure of an industrial machine (113) according to claim 15, wherein the machine controller is caused to assume a mode for predicting the time until a failure that occurs at the latest in the industrial machine.

17. Use of a method for predicting a failure of an industrial machine (113) according to claim 15, wherein the machine controller is caused to assume a mode in which the time until maintenance of the machine occurs at the latest in the industrial machine.

18. An industrial machine (113) adapted to provide machine data ({{X...}}N) to a computer adapted to execute the method according to claim 1 and to receive prediction data ({Z...}) from the computer, and associated with a machine controller for switching the operating mode of the industrial machine according to a predefined optimum target.

19. The predefined optimum target is selected from avoiding maintenance as much as possible and operating in a mode in which a failure is predicted to occur at the latest, for the industrial machine (113) according to claim 18.

20. The industrial machine (113) according to claim 18 or 19, selected from a chemical reactor, a metallurgical furnace, a container, a pump, a motor, and an engine.

21. A method (702 / 802) implemented on a computer for training a module structure (372) having first, second, and third dependent processing modules (312, 322, 332) coupled to an output processing module (362), so that the module structure (372) can provide a failure indicator ({Z...}) with failure prediction for an industrial machine. Training the subordinate processing modules (312, 322, 332), then operating the trained subordinate processing modules, and subsequently including the application of cascade training for training the output processing module, The cascade training includes: Training the third subordinate processing module (333) with the history machine data ({{X...}}N); Executing the trained third subordinate processing module (333) (742) to obtain the history mode index (3{Y...}) by processing the history machine data ({{X...}}N); Training the first and second subordinate processing modules (312, 322) with the history machine data ({{X...}}N)) and the history mode index (3{Y...}) (722, 822); Executing the trained first and second subordinate processing modules (312, 322) (762, 862) to obtain the first and second intermediate status indices (1{Y...}, 2{Y...}) by processing the history machine data {{X...}N}; and Training the output processing module (362) with the history mode index, with the history machine data, and with the history failure data ({Q...}) (732, 832), a method implemented by a computer including this.

22. A computer program product that, when loaded into the memory of a computer system and implemented by at least one processor of the computer system, causes the computer system to execute the steps of the computer-implemented method according to claim 1 or claim 21.

23. A computer system including a plurality of processing modules that execute the steps of the computer-implemented method according to claim 1 or claim 21 when executed by the computer system.

24. A computer adapted to process machine data ({{X...}}N) by implementing the method according to claim 1 and further adapted to provide prediction data ({Z...}), The computer selects and switches the operating mode of the industrial machine (113) according to the prediction data and in accordance with a predefined optimal target, from operating in a mode that avoids maintenance as much as possible and in which a failure is predicted to occur at the latest,