Providing explanations of anomalies in industrial machines or plants through language sentences

The method uses a Large Language Model to convert sensor data into natural language explanations of anomalies in industrial machines, addressing user-friendliness and adaptability issues in existing systems, providing comprehensive insights into anomalies.

WO2025153725A1PCT designated stage expired Publication Date: 2025-07-24NUOVO PIGNONE TECH SRL
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/051224
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-18
Filing Date
2025-01-17
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing anomaly detection systems in industrial machines and plants provide explanations that are not user-friendly and require extensive lookup tables or predefined clusters, failing to offer detailed, natural language explanations that cater to the knowledge level of the user.

Method used

A computer-implemented method using a Large Language Model (LLM) to encode time series sequences from sensors, map them to sentence embeddings, and decode them into natural language sentences, enabling user-friendly explanations of anomalies without relying on predefined clusters.

Benefits of technology

Generates detailed, human-understandable explanations of anomalies, adaptable to user knowledge levels, even with limited training data, and capable of explaining previously unseen anomalies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025051224_24072025_PF_FP_ABST
    Figure EP2025051224_24072025_PF_FP_ABST
Patent Text Reader

Abstract

The computer-implemented method (2000) serves for providing explanatory information regarding anomalies in an industrial machine or plant in the form of natural language sentences; the method comprises the steps of: a) receiving (2100) a plurality of time series sequences deriving from a corresponding plurality of sensors of said industrial machine or plant, b) encoding (2200) said plurality of time series sequences thereby generating a corresponding plurality of time series embedding features, c) mapping (2300) said plurality of time series embedding features thereby generating a plurality of sentence embedding features, d) decoding (2400) said plurality of sentence embedding features thereby generating a natural language sentence, and e) transmitting (2500) said natural language sentence.
Need to check novelty before this filing date? Find Prior Art

Description

TITLEProviding Explanations of Anomalies inIndustrial Machines or Plants through Language SentencesDESCRIPTIONTECHNICAL FIELD

[0001] The subject matter disclosed herein relates to a method and a system for providing explanations of anomalies in an industrial machine or plant through text messages, as well as their applications.BACKGROUND ART

[0002] An “anomaly” in an industrial machine or plant is a situation in the machine or plant (during its operation) different from expected, i.e. an “anomalous situation”. It may be, for example, a temperature at a position of the machine or plant higher than a normal value. Typically, the normal value derives from the design of the machine or plant. The normal value may depend on the current operating condition of the machine or plant, for example the current ambient temperature, and / or on the current operating mode of the machine or plant, for example full-power or shut-down.

[0003] An “anomaly” may be more or less serious. For example, in a certain machine, in a certain operating condition, in a certain operating mode, if a temperature at a certain position has a value 50% higher than its normal value there is a problem (i.e. there is a “serious anomaly”), while if this temperature is only 5% higher than its normal value, the issue might be tolerated (i.e. there is a “mild anomaly”) and / or might even not be considered an anomaly.

[0004] When talking about an “anomaly” in an industrial machine or plant, reference may also be made to time, i.e. to the duration of an “anomalous situation”. For example, in a certain machine, in a certain operating condition, in a certain operating mode, if a temperature at a certain position maintains a value 50% higher than its normal value for e.g. 1 minute there is an “anomaly”, while if this temperature maintains a value 50% higher than its normal value only for e.g. 1 second there is no “anomaly”.

[0005] The desire to automatically detect anomalies in industrial machines and plants is well known. Specific solutions have already been offered both in patent literature and in scientific literature. For example, US 2015 / 267591 Al discloses a method of monitoring combustion anomalies in a gas turbomachine.

[0006] A further desire is to automatically determine the “cause” or “root cause” of an “anomaly” once detected in an industrial machine or plant. Few specific solutions have already been offered both in patent literature and in scientific literature. For example, US 2014 / 163926 Al discloses automated root cause analysis, US 2018 / 348747 Al discloses a system and a method for unsupervised root cause analysis of machine failures, and US2023205161 discloses a method and apparatus for monitoring industrial devices.SUMMARY

[0007] It would be desirable to explain a detected anomaly. This is specifically true if the explanation is for the benefit of the user of the industrial machine or plant, that sometimes is also its manager and / or its owner. In fact, the technicians who designed the industrial machine or plant has a much higher knowledge and understanding thereof.

[0008] As better explained in the following, an explanation of an anomaly may be reply information to one or more of the following questions: what is the anomaly ? what part(s) of the machine or plant is / are involved in theanomaly ? what part(s) of the machine or plant will be influenced the anomaly ? where does the anomaly comes from ? what is a cause of the anomaly ? what is the root cause of the anomaly ?

[0009] US 2022 / 207326 Al is a document entitled “anomaly detection, data prediction and generation of human-interpretable explanations of anomalies”. However, even if its title refers to “human-interpretable explanations”, it discloses a system that is only able to provide so-called “Shapley additive explanations”, i.e. SHAP; in other words, the explanation consists in the relationship between the input features and the output feature and the system presents to the user SHAP values as explanation (see paragraph

[0083] ). Furthermore, “The system may be configured to display the top 20 features with the largest SHAP values. In this manner, the system may arrange, organize, filter, or otherwise process the features' SHAP values in any suitable manner to explain to the user outputs of the system (such as which features most contribute to an anomaly or predicted data points).”. Therefore, such kind of explanation may be used only by a high-level technician and is far from being user-friendly.

[0010] Considering especially the user of the machine or plant, the best thing would be to provide explanations through natural language sentences. Possibly, such sentences may be turned into voices messages.

[0011] A trivial solution would be to detect an anomaly based on data collected from a plurality of sensors, to identify specifically what is the anomaly, and to use a LUT (= Look Up Table) for extracting a text message associated to the identified anomaly. However, in this case, the anomaly has to be identified in a very specific manner so that it can be used as an entry to the LUT. Furthermore, the size of the LUT should be enormous if a reasonably detailed explanation is expected.

[0012] The use of a lookup connected to anomaly detection is disclosed in US 2017 / 284896 Al; such solution provides the steps of: receiving time-series data associated with a piece of machinery, automatically determining an anomaly associated with the piece of machinery by comparing the time-series data with a model associated with the piece of machinery, automatically determining that the anomaly is not a known fault based on performing a lookup of known failure modes, and transmitting an alert associated with an unknown failure mode.

[0013] US 2014 / 163926 Al discloses the use of a matrix, called “correlation matrix”, in order to present to the user a list of the most likely explanations to a misbehavior; such solution is based on calculating distances between a detected misbehavior and a plurality of known misbehaviors stored in at least one database.

[0014] According to a first aspect, the subject matter discloses herein relates to a computer-implemented method for providing explanatory information regarding anomalies in an industrial machine or plant, the explanatory information being in the form of natural language sentences; the method comprises the steps of: a) receiving a plurality of time series sequences deriving from a corresponding plurality of sensors of said industrial machine or plant, b) encoding said plurality of time series sequences thereby generating a corresponding plurality of time series embedding features, c) mapping said plurality of time series embedding features thereby generating a plurality of sentence embedding features, d) decoding said plurality of sentence embedding features thereby generating a natural language sentence, and e) transmitting said natural language sentence.

[0015] According to a second aspect, the subject matter discloses herein relates to a computer-based system configured to carry out the above method; in the system, the steps “a”, “b”, “c”, “d”, “e” of the method are carried out byone or more electronic processing units, for example a computer, in particular a personal computer or a workstation, provided with a software program including software sets of instructions for these steps.

[0016] According to further aspects, the subject matter discloses herein relates to industrial machines and industrial plants able to provide to their users explanatory information in the form of natural language sentences regarding any anomaly occurring inside them.

[0017] The above method and the above system are of the type so-called Al (= Artificial Intelligence) method and system. Therefore, the quality of their performance depends on the preliminary training and the data used therefor.

[0018] The above method and the above system generates natural language sentences not directly from the time series sequences coming from sensors of an industrial machine or plant but passing through “embedding” features of the time series and “embedding” features of the sentence. In this way, the task is much simplified.

[0019] US2023205161 discloses a system that uses an autoencoder to detect anomalous behaviors in signals from industrial sensors and discloses a method to detect anomalies using an autoencoder-based approach. In particular, a textual report includes free-text information to classify the anomaly. The classification is done as follows: 1) extract the most important words in the text based on grammatical rules (e.g., remove “and,” “.” and other less- relevant words while keeping keywords). 2) use the extracted “bag-of-words” for clustering anomaly signals based on their descriptions. 3) cluster anomalies by first converting words to numbers (word2vec or similar) and then using k- means or t-SNE or other unsupervised clustering techniques. This approach “labels” anomalies based on unsupervised clustering: i.e., it only associates new anomalies with similar anomalies encountered in the past but does notspecify in what way they are similar.

[0020] It is to be noted that, according to the subject matter disclosed herein, the method based on the Large Language Model (LLM) approach does not need pre-defined clusters or categories, but the model can describe the signal more naturally, in human language, using an LLM to check the signal and describe (explain) the anomaly with text.

[0021] It is to be noted that, according to the subject matter disclosed herein, anomaly explanations may be generated even for anomalies not considered during preliminary training by leveraging on the acquired knowledge.

[0022] It is also to be noted that solutions according to the subject matter disclosed herein may be implemented even if not many (annotated) data regarding anomalies are available for preliminary training. However, it is important that sufficient data are available relating to normal operation of the machine or plant to be monitored and used for preliminary training.

[0023] Thanks to a Large Language Model that may be used according to the subject matter disclosed herein, explanations can go deeper if the user provides further so-called “prompt” to the Al system (“prompting” means “the act of trying to make someone say something”) after having received an initial explanation.

[0024] The Large Language Model that may be used according to the subject matter disclosed herein, may be an already-trained commercially-available one and is advantageously fine-tuned for anomaly detection and / or explanation.BRIEF DESCRIPTION OF THE DRAWINGS.

[0025] A more complete appreciation of the disclosed embodiments of the invention and many of the attendant advantages thereof will be readilyobtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings, wherein:Fig. 1 shows a block diagram of a first embodiment of an innovative explanation system,Fig. 2 shows a block diagram of a second embodiment of an innovative explanation system,Fig. 3 shows a block diagram of the first embodiment of the innovative explanation system including also components used only during training, andFig. 4 shows a flowchart of an embodiment of an innovative explanation method corresponding to the first embodiment of an innovative explanation system.DETAILED DESCRIPTION OF EMBODIMENTS

[0026] The subject matter disclosed herein falls in the field of Al (= Artificial Intelligence) and the terminology used herein corresponds to the one familiar to the person skilled in this art.

[0027] Industrial machines and plants are subject to anomalies, i.e. may operate differently from expected. In this case, its user may wonder: what is the anomaly ? and / or what part(s) of the machine or plant is / are involved in the anomaly ? and / or what part(s) of the machine or plant will be influenced by the anomaly ? and / or where does the anomaly comes from ? and / or what is a cause of the anomaly ? and / or what is the root cause of the anomaly ? Any answer to these questions may be considered an “explanation” of the anomaly.

[0028] Even if the machine or plant includes several sensors that provide a lot of data, it is not easy for a user to answer such questions. Al may help.

[0029] According to the subject matter disclosed herein, natural language sentences explaining the anomaly are generated for being then output to a user. Such sentences are not generated directly from the sensor data, but by mapping “embedding” features of explanation sentences with “embedding” features of sensor data.

[0030] Simply said, the word “embedding” in Al means representing in a numerical compressed way. More precisely, the key idea behind embeddings is to transform discrete and often high-dimensional data into vector (typically of real numbers) representations that can be easily understood and processed by machine learning algorithms. This facilitates the extraction of meaningful patterns and relationships in the data, allowing Al models to better understand and generalize from the input information.

[0031] In order to obtain explanations of anomalies in an industrial system or plant, a plurality of time series sequences are necessary. They derive from a corresponding plurality of sensors in the industrial machine or plant that detect or measure repeatedly, usually periodically, physical variables or quantities (called “features” in Al) such as temperatures, pressures, volumetric and mass flows, displacements, speeds (e.g. rotation speeds), accelerations, vibrations, valve opening levels, IGV set angular positions, IGV detected angular positions, gas compositions, burner statuses. Typically, sensors are “real” sensors, i.e. devices that repeatedly or continuously perform a measurement inside the machine or plant and repeatedly or continuously determine / output a (analog or digital) signal whose amplitude corresponds to the measured value. Alternatively, according to the subject matter disclosed herein, one or more of the sensors may be a so-called “virtual” sensor; as known, a “virtual” sensor is a piece of software running in a computer (it maybe the same computer carrying out the innovative method) that repeatedly calculates e.g. a formula using as input data from one or more “real” sensors and producing as output data of the “virtual” sensor as if a machine or plant would have a “real” sensor on-board instead of the “virtual” sensor.

[0032] Such time series sequences may derive from measurements (or calculations) performed “offline”, for example 1 hour or 1 day or 1 week or 1 month before being processed according to the present innovative method or by the present innovative system, or “in real time”. Even in the case of real time processing, a small delay usually exists between the measurement time and the processing time, for example 1 second or 1 minute. In any case, as far as the subject matter disclosed herein concerns, such time series sequences may be considered to be received from an external system having the function of collecting measurement data; this is schematically represented by the black arrow on the left of Fig. 1 and Fig. 2.

[0033] It is to be noted that a same electronic processing unit, for example a personal computer or a workstation, may integrate both the data receiving function and the processing functions according to the subject matter disclosed herein; for example the first function is performed by a first software task and the second function(s) is performed by a second software task.

[0034] In order to output natural language sentences to a user, the sentences have to be previously generated according to the subject matter disclosed herein (and usually stored in a memory). Thereafter, outputting to a user is usually performed by an external system, that may be called HI (= Human Interface), having specific abilities for user interaction, for example a screen or a speaker. Such sentences may be considered to be transmitted to the external system; this is schematically represented by the black arrow on the right of Fig. 1 and Fig. 2.

[0035] It is to be noted that a same electronic processing unit, for example a personal computer or a workstation, may integrate both the sentence transmitting function and the processing functions according to the subject matter disclosed herein; for example the first function is performed by a first software task and the second function(s) is performed by a second software task.

[0036] The block diagram of Fig. 1 shows the main components of an innovative system 1000 for providing explanatory information regarding anomalies in an industrial machine or plant, and corresponds to the flowchart of Fig. 4 of an innovative method 2000 for providing explanatory information regarding anomalies in an industrial machine or plant; Fig. 4 shows the main steps of method 2000.

[0037] Each component of system 1000 may be an electronic processing unit, for example with suitable software program for performing the different functions, in particular the functions of the steps of method 2000. It is to be noted that a same electronic processing unit may perform more than one function. A single electronic processing unit, for example a personal computer or a workstation, may perform some or all of the functions of the steps of method 2000 if suitably programmed.

[0038] System 1000 comprises: a receiver 100, a decoder 200, a mapper 300, a decoder 400, and a transmitter 500. As shown in Fig. 1, all these components are connected in series, i.e. the output of a previous block is coupled to the input of the following block. Receiver 100 receives sensor data, more precisely time series sequences (see black arrow on the left of Fig. 1 and Fig. 2), to be processed and transmitter 500 transmits natural language sentences for users (see black arrow on the right of Fig. 1 and Fig. 2). Therefore, decoder 200, mapper 300, and decoder 400 are the essential components of system 1000; the core of system 1000 is mapper 300, as will be apparent from the following.

[0039] Method 2000 comprises: receiving step 2100, encoding step 2200, mapping step 2300, decoding step 2400, and transmitting step 2500. Correspondingly to system 1000, encoding step 2200, mapping step 2300, and decoding step 2400 are the essential components of method 2000; the core of method 2000 is mapping step 2300, as will be apparent from the following. In Fig. 4, a BEGIN block and an END block are shown for clarity reasons. A (outer) loop control flow is also shown as, in general, steps 2100-2500 are repeated several times on different data, i.e. on different pluralities of time series sequences; typically, different pluralities relate to different time frames but derive from the same plurality of sensors of the industrial machine or plant; two different time frames may be overlapping or non-overlapping in time.

[0040] According to the embodiment of Fig. 1, computer-implemented method 2000 provides explanatory information regarding anomalies in an industrial machine or plant in the form of natural language sentences. It comprises the steps of: a) receiving (step 2100) a plurality of time series sequences deriving from a corresponding plurality of sensors of the industrial machine or plant, b) encoding (step 2200) the received plurality of time series sequences thereby generating a corresponding plurality of time series embedding features, c) mapping (step 2300) the encoded plurality of time series embedding features thereby generating a plurality of sentence embedding features, d) decoding (step 2400) the mapped plurality of sentence embedding features thereby generating a natural language sentence, and e) transmitting (step 2500) the generated natural language sentence.

[0041] The innovative method according to the subject matter disclosed herein, including method 2000 of Fig. 4, is not only a computer-implemented method but also an artificial intelligence method; therefore, typically, its processing (especially steps 2200, 2300 and 2400 in the case of method 2000) requires preliminary training of the electronic processing unit or units (see e.g. Fig. 1) configured to perform its function.

[0042] Herein, the word “explanation” is not exactly defined as it may correspond to different kind of understanding associated to an “anomaly”. For example, as already said, it may be an answer to any of the following questions: what is the anomaly ? or what part(s) of the machine or plant is / are involved in the anomaly ? or what part(s) of the machine or plant will be influenced by the anomaly ? or where does the anomaly comes from ? or what is a cause of the anomaly ? or what is the root cause of the anomaly ? An exact definition is not necessary as, according to any specific different realization of the invention, what is meant by “explanation” depends on the “explanations” offered to the system during preliminary training, and on what is learned by the system.

[0043] According to some embodiments, the natural language sentence may be an explanation of an anomaly or an indication of no anomaly if the processed plurality of time series sequences do contain an anomaly. In this case, the method incorporates implicitly also an anomaly detection function.

[0044] According to other embodiments, the natural language sentence corresponds always to an explanation of an anomaly. In this case, the method provides for an anomaly detection function additional to the functions shown for example in Fig. 4. Such additional function may be performed just before encoding step “b” (before step 2200 in e.g. Fig. 4) or just before mapping step “c” (before step 2300 in e.g. Fig. 4). Depending on the result of such additional function, i.e. whether the plurality of time series sequences contains ananomaly or not, further processing is performed or not.

[0045] According to some embodiments, step “d” is based on a predetermined so-called “prompt” (that could be considered implicit in such embodiments). This means that the decoding function is performed as if a same and predetermined request or question is input by a user; for example, a request may be: please provide an explanation of the anomaly; for example, a question may be: what is the anomaly ? or what part(s) of the machine or plant is / are involved in the anomaly ? or what part(s) of the machine or plant will be influenced by the anomaly ? or where does the anomaly comes from ? and / or what is a cause of the anomaly ? or what is the root cause of the anomaly ?

[0046] According to other embodiments, step “d” is based on a so-called “prompt” previously received from a user (see e.g. the system embodiment of Fig. 2). The “prompt” may change from iteration to iteration (see the outer and inner loops in Fig. 4) or from time to time or from user to user. In this case, preliminary training should preferably take into account the possibility of different “prompt”; in other words, there is be enough training data for each different “prompt”. This possibility allows the innovative system to adapt automatically to the needs of different users (for example, users with a different technical background knowledge) and / or to the current needs of the same user (for example, at a certain time the user may be interested in a simple explanation while at a different time the same user may be interested in a more complex explanation). In fact, system embodiment 1000’ of Fig. 2 differs from the system embodiment 1000 of Fig. 1 only in that decoder 400’ is configured to receive “prompt” from a user.

[0047] The following description refers both to Fig. 1 and Fig. 4. Reference will be made to Fig. 3 when describing components of system 100 used only during its training.

[0048] Step “b” is preferably carried out through an encoder section (see block 210 in Fig. 3) of an autoencoder (see block 200 in Fig. 1). The encoder section comprises preferably several layers, and the encoder section has a number of inputs and a number of outputs, the number of outputs being smaller than the number of inputs. In this way, the complexity of the following mapping step is reduced as the number of input features is lower. The size of the input data to the encoder section may be the number of sensors multiplied by the number of samples for each sensor multiplied by the number of neurons in the first layer. The size of the output data from the encoder section may be the number of sensors multiplied by the number of neurons in the last layer (that is well smaller than the number of neurons in the first layer). Autoencoder is a known architecture consisting of the connection of an encoder and a decoder, and is popular for reducing the dimensionality of dataset. For example, Chapter 7 of the article by M. Sewak et al. entitled “An Overview of Deep Learning Architecture of Deep Neural Networks and Autoencoders” in the Journal of Computational and Theoretical Nanoscience deals with this topic. During training of an autoencoder, hyperparameters are also learnt corresponding to the number of neurons in the various layers (see e.g. the article by H.J.P. Weerts et al. entitled “Importance of Tuning Hyperparameters of Machine Learning Algorithms” - arXiv:2007.07588), in particular the first layer of encoder section 210 and the last layer of encoding section 210.

[0049] Autoencoder 200 is typically pre-trained through an unsupervised training. Fig. 3 shows a block 250 configured to supervise the training activity on autoencoder 200 (see arrow pointing upward). Preferably the training of encoder section 210 of autoencoder 200 is performed using also decoder section 220 autoencoder 200.

[0050] Step “c” is preferably carried out through a first electronic processing unit, in particular an estimator (see block 300 in Fig. 1), configured toimplement a regression model. The number of inputs of estimator 300 corresponds to the number of outputs of encoder section 210 of autoencoder 200. The number of outputs of estimator 300 depends on the number of inputs of decoder 400 and will be explained afterwards.

[0051] The first electronic processing unit may comprises a neural network, in particular a recurrent neural network.

[0052] The first electronic processing unit is typically pre-trained through a supervised training. Fig. 3 shows a block 350 configured to supervise the training activity on estimator 300 (see arrow pointing downward). Such training aims at learning mappings from the embeddings generated from the trained autoencoder on time series sequences to embeddings generated from the trained autoencoder on sentences describing the sample / samples. This mapping can be learnt by training an estimator that is able to do such mapping; it may be a supervised regression problem. A supervised regression model that can handle the mapping from a multivariate sequential input to a multivariate sequential output can be used. An example of such models is a neural network of the RNN (= Recurrent Neural Networks) family. A RNN is an extension of a conventional Feed-Forward neural network with the ability of managing variable-length sequence inputs. Unlike the conventional Feed-Forward neural network, which are not generally able to handle sequential inputs and all the inputs (and outputs) dependent on each other, a RNN model provides some gates to store the previous inputs and leverages sequential information of the previous inputs. Such “special” memory is called recurrent hidden states and gives RNN the ability to predict what input is coming next in the sequence of input data. More detailed theoretical information about the RNNs can be found for example in the article of Alex Sherstinsky entitled “Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) Network” (arXiv: 1808.03314vl0).

[0053] Step “d” is carried out preferably through a second electronic processing unit (see block 400 in Fig. 1), in particular a decoder, configured to implement a transformer-based model, in particular a so-called LLM (= “Large Language Model”). Decoder 400 is configured to receive sentence embeddings from estimator 300 and to use for performing its function also “prompt” embeddings (whether implicit or received from a user). An LLM has several inputs, typically hundreds of inputs, (for example 384 or 512, depending on the embedding model chosen) and at least one output sentence, being the sentence a concatenation of so-called “tokens”. It is not to be excluded that, according to some embodiments, decoder 400 may have a fewer number of inputs, but typically at least some dozens. LLMs are nowadays generally quite well known in the field of Al; reference may be may for example to the article by Humza Naveed et al. entitled “A Comprehensive Overview of Large Language Models” (arXiv: 2307.06435v5).

[0054] The second electronic processing unit is typically pre-trained through a supervised training. Fig. 3 shows a block 450 configured to supervise the training activity on decoder 400 (see arrow pointing upward). The Large Language Model may be an already-trained commercially-available general- purpose one.

[0055] The second electronic processing unit is advantageously fine-tuned. In fact, if the Large Language Model is an already-trained commercially- available general-purpose one, it is advantageously to fine-tune it for anomaly detection and / or explanation based on specific training data. Block 450 in Fig. 3 shows a double function: training and fine-tuning; this is represented schematically by the crossing dashed line; however, in case the Large Language Model is an already-trained commercially-available one, only the fine-tuning function may be present in block 450.

[0056] The second electronic processing unit (block 400 in Fig. 1 and block400’ in Fig. 3) may be configured to transmit another natural language sentence based on the natural language sentence transmitted at step “e” and “prompt” from a user. In particular, if a Large Language Model, explanations can go deeper if the user provides to the Al system further “prompt” after having received an initial explanation.

[0057] As already explained, steps 2100-2500 may be are repeated several times on different data; this is represented schematically by the outer loop control flow in Fig. 4. Furthermore, steps 2200-2400 may be repeated even on the same plurality of time series sequence (see inner loop control flow in Fig. 4); in fact, considering a moving window, several sets of data may be extracted from a sequence of data by shifting progressively the window of one sample at a time.

[0058] As already explained, the innovative system is a computer-based system configured to carry out the innovative method. Each of steps “a”, “b”, “c”, “d”, “e” are carried out by one or more electronic processing units; in the system embodiment 1000 of Fig. 1, there is a unit for each step of the method embodiment 2000 of Fig. 2.

[0059] According to some embodiments, the innovative system may comprise also an anomaly detector, preferably an LLM based anomaly detection system: this is not shown in any figure. In this way, the system transmits explanatory information only in case of anomalies, itdetects generic anomalies, which do not fall in pre-defined categories. The LLM describes the anomalies in textual form and an agentic system can make decision and take actions based on the textual description of the anomalies, In particular, the agentic system can make a decision based on the response of the LLM (e.g., it could say 'there is an anomaly with these characteristics).

[0060] An agentic system can significantly enhance the functionality of LargeLanguage Models (LLMs) by enabling goal-oriented decision-making, autonomous learning, contextual understanding, and self-improvement. For example, it can decide to check the functioning of another sensor or read other log files before making a decision on how to handle a situation to achieve specific goals such as safety, speed, or reliability. The agentic system can make decisions based on the LLM’s responses, such as identifying an anomaly and determining the need to send an alarm. This approach allows LLMs to become more proactive by anticipating and responding to user needs, adaptive by adjusting to changing contexts and user preferences, and personalized by tailoring responses to individual users and their goals. In essence, agentic systems empower LLMs to operate more effectively and efficiently, ultimately leading to more focused, relevant, and empathetic interactions.

[0061] As the innovative system is typically an artificial intelligence system, it is subject to preliminary training. Considering the embodiments of Fig. 1, Fig. 2 and Fig. 3, units 200, 300 and 400 are typically trained. Such training may be largely independent. It is to be noted that training of unit 200 may be largely carried out on non-anomalous data while training of unit 300 is largely or only carried out on anomalous data.

[0062] According to some embodiments, the innovative system is configured to carry out the innovative method during (or only during) operation of the industrial machine or plant. In general, it is possible that an innovative system may have an “offline” operating mode and a “real-time” operating mode.

[0063] According to some embodiments, the innovative system is configured to trigger an alert / alarm based on the natural language sentence generated through the innovative method. In other words, the alert / alarm is triggered based on the generated explanation of the anomaly and not on simple detection of the anomaly.

[0064] According to some embodiments, the innovative system is configured to prompt a user to take an action on the industrial machine or plant or on one or more sub-systems of the industrial machine or plant or coupled to the industrial machine or plant based on the text message generated through the method. In other words, the explanation natural language sentences may contain recommendations to the user to take an action (if necessary) or to go deeper into the understanding of the anomaly so that he can take an action (if necessary).

[0065] An innovative system may be integrated into an industrial machine. Alternatively, it may be associated to an industrial machine and may be located for example remotely from the machine.

[0066] An innovative system may be integrated into an industrial plant. Alternatively, it may be associated to an industrial plant and may be located for example remotely from the plant.

[0067] In case of offline operation, what is necessary is that the innovative system receives the data generated by the sensors of the machine or plant.

[0068] Anomaly explanation may be useful at the place where the machine or plant is installed or remotely from the installation place for example at a maintenance data center.

[0069] The subject matter described herein can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structural means disclosed in this specification and structural equivalents thereof, or in combinations of them. The subject matter described herein can be implemented as one or more computer program products, such as one or more computer programs tangibly embodied in an information carrier (e.g., in a machine readable storage device), or embodied in a propagated signal, for execution by, or to control the operation of, data processing apparatus (e.g., aprogrammable processor, a computer, or multiple computers). A computer program (also known as a program, software, software application, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file. A program can be stored in a portion of a file that holds other programs or data, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network.

[0070] The processes and logic flows described in this specification, including the method steps of the subject matter described herein, can be performed by one or more programmable processors executing one or more computer programs to perform functions of the subject matter described herein by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus of the subject matter described herein can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0071] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processor of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled toreceive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. Information carriers suitable for embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices, (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks, (e.g., internal hard disks or removable disks); magneto optical disks; and optical disks (e.g., CD and DVD disks). The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0072] To provide for interaction with a user, the subject matter described herein can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, (e.g., a mouse or a trackball), by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0073] The techniques described herein can be implemented using one or more modules. As used herein, the term “module” refers to computing software, firmware, hardware, and / or various combinations thereof. At a minimum, however, modules are not to be interpreted as software that is not implemented on hardware, firmware, or recorded on a non-transitory processor readable recordable storage medium (i.e., modules are not software per se). Indeed “module” is to be interpreted to always include at least some physical, non-transitory hardware such as a part of a processor or computer. Two different modules can share the same physical hardware (e.g., two differentmodules can use the same processor and network interface). The modules described herein can be combined, integrated, separated, and / or duplicated to support various applications. Also, a function described herein as being performed at a particular module can be performed at one or more other modules and / or by one or more other devices instead of or in addition to the function performed at the particular module. Further, the modules can be implemented across multiple devices and / or other components local or remote to one another. Additionally, the modules can be moved from one device and added to another device, and / or can be included in both devices.

[0074] The subject matter described herein can be implemented in a computing system that includes a back end component (e.g., a data server), a middleware component (e.g., an application server), or a front end component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described herein), or any combination of such back end, middleware, and front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), e.g., the Internet.

[0075] Although some specific embodiments have been described herein, other embodiments are within the scope of the subject matter as per the annexed claims.

[0076] It is noted that one or more references are incorporated herein. To the extent that any of the incorporated material is inconsistent with the present disclosure, the present disclosure shall control. Furthermore, to the extent necessary, material incorporated by reference herein should be disregarded if necessary to preserve the validity of the claims.

Claims

CLAIMS1. A computer-implemented method for providing explanatory information regarding anomalies in an industrial machine or plant, the explanatory information being in the form of natural language sentences, wherein said natural language sentence is an explanation of an anomaly or an indication of no anomaly. wherein the method comprises the steps of: a) receiving (2100, 100) a plurality of time series sequences deriving from a corresponding plurality of sensors of said industrial machine or plant, b) encoding (2200, 200) said plurality of time series sequences thereby generating a corresponding plurality of time series embedding features, is characterized in that said method based on Large Language Model (LLM) approach, further comprising the steps of c) mapping (2300, 300) said plurality of time series embedding features thereby generating a plurality of sentence embedding features, d) decoding (2400, 400) said plurality of sentence embedding features thereby generating a natural language sentence, e) transmitting (2500, 500) said natural language sentence.

2. The method of claim 1, wherein step “d” is based on a predetermined “prompt”.

3. The method of claim 1, wherein step “d” is based on a “prompt” previously received from a user.

4. The method of claim 1, wherein step “b” is carried out through an encoder section (210) of an autoencoder (200),wherein said encoder section (210) comprises preferably several layers, wherein said encoder section (210) has a number of inputs and a number of outputs, said number of outputs being smaller than said number of inputs.

5. The method of claim 4, wherein said autoencoder (200) is pre-trained, wherein training of said autoencoder (200) is unsupervised.

6. The method of claim 1, wherein step “c” is carried out through a first electronic processing unit, in particular an estimator, configured to implement a regression model.

7. The method of claim 6, wherein said first electronic processing unit comprises a neural network, in particular a recurrent neural network.

8. The method of claim 6 or 7, wherein said first electronic processing unit is pre-trained, wherein training of said first electronic processing unit is supervised.

9. The method of claim 1, wherein step “d” is carried out through a second electronic processing unit, in particular a decoder, configured to implement a transformer-based model, in particular a Large Language Model.

10. The method of claim 9, wherein said second electronic processing unit is pre-trained, wherein training of said second electronic processing unit is semi-supervised or self-supervised.

11. The method of claim 10, wherein said second electronic processing unit is fine-tuned.

12. The method of claim 10, wherein said second electronic processing unit is configured to transmit another natural language sentence based on the natural language sentence transmitted at step “e” and “prompt” from a user.

13. The method of any preceding claims, wherein at least steps “b”, “c”, “d” and possibly step “e” are repeated for all samples in said time series sequences.

14. A computer-based system configured to carry out the method according claim 1, wherein each of steps “a”, “b”, “c”, “d”, “e” are carried out by one or more electronic processing units.

15. The computer-based system of claim 14, wherein it is configured to trigger an alert / alarm based on the natural language sentence generated through the method.

16. The computer-based system of claim 14, wherein it is further configured to prompt a user to take an action on the industrial machine or plant or on one or more sub-systems of the industrial machine or plant or coupled to the industrial machine or plant based on the text message generated through the method.

17. The computer-based system of claim 14, comprising further an anomaly detector, preferably an LLM based anomaly detection system to detects generic anomalies, which do not fall in pre-defined categories.

18. The computer-based system of claims 14 and 17, wherein the anomaly detection system is configurated for monitoring data from an industrial machine, detecteding at least an anomaly the LLM describes the at least an anomaly in textual form and an agentic system makes at least a decision and take at least an action based on the textual description of the at least an anomaly detected.

19. The computer-based system of claim 14, wherein an industrial machine comprising said computer-based system.

20. The computer-based system of claim 14, wherein an industrial plant comprising said computer-based system.

Citation Information

Patent Citations

  • Automated root cause analysis

    US20140163926A1

  • Method of monitoring for combustion anomalies in a gas turbomachine and a gas turbomachine including a combustion anomaly detection system

    US20150267591A1

  • System and method for unsupervised anomaly detection on industrial time-series data

    US20170284896A1

  • System and method for unsupervised root cause analysis of machine failures

    US20180348747A1

  • Anomaly detection, data prediction, and generation of human-interpretable explanations of anomalies

    US20220207326A1