Artificial intelligence (AI) collaborator for predictive efficiency

The integration of trained LLMs for expert agents in a debating strategy addresses the inefficiencies of human-dependent predictive maintenance, improving accuracy and reliability by merging data-driven analytics with domain knowledge.

WO2026035277A1PCT designated stage Publication Date: 2026-02-12GE INFRASTRUCTURE TECH LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/041698
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing predictive maintenance techniques heavily rely on human intervention, lacking adaptability and requiring manual feature engineering or sophisticated neural networks, which are time-consuming and lack intuitive decision explanations, leading to inefficiencies and trustworthiness issues.

Method used

A system integrating trained large language models (LLMs) for expert agents to engage in a self-exclusionary debating strategy to identify the most likely root cause of anomalies, seamlessly merging data-driven analytics with domain knowledge to enhance reliability and stability.

Benefits of technology

This approach reduces human intervention, enhances adaptability, and improves the accuracy and reliability of predictive maintenance by leveraging automated dialogue and cooperation among multiple human-like agents, optimizing maintenance capabilities and reducing downtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024041698_12022026_PF_FP_ABST
    Figure US2024041698_12022026_PF_FP_ABST
Patent Text Reader

Abstract

Briefly, embodiments are direct to a process, system, and article for receiving an initial ranked list of potential root causes of at least one anomaly of an industrial asset being monitored by a set of environmental sensors. Expert agents may identify second ranked lists of potential root causes of the at least one anomaly based at least in part on the initial ranked list, at least a portion of the set of expert agents comprising trained large language models (LLMs). The second ranked lists from each expert agent are processed to determine whether there is a consensus in corresponding rankings of the second ranked lists regarding the potential root causes. In response to a determination that a consensus has not been reached in the corresponding rankings of the second ranked lists, each expert agent of the set of expert agents engages in a self-exclusionary debating strategy to identify the most likely correct ranked list from the second ranked lists until a consensus among each expert agents is reached regarding the potential root causes.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No.: 701017-WO-l (G30.435PCT)ARTIFICIAL INTELLIGENCE (Al) COLLABORATOR FOR PREDICTIVE EFFICIENCYBACKGROUND

[0001] Reducing manufacturing and operational costs, improving production reliability and efficiency, and addressing sustainability and safety concerns are paramount priorities across industries. Predictive maintenance has emerged as a crucial solution strategy, harnessing extensive real-time and historical data from the shop floor. An objective of predictive maintenance strategies is to detect early anomalies in manufacturing or equipment behavior, predict the future health of equipment, and formulate maintenance plans to eliminate or mitigate the impact of predicted failures.

[0002] However, prevailing predictive maintenance techniques, though predominantly data- driven, heavily rely on human intervention. This dependency is notable in methods requiring manual feature engineering by domain experts or sophisticated neural networks lacking intuitive decision explanations. The former lacks adaptability, necessitating repetitive reconfiguration when applied to new applications, while the latter demands additional human time for comprehension, potentially eroding trust in analytics if findings lack physical coherence.

[0003] Moreover, applying maintenance knowledge and experience encounters constraints, leading to impractical scenarios for manual analysis. For example, manual scanning of an analytic output of an asset being monitored consumes significant human labor. Moreover, not all required knowledge may be available, especially for new assets, in which case a human may not have the requisite knowledge base to detect potential maintenance issues. Additionally, experienced human experts may not be available promptly, particularly in emergencies.SUMMARY

[0004] According to an aspect of an example embodiment, a process may include receiving an initial ranked list of potential root causes of at least one anomaly of an industrial asset being monitored by a set of environmental sensors, where the at least one anomaly is representative of a failure or predicted failure of at least a portion of the industrial asset. Each expert agent of a set of expert agents may be instructed to identify second ranked lists of the potential root causesDocket No.: 701017-WO-l (G30.435PCT) of the at least one anomaly based at least in part on the initial ranked list, at least a portion of the set of expert agents comprising trained large language models (LLMs). The second ranked lists may be received from each expert agent and a determination may be made whether there is a consensus in corresponding rankings of the second ranked lists regarding the potential root causes of the at least one anomaly. In response to a determination that a consensus has not been reached in the corresponding rankings of the second ranked lists, each expert agent of the set of expert agents may be instructed to engage in a self-exclusionary debating strategy to identify the most likely correct ranked list from the second ranked lists until a consensus among each expert agent of the set of expert agents is reached regarding the potential root causes of the at least one anomaly.

[0005] According to an aspect of another example embodiment, a system may include a set of environmental sensors to measure operating conditions of an industrial asset. An analytics agent group may implement one or more data-driven artificial intelligence (Al) model modules to process the measured operating conditions of the industrial asset from the set of environmental sensors and determine an initial ranked list of potential root causes of at least one anomaly of the industrial asset, where at least one anomaly is representative of a failure or predicted failure of at least a portion of the industrial asset. An expert agents group of expert agents may: (a) identify second ranked lists of the potential root causes of the at least one anomaly based at least in part on the initial ranked list, at least a portion of the set of expert agents comprising trained large language models (LLMs), (b) determine whether there is a consensus in corresponding rankings of the second ranked lists regarding the potential root causes of the at least one anomaly, (c) in response to a determination that a consensus has not been reached in the corresponding rankings of the second ranked lists, each expert agent of the set of expert agents may engage in a self- exclusionary debating strategy to identify the most likely correct ranked list from the second ranked lists until a consensus among each expert agent of the set of expert agents is reached regarding the potential root causes of the at least one anomaly, and (d) identify actionable suggestions to address the at least one anomaly in response to the set of expert agents reaching the consensus.

[0006] According to an aspect of another example embodiment, an article may comprise a non- transitory storage medium comprising machine-readable instructions executable by one or more processors to perform one or more operations. For example, the one or more processors mayDocket No.: 701017-WO-l (G30.435PCT) process a received initial ranked list of potential root causes of at least one anomaly of an industrial asset being monitored by a set of environmental sensors, where the at least one anomaly is representative of a failure or predicted failure of at least a portion of the industrial asset. The one or more processors may also instruct each expert agent of a set of expert agents to identify second ranked lists of the potential root causes of the at least one anomaly based at least in part on the initial ranked list, at least a portion of the set of expert agents comprising trained large language models (LLMs). The one or more processors may additionally process received second ranked lists from each expert agent and determine whether there is a consensus in corresponding rankings of the second ranked lists regarding the potential root causes of the at least one anomaly. The one or more processors may further, responsive to a determination that a consensus has not been reached in the corresponding rankings of the second ranked lists, instruct each expert agent of the set of expert agents to engage in a self-exclusionary debating strategy to identify the most likely correct ranked list from the second ranked lists until a consensus among each expert agent of the set of expert agents is reached regarding the potential root causes of the at least one anomaly.

[0007] Other features and aspects may be apparent from the following detailed description taken in conjunction with the drawings and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Features and advantages of the example embodiments, and the manner in which the same are accomplished, will become more readily apparent with reference to the following detailed description taken in conjunction with the accompanying drawings.

[0009] FIG. 1 illustrates an embodiment of a pressurized water reactor (PWR) for which anomalies have been detected.

[0010] FIG. 2 illustrates an embodiment of a process utilizing a Root Cause Analysis (RA) Framework for determining root causes of a potential failure in accordance with an embodiment.

[0011] FIG. 3 illustrates an embodiment of a process utilizing a Deep Root Cause Analysis (DRA) Framework for determining root causes of a potential failure in accordance with an embodiment.Docket No.: 701017-WO-l (G30.435PCT)

[0012] FIG. 4 illustrates an embodiment of a process utilizing a Large Language Model (LLM) - Enhanded Root Cause Analysis (LDRA) Framework for determining fundamental root causes of a potential failure in accordance with an embodiment.

[0013] FIG. 5 illustrates an embodiment of a flow diagram of a process for an LDRA Framework to process training data and determine a fundamental root cause of a fault of an asset.

[0014] FIG. 6 illustrates a diagram of an embodiment of a system for determining a fundamental root cause of an anomaly or fault experienced by an asset.

[0015] FIG. 7 illustrates an embodiment of an LLM module which includes multiple LLMs according to an embodiment.

[0016] FIG. 8 is an embodiment of a flow diagram of a process for a multi-expert agent debating strategy.

[0017] FIGS. 9A-D illustrate an embodiment of a multi-expert agent debating strategy.

[0018] FIG. 10 illustrates a graphical representation of results of a multi-expert agent debating strategy in accordance with an embodiment.

[0019] FIG. 11 illustrates a computing device according to an embodiment.

[0020] Throughout the drawings and the detailed description, unless otherwise described, the same drawing reference numerals will be understood to refer to the same elements, features, and structures. The relative size and depiction of these elements may be exaggerated or adjusted for clarity, illustration, and / or convenience.DETAILED DESCRIPTION

[0021] In the following description, specific details are set forth in order to provide a thorough understanding of the various example embodiments. It should be appreciated that various modifications to the embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosure. Moreover, in the following description, numerous details are set forth for the purpose of explanation. However, one of ordinary skill in the art should understand that embodiments may be practiced without the use of these specificDocket No.: 701017-WO-l (G30.435PCT) details. In other instances, well-known structures and processes are not shown or described in order not to obscure the description with unnecessary detail. Thus, the present disclosure is not intended to be limited to the embodiments shown but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0022] Data-driven tools for asset health management face significant challenges, including a lack of understanding of physical principles, difficulty incorporating domain experts’ experiences, and consequently low detection accuracy, leading to trustworthiness issues. Automatically integrating data-driven analysis with human knowledge and experience, as found in literature and maintenance logs, is of critical importance. Recent progress in large language models (LLMs) offers opportunities to achieve this goal.

[0023] One or more embodiments, as discussed herein, present a framework that integrates a data-driven Artificial Intelligence (Al) approach with pretrained LLMs to address root cause detection in industrial failure analysis. Such a framework may also be used to determine how to construct or build an industrial asset in a way which minimizes the risk pf potential faults or failures. Such a framework may employ LLMs to analyze outputs from data-driven root cause analysis models, filtering out less relevant results and prioritizing those that align closely with physical principles and domain expertise. Such an approach leverages advanced data-driven analytics and a multi-LLM debate strategy for collaborative decision-making, seamlessly merging data-driven insights with domain knowledge. For example, through the use of self- exclusionary debates among multiple LLMs, biases inherent in single-LLM systems may be effectively mitigated, thereby enhancing reliability and stability. Crucially, a framework of one or more embodiments bridges a gap between data- driven models and physics-informed LLMs, accelerating the interaction between data and knowledge for more informed and realistic decision-making processes.

[0024] Reducing manufacturing and operational costs, improving production reliability and efficiency, and addressing sustainability and safety concerns are paramount priorities across industries. Predictive Maintenance emerges as a crucial solution, harnessing extensive real-time and historical data from the shop floor. Its objective is to detect early anomalies in manufacturing or equipment behavior, predict the future health of equipment, and formulate maintenance plans to eliminate or mitigate the impact of predicted failures.Docket No.: 701017-WO-l (G30.435PCT)

[0025] However, prevailing predictive maintenance techniques, though predominantly data- driven, heavily rely on human intervention. This dependency is notable in methods requiring manual feature engineering by domain experts or sophisticated neural networks lacking intuitive decision explanations. The former lacks adaptability, necessitating repetitive reconfiguration when applied to new applications, while the latter demands additional human time for comprehension, potentially eroding trust in analytics if findings lack physical coherence.

[0026] Moreover, applying maintenance knowledge and experience encounters constraints, leading to impractical scenarios for manual analysis. For example, the manual scanning of analytic outputs for an asset or system being monitored consumes significant human labor. Moreover, in a manual scanning approach, there are scenarios in which not all of the requisite knowledge may be available, particularly for relatively new systems or assets for which there may be few human experts. Experienced human experts also might not be available promptly, particularly in emergencies. Accordingly, one or more embodiments as discussed herein provide a technique to leverage prior knowledge from extensive documents such as textbooks, maintenance logs, and shop orders, therefore reducing human intervention in predictive maintenance systems and enhancing their adaptability.

[0027] In accordance with an embodiment, a system and process revolutionize predictive maintenance by seamlessly integrating Al analytics and human-like agents. Through an automated dialogue and cooperation among multiple human-like agents, a system and method may optimize maintenance capabilities by continuously encoding knowledge in a self-learning loop. Such an embodiment may enhance operational efficiency, reduce costs, and address sustainability and safety concerns in various industries.

[0028] In asset health management, fault detection and root cause analysis (RCA) refers to a process of identifying and diagnosing anomalies or malfunctions in equipment or processes to prevent failures and maintain efficiency. For example, RCA complements fault detection by uncovering fundamental causes of a predicted fault or anomaly, enabling targeted corrective actions and facilitating predictive maintenance strategies. An “anomaly,” as used herein, refers to an unusual behavior of a system, machine, or industrial asset and may be a signal of a malfunction.Docket No.: 701017-WO-l (G30.435PCT)

[0029] In recent years, a surge in sensor technologies has led to an unprecedented volume of time series data across diverse sectors, presenting both opportunities and challenges. Al models, particularly those designed to operate on time series data, have become crucial for autonomously identifying the underlying root causes of failures.

[0030] With the success and continuous improvement of LLMs, an opportunity has arisen to design human-like agents that leverage knowledge from vast textual resources. These agents may function as domain experts, digesting and summarizing findings from Al analytics. The use of such agents enables autonomous interaction between real-time data analysis and domain knowledge, facilitating optimized and automatic human-machine collaboration by combining computational capabilities from Al and knowledge from humans.

[0031] In accordance with one or more embodiments, a comprehensive digital twins system is provided for predictive maintenance, incorporating Al analytics and human-like agents for seamless communication between analytic findings and domain knowledge. Such a system integrates multiple agents responsible for data processing, Al analytics, human expertise, and system control. By establishing an automated dialogue and communication between human-like agents, a system may continuously encode knowledge in a self-learning loop, optimizing its capability in predictive maintenance - a process akin to the self-accumulation of knowledge.

[0032] These human-like agents play a crucial role in bridging the gap between data-driven analytics and human expertise, enhancing decision-making in predictive maintenance scenarios. Such an arrangement may gain insights into the emergent capabilities and limitations of these collaborative teams in proxy operational settings, specifically within the context of predictive maintenance. This process mirrors the self-accumulation of knowledge, with each side benefiting from and enhancing the intelligence of the other, thereby contributing to an efficient and effective autonomous predictive maintenance system. A system as discussed herein holds significant potential to enhance operational efficiency, reliability, and safety across various industries. By reducing downtime, optimizing maintenance schedules, and improving overall equipment effectiveness, it offers tangible benefits for businesses and organizations.

[0033] FIG. 1 illustrates an embodiment of a pressurized water reactor (PWR) 100 for which anomalies have been detected. Such anomalies may be indicative of an impending fault or failure of the PWR 100 or at least a portion thereof PWR 100 may include various componentsDocket No.: 701017-WO-l (G30.435PCT) within a reinforced concrete containment and shield 105. A steam generator 1 10 may receive water as an input and may convert the input water into steam. Steam generator 110 may provide coolant to a steel pressure vessel 115. Steel pressure vessel 115 may include control rods 120 as well as fuel elements 125. A pressure level of the coolant may be controlled be pressurizer 130 and provided back to steam generator 110.

[0034] PWR 100 may include various sensors disposed at various points to measure various changes in environmental conditions, such as temperature, pressure, liquid flow rate, gas flow rate, vibration levels, and humidity levels, to name just a few examples among many. If a predictive maintenance system detects certain anomalies such as a high pressure measurement, high temperature, and / or high flow rate, the predictive maintenance system may generate a warning to indicate that the detected anomalies may result in a system fault which might require downtime. In this example, a root cause of the high sensor value readings is a malfunctioning cooling system for the PWR 100. A “root cause,” as used herein, refers to a fundamental reason for the occurrence of an anomaly, problem, and / or predicted failure. However, if a human operator is tasked with determining why the high sensor measurements are detected, the human operator may have difficulty determining a root cause of the anomalies and may, in some circumstances, improperly diagnose the root cause of the anomalies, potentially resulting in catastrophic damage to the PWR 100 and / or a potentially unavoidable extended period of downtime.

[0035] There are different ways in which to detect anomalies. One way is for a human operator to manually scan a signature from sensors and check each individual sensor reading before a failure one at a time to try to identify the root cause of the anomaly. A “signature,” as used herein, refers to is a visual representation of everything that happened to an asset or a component of the asset during operation or use of the asset. For example, various functions performed by the asset may be associated with a repeatable signature or digital “fingerprint” when functioning properly.

[0036] Another approach to identify a root cause of an anomaly is by using data driven methods for which there is no physics knowledge or no domain expert charming. Such a data driven approach may involve a data driven model running a scan of sensor data and then detecting an anomaly and root cause. In a traditional data driven method, a time series may be obtained fromDocket No.: 701017-WO-l (G30.435PCT) various sensor readings and a prediction model may predict a residual value and a corresponding root cause. A “residual,” “prediction residual,” “residual value,” or “prediction residual value,” as used herein, in a time series model refers to what is left over after fitting a model. For example, a residual may comprise a difference between actual sensor readings and corresponding expected values for the sensor readings if a monitored system or component is functioning properly or as expected.

[0037] Conventional data-driven root cause analysis approaches often rely on identifying deviations between a detected sensor signature and an expected sensor signature through reconstruction or prediction, leveraging Al techniques such as autoencoders or long short-term memory (LSTM) networks. A signature may be associated with various channels such as data streams associated with various sensors, such as temperature and pressure sensors, for example. A channel for a temperature sensor may comprise a stream of temperature measurements determined by the temperature sensor, whereas a different channel for a pressure sensor may comprise a stream of pressure measurements determined by a pressure sensor in accordance with an embodiment. However, pinpointing the most deviated channels might emphasize downstream effects rather than direct causes. For example, in the scenario described above with respect to the PWR 100 of FIG. 1, a malfunction in the cooling system, leading to spikes in pressure, temperature, and flow rate within the control rods might be flags within conventional approaches as being causes as a result of to their substantial residuals in prediction. However, in reality, as discussed above, such elevated measurements may instead serve more as indications of the underlying cooling system problem rather than being the immediate cause of the malfunctions.

[0038] If an anomaly or failure occurs or is predicted to occur, there may be multiple reasons why the anomaly or failure happened. In order for an asset owner or operator to take an appropriate corrective action to extend the usable lifetime of the asset, it may be important to know what the actual root cause of the predicted or actual anomaly or failure is. For example, once a root cause has been determined, an asset owner or operator may be in a position to undertake corrective action and make corresponding repairs such as changing a part or component, or performing a software upgrade, to name just a couple examples of corrective actions among many. Identifying the root cause may take a relatively long time if performedDocket No.: 701017-WO-l (G30.435PCT) manually by a human operator. To improve a speed of root cause identification, data driven methods may be applied.

[0039] FIG. 2 illustrates an embodiment 200 of a process utilizing a Root Cause Analysis (RA) Framework 205 for determining root causes of a potential failure in accordance with an embodiment. In embodiment 200, an asset 210, such as an industrial asset, may be monitored to identify anomalies and predict impending failures. A time series of measurements may be generated by sensors 215 monitoring environmental conditions observed for asset 210. For example, the time series may comprise a multivariate time series collected from industrial sensors measuring items such as temperature, pressure, stress levels, humidity, vibration, and thermal expansion, to name just a few examples among many.

[0040] RA Framework 205 may include a prediction model module 220. Measurements from sensors 215 may be provided to prediction model module 220. A “module,” as used herein, refers to a component for performing one or more specific functions. For example, prediction model module 220 may be configured or otherwise designed to implement one or more prediction models to generate various prediction residuals as an output.

[0041] In accordance with embodiment 200, fault detection and root cause analysis may be performed on multivariate time series data collected from asset 210. A time series may be denoted as X Emxn, where m represents the number of input channels and n denotes the number of time steps. The value of the z-th channel at time step t may be denoted as Xt,t. A sliding window may be used to scan each time series, where Xt-e+i:t GRinz£represents the window containing the latest 6 time steps. RCA performed in accordance with one or more embodiments is built upon anomaly (fault) detection. Accordingly, an objective is to build a regression model for anomaly detection that takes a sliding window as input to predict the next time step and captures anomalies with high prediction residuals. Simultaneously, a second objective is to identify root cause channels for each detected faulty time series. Formally, the dual objectives are to first perform anomaly detection upon time series prediction:. For example, to predict the values of Xt+1 with each input window, Relation 1 may be performed:'Xt+1 = f(Xt-f+l :t). [Relation 1]Docket No.: 701017-WO-l (G30.435PCT)

[0042] An abnormal segment Xt-£+l :t may be identified if the absolute residual | "Xt+1 - Xt+11 is high. The entire time series X is labeled as an anomaly if the cumulative absolute residual | "X -X| surpasses a predefined threshold.

[0043] Prediction model module 220 may detect anomalies using a prediction model f(*) (e.g., Relation 1) trained on normal data. In an inference stage, anomalies may be identified by high prediction residuals, with channels exhibiting the highest residuals referred to as symptom signals.

[0044] Based on the prediction residual, a root cause may be identified. For example, a list of several potential root causes may be determined, where the potential root causes are ranked by likelihood of being the true or correct root cause in some implementations. A root cause can be said to be a factor that initiates a sequence of events leading to an outcome. Once a root cause is removed from a sequence of events that cause a problem, then the problem should not recur as the cause or sequence has been addressed. For example, different residual values may be associated with different potential root causes. In this case, the use of the prediction model and the output prediction residual may identify symptom channels. As used herein, a “symptom channel” refers to a channel of data which is deviating from its normal distribution or baseline profile by at least a threshold amount. For example, if temperature measurement values of a data stream from a temperature sensor are expected to have a baseline profile or particular distribution of temperature measurement values, but the actual temperature measurement values have different values, such as higher than expected temperature measurements, the elevated temperature measurements may comprise observable and / or visible indications of an issue or problem, such as an impending failure. In other words, a symptom channel refers to a channel of data values which is outside of a baseline profile or expected range of values and is indicative of an issue or problem with a device or system being monitored, for example.

[0045] In embodiment 200, a result of a prediction residual may be utilized to detect a root cause of an anomaly or predicted failure of asset 210. The prediction model module 220 of RA Framework 205 may detect anomalies with high prediction residuals. For example, prediction model module 220 may implement deviation detection, considering deviations from the norm as anomalies or faults. A prediction-residual based method may identify faults and root causes byDocket No.: 701017-WO-l (G30.435PCT) measuring prediction or reconstruction errors. Root causes are typically detected by high residuals between actual and predicted values across all observable channels.

[0046] However, the determination of the root cause based on a prediction residual determined via use of RA Framework 205 may be inaccurate or otherwise imprecise. Accordingly, further processing may be performed to identify a root cause with a greater level of confidence.

[0047] FIG. 3 illustrates an embodiment 300 of a process utilizing a Deep Root Cause Analysis (DRA) Framework 230 for determining root causes of a potential failure in accordance with an embodiment. In embodiment 300, DRA Framework 230 incorporates the RA Framework 205 of embodiment 200 shown in FIG. 2 with the inclusion of one or more additional processing model modules.

[0048] In embodiment 300, the prediction residual determined by prediction model module 220 may be provided to a root cause analysis module 235 which may perform processing or calculations to generate a saliency map. For example, root cause analysis model module 235 may regress residuals determined by prediction model module 220 and may derive one or more saliency maps. A “saliency map,” as used herein, refers to a topographically arranged map that represents visual saliency of a corresponding visual scene. For example, a saliency map may provide a visual representation of potential root causes and their associated relative importances to a detected anomaly or impending failure of asset 205.

[0049] Saliency maps derived by root cause analysis model module 235 may highlight channels potentially responsible for the high prediction residuals determined by prediction model module 220. Such a hierarchical structure involving both prediction model module 220 and root cause analysis model module 235 may enhance interpretability, providing detailed insights into root causes.

[0050] One or more root causes may be determined based on the saliency map generated by root causes analysis module 235 with a relatively greater degree of confidence or accuracy than may be possible if the root cause is determined solely based on the prediction residual output of prediction model module 220. For example, the root cause analysis module 235 may regress residuals from faulty time series. Saliency maps from the root cause analysis module 235 may highlight channels potentially responsible for the high prediction residuals identified by prediction model module 220, offering transparency in root cause location.Docket No.: 701017-WO-l (G30.435PCT)

[0051] Root causes analysis module 235 may implement a regression model regressing residuals from prediction model module 220. Saliency maps derived from root causes analysis module 235 may provide potential root causes for detected anomalies, extending their application beyond their origin in computer vision

[0052] Although the potential root cause channels identified by prediction model module 220 or root causes analysis module 235 may often include the fundamental or true root cause, there may be numerous scenarios in which the fundamental root causes of an impending failure of asset 205 are not always among the highest ranking determined root causes, leading to trustworthiness issues.

[0053] Although a structure of a DRA Framework 230 offers greater accuracy and promise than an RA Framework 205, particularly in scenarios where traditional approaches mistakenly focus on symptoms and downstream effects rather than root causes, a DRA Framework 230 remains purely data-driven and lacks integration with physical principles and validation by domain experts’ experiences. To further increase detection accuracy and trustworthiness, pre-trained LLMs may be utilized to analyze the output of the DRA Framework 230.

[0054] FIG. 4 illustrates an embodiment 400 of a process utilizing an LLM- Enhanded Root Cause Analysis (LDRA) Framework 250 for determining fundamental root causes of a potential failure in accordance with an embodiment. In embodiment 400, LDRA Framework 250 incorporates the DRA Framework 230 of embodiment 300 shown in FIG. 3 with the inclusion of an LLM module 255 to perform additional processing. In accordance with an embodiment, root causes determined based on saliency map after processing by root cause analysis model module 235 may be provided to LLM module 255. LLM module 255 may apply one or more LLMs to a list of ranked root causes determined based on the saliency map. For example, physical principles and domain experts’ experiences may be integrated into a root cause analysis framework by utilizing pretrained LLMs to identify the most relevant channels from the potential root causes identified by prediction model module 230 or root cause analysis module 235 according to their impact to symptom channels (those with high prediction residuals).

[0055] If LLM module 255 includes multiple LLMs, LDRA Framework 250 may include or otherwise implement a multi-LLM debating system to enhance the accuracy of conclusions, mitigating potential bias from a single LLM. A debating strategy may be implemented fromDocket No.: 701017-WO-l (G30.435PCT) among multiple LLMs, where the strategy involves multiple rounds of debating and self- exclusionary voting.

[0056] LLMs may be applied by LLM module 225 to determine a fundamental root cause of an impending failure of asset 205 with a greater level of precision than is possible based solely on the prediction residual from prediction model module 230 or the determined saliency map by root cause analysis model module 235. For example, determining a fundamental root cause via use of LDRA Framework 250 may provide an increase in accuracy in determining the true fundamental root cause be at least 100% versus a determination by DRA Framework 230 in some implementations.

[0057] Although a human operator may be capable of identifying many fundamental root causes of impending failures, a human operator may take a relatively long amount of time to identify the fundamental root cause and may lack the requisite knowledge base to determine the fundamental root cause in certain situations. For example, some rule-based systems which rely on predefined rules and expert knowledge may provide valuable insights but may lack adaptability and precision. Moreover, many devices accumulate vast quantities of high-frequency time series data over time, making manual analysis and creation of new rules impractical.

[0058] Although LLMs have been effectively used for certain scientific applications, their potential in prognostics and health management (PHM) remains underexplored. A significant challenge impeding the adoption of LLMs in industrial applications is the prevalence of biases and response variability. For example, in a system which employs multiple LLMs, each LLM may contain biases and may produce different answers to the same inputs, raising trustworthiness concerns. This discrepancy arises from several factors. First, different LLMs may be trained on diverse datasets or sources, leading to inherent biases or subjective interpretations. Second, differing perspectives and interpretations of the same context may cause variations in the responses generated by different LLMs. For example, even if several LLMs are trained on the same dataset with the same model structure, the patterns learned during pretraining of the LLMs may not fully generalize to all possible inputs. Consequently, LLMs may exhibit uncertainty or ambiguity in their predictions for certain inputs. When utilizing LLMs in PHM, it may be crucial to address and mitigate such biases and variabilities to ensure accurate and robust results.Docket No.: 701017-WO-l (G30.435PCT)

[0059] FIG. 5 illustrates an embodiment 500 of a flow diagram of a process for an LDRA Framework to process training data and determine a fundamental root cause of a fault of an asset. Embodiments in accordance with claimed subject matter may include all of, less than, or more than operations 505 through 545. Also, the order of operations 505 through 545 is merely an example order. For example, a method in accordance with process 500 may be performed by a computing device having one or more processors.

[0060] At operation 505, a first dataset of training data may be received, where the first dataset comprises data comprises “normal data.” Normal data comprises measurements received from sensors monitoring the asset which indicate what data streams should look like if a device or system being monitored is operating properly. At operation 510, prediction model processing may be performed on the first dataset. For example, the prediction model processing may correspond to processing performed by prediction model module 220 of embodiment 400 shown in FIG. 4.

[0061] At operation 515, prediction residual results may be determined which comprise anomalies with high residual values. At operation 520, channels with high residual values may be ranked as symptoms. At operation 525, normal data may be combined with second training data into a second dataset, where the second training data comprises the detected anomalies determined from the prediction model processing.

[0062] At operation 530, root cause analysis model processing may be performed on the second dataset. For example, the root cause analysis model processing may correspond to processing performed by root causes analysis model module 235 of embodiment 400 shown in FIG. 4.

[0063] At operation 535, regression may be performed on the predicted residual results determined at operation 515 and a saliency map be determined, where the saliency map comprises a map of a detected anomaly. At operation 540, a ranked list of potential root causes may also be determined based on the saliency map.

[0064] At operation 545, LLM model processing may be performed on the ranked link of potential root causes to identify a fundamental root cause of the anomaly experienced by the asset. The LLM model processing may include multi-LLM debating to ensure that the results of application of different LLM models are in agreement with the determination of the fundamental root cause.Docket No.: 701017-WO-l (G30.435PCT)

[0065] FIG. 6 illustrates a diagram of an embodiment 600 of a system for determining a fundamental root cause of an anomaly or fault experienced by an asset. The asset may comprise a digital twin of a real world asset, such as an industrial asset. A “digital twin,” as used herein, refers to a digital representation of a physical object, person, or process, contextualized in a digital version of its environment. Digital twins may assist in simulating real operational situations and associated outcomes, ultimately a preventive maintenance system to make better decisions.

[0066] Embodiment 600 may include various entities, such as an environment 605, data process agent 610, analytics agent group 615, experts agent group 620, control agent 625. A human operator 630 may be in communication with the environment 605 and the control agent 625.

[0067] Environment 605 may include a monitored or target industrial asset or system. Environment 600 may also include monitoring sensors, robots, cameras, or other monitoring devices or components for collecting data relating to the operation of the targeted industrial asset or system. Measurements determined or obtained by environment 605 may be provided to data process agent 610.

[0068] Data process agent 610 may perform data collection, digitization, and preprocessing based on the measurements obtained from environment 605. For example, data process agent 610 may determine a preprocessed dataset. Data process agent 610 may collect data from various sources in the environment 605 or monitored system and may also perform data cleaning, imputation, smoothing, and normalization with multiple options for data preprocessing, to name just a few examples among many. Data process agent 610 may also utilize or implement a reward mechanism within the entire digital-twin system to intelligently select the most suitable preprocessing option.

[0069] The preprocessed dataset may be provided to analytics agent group 615. Analytics agent group 615 may include or may implement various data-driven Al models to analyze or otherwise process the preprocessed dataset to detect any ongoing issues such as impending faults or anomalies based on residuals values determined from deviated signals. A ranked list of potential root causes of an anomaly or impending failure may be determined and prioritized based on the application of various data-driven rule-based models. Analytics agents group may, for example,Docket No.: 701017-WO-l (G30.435PCT) implement a DRA Framework such as DRA Framework 230 shown in embodiment 300 of FIG. 3.

[0070] Each agent of analytics agent group 615 may learn from the data collected and processed by the data process agent 610 to monitor and assess the current health status of a monitored or targeted system. Analytics agent group 615 may detect ongoing issues that may lead to system failure for the asset and may evaluate various anomalies. Analytics agent group 615 may identify the most deviated signals and potential root causes associated with an anomaly or impending failure. Analytics agent group may enable interaction and information exchange among agents to collectively arrive at conclusions or prioritize results. Consistency among agents’ results, individual confidence levels, and rewards from previous learning processes influence the decision-making process. For example, a rewards process may be implemented to improve the functioning of analytics agent group 615. If there are five different Al process models implemented within analytics agents group 615, but one of the Al process models has a recent history of more accurately identifying a true root cause of an anomaly or impending failure, control agent 625 may effectively reward the Al process model with the best proven track record by increasing a relatively weighting assigned to results determined by the rewarded Al process model while simultaneously reducing a relatively weighting of other Al process models associated with less accurate predictions. By implementing a rewards process in a feedback loop, the analytics agent group may improve its root cause identification ability over time.

[0071] A ranked list of potential root causes determined by analytics agent group 615 may be provided to experts agent group 620. Expert agents group 620 may include a plurality of different expert agents. For example, each of the expert agents may comprise a different LLM, where each LLM is trained in a different area of expertise. Each LLM may be trained on documents involving domain knowledge (e.g., textbooks, manuscripts, technical design documents, and papers) and human experiences (e.g., physics analytic records and papers), terming it as human-like expert agents. In one implementation, a first LLM may comprise a knowledge base in the field of physics, a second LLM may comprise a knowledge base in the field of thermodynamics, a third LLM may comprise a knowledge base in the field of chemistry, to name just a few examples among many. In some implementations, one or more of the expertDocket No.: 701017-WO-l (G30.435PCT) agents may comprise a human operator, such as a human being experienced or otherwise trained to monitor the industrial asset or system being monitored.

[0072] Expert agents group 620 may act as human experts, learn from domain knowledge and historical documents, and engage in debates on analytics agents outputs to determine a fundamental root cause of an impending fault or anomaly. Experts agent group 620 may also determine actionable suggestions as to how to correct or fix the anomaly or failure of the industrial asst or system, for example. Examples of actionable suggestions include recommendations for fixing operations of the monitored industrial asset, such as altering a flow rate, reducing a temperature of operation, or cleaning a valve, to name just a few examples among many. The actionable suggestions examples may also include suggestions to perform further examination, such as by collecting additional measurements from the industrial asset or system, The actional suggestions may additionally include initiating calls for maintenance and / or to perform an emergency shutdown of the industrial asset or system in some scenarios.

[0073] Each expert agent of experts agent group 620 may act or behave as a human expert, learning from domain knowledge and historical system documents, such as maintenance logs and shop order logs. Expert agents may be trained by an LLM, showcasing proficiency in generating coherent and contextually relevant text. An expert agent may analyze inputs containing an anomaly degree, most deviated signals, and a ranked list of potential root causes received from analytics agent group 620. Expert agents may engage in debates considering possible causes of detected anomalies and effects, and may provide actionable suggestions, such as recommending fixing operations or initiating calls for maintenance or emergency shutdown.

[0074] The actionable suggestions may be provided to a control agent 625. Control agent 625 may have access to suggested actions from the expert agent group 620, analysis reports, and relevant information from the analytics agent group 615 and expert agents group 620. Control agent 625 may also engage in direct interaction with real users such as human operator 630, providing feedback and receiving queries and commands. Control agent 625 may directly execute actions on the environment 605 or targeted system, eliciting responses detailing the sequence of these actions. Control agent 625 may additionally assign rewards to each agent based on the positive impact of their output or reaction in reaching a proper conclusion to enhance the status of the targeted system or accurately revealing its current state, as discussedDocket No.: 701017-WO-l (G30.435PCT) above. Based on user commands from human user 630 and / or subsequent reports from the environment after previous actions, control agent 625 may configure or tune settings of a data process agent 610 or analytics agent group 615, and / or provide new prompts to the experts agent group 625 for further instruction.

[0075] Expert agents group 620 may include the use of multiple LLMs. Utilizing multiple LLMs instead of a single LLM may provide various advantages. Using only one LLM via a single agent strategy to analyze the output from analytics agent group 615 may be potentially risky for several reasons. First, a single LLM may have inherent biases or subjective interpretations of the data, leading to potentially skewed or inaccurate results. A single LLM may therefore be associated with a limited perspective because different LLMs may have been trained on different datasets or sources of information, resulting in varying perspectives and interpretations of the same data. LLMs may also inherently have uncertainties in their predictions, and using only a single LLM may not capture the full range of possible interpretations or explanations for the data. Moreover, if a single LLM makes an incorrect interpretation or analysis, there is no mechanism in place to cross-check or validate its findings, potentially leading to erroneous conclusions.

[0076] To address these risks and improve the accuracy and reliability of the analysis, employing multiple expert agents such as LLMs for debating purposes may be beneficial. The use of multiple expert agents may provide a diversity of perspectives. For example, each expert agent may bring a unique perspective and interpretation to the analysis, increasing the likelihood of capturing diverse insights and reducing the impact of individual biases. The use of multiple expert agents may also provide a cross-validation benefit. For example, multiple expert agents such as LLMs may cross-validate each other’s findings, helping to identify and correct errors or inconsistencies in the analysis. The use of multiple expert agents may be useful in building a consensus as to the true or fundamental root cause of an anomaly or impending or predicted failure. Through rigorous debate and consensus-building among multiple expert agents, more robust and reliable conclusions can be reached, enhancing the overall confidence in the analysis. An enhanced level of accuracy may also be achieved. For example, by leveraging the collective intelligence of multiple expert agents, the likelihood of accurately identifying potential root causes and making informed decisions is increased. Overall, employing a multi-expert agent debating strategy may help mitigate the risks associated with using a single expert agent or LLMDocket No.: 701017-WO-l (G30.435PCT) for analysis and improve the accuracy and reliability of the output from the analytics agent group 615.

[0077] Moreover, requiring a consensus among LLMs or expert agents provides advantages relative to a strategy in which a single judge determines which root cause or list of ranked root causes most accurately describes the root cause of an anomaly. For a example, a single judge may exhibit certain biases, whereas a decentralized architecture in which a consensus among LLMs or expert agents is reached reduces or eliminates such biases.

[0078] FIG. 7 illustrates an embodiment 700 of an LLM module 702 which includes multiple LLMs according to an embodiment. As illustrated, LLM module 702 may include a first expert agent 705, a second expert agent 710, a third expert agent 715, and a clerk 720. Only three expert agents are shown in embodiment 700 for the sake of simplicity, but it should be appreciated that in some embodiments, more or fewer than three expert agents may be employed. Each of the expert agents may comprise a trained LLM. In some implementations, one or more of the expert agents may comprise a human being with expertise in the field of operation of an asset being monitored. Each of first expert agent 705, second expert agent 710, and third expert agent 715 may process a ranked list of root causes identified by a data-drive approach, such as via analytics agent group 615 in embodiment 600, and may determine its own ranked list of potential root causes of an anomaly or impending failure of a monitored asset. There is a multistage debating process conducted between the expert agents until a consensus has been reached among the expert agents where they are all in agreement as to the fundamental root cause of the anomaly or predicted failure. Clerk 720 may serve to keep a record of the results at each stage of the debate, such the ranked list of potential root causes determined by each expert agent and the associated reason for the anomaly or predicted failure. For example, the record maintained by clerk 720 may be stored in a memory or storage device such as in logs for later inspection or subsequent processing.

[0079] A multi-stage or multi-round debating process among a group of expert agents may be initiated when each of the expert agents receives a ranked list of potential root causes of an anomaly from a data-drive DRA framework. Each of the expert agents may be tasked with ranking potential root causes and providing a rational for the ranking. As discussed above, basedDocket No.: 701017-WO-l (G30.435PCT) on the knowledge base for each expert agent, the results of ranking potential root causes may differ among the expert agents.

[0080] FIG. 8 is an embodiment 800 of a flow diagram of a process for a multi-expert agent debating strategy. Embodiments in accordance with claimed subject matter may include all of, less than, or more than operations 805 through 825. Also, the order of operations 805 through 825 is merely an example order. For example, a method in accordance with process 800 may be performed by a computing device having one or more processors.

[0081] Embodiment 800 illustrates a process which may be performed by an expert agents group 620, which may include a plurality of LLMs, after receiving a ranked list of potential root causes from an analytics agents group 615 of Al models in response to an anomaly or impending failure of an industrial device or other system, such as is shown in embodiment 600 of FIG. 6. At operation 805 of embodiment 800, an initial ranked list of root causes may be received by an expert agents group from an analytics agents group of a DRA framework.

[0082] At operation 810, each expert agent, such as an LLM or a human expert in certain implementations, may be tasked with identifying a second ranked list of potential root causes of a detected anomaly based on the initial ranked list. For example, each expert agent may determine a ranking of likely root causes and corresponding explanations for the ranking, e.g., to give context as to which a particular ranking was determined.

[0083] At operation 815, a determination is made as to whether there is a consensus among the expert agents regarding the most likely root cause of the anomaly. For example, reaching a consensus may require that each of the expert agents identify the same root cause of the anomaly or impending failure of the asset being monitored.

[0084] If “no” at operation 815, processing proceeds to operation 820, where each expert agent is tasked with identifying the most likely correct ranked list from all of the ranked lists determined by the expert agents with a self-exclusion policy. For example, if there are three expert agents, a first expert agent, a second expert agent and a third expert agent, the first expert agent cannot select the list previously determined by the first expert agent at this operation. Instead, the first expert agent has to select the most likely root cause as being the ranked list determined by either the second expert agent or the third expert agent. The second and third expert agents similarly have to select the most likely ranked list as determined by the otherDocket No.: 701017-WO-l (G30.435PCT) expert agents again with the condition that each of these expert agents cannot select the ranked list of root causes they previously determined.

[0085] An advantage of self-exclusionary voting is promoting fairness and impartiality. For example, by preventing individual expert agents from voting for themselves, biases are reduced, leading to more objective decision-making. Exclusionary voting also encourages candidates to consider their peers’ qualifications and merits, fostering cooperation and collaboration within the voting community.

[0086] After the expert agents have determined their respective ranked lists from this second round, processing returns to operation 815, where a determination is made as to whether a consensus has been reached at the second round. If “no,” processing proceeds again to operation 820 for a third round of debating among the expert agents. However, if a determination is made that “yes,” a consensus has been reached at operation 815 as to the fundamental root cause of the detected anomaly or impending failure, and processing advances to operation 825, at which point actionable suggestions are identified to address the anomaly based on the determined fundamental root cause. The actional suggestions may also be transmitted to a control agent at operation 825 to either present the actionable suggestions to a user, such as via a user interface, or to automatically perform a suggested action in certain circumstances.

[0087] FIGS. 9A-D illustrate an embodiment 900 of a multi-expert agent debating strategy. Embodiment 900 illustrates a process which may be performed by an expert agents group 620, which includes a plurality of LLMs in this example, after receiving a ranked list of potential root causes from an analytics agents group 615 of Al models in response to an anomaly or impending failure of an industrial device or other system, such as is shown in embodiment 600 of FIG. 6. Embodiment 900 illustrates a three round debate among four different expert agents of an expert agents group. However, it should be appreciated that in some implementations, more or few than three rounds of debating may be performed among more or fewer than four expert agents. One or more of the expert agents may comprise a human expert or operator in some implementations.

[0088] During a first round 905, a prompt may be received and provided to each of the expert agents which indicates abnormal channels with high residual values and also indicates a ranked list of root causes as determined by a DRA framework, which may include an analytics agentDocket No.: 701017-WO-l (G30.435PCT) group which may apply one or more Al processing models. In the example depicted in embodiment 900, the prompt at Round 1 recites:“Our Anomaly Detector has identified ongoing issues with the following symptom signals: ['XMV_10-Reactor _cooling_ water_flow_valve’, 'XMEAS_ 41-Composition_ of_ H_producf , 'XMEAS_ 40-Composition_ of_ G_product’]The prompt further recites:“The potential root causes identified through our Deep Root Cause Analysis are: ['XMEAS_9-Reactor_temp', XMEAS_6-Reactor_feed_rate; 'XMEAS_24- Composition of B rxtr feed’, ‘XMEAS 17-Stripper_underflow’,'XM V_3 _feed_fl ow_val ve' ]The prompt additionally recites the following instructions for each of the expert agents:“In the context of Tennessee Eastman Process, please prioritize the potential root causes channels based on their impact on the symptom signals.”

[0089] In response to the prompt, each of the four expert agents determines its own ranking of the five ranked root causes received from the DRA framework as well as a corresponding explanation for the ranking. For example, in embodiment 900, Agent 1 determines the following in first round 905:Agent 1: 1 would rank the five channels as follows: XMEAS_ 6, XMJJ3, XMEASJ, XMEAS 24, and finally XMEASJ 7.This is because the reactor feed rate (XMEAS_ 6) directly affects the amount of reactants entering the system, impacting the temperature (XMEASJ), composition {XMEAS 24), and underflow (XMEAS 17). The A feed flow valve (XMV 3) controls the rate at which component A enters the reactor, further influencing these signals. Lastly, the stripper underflow (XMEASJ7) is influenced by all the previous channels and provides feedback on the process functionality.

[0090] Agent 2 determines the following in first round 905 of embodiment 900:Agent 2: 1 would rank the five channels as follows: XMEAS_ 6 >XMV 3 > XMEASJ > XMEAS 24 XMEASJ 7.Docket No.: 701017-WO-l (G30.435PCT)This is because: 1) Reactor feed rate (XMEAS 6) impacts the amount of reactants entering the reactor, affecting process efficiency. 2) A feed flow valve (XME 3) controls raw material flow into the reactor, influencing reaction rates and product composition.3) Reactor temperature (XMEAS _9) is crucial for controlling reaction rates and ensuring optimal performance. 4) The composition of B (XMEAS 24) affects process efficiency, selectivity, and yield.

[0091] Agent 3 determines the following in first round 905 of embodiment 900:Agent 3: 1 would rank the five channels as follows: XMEAS 9, XMEAS 6, XMV 3, XMEAS 24, and finally XMEAS _17.This is because 1) Reactor temperature (XMEAS _9) impacts the reaction rate, influencing the feed rate (XMEAS _6). 2) A feed flow valve (XMV 3) controls reactant flow, affecting both temperature and feed rate. 3) Composition ofB In reactor feed (XMEAS 24) affects the overall reaction balance. 4) Stripper underflow (XMEAS 17) is influenced by other factors but doesn't directly affect input signals like temperature, feed rate, or reactant flow.

[0092] Agent 4 determines the following in first round 905 of embodiment 900:Agent 4: 1 would rank the five channels as follows: XMEAS _9, XMEAS 6, XXIV 3, XMEAS 24, and finally XMEAS 17.This is because 1} Reactor temperature (XMEAS 9) controls reactions. 2) Reactor feed rate (XMEAS 6) affects production rates. 3) B-feed composition (XMEAS 24) influences product selectivity. 4) A-feedflow valve (XMV_ 3) controls reactant entry. 5) Stripper underflow (XMEAS 17) affects product purity. These rankings consider their impact on reactor cooling water flow valve and product compositions.

[0093] A clerk receives the ranked lists of potential root causes from each of the four expert agents and determines whether there is a consensus among the expert agents. In this example, the clerk determines that no consensus was reached among the expert agents. Since no consensus was reached during the first round 905 debate, a second round 910 is initiated during each of the expert agents is tasked to determine which of the ranked lists of potential root causes is most likely the true root causes based on the corresponding explanations provided with theDocket No.: 701017-WO-l (G30.435PCT) caveat that an expert agent cannot vote for the ranked list and explanation determined by the same expert agent in the previous round. In other words, a self-exclusionary rule is applied during the voting / debating process.

[0094] In embodiment 900, the following prompt is presented to each of the LLMs during second round 910:Each LLM need to evaluates all other LLMs opinion from the previous round and select the best one. A LLM cannot choose its own opinion unless at least one other LLM shared the same opinion in the previous round.

[0095] Agent 1 determines the following in second round 910 of embodiment 900:Agent 1: Agent 2 provides a better answer as it considers the direct impact of each channel on the process audits efficiency, starting with reactor feed rate which has the most significant effect on reaction temperature and overall production speed.

[0096] Agent 2 determines the following in second round 910 of embodiment 900:Agent 2: Agent 3 provides a better answer as it prioritizes temperature first because it has a significant impact on reaction rates, followed by feed rate which affects conversion levels and product distributions. Additionally, Agent 3 considers the control valve for A feed law and composition of B in the reactor feed, both of which have direct effects on product compositions and selectivity.

[0097] Agent 3 determines the following in second round 910 of embodiment 900:Agent 3: Agent 2 provides a better answer because it considers adjusting the composition of reactant feed (B) to be more influential than controlling the reactor temperature or stripper underflow rote, which aligns with the given symptom signals XMEAS_ 41 and XMEAS 40.

[0098] Agent 4 determines the following in second round 910 of embodiment 900:Agent 4: Agent 2 provides a better answer because its ranking takes into account the direct effects of each variable on the reaction rotes, product compositions, and process efficiency, ultimately leading to more accurate decision-making for optimizing the Tennessee Eastman Process.Docket No.: 701017-WO-l (G30.435PCT)

[0099] A clerk receives the ranked lists of potential root causes from each of the four expert agents determined in second round 910 and determines whether there is a consensus among the expert agents. In this example, the clerk determines that no consensus was reached among the expert agents during second round 910. Since no consensus was reached during the Round 2 debate, a third round 915 is initiated during each of the expert agents is tasked to determine which of the ranked lists of potential root causes is most likely the true root causes.[00100J In embodiment 900, the following prompt is presented to each of the LLMs during third round 915:Each LLM need to evaluates all other LLMs opinion from the previous round and select the best one. A LLM cannot choose its own opinion unless at least one other LLM shared the same opinion in the previous round.

[0101] Agent 1 determines the following in third round 915 of embodiment 900:Agent 1: Agent 2's opinion is slightly more comprehensive as it considers that the reaction rote and thus feed rate (XMEA5 6), directly impacts the reactor temperature (XMEA5 9), while also considering how the A-feed flow valve (XMV 3) affects multiple channels including temperature and feed rate. Therefore, I would select Agent 2.

[0102] Agent 2 determines the following in third round 915 of embodiment 900:Agent 2: Agent 2 offers a deeper analysis of how different factors influence each other in chemical reactor dynamics. Unlike other agents, Agent Z emphasizes the cascading impact from reactor feed rotes to reactant flows, reactor temperature, and composition balance, down to stripper underflow effects. This nuanced understanding of process dynamics makes Agent 2 the most suitable choice.

[0103] Agent 3 determines the following in third round 915 of embodiment 900:Agent 3: Agent 2 provides a more comprehensive explanation as it mentions how each channel impacts other channels In order, which Is crucial for understanding the causality of events within the system.

[0104] Agent 4 determines the following in third round 915 of embodiment 900:Docket No.: 701017-WO-l (G30.435PCT)Agent 4: Agent 2 provides a better answer as it takes into account more nuanced considerations such as reaction rates, output quality, efficiency etc., making their reasoning more comprehensive and well-rounded. Their response seems to demonstrate greater knowledge about intricate relationships within a complex industrial setting like the Tennessee Eastman Process.

[0105] The clerk receives the ranked lists of potential root causes from each of the four expert agents determined in third round 915 and determines that a consensus has been reached among the expert agents. Specifically, each of the expert agents is in agreement that the opinion of Agent 2 is most likely and that the ranking of potential root causes determined by Agent 2 in the first round 905 indicates the fundamental or true root cause, where XMEAS_ 6 >XMV_3 > XMEAS 9 > XMEAS 24 > XMEAS 17.

[0106] Since a consensus was reached, actionable suggestions are subsequently determined by the experts agents group and are transmitted to a control agent, such as is discussed above with respect to operation 825 of embodiment 800 of FIG. 8.

[0107] FIG. 10 illustrates a graphical representation 1000 of results of a multi-expert agent debating strategy in accordance with an embodiment. For example, graphical representation 1000 illustrates the results of each round of the debating strategy discussed above with respect to embedment 900 of FIG. 9. For example, graphical representation 1000 shows that during a first round, each expert agent determines a different ranked list of potential root causes of an anomaly and a corresponding explanations. During a second round, each expert agent selects the best ranked list determined by the other expert agents during the first round. During the third round of debating, a consensus is reached that the ranked list of corresponding root causes and the associated explanation initially provided by the second expert agent is correct and indicates that fundamental root cause of the anomaly.

[0108] FIG. 11 illustrates a computing device 1100 according to an embodiment. Computing device 1100 may implement a module, service, or any other device in accordance with the teachings discussed herein. Computing device 1100 may include a processor 1105. Processor 1105 may be utilized to execute an application, such as an application which implements a module to implement a processing model, LLM, or perform a multi-LLM debating strategy, for example. Computing device 1100 may include additional components, such as aDocket No.: 701017-WO-l (G30.435PCT) memory 11 10 or other type of storage device, a receiver 1 115, a transmitter 1120, and an Input / Output (I / O) port 1125. Processor 1105 may execute computer-executable code stored in memory 1110 which may be related to an application. For example, computing device 1100 may communicate via receiver 1115, transmitter 1120, and / or I / O port 1125.

[0109] One or more embodiments as discussed above introduces a comprehensive digital twins system tailored for predictive maintenance, amalgamating Al analytics and human-like agents to facilitate seamless communication between analytic findings and domain knowledge. Such a system may comprise various components, including a data process agent, a group of Al analytics agents, a group of human-like expert agents, and a control or interface agent, each of which play a vital role in data processing, analysis, human expertise integration, and system control. By integrating these components, the digital twins system may optimize maintenance operations through autonomous interaction, self-learning loops, and collaborative decisionmaking between Al and human expertise.

[0110] Moreover, employing multiple LLMs with an iterative and self-exclusionary debating strategy helps mitigate biases, enhances accuracy, and fosters consensus-building in identifying potential root causes of failures. Such a multi-agent and referendum / consensus approach introduces a sophisticated technique for debating and voting among LLMs, promoting fairness, impartiality, and cooperation within the decision-making process.

[0111] Through these advancements, the digital twins system offers promising prospects for predictive maintenance, addressing challenges associated with human intervention, adaptability, and knowledge availability. By continuously encoding knowledge and leveraging the collective intelligence of Al and human expertise, the system may enhance operational efficiency, reliability, and safety across various industries. Future research directions may focus on refining the system’s performance metrics, scalability, and generalizability to accommodate diverse predictive maintenance scenarios, further advancing the field towards fully autonomous and effective maintenance solutions.

[0112] As will be appreciated based on the foregoing specification, the above-described examples of the disclosure may be implemented using computer programming or engineering techniques including computer software, firmware, hardware or any combination or subset thereof. Any such resulting program, having computer-readable code, may be embodied orDocket No.: 701017-WO-l (G30.435PCT) provided within one or more non-transitory computer readable media, thereby making a computer program product, i.e., an article of manufacture, according to the discussed examples of the disclosure. For example, the non-transitory computer-readable media may be, but is not limited to, a fixed drive, diskette, optical disk, magnetic tape, flash memory, semiconductor memory such as read-only memory (ROM), and / or any transmitting / receiving medium such as the Internet, cloud storage, the internet of things, or other communication network or link. The article of manufacture containing the computer code may be made and / or used by executing the code directly from one medium, by copying the code from one medium to another medium, or by transmitting the code over a network.

[0113] The computer programs (also referred to as programs, software, software applications, “apps”, or code) may include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, cloud storage, internet of things, and / or device (e.g., magnetic discs, optical disks, memory, programmable logic devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The “machine-readable medium” and “computer- readable medium,” however, do not include transitory signals. The term “machine-readable signal” refers to any signal that may be used to provide machine instructions and / or any other kind of data to a programmable processor.

[0114] The above descriptions and illustrations of processes herein should not be considered to imply a fixed order for performing the process steps. Rather, the process steps may be performed in any order that is practicable, including simultaneous performance of at least some steps. Although the disclosure has been described in connection with specific examples, it should be understood that various changes, substitutions, and alterations apparent to those skilled in the art can be made to the disclosed embodiments without departing from the spirit and scope of the disclosure as set forth in the appended claims.

[0115] Some portions of the detailed description are presented herein in terms of algorithms or symbolic representations of operations on binary digital signals stored within aDocket No.: 701017-WO-l (G30.435PCT) memory of a specific apparatus or special purpose computing device or platform. In the context of this particular specification, the term specific apparatus or the like includes a general-purpose computer once it is programmed to perform particular functions pursuant to instructions from program software. Algorithmic descriptions or symbolic representations are examples of techniques used by those of ordinary skill in the signal processing or related arts to convey the substance of their work to others skilled in the art. An algorithm is here, and generally, considered to be a self-consistent sequence of operations or similar signal processing leading to a desired result. In this context, operations or processing involve physical manipulation of physical quantities. Typically, although not necessarily, such quantities may take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared or otherwise manipulated.

[0116] It has proven convenient at times, principally for reasons of common usage, to refer to such signals as bits, data, values, elements, symbols, characters, terms, numbers, numerals or the like. It should be understood, however, that all of these or similar terms are to be associated with appropriate physical quantities and are merely convenient labels. Unless specifically stated otherwise, as apparent from the following discussion, it is appreciated that throughout this specification discussions utilizing terms such as "processing," "computing," "calculating," "determining" or the like refer to actions or processes of a specific apparatus, such as a special purpose computer or a similar special purpose electronic computing device. In the context of this specification, therefore, a special purpose computer or a similar special purpose electronic computing device is capable of manipulating or transforming signals, typically represented as physical electronic or magnetic quantities within memories, registers, or other information storage devices, transmission devices, or display devices of the special purpose computer or similar special purpose electronic computing device.

[0117] It should be understood that for ease of description, a network device (also referred to as a networking device) may be embodied and / or described in terms of a computing device. However, it should further be understood that this description should in no way be construed that claimed subject matter is limited to one embodiment, such as a computing device and / or a network device, and, instead, may be embodied as a variety of devices or combinations thereof, including, for example, one or more illustrative examples.Docket No.: 701017-WO-l (G30.435PCT)

[0118] The terms, “and”, “or”, “and / or” and / or similar terms, as used herein, include a variety of meanings that also are expected to depend at least in part upon the particular context in which such terms are used. Typically, “or” if used to associate a list, such as A, B or C, is intended to mean A, B, and C, here used in the inclusive sense, as well as A, B or C, here used in the exclusive sense. In addition, the term “one or more” and / or similar terms is used to describe any feature, structure, and / or characteristic in the singular and / or is also used to describe a plurality and / or some other combination of features, structures and / or characteristics. Likewise, the term “based on” and / or similar terms are understood as not necessarily intending to convey an exclusive set of factors, but to allow for existence of additional factors not necessarily expressly described. Of course, for all of the foregoing, particular context of description and / or usage provides helpful guidance regarding inferences to be drawn. It should be noted that the following description merely provides one or more illustrative examples and claimed subject matter is not limited to these one or more illustrative examples; however, again, particular context of description and / or usage provides helpful guidance regarding inferences to be drawn.

[0119] A network may also include now known, and / or to be later developed arrangements, derivatives, and / or improvements, including, for example, past, present and / or future mass storage, such as network attached storage (NAS), a storage area network (SAN), and / or other forms of computing and / or device readable media, for example. A network may include a portion of the Internet, one or more local area networks (LANs), one or more wide area networks (WANs), wire-line type connections, wireless type connections, other connections, or any combination thereof. Thus, a network may be worldwide in scope and / or extent. Likewise, sub-networks, such as may employ differing architectures and / or may be substantially compliant and / or substantially compatible with differing protocols, such as computing and / or communication protocols (e.g., network protocols), may interoperate within a larger network. In this context, the term sub-network and / or similar terms, if used, for example, with respect to a network, refers to the network and / or a part thereof. Sub-networks may also comprise links, such as physical links, connecting and / or coupling nodes, such as to be capable to transmit signal packets and / or frames between devices of particular nodes, including wired links, wireless links, or combinations thereof. Various types of devices, such as network devices and / or computing devices, may be made available so that device interoperability is enabled and / or, in at least some instances, may be transparent to the devices. In this context, the term transparent refers toDocket No.: 701017-WO-l (G30.435PCT) devices, such as network devices and / or computing devices, communicating via a network in which the devices are able to communicate via intermediate devices of a node, but without the communicating devices necessarily specifying one or more intermediate devices of one or more nodes and / or may include communicating as if intermediate devices of intermediate nodes are not necessarily involved in communication transmissions. For example, a router may provide a link and / or connection between otherwise separate and / or independent LANs. In this context, a private network refers to a particular, limited set of network devices able to communicate with other network devices in the particular, limited set, such as via signal packet and / or frame transmissions, for example, without a need for re-routing and / or redirecting transmissions. A private network may comprise a stand-alone network; however, a private network may also comprise a subset of a larger network, such as, for example, without limitation, all or a portion of the Internet. Thus, for example, a private network “in the cloud” may refer to a private network that comprises a subset of the Internet, for example. Although signal packet and / or frame transmissions may employ intermediate devices of intermediate nodes to exchange signal packet and / or frame transmissions, those intermediate devices may not necessarily be included in the private network by not being a source or destination for one or more signal packet and / or frame transmissions, for example. It is understood in this context that a private network may provide outgoing network communications to devices not in the private network, but devices outside the private network may not necessarily be able to direct inbound network communications to devices included in the private network.

[0120] While certain exemplary techniques have been described and shown herein using various methods and systems, it should be understood by those skilled in the art that various other modifications may be made, and equivalents may be substituted, without departing from claimed subject matter. Additionally, many modifications may be made to adapt a particular situation to the teachings of claimed subject matter without departing from the central concept described herein. Therefore, it is intended that claimed subject matter not be limited to the particular examples disclosed, but that such claimed subject matter may also include all implementations falling within the scope of the appended claims, and equivalents thereof.

Claims

Docket No.: 701017-WO-l (G30.435PCT)WHAT IS CLAIMED IS:

1. A process comprising: receiving an initial ranked list of potential root causes of at least one anomaly of an industrial asset being monitored by a set of environmental sensors, the at least one anomaly being representative of a failure or predicted failure of at least a portion of the industrial asset; instructing each expert agent of a set of expert agents to identify second ranked lists of the potential root causes of the at least one anomaly based at least in part on the initial ranked list, at least a portion of the set of expert agents comprising trained large language models (LLMs); receiving the second ranked lists from each expert agent and determining whether there is a consensus in corresponding rankings of the second ranked lists regarding the potential root causes of the at least one anomaly; in response to a determination that a consensus has not been reached in the corresponding rankings of the second ranked lists, instructing each expert agent of the set of expert agents to engage in a self-exclusionary debating strategy to identify the most likely correct ranked list from the second ranked lists until a consensus among each expert agent of the set of expert agents is reached regarding the potential root causes of the at least one anomaly.

2. The process of claim 1, further comprising identifying actionable suggestions to address the at least one anomaly in response to the set of expert agents reaching the consensus.

3. The process of claim 2, further comprising transmitting one or more messages containing the actionable suggestions to a control agent.

4. The process of claim 3, wherein the control agent is to perform at least one of: presenting one or more of the actionable suggestions to a user via a user interface, or automatically perform at least one of the one or more actionable suggestions.

5. The process of claim 1, wherein at least one expert agent of the set of expert agents comprises a human expert.Docket No.: 701017-WO-l (G30.435PCT)6. The process of claim 1, wherein the industrial asset is a physical system being operated.

7. The process of claim 1, wherein the industrial asset is a physical system being designed.

8. The process of claim 1, wherein the initial ranked list of potential root causes is received from a group of one or more data-driven artificial intelligence (Al) model modules.

9. The process of claim 1, wherein the second ranked lists include corresponding explanations for the ranking of the potential root causes.

10. A system comprising: a set of environmental sensors to measure operating conditions of an industrial asset; an analytics agent group to implement one or more data-driven artificial intelligence (Al) model modules to process the measured operating conditions of the industrial asset from the set of environmental sensors and determine an initial ranked list of potential root causes of at least one anomaly of the industrial asset, the at least one anomaly being representative of a failure or predicted failure of at least a portion of the industrial asset; an expert agents group of expert agents to: identify second ranked lists of the potential root causes of the at least one anomaly based at least in part on the initial ranked list, at least a portion of the set of expert agents comprising trained large language models (LLMs), determine whether there is a consensus in corresponding rankings of the second ranked lists regarding the potential root causes of the at least one anomaly, in response to a determination that a consensus has not been reached in the corresponding rankings of the second ranked lists, each expert agent of the set of expert agents is to engage in a self-exclusionary debating strategy to identify the most likely correct ranked list from the second ranked lists until a consensus among each expert agent of the set of expert agents is reached regarding the potential root causes of the at least one anomaly, andDocket No.: 701017-WO-l (G30.435PCT) identify actionable suggestions to address the at least one anomaly in response to the set of expert agents reaching the consensus.

11. The system of claim 10, wherein the expert agents group is to transmit one or more messages containing the actionable suggestions to a control agent.

12. The system of claim 11, wherein the control agent is to perform at least one of: presenting one or more of the actionable suggestions to a user via a user interface, or automatically perform at least one of the one or more actionable suggestions.

13. The system of claim 10, wherein at least one expert agent of the set of expert agents comprises a human expert.

14. The system of claim 10, wherein the industrial asset is a physical system being operated.

15. The system of claim 10, wherein the industrial asset is a physical system being designed.

16. An article, comprising: a non-transitory storage medium comprising machine-readable instructions executable by a processor to perform: receiving an initial ranked list of potential root causes of at least one anomaly of an industrial asset being monitored by a set of environmental sensors, the at least one anomaly being representative of a failure or predicted failure of at least a portion of the industrial asset; instructing each expert agent of a set of expert agents to identify second ranked lists of the potential root causes of the at least one anomaly based at least in part on the initial ranked list, at least a portion of the set of expert agents comprising trained large language models (LLMs); receiving the second ranked lists from each expert agent and determining whether there is a consensus in corresponding rankings of the second ranked lists regarding the potential root causes of the at least one anomaly;Docket No.: 701017-WO-l (G30.435PCT) in response to a determination that a consensus has not been reached in the corresponding rankings of the second ranked lists, instructing each expert agent of the set of expert agents to engage in a self-exclusionary debating strategy to identify the most likely correct ranked list from the second ranked lists until a consensus among each expert agent of the set of expert agents is reached regarding the potential root causes of the at least one anomaly.

17. The article of claim 16, wherein the machine-readable instructions are further executable by the processor to identify actionable suggestions to address the at least one anomaly in response to the set of expert agents reaching the consensus.

18. The article of claim 17, wherein the machine-readable instructions are further executable by the processor to initiate transmission of one or more messages containing the actionable suggestions to a control agent.

19. The article of claim 16, wherein at least one expert agent of the set of expert agents comprises a human expert.

20. The article of claim 16, wherein the initial ranked list of potential root causes is received from a group of one or more data-driven artificial intelligence (Al) model modules.

Citation Information

Patent Citations

  • Container cloud micro-service performance anomaly root cause positioning method based on large language model

    CN118051370A

  • Identifying a root cause of an error

    US11797366B1