DCS software troubleshooting assistant
By monitoring data from DCS and automated equipment, and utilizing anomaly detection and Large Language Model (LLM) for troubleshooting, the problems of diagnostic latency and poor user experience in large distributed systems have been solved, enabling faster and more efficient troubleshooting.
Patent Information
- Application Number
- CN202510643917.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-21
- Filing Date
- 2025-05-19
- Publication Date
- 2025-11-21
AI Technical Summary
In large-scale distributed systems, there are latency and poor user experience issues when diagnosing software and hardware problems, especially in complex systems that span real-time controllers, field servers, and cloud host servers. Existing observability and diagnostic tools struggle to effectively distinguish between hardware and software failures.
By monitoring data from DCS and automated equipment, applying predetermined anomaly detection rules, performing similarity searches, and utilizing Large Language Models (LLM) for troubleshooting diagnosis and suggestions, a user interface is provided to improve diagnostic efficiency.
It reduces troubleshooting delays, improves fault tracking capabilities and the effectiveness of the monitoring process, reduces unnecessary downtime and improves user experience, and provides more comprehensive anomaly detection and diagnostic support.
Smart Images

Figure CN120994424A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a DCS software troubleshooting assistant. BACKGROUND
[0002] Diagnosing software problems and hardware problems in large distributed systems spanning real-time controllers, field servers and cloud host servers is challenging due to the complex interplay of thousands of hardware and software components. The observability and diagnostic software used in consumer-facing applications typically only focuses on logs, metrics and trace data from software components. Thus, if the root cause is not in the software components, it can be difficult for the user to diagnose the hardware / software problem. This can severely delay troubleshooting and can lead to unnecessary extended downtime and give a poor user experience.
[0003] Thus, there are several drawbacks with respect to diagnosing software problems and hardware problems in large distributed systems. Therefore, there is room for improvement and a need. SUMMARY
[0004] In view of the above, it is an object of the present disclosure to overcome at least part of the drawbacks that can arise when diagnosing software problems and hardware problems in large distributed systems.
[0005] Thus, in order to address one or more of these drawbacks, in a first aspect, there is provided a method for supporting troubleshooting of software and hardware problems in a Distributed Control System, DCS, associated with automation equipment in an industrial plant. The method comprises monitoring data from the DCS and / or monitoring data from the automation equipment. The method further comprises detecting an anomaly in the monitored data based on predetermined anomaly detection rules. The method further comprises, based on a result of the detecting, performing a similarity search for the detected anomaly against historical anomaly data associated with the DCS and / or the automation equipment. The method further comprises, based on a result of the performed similarity search, interrogating a Large Language Model, LLM, for a diagnosis and / or a suggestion for troubleshooting of the detected anomaly. The method further comprises obtaining, based on the interrogating, an output from the LLM, wherein the output indicates the diagnosis and / or the suggestion for troubleshooting of the detected anomaly. The method further comprises providing the output to a user.
[0006] It should be noted that the historical anomaly data can be understood to comprise data indicating past detected anomalies detected at the DCS and / or the automation equipment. Further, troubleshooting of the detected anomaly can be understood as eliminating or correcting the detected anomaly.
[0007] As a simple example to enhance understandability only, the diagnosis can include detecting an anomaly of a component of the automation equipment, e.g. a temperature in a boiler being above a predetermined temperature threshold. The data monitored from the DCS can not provide any hint why the predetermined temperature threshold can be exceeded. For example, no user logged in, who can have been allowed to override the preset temperature control. However, a similarity search can find a similar past case, according to which a heating element in the boiler was defective and thus exceeded the predetermined temperature threshold. Then, the diagnosis can include that the boiler comprises a defective heating element. Further, data related to the heating element of the boiler can be checked before providing the diagnosis and the defective heating element can be identified, so that the defective heating element can be indicated in the diagnosis. Based thereon, the recommendation can include repairing, replacing or deactivating the defective heating element.
[0008] The method according to the first aspect is advantageous in that it can contribute to enable a reduction of delay in troubleshooting since information about the operational technology (OT) equipment can not need to be manually related to information about the IT equipment. Thus, hardware / software problems can be diagnosed even if the root cause is in the production process or automation equipment. Thus, at least unnecessary prolonged downtime and poor user experience are reduced. Furthermore, more troubleshooting capabilities and more efficient monitoring processes are provided. Furthermore, troubleshooting sessions can be collected for further analysis and long-term product improvement.
[0009] According to some examples of the present disclosure, monitoring data from the DCS can include monitoring metrics, logs and traces of software components and / or hardware components from a server of the DCS. Monitoring data from the automation equipment can include monitoring production process data associated with a production process at the automation equipment.
[0010] It is noted that the production process data can be data from the automation equipment, e.g. in case the production process runs on the automation equipment.
[0011] Thus, a broad and comprehensive monitoring can be enabled. This can provide a basis for a broad and comprehensive anomaly detection.
[0012] According to several examples of the present disclosure, monitoring the metrics, logs and traces can include monitoring the metrics, logs and traces of at least one of the software components and / or hardware components based on a component-specific predetermined anomaly detection rule for the at least one of the software components and / or hardware components.
[0013] Thus, the monitoring can be made more independent, thus more detailed and more reliable. This can provide a basis for a further improved broad and comprehensive anomaly detection.
[0014] According to several examples of the present disclosure, the method can further comprise obtaining, from the monitoring of data from the DCS, first data indicative of first monitoring data. The method can further comprise obtaining, from the monitoring of data from the automation plant, second data indicative of second monitoring data. The method can further comprise jointly analyzing the first data and the second data based on correlating at least a part of the first data with at least a part of the second data and / or based on correlating at least a part of the second data with at least a part of the first data. The detecting can comprise detecting the anomaly further based on a result of the joint analysis.
[0015] It is noted that the joint analysis can also be understood as a combined analysis. The joint analysis can also be understood as a separate or sequential analysis of the first data and the second data, but a comparative or cross analysis of results from the separate or sequential analysis. The summary of the invention is that the first data is not analyzed separately or exclusively, and the second data is not analyzed separately or exclusively.
[0016] Merely as an example to improve understandability, an anomaly of a component of the automation plant can be detected, e.g. a temperature in a boiler is above a predetermined temperature threshold. The data monitored from the automation plant can not provide any hint why the predetermined temperature threshold can be exceeded. For example, no defective heating element can be identified. The data monitored from the DCS can not provide any hint why the predetermined temperature threshold can be exceeded when analyzed separately. However, for example, the data monitored from the DCS can indicate that a system user who is allowed to override a preset temperature control has logged into the system. The joint analysis can allow to identify a link between the user login and the exceeded predetermined temperature threshold. The similarity search can allow to diagnose that the user has (possibly simultaneously) overridden the preset temperature control, so that the result exceeded the predetermined temperature threshold. Based thereon, a recommendation can comprise readjusting the temperature or deactivating the permission of the user to override the preset temperature control.
[0017] Thus, further insights are obtained. Thus, the root cause identification of hardware / software problems is improved.
[0018] According to several examples of the present disclosure, the method can further comprise receiving, via a user interface, user feedback for the output; and performing at least one of:
[0019] - training or fine-tuning the LLM based on the received user feedback, and
[0020] - providing the received user feedback to the LLM to cause the LLM to process the received user feedback in a conversation between the user and the LLM.
[0021] It should be noted that since the method takes into account conversational user interfaces, the user can interact with the LLM (or the software agent communicatively connected to the LLM) in a dialogue or conversation to improve the diagnosis and / or recommendations obtained as output from the LLM regarding one or more detected anomalies. Thus, the user feedback can also be used immediately for the conversation with the LLM, since the user feedback can be directly processed by the LLM.
[0022] Thus, the accuracy, speed and quality of anomaly detection is further improved.
[0023] According to several examples of the present disclosure, the method can further comprise storing historical data indicative of at least one of:
[0024] - a plurality of past diagnoses regarding the detected anomaly,
[0025] - a plurality of past recommendations regarding the detected anomaly, and
[0026] - a plurality of user conversations regarding the detected anomaly, the conversations being received via the user interface.
[0027] The method can further comprise training or fine-tuning the LLM based on the stored historical data.
[0028] Thus, an extended knowledge base for training or fine-tuning the LLM can be obtained. Thus, the performance of the LLM for anomaly detection can be further improved.
[0029] According to several examples of the present disclosure, the interrogating can comprise iteratively interrogating the LLM. Additionally or alternatively, the output can be indicative of a plurality of diagnoses and / or a plurality of recommendations. Additionally or alternatively, obtaining the output can comprise obtaining a ranking of diagnosis alternatives and / or recommendation alternatives.
[0030] It should be noted that iteratively interrogating the LLM can be understood to also comprise repeatedly interrogating. For example, in a conversation between the LLM and the user via the software agent, the user can respond to the output provided by the LLM via the software agent. The LLM can then provide further output in response to the user's response. The user can then further respond and continue the interaction between the LLM and the user via the software agent, i.e. the user repeatedly responds to the output provided by the LLM, where each output is based on the latest user response received. Thus, it can be understood that the software agent iteratively interrogates the LLM.
[0031] Thus, the reliability, applicability and / or accuracy of the output of the LLM, i.e. the output diagnoses and / or recommendations, can be further improved in an easy, efficient, intuitive and user-friendly manner.
[0032] According to several examples of the present disclosure, the ranking can indicate a probability that the diagnosed correct and / or recommended application will be successful, and wherein the ranking can be based on training data used to train the LLM.
[0033] It is noted that the more training data related to anomalies is available for training the LLM, the more accurate the LLM can be in diagnosing detected anomalies and / or can provide applicable recommendations for detected anomalies. The training data can be historical data and / or simulated data. Thus, based on the increased knowledge base of the LLM, the LLM can more reliably know whether a diagnosis or recommendation is correct. Based thereon, the LLM can assign a probability to its output diagnosis or recommendation, wherein the probability can indicate a potential correctness of the diagnosis and / or can indicate a potential success in applying the recommended measures to eliminate the detected anomaly.
[0034] Thus, the user can be provided with several ranked alternatives, which can enable the user to more comprehensively evaluate the detected anomaly, and thus can provide further improved and more efficient troubleshooting.
[0035] According to several examples of the present disclosure, the method can further comprise automatically or autonomously taking measures for troubleshooting the detected anomaly based on predetermined rules for automatic or autonomous anomaly troubleshooting based on obtaining the result of the output.
[0036] For example, with reference to the example outlined above, the measures that can be taken automatically or autonomously by the software agent can be deactivating the defective heating element in the boiler.
[0037] Thus, the user is relieved and thus more focused on the detected anomaly, which can require the user to take action.
[0038] According to a second aspect, there is provided a control device or data processing device for supporting troubleshooting of software and hardware problems in a DCS associated with automation equipment in an industrial plant. The data processing device comprises a processor configured to perform the method of the first aspect.
[0039] An advantage of the data processing device according to the second aspect is that since information about OT equipment can not need to be manually related to information about IT equipment, it can contribute to a reduction in the delay in troubleshooting. Thus, it is possible to diagnose hardware / software problems even if the root cause is in the production process or automation equipment. Thus, at least unnecessary prolonged downtime and poor user experience is reduced. Furthermore, more troubleshooting capabilities and more efficient monitoring processes are provided. Furthermore, troubleshooting sessions can be collected for further analysis and long-term product improvements.
[0040] According to a third aspect, there is provided a data processing system or a security system for supporting troubleshooting of software and hardware problems in a DCS associated with automation equipment in an industrial plant. The data processing system comprises the data processing apparatus of the second aspect, configured to perform the method of the first aspect. Additionally or alternatively, the system comprises means for performing the method of the first aspect.
[0041] An advantage of the data processing system according to the third aspect is that since information about OT equipment can not need to be manually related to information about IT equipment, it can contribute to a reduction in the delay in troubleshooting. Thus, hardware / software problems can be diagnosed even if the root cause is in the production process or automation equipment. Thus, at least unnecessary extended downtime and poor user experience is reduced. Furthermore, more troubleshooting capabilities and more efficient monitoring processes are provided. Furthermore, troubleshooting sessions can be collected for further analysis and long-term product improvements.
[0042] According to some examples of the present disclosure, the data processing system can comprise a DCS software troubleshooting assistant, wherein the DCS software troubleshooting assistant can comprise:
[0043] - an anomaly configurator;
[0044] - an anomaly detector in communicative connection with the anomaly configurator;
[0045] - a diagnostic intelligence retriever in communicative connection with the anomaly detector and a conversational user interface; and
[0046] - the conversational user interface.
[0047] The anomaly detector can comprise an interface to a database comprising a software service log database and a hardware and / or software metrics database. The diagnostic intelligence retriever can comprise an interface to a plant history database and an LLM. The conversational user interface can comprise an interface for communication with a user. The anomaly detector can be configured to perform the monitoring of the method according to the first aspect via the interface to the database. The anomaly detector can further be configured to perform the detecting of the method according to the first aspect based on predetermined anomaly detection rules obtained from the anomaly configurator. The diagnostic intelligence retriever can be configured to perform the similarity search of the method according to the first aspect by using the interface to the plant history database. The diagnostic intelligence retriever can further be configured to perform the interrogation of the method according to the first aspect via the interface to the LLM. The conversational user interface can be configured to perform the providing of the method according to the first aspect to a user via the interface.
[0048] Thus, a system for efficient anomaly troubleshooting is provided.
[0049] According to several examples of the present disclosure, the DCS software fault diagnosis assistant can further comprise a fault diagnosis session saver in communication connection with the diagnostic intelligence retriever and the conversational user interface. The fault resolution session saver can comprise an interface to a fault resolution session database. The fault resolution session saver can be configured to store user feedback received via the conversational user interface and perform storage according to the method of the first aspect.
[0050] Thus, it is allowed to take into account user feedback and / or user sessions.
[0051] According to a fourth aspect, there is provided an industrial plant comprising the data processing apparatus of the second aspect and / or the data processing system of the third aspect, the data processing apparatus being configured to perform the method of the first aspect.
[0052] According to several examples, the “industrial apparatus” can refer to an industrial apparatus or industrial production apparatus comprising one or more pipelines, production lines and / or assembly lines for converting one or more educts into a product and / or for assembling one or more components into an end product. According to several examples, it can refer to an industrial apparatus in the oil industry, the natural gas industry or the chemical industry.
[0053] The industrial plant according to the fourth aspect is advantageous in that it can contribute to enable a reduction of the delay in fault resolution since information about OT devices can not need to be manually related to information about IT devices. Thus, hardware / software problems can be diagnosed even if the root cause is in the production process or automation equipment. Thus, at least unnecessary prolonged downtime and poor user experience is reduced. Furthermore, more fault tracking capabilities and more efficient monitoring processes are provided. Furthermore, fault resolution sessions can be collected for further analysis and long-term product improvement.
[0054] According to a fifth aspect, there is provided a computer readable medium comprising instructions which, when executed by a computing system, cause the computing system to perform the method of the first aspect. The computer readable medium can be transitory or non-transitory, volatile or non-volatile.
[0055] The computer readable medium according to the fifth aspect is advantageous in that it can contribute to enable a reduction of the delay in fault resolution since information about OT devices can not need to be manually related to information about IT devices. Thus, hardware / software problems can be diagnosed even if the root cause is in the production process or automation equipment. Thus, at least unnecessary prolonged downtime and poor user experience is reduced. Furthermore, more fault tracking capabilities and more efficient monitoring processes are provided. Furthermore, fault resolution sessions can be collected for further analysis and long-term product improvement.
[0056] According to a sixth aspect, there is provided a computer program product comprising instructions which, when executed by a computing system, enable the computing system or cause the computing system to perform the method of the first aspect. The computer program product can comprise a computer-readable medium comprising the instructions of the computer program product.
[0057] The computer program product according to the sixth aspect is advantageous in that it can contribute to enabling a reduction of the delay of troubleshooting since information about OT devices can not need to be manually related to information about IT devices. Thus, hardware / software problems can be diagnosed even if the root cause is in the production process or automation equipment. Thus, at least unnecessary prolonged downtime and poor user experience is reduced. Furthermore, more troubleshooting capabilities and more efficient monitoring processes are provided. Moreover, troubleshooting sessions can be collected for further analysis and long-term product improvements.
[0058] According to a seventh aspect, there is provided the use of the data processing apparatus of the second aspect, and / or the data processing system of the third aspect, and / or the industrial equipment of the fourth aspect.
[0059] The use according to the seventh aspect is advantageous in that it can contribute to allowing a reduction of the delay of troubleshooting since information about OT devices can not need to be manually related to information about IT devices. Thus, hardware / software problems can be diagnosed even if the root cause is in the production process or automation equipment. Thus, at least unnecessary prolonged downtime and poor user experience is reduced. Furthermore, more troubleshooting capabilities and more efficient monitoring processes are provided. Moreover, troubleshooting sessions can be collected for further analysis and long-term product improvements.
[0060] The method of the first aspect can be computer-implemented.
[0061] Optional features of the first aspect can form part of any of the second to seventh aspects mutatis mutandis.
[0062] As used herein, the term “obtaining” can include, for example, receiving from another system, device, or process; receiving via an interaction with a user; loading or retrieving from storage or memory; measuring or capturing using a sensor or other data acquisition device.
[0063] As used herein, the term “determining” encompasses a wide variety of actions and can include, for example, calculating, computing, processing, deriving, investigating, looking up (such as, for example, looking up in a table, a database or another data structure), ascertaining and the like. Also, “determining” can include receiving (such as, for example, receiving information), accessing (such as, for example, accessing data in a memory) and the like. Also, “determining” can include resolving, selecting, choosing, establishing and the like.
[0064] The indefinite article "a" or "an" does not exclude a plurality. Furthermore, "a" and "an" as used herein are used generically and are thus intended to include both the singular and the plural, unless the context clearly indicates otherwise.
[0065] The phrases "one or more of A, B and C", "at least one of A, B and C" and "A, B and / or C" as used herein are intended to mean any permutation of one or more of the listed items. That is, the phrase "A and / or B" means (A), (B) or (A and B), and the phrase "A, B and / or C" means (A), (B), (C), (A and B), (A and C), (B and C) or (A, B and C).
[0066] The term "comprising", used in the detailed description and throughout the claims, does not exclude other elements or steps. Furthermore, "comprising" as used herein is used generically that also covers "consisting only of".
[0067] The application can include one or more aspects, examples, or features alone or in combination. Any optional feature of one aspect of the above description can be applied to any other aspect, mutatis mutandis.
[0068] The above aspects will become apparent and elucidated from the detailed description provided hereinafter, reference being made to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0069] The detailed description will now be made with reference to the drawings, wherein:
[0070] Figure 1 Components and contextual views of a DCS software troubleshooting assistant are shown, in accordance with some examples of the present disclosure;
[0071] Figure 2 Techniques and architecture for implementing a DCS troubleshooting assistant are shown, in accordance with some examples of the present disclosure;
[0072] Figure 3 Control flow through a DCS software troubleshooting assistant is shown, in accordance with some examples of the present disclosure;
[0073] Figure 4 Flowcharts representing methods, in accordance with some examples of the present disclosure, are shown; and
[0074] Figure 5 A block diagram schematically illustrating a data processing apparatus, in accordance with some examples of the present disclosure, is shown. DETAILED DESCRIPTION
[0075] According to some examples of the present disclosure, a software agent, namely a DCS software troubleshooting assistant, is provided that can continuously monitor anomalies from DCS server hardware / software components, logs, and traces and can simultaneously correlate this data with production process data (e.g., time-series sensor data) from automation equipment. Thus, application domain-specific data from the production process of the automation equipment is not neglected. Thus, a method of diagnosing hardware / software problems is provided even if the root cause is in the production process or automation equipment. Thus, at least a severe delay in troubleshooting is reduced because information about operational technology (OT) equipment can not need to be manually correlated with information about information technology (IT) equipment. Thus, at least unnecessary extended downtime and poor user experience is reduced.
[0076] Further, according to some examples of the present disclosure, the software agent can utilize a large language model (LLM) to perform interrogations on summarized plant OT data, e.g., to generate appropriate and easy-to-understand diagnoses and / or to recommend suggestions for troubleshooting to a user. Further, the software agent can feature a conversational user interface so that a user can interact with the software agent in a conversation to improve diagnoses and recommendations.
[0077] Thus, to diagnose and troubleshoot complex hardware / software problems in a DCS, according to some examples of the present disclosure, a DCS software troubleshooting assistant is provided. A method of combining IT-related data and OT-related data to diagnose hardware / software problems and utilizing a large language model (LLM) to summarize and analyze the data is also provided. When an anomaly is encountered in a DCS, the DCS software troubleshooting assistant can autonomously create an analysis of the situation and can ask a user interrogatories for a hierarchical diagnosis, recommendation, or solution alternative. The user can enter into a conversation with the DCS software troubleshooting assistant to improve the situation analysis and suggested solution steps.
[0078] The DCS software troubleshooting assistant can be a tailored observability and troubleshooting system for a DCS. It can include an anomaly detector, a diagnostic intelligence retriever, a conversational user interface (UI), and a troubleshooting session saver. The system can interact with a metrics and logs database and a plant historian and a vector database with contextual information. To analyze a given anomaly and find a solution alternative, i.e., a suggestion, the DCS software troubleshooting assistant can iteratively interrogate an LLM, which can convert augmented textual prompts into textual and graphical analyses and can provide a hierarchical list of solution alternatives. With this system, a user can troubleshoot hardware / software problems faster and with higher quality. In many cases, the user does not need to interrogate many databases for diagnostic information but can solve the problem in an assisted manner in a streamlined user interface. This user interface can manage the complexity of the underlying IT infrastructure for the user.
[0079] According to some examples of the present disclosure, a process and coordination of IT related data for troubleshooting can be provided. A DCS specific context can also be provided to the LLM for timely and accurate troubleshooting assistant. Furthermore, a continuous learning approach can be implemented by saving past alarm solutions and considering them when handling newly occurring faults. Moreover, it is made possible to customize the features and thresholds to be monitored based on critical levels of components of the DCS. Thus, monitoring can be adapted individually for a single component.
[0080] According to some examples of the present disclosure, the purpose of the proposed DCS troubleshooting method disclosed herein can be to support a DCS user in solving software problems and hardware problems that can affect the functionality and / or performance of the DCS. For example, if a software component fails, the user can use the system, i.e. the DCS, to analyze log data written by the component before the crash and the proposed DCS software troubleshooting assistant. The DCS software troubleshooting assistant supports the analysis by referring to logs, metrics and traces available from the system, by summarizing process data and by referring to previous similar troubleshooting sessions. The DCS software troubleshooting assistant can suggest different alternatives to solve the fault, e.g. restart the component, reconfigure the component, reconfigure the component or update the component, and can provide step-by-step instructions to the user. The user can still decide which alternative of the solution approach to execute, possibly assisted by the system with a predicted probability of success and / or effort estimate. The predicted probability of success and / or effort estimate can be provided by the DCS software troubleshooting assistant.
[0081] As another example, a server node in a DCS cluster can fail due to a hardware problem. The DCS system will automatically try to restart the faulty component on other server nodes in the cluster. But in case the cluster capacity is not sufficient to restart all components, the troubleshooting assistant can ask the user to select non-critical components not to restart to keep critical parts of the system running. The DCS software troubleshooting assistant utilizes the LLM to summarize historical plant data, e.g. time series data of sensor values, and selects similar previous troubleshooting sessions from the database and their found solutions.
[0082] Reference is now made to Figure 1 , Figure 1 The system 100 is schematically illustrated, which comprises a DCS software troubleshooting assistant 110. Figure 1 The illustrated system 100 further comprises databases 120, 130, 150, 170, a DCS 140 and a LLM 160. In more detail, Figure 1The components and context view of a DCS software troubleshooting assistant 110 according to some examples of this disclosure are shown. The DCS software troubleshooting assistant 110 may include an anomaly detector 112, an anomaly configurator 111, a diagnostic smart retrieval unit 113, a troubleshooting session saver 115, and a session user interface 114 for interaction with a user 180. Each component can be implemented as several interactive software processes. The entire DCS software troubleshooting assistant software can run as a continuously running software agent within a DCS cluster. The responsibilities of the system components are as follows:
[0083] Anomaly Detector 112: Anomaly Detector 112 can run continuously (see...) Figure 1 Step 1) of the process can monitor the software service log database 120 and the hardware / software metrics database 130. Logs may contain structured log data with timestamps and events or simple text messages generated by various software components in the DCS. This data may need to be stored in an appropriate database, such as a document-oriented database, to allow for rapid querying. Metrics provide telemetry data on hardware and software, such as CPU utilization, memory consumption, network traffic, execution time, latency, etc. Metrics in IT systems are typically stored in efficient time-series databases to save storage space and support efficient querying. Anomaly detector 112 can be configured to detect outliers or unexpected patterns in log and metric data. Once detected, anomaly detector 112 can notify the diagnostic smart retriever 113 of a summary of the event that occurred. The implementation of anomaly detector 112 can be based on a log aggregator, such as Kibana or Grafana Loki.
[0084] Anomaly Configurator 111: Anomaly Configurator 111 can define rules for Anomaly Detector 112, which encodes thresholds and patterns for identification in monitored logs and metrics. Therefore, Anomaly Configurator 111 can provide predefined anomaly detection rules to Anomaly Detector 112 for application. For example, a rule might declare that more than 20 log messages from a specific component within one minute indicates an anomaly. Another rule might declare that more than 90% CPU load for more than 5 minutes is an anomaly. Yet another rule might declare that the word "error" appearing in certain log messages indicates an anomaly. Many of these rules can be generic and reused across different DCS appliances. However, there may be plant-specific rules that need to be added for each project. For example, if a production plant uses a batch management application in a specific way, custom rules regarding anomalies around the batch management component must be defined. The implementation of Anomaly Configurator 111 can be based on a key-value store or directly embedded into Kubernetes as a custom configuration map.
[0085] Diagnostic intelligence retriever 113: The responsibility of the diagnostic intelligence retriever 113 can be to formulate prompts to ask the LLM 160 using available sources of information, improve the quality of the LLM output by iteratively refining the prompts, and then pass the output results to the conversational user interface 114. The diagnostic intelligence retriever 113 can obtain a summary of the occurring anomaly from the anomaly detector 112 and can use the summary to perform information retrieval on the plant history database 150 and the troubleshooting session database 170. Both databases can be extended with an embedded vector database to store text embeddings of their data. Text embeddings are encoding of text data into floating point vectors. These can allow performing efficient similarity searches on the anomaly summary with the text embeddings. With the information retrieved from these databases, the diagnostic intelligence retriever 113 can augment the prompts of the LLM 160 that request a diagnosis of the occurring anomaly and possible remedies. Because the prompts can be augmented with plant-specific data, the LLM 160 can provide more precise and appropriate answers for the context. The implementation of the diagnostic intelligence retriever 113 can be based on the LangChain framework.
[0086] Conversational user interface 114: The output of the LLM 160 can be an explanation text of the anomaly and is passed by the diagnostic intelligence retriever 113 to the conversational user interface 114. The conversational user interface 114 informs the user and displays the obtained information in an easy-to-process way. The conversational user interface 114 can for example simply present the information as an explanation text, show a line plot of the threshold violation, or even overlay on a topology map of the DCS cluster to enable a quick user understanding and resolution strategy of the notification. The output can have provided possible anomaly resolution alternatives. The conversational user interface 114 can prompt the user 180 to obtain more information, for example information not represented in the system, or even simply select one of the suggested solution alternatives or recommended alternatives. Some solution alternatives can be executed automatically, for example by running a script or issuing a command to the Kubernetes API, but this can not be the responsibility of the DCS software troubleshooting assistant described here. The user interaction with the conversational user interface 114 can be logged, for example in JSON format, for later reuse. The implementation of the conversational user interface 114 can be based on Streamlit.
[0087] Troubleshooting session saver 115: The troubleshooting session saver 115 can create text embeddings of user sessions, e.g., a series of questions and answers. The embeddings can also be floating point vectors that allow for similarity searchers to be effective in the future. These embeddings can be stored in a troubleshooting session database 170, which can be, for example, a vector database. Possibly, the contents of these databases can even be shared between different DCS systems, so that users in different systems can learn from the experiences of other users. However, privacy requirements need to be taken into account, so the shared troubleshooting session database 170 should feature some form of anonymization and obfuscation of details that can be sensitive intellectual property of the production plant.
[0088] Reference is now made to Figure 2 , Figure 2 Different software technologies, components, and databases are depicted that can be used in the system 100 to implement the DCS software fault diagnostic assistant 110 according to some examples of the present disclosure. The present disclosure itself is high-level and can not rely on these technologies according to several examples of the present disclosure. Each of the cited technologies needs to be configured and / or implemented and cannot be used as-is. However, these technologies can provide a basis for quickly creating possible implementations. These technologies are relevant at the time of writing and emphasize the technology of the present disclosure. In the future, they can be replaced by more advanced counterparts, while keeping the core idea of the present disclosure intact and effective.
[0089] It should be noted that Figure 2 Only possible implementation technologies are shown, not mandatory implementation technologies. These shown implementation technologies are merely examples and should not be understood to limit the present disclosure to these technologies. Rather, the present disclosure covers any (now and / or future available) possible implementation technology for implementing at least part of the system 100 according to Figure 1 .
[0090] Reference is now made to Figure 3 , Figure 3 A possible control flow of the method through the DCS software fault diagnostic assistant according to some examples of the present disclosure is shown. Figure 3 Twelve steps are shown, steps 1 to 12, where these steps are also shown in Figure 1 and 2 to more easily match the method steps with the system.
[0091] Step 1 of the method can have been executed before any event during the start of the production process. The anomaly detector 112 can continuously monitor the logs and metrics generated by the system. For the purpose of illustration, the following can be considered. A leak in a tank in plant area C has caused a flood of alarms in the DCS, as the reduced pressure in the tank has a cascading effect on the feed pumps and heat exchangers, all of which raise many alarms. This not only affects the automation equipment, but also the software service “B” that handles the alarm filtering now crashes due to overload.
[0092] In step 2 of the method, continuing the example of step 1, multiple anomaly rules for high CPU load and high logarithmic rate have been triggered and the anomaly detector 112 has identified the anomaly. At this point, the root cause, i.e. the tank leak, is unknown, as this is not reflected in the well logging data.
[0093] In step 3, the anomaly detector 112 extracts relevant logs and metrics from the available data and sends them to the diagnostic intelligence retriever 113.
[0094] In step 4, the diagnostic intelligence retriever 113 builds general troubleshooting hints for the LLM 160 and then performs a similarity search on the latest plant history data using the data from the anomaly detector 112. Because the anomaly detector 112 reports the time stamp of the event, the similarity search finds sensor readings of the same time stamp in the plant history data and can identify the affected plant area.
[0095] In step 5, the search results already contain a sharp drop in pressure readings in the affected tank.
[0096] In step 6, the diagnostic intelligence retriever 113 therefore enriches the general hints with a summary of the pressure readings and hints the LLM 160 for possible resolutions.
[0097] In step 7, the LLM 160 can now propose possible resolutions to the situation (i.e. diagnoses and / or recommendations) with its large amount of trained domain and event knowledge, as well as technical knowledge about various hardware and software components. For example, the LLM 160 suggests restarting the software service “B” after the tank has been repaired and the method proceeds to step 8. Otherwise, for example in case the LLM 160 can not have obtained any possible or sufficiently reliable resolution (e.g. the reliability, the applicability and / or the success rate do not reach a predetermined minimum reliability, a predetermined minimum applicability and / or a predetermined minimum success rate), the LLM 160 can inform the diagnostic intelligence retriever 113 accordingly and the method returns to step 6.
[0098] In step 8, the diagnostic smart retriever 113 sends the LLM output to the conversational user interface 114, where the LLM output is displayed to the user 180. The conversational user interface 114 can contain graphs, charts and textual explanations for the detailed diagnosis of the situation.
[0099] In step 9, the diagnostic smart retriever 113 can also send the LLM output obtained from the troubleshooting session database 170 via retrieval augmentation to the conversational user interface 114.
[0100] In step 10, the user 180 can now request additional information, e.g. retrieve more detailed log data, and judge the resolution alternatives provided. In this case, once the tank is repaired and the session is finally ended, the user can decide to follow the recommendation to restart service B.
[0101] In step 11, the questions and answers in this user session are extracted, possibly anonymized and filtered for unneeded details, and then, in step 12, stored by the troubleshooting session saver 115 for future reference.
[0102] Reference is now made to Figure 4 , Figure 4 A flowchart representing a method according to some examples of the present disclosure is shown. The method is a method for supporting troubleshooting of software and hardware problems in a DCS associated with automation equipment in an industrial plant.
[0103] The method according to Figure 4 may be applied by the DCS software troubleshooting assistant 110 as described above with reference to Figure 1 and 2 .
[0104] The method starts in S400.
[0105] In S410, the method comprises monitoring data from the DCS and / or monitoring data from the automation equipment.
[0106] In S420, the method comprises detecting an anomaly in the monitored data based on predetermined anomaly detection rules.
[0107] In S430, the method comprises, based on the detection result, performing a similarity search for the detected anomaly against historical anomaly data associated with the DCS and / or the automation equipment.
[0108] In S440, the method comprises, based on the result of the performed similarity search, querying the LLM 160 for a diagnosis and / or a recommendation for troubleshooting the detected anomaly.
[0109] In S450, the method includes: obtaining an output from LLM160 based on the inquiry, wherein the output indicates a diagnosis and / or suggestion for troubleshooting the detected anomaly.
[0110] In S460, the method includes providing output to user 180.
[0111] This method ends at S470.
[0112] According to some examples of this disclosure, a data processing apparatus is provided for supporting troubleshooting of software and hardware problems in a DCS associated with automated equipment in an industrial plant. The data processing apparatus can be configured to, i.e., include components configured to perform... Figure 3 Methods and / or execution Figure 4 The processor of the method. The data processing apparatus can represent and / or can be used as referenced above. Figure 1 and 2 The DCS software troubleshooting assistant 110.
[0113] More specifically, based on various examples, it is configured to execute Figure 3 and / or Figure 4 The data processing apparatus of the method can, for example, refer to Figure 5 , Figure 5 A block diagram schematically illustrating a data processing apparatus 500 according to some examples of the present disclosure is shown. The data processing apparatus 500 includes processing circuitry, processing functions, processing means, or a processor 501, enabling the data processing apparatus 500 to participate in supporting troubleshooting of software and hardware problems in a DCS associated with automated equipment in an industrial plant. The processor 501 may include one or more processing portions or functions, wherein the processing portions or functions may be provided as one or more physical or virtual entities. The data processing apparatus 500 may include one or more communication interfaces 502. The data processing apparatus 500 may also include a memory or storage unit 503 for storing data, programs, and / or instructions to be executed by the processing unit 501. The memory 503 may be internal to the data processing apparatus 500 or external to the data processing apparatus 500, such as at a cloud server. The processor 501 may include one or more portions enabling the data processing apparatus 500 to perform, for example... Figure 3 and / or Figure 4 The method. Based on some examples of this disclosure, refer to those provided as examples. Figure 5 The monitoring section 510 can be configured to... Figure 4 The S410 performs such monitoring, and the detection section 520 can be configured to perform monitoring based on... Figure 4 The S420 performs such detection, and the execution part 530 can be configured to perform detection based on... Figure 4S430, the interrogation part 540 can be configured to perform the interrogation according to Figure 4 S440, the obtaining part 550 can be configured to perform the obtaining according to Figure 4 S450, and the providing part 560 can be configured to perform the providing according to Figure 4 S460.
[0114] For example, the parts of the data processing apparatus 500 can also be implemented by means for performing certain functions. For example, the data processing apparatus 500 can comprise means for performing the method according to Figure 3 and / or Figure 4 S440.
[0115] According to some examples of the present disclosure, there is provided a data processing system for supporting troubleshooting of software and hardware problems in a DCS associated with automation equipment in an industrial plant. The data processing system can comprise a data processing apparatus as described above, configured to perform the method according to Figure 3 and / or to perform the method according to Figure 4 Additionally or alternatively, the data processing system can comprise means for performing the method according to Figure 3 and / or means for performing the method according to Figure 4 The data processing system can represent and / or function as the system 100 as described above with reference to Figure 1 and 2 .
[0116] According to some examples of the present disclosure, there is provided an industrial plant comprising a data processing apparatus as described above and / or a data processing system as described above.
[0117] According to several examples of the present disclosure, there is provided a computer-readable medium comprising instructions which, when executed by a computing system, cause the computing system to perform the method according to Figure 3 and / or to perform the method according to Figure 4 The computer-readable medium can be transitory or non-transitory, volatile or non-volatile.
[0118] According to several examples of the present disclosure, there is provided a computer program product comprising instructions which, when executed by a computing system, enable or cause the computing system to perform the method according to Figure 3 and / or to perform the method according to Figure 4 The computer program product can comprise a computer-readable medium comprising the instructions of the computer program product.
[0119] According to some examples of the present disclosure, there is provided use of a data processing apparatus as described above, and / or of a data processing system as described above, and / or of an industrial plant as described above.
[0120] The method according to Figure 3 and / or Figure 4 may be computer-implemented.
[0121] The optional features of the method according to Figure 3 and / or Figure 4 may form, mutatis mutandis, part of any data processing apparatus, data processing system, industrial factory, computer-readable medium, computer program product and use.
[0122] Any of the units, modules, circuits or methods described herein can be implemented using hardware, software and / or firmware configured to perform any of the operations described herein. The hardware can include one or more processor cores, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc. The software can be implemented as software packages, codes, instructions, instruction sets and / or data recorded on at least one transitory or non-transitory computer-readable storage medium. The firmware can be implemented as code, instructions or instruction sets and / or hard-coded data in a memory device (e.g., a non-volatile memory device).
[0123] If implemented in software, the functions can be stored or transmitted over as one or more instructions or code on a computer-readable medium. Computer- readable media include computer-readable storage media. A computer-readable storage medium can be any available storage media that can be accessed by a computer. By way of example, and not limitation, such computer-readable storage media can comprise FLASH memory, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Disk and disc, as used herein, include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs (BDs), where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0124] Applicant hereby discloses the subject matter of this patent application corresponding to the disclosure of the following concurrently filed patent applications: 15 / ---, ---, --- and ---, each of which is herein incorporated by reference in its entirety.
[0125] It must be noted that embodiments of the application are described with reference to different classes of subject matter. In particular, some examples are described with reference to methods, while other examples are described with reference to devices. However, a person skilled in the art will gather from the descriptions herein that, unless otherwise stated, any combination of features belonging to one class of subject matter are also considered to be disclosed in relation to the other class of subject matter. However, all features can be combined in order to provide synergistic effects that are more than the simple sum of the features.
[0126] While the application has been illustrated and described in detail in the drawings and foregoing description, such illustration and description is to be considered exemplary and not restrictive in character. The application is not limited to the disclosed embodiments. Other variations to the disclosed embodiments can be understood and effected by persons skilled in the art from study of the drawings, the disclosure, and the appended claims.
[0127] The mere fact that measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0128] Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. A method for supporting troubleshooting software and hardware problems in a distributed control system (DCS), said DCS being associated with automated equipment in an industrial plant, the method comprising: Monitoring (S410) data from the DCS and / or monitoring (S410) data from the automated equipment; Anomalies in the monitored data are detected based on predetermined anomaly detection rules (S420); Based on the detection results, for the detected anomalies, a similarity search is performed (S430) on the historical anomaly data associated with the DCS and / or the automation equipment; Based on the results of the similarity search performed, the large language model LLM (160) is queried (S440) for diagnostic and / or suggestions for troubleshooting the detected anomalies; Based on the inquiry, an output (S450) is obtained from the LLM (160), wherein the output indicates a diagnosis and / or suggestion for fault diagnosis of the detected anomaly; and The output described in (S460) is provided to the user (180).
2. The method according to claim 1, The data monitored from the DCS includes: Monitor metrics, logs, and traces of the software and / or hardware components of the servers from the DCS; as well as The monitoring of data from the automated equipment includes: monitoring production process data associated with the production process at the automated equipment.
3. The method of claim 2, wherein monitoring the indicators, logs, and traces comprises: The metrics, logs, and traces of the at least one software component and / or hardware component are monitored based on component-specific predetermined anomaly detection rules specific to at least one of the software components and / or hardware components.
4. The method according to any one of claims 1 to 3, further comprising: First data indicating the first monitoring data is obtained from monitoring the data from the DCS; Second data indicating second monitoring data is obtained from monitoring the data from the automated equipment; as well as Joint analysis is performed on the first data and the second data based on associating at least a portion of the first data with at least a portion of the second data and / or based on associating at least a portion of the second data with at least a portion of the first data. The detection includes detecting anomalies based on the results of the joint analysis.
5. The method according to any one of claims 1 to 4, further comprising: Receive user feedback on the output via the user interface (114) and perform at least one of the following: The LLM(160) is trained or fine-tuned based on the received user feedback, and The received user feedback is provided to the LLM (160) so that the LLM (160) processes the received user feedback in a session between the user (180) and the LLM (160).
6. The method according to any one of claims 1 to 5, further comprising: Store historical data indicating multiple past diagnoses and / or multiple past recommendations and / or multiple user sessions regarding detected anomalies received via the user interface (114); and The LLM(160) is trained based on the stored historical data.
7. The method according to any one of claims 1 to 6, The query includes iteratively querying the LLM(160); and / or The output indicates multiple diagnoses and / or multiple recommendations; and / or The output is obtained by: Obtain a ranking of diagnostic alternatives and / or recommended alternatives.
8. The method of claim 7, wherein the ranking indicates the probability that the correct diagnosis and / or recommended application will be successful, and wherein the ranking is based on training data used to train the LLM(160).
9. The method according to any one of claims 1 to 8, further comprising: Based on the obtained output, and according to predetermined rules for autonomous anomaly troubleshooting, measures are automatically taken to troubleshoot the detected anomalies.
10. A data processing apparatus (110) for supporting troubleshooting software and hardware problems in a distributed control system (DCS) associated with automated equipment in an industrial plant, the data processing apparatus (110) including a processor configured to perform the method according to any one of claims 1 to 9.
11. A data processing system (100) for supporting troubleshooting software and hardware problems in a distributed control system (DCS) associated with automated equipment in an industrial plant, the data processing system (100) comprising a data processing apparatus (110) according to claim 10, and / or the data processing system (100) comprising components for performing the method according to any one of claims 1 to 9.
12. The data processing system (100) according to claim 11, wherein the data processing system (100) includes a DCS software troubleshooting assistant (110), the DCS software troubleshooting assistant (110) comprising: Exception Configurator (111); An anomaly detector (112) is communicatively connected to the anomaly configurator (111); A diagnostic intelligent retrieval unit (113) is communicatively connected to the anomaly detector (112) and the conversational user interface (114); and The conversational user interface (114); in The anomaly detector (112) includes an interface to a database, which includes a software service log database (120) and a hardware and / or software metrics database (130). The diagnostic smart retriever (113) includes interfaces to the factory history database (150) and to the LLM (160); and The conversational user interface (114) includes an interface for communicating with a user (180); And among them The anomaly detector (112) is configured to perform monitoring according to any one of claims 1 to 9 via an interface to the database (120; 130); The anomaly detector (112) is further configured to perform detection according to any one of claims 1 to 9 based on the predetermined anomaly detection rules obtained from the anomaly configurator (111); The diagnostic smart retrieval unit (113) is configured to perform a similarity search according to any one of claims 1 to 9 by using the interface to the factory history database (150); The diagnostic smart retrieval unit (113) is also configured to perform an interrogation according to any one of claims 1 to 9 via an interface to the LLM (160); and The conversational user interface (114) is configured to provide services to the user via the interface (180) according to any one of claims 1 to 9.
13. The data processing system (100) according to claim 12, wherein the DCS software troubleshooting assistant (110) further includes: A troubleshooting session saver (115) is communicatively connected to the diagnostic smart retrieval unit (113) and the session user interface (114); The troubleshooting session saver (115) includes an interface to a troubleshooting session database (170); And the troubleshooting session saver (115) is configured to store user feedback received via the session user interface (114) and perform the storage as described in claim 6.
14. A computer-readable medium comprising instructions that, when executed by a computing system, cause the computing system to perform the method according to any one of claims 1 to 9.
15. A computer program product comprising instructions that, when executed by a computing system, enable and / or cause the computing system to perform the method according to any one of claims 1 to 9.