Intelligent exception handling in process control and other systems
Patent Information
- Application Number
- CN202480088943.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2026-09-25
AI Technical Summary
然而,当前的技术对于非数据科学专家而言不易使用,并且无法识别所检测到的异常的原因或解决方案
[0008]在着手进行下文的具体实施方式之前,阐述本专利文件通篇使用的某些词语或短语的定义可能是有利的:术语“包括(include)”和“包含(comprise)”及其派生词意指无限制的包含;术语“或(or)”是包含性的,意指和/或;短语“与……相关联(associated with)”和“与之相关联(associated therewith)”及其派生词能够是指包括、被包括在……内、与……互连、包含、被包含在……内、连接到或与……连接、耦接到或与……耦接、可与……通信、与……协作、交织、并置、接近、绑定到或与……绑定、具有、具有……的性质等;并且术语“控制器(controller)”是指控制至少一个操作的任何设备、系统或其部分,无论此类设备是以硬件、固件、软件还是其至少两种的某种组合来实现。应当注意,与任何特定控制器相关联的功能能够是集中式或分布式的,无论是本地还是远程。本专利文件通篇提供了某些词语和短语的定义,并且本领域普通技术人员将理解,此类定义在许多(如果不是大多数)情况下适用于此类所定义词语和短语的先前以及未来的使用。虽然一些术语能够包括各种各样的实施方式,但所附权利要求能够将这些术语明确限制到特定实施方式。
Smart Images

Figure CN122826558A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to using artificial intelligence and machine learning techniques to detect anomalies in process control systems. Background Technology
[0002] Artificial intelligence (AI) and machine learning (ML) capabilities are increasingly being utilized across various industries, with anomaly detection being a key application area. Examples of such anomaly detection include quality assurance in manufacturing, building management, energy management, cybersecurity, and other process control systems. However, current technologies are not easily used by non-data science experts and cannot identify the causes or solutions for detected anomalies. Improved systems are desired. Summary of the Invention
[0003] Various disclosed implementations include methods for identifying and handling anomalies in AI and ML models, as well as corresponding systems and computer-readable media. One method includes receiving evidence data, including model metric data, historical application domain data, and causal chain graphs. The method includes identifying anomalies in the AI or ML model. The method includes identifying the application domain context corresponding to the identified anomaly, and assembling a causal chain corresponding to the identified anomaly and the identified application domain context. The method includes outputting the causal chain, including a response action.
[0004] Various implementations also include the computer system performing the response action. In various implementations, the anomaly is identified based on the model metric data. In various implementations, the anomaly is identified by comparing the model metric data to one or more thresholds to determine whether the model metric data exceeds one or more thresholds. In various implementations, the application domain context is identified based on historical application domain data.
[0005] In various implementations, the causal chain includes the identified anomaly, the root cause corresponding to the anomaly, and the response action. In various implementations, the output is generated to the presentation layer of the computer system for display to the user. In various implementations, the output is stored in a non-transitory medium. In various implementations, the method is performed by an interpretability layer implemented by the computer system.
[0006] Various embodiments include a computer system having a processor and accessible memory, configured to perform the processes as described herein. Various embodiments also include a non-transitory computer-readable medium encoded with executable instructions that, when executed by one or more computers, cause those computers to perform the processes as described herein.
[0007] The features and technical advantages of this disclosure have been summarized quite extensively above to enable those skilled in the art to better understand the following specific embodiments. Additional features and advantages of this disclosure will be described below, forming the subject matter of the claims. Those skilled in the art will understand that they can readily use the disclosed concepts and specific embodiments as the basis for modifying or designing other structures for carrying out the same purposes of this disclosure. Those skilled in the art will also recognize that such equivalent constructions do not depart from the spirit and scope of this disclosure in its broadest form.
[0008] Before proceeding with the specific implementation described below, it may be advantageous to define certain words or phrases used throughout this patent document: the terms “include” and “comprise” and their derivatives mean unlimited inclusion; the term “or” is inclusive, meaning and / or; the phrases “associated with” and “associated therewith” and their derivatives can mean including, being included in, interconnected with, containing, being contained within, connected to or connected to, coupled to or coupled to, able to communicate with, cooperate with, intertwine, juxtapose, approach, bind to or bind to, have, have the properties of, etc.; and the term “controller” means any device, system or part thereof that controls at least one operation, whether such device is implemented in hardware, firmware, software or some combination of at least two of them. It should be noted that the functionality associated with any particular controller can be centralized or distributed, whether local or remote. This patent document provides definitions for certain words and phrases throughout, and those skilled in the art will understand that such definitions apply in many (if not most) cases to the prior and future uses of such defined words and phrases. While some terms can encompass a wide variety of embodiments, the appended claims are intended to explicitly limit these terms to specific embodiments. Attached Figure Description
[0009] To gain a more complete understanding of this disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, wherein the same numerals denote the same objects, and in the drawings: Figure 1 A block diagram of a computer system in which implementation methods can be carried out is shown; Figure 2 Examples of layers of an AI / ML system implemented in one or more computer systems according to the disclosed embodiments are shown; Figure 3 An example of a causal chain according to the disclosed implementation is shown; Figure 4An example of an interpretability layer according to the disclosed implementation is shown; Figure 5 A flowchart depicting the process according to the disclosed implementation method; Figure 6 An example of a measurement threshold table according to the disclosed implementation is shown; Figure 7 An example of an application domain context table according to the disclosed implementation is shown; and Figures 8 to 13 Exemplary output of an interpretability layer or computer system according to the disclosed implementation is shown. Detailed Implementation
[0010] The following discussion Figures 1 to 13 The various embodiments used to describe the principles of this disclosure are illustrative only and should not be construed as limiting the scope of this disclosure in any way. Those skilled in the art will understand that the principles of this disclosure can be implemented in any suitably arranged device. Numerous innovative teachings of this application will be described with reference to exemplary, non-limiting embodiments.
[0011] As mentioned above, AI and ML tools can be useful in identifying anomalies in the operation or output of many process control systems and other systems. Anomalies detected by AI and machine learning models are often quantified and visualized using various methods for more effective interpretation. These can include metrics displayed as time series plots, which allow for easy identification of trends and patterns within a given time period. Alternatively, significant anomalies can be isolated and represented as individual data points, thus focusing on a single event of interest.
[0012] However, such representations often require a data science background to be fully understood, creating a barrier for operators who need to respond to these detected anomalies but lack the necessary expertise. Most current operators are trained in their application domains (e.g., manufacturing, construction management, energy management, cybersecurity), not AI. Therefore, the displayed AI model metrics are ineffective for current operators, and effective analysis requires companies to hire data scientists. This option is not feasible due to the limited availability and high salaries of data scientists in the job market.
[0013] Another problem with the current system is that the detected anomalies, as measured by the AI model, do not reveal their root cause. The detected anomalies are merely an indication. Yet another problem is that the detected anomalies, as measured by the AI model, do not reveal how to fix them (preventive or corrective actions, referred to here as "response actions").
[0014] The disclosed implementation overcomes the following and other problems: detected anomalies displayed by opaque AI model metrics, detected anomalies whose root causes are not identified, and detected anomalies whose response actions are not identified.
[0015] The disclosed implementation includes a more intuitive and user-friendly visualization of anomalies detected by AI, which can be easily interpreted across different levels of expertise.
[0016] Under normal circumstances, data scientists are responsible for detecting and handling anomalies. These include situations where the ML / AI model is offline and situations where the ML / AI model is online. When the model is offline, data scientists can analyze model metrics during model training, testing, and execution, and identify deviations if any exist. When the model is online, model metrics are continuously reported, and data scientists have access to these metrics. If a deviation occurs, the model metrics become available to data scientists, who can then identify the deviation. In both cases, data scientists need to communicate with operators in the application domain to identify the root cause of the deviation and the appropriate response. Today, it can take up to several days to identify the root cause and initiate a response after an anomaly has been detected.
[0017] To illustrate the exemplary hardware environment, Figure 1 A block diagram of a computer system in which implementation methods can be carried out is shown, for example as a computer system specifically configured by software or other means to perform the processes described herein, and particularly as each of the plurality of interconnect and communication systems described herein. The depicted computer system includes a processor 102 connected to a secondary cache / bridge 104, which in turn is connected to a local system bus 106. The local system bus 106 can be, for example, a peripheral component interconnect (PCI) architecture bus. In the depicted example, main memory 108 and a graphics adapter 110 are also connected to the local system bus. The graphics adapter 110 can be connected to a display 111.
[0018] Other peripheral devices, such as LAN / WAN / wireless (e.g., WiFi) adapter 112, can also be connected to the local system bus 106. An expansion bus interface 114 connects the local system bus 106 to the input / output (I / O) bus 116. The I / O bus 116 connects to a keyboard / mouse adapter 118, a disk controller 120, and an I / O adapter 122. The disk controller 120 can be connected to a storage space 126, which can be any suitable machine-usable or machine-readable storage medium, including but not limited to non-volatile hard-coded media such as read-only memory (ROM) or erasable electrically programmable read-only memory (EEPROM), magnetic tape storage devices, and user-recordable media such as floppy disks, hard disk drives, and optical disc read-only memory (CD-ROM) or digital versatile discs (DVDs), as well as other known optical, electrical, or magnetic storage devices.
[0019] Storage space 126 is capable of storing any data necessary or useful for performing the processes described herein, including executable code 150, AI / ML models 152, model metric data 154, historical application domain data 156, causal chain graphs 158, causal chains 160, anomalies 162, root causes 164, response actions 166, functional components 168, evidence data 170, tables 172, application domain context 174, and other data 176.
[0020] In the example shown, audio adapter 124 is also connected to I / O bus 116, and a speaker (not shown) can be connected to the audio adapter to play sound. Keyboard / mouse adapter 118 provides connectivity for pointing devices (not shown), such as a mouse, trackball, track indicator, touchscreen, etc.
[0021] Those skilled in the art will understand that Figure 1 The hardware depicted may vary depending on the specific implementation. For example, other peripheral devices, such as optical disc drives, may be used as supplements or alternatives to the depicted hardware. The examples depicted are provided for illustrative purposes only and are not intended to impose architectural limitations on this disclosure.
[0022] A computer system according to embodiments of this disclosure includes an operating system employing a graphical user interface. The operating system allows multiple display windows to be simultaneously presented in the graphical user interface, each display window providing an interface to different applications or different instances of the same application. A cursor in the graphical user interface can be manipulated by a user using a pointing device. The cursor position can be changed, and / or events (e.g., clicking a mouse button) can be generated to trigger a desired response.
[0023] It is possible to adopt and appropriately modify one of various commercial operating systems, such as a version of Microsoft Windows™, a product of Microsoft Corporation located in Redmond, Washington. This operating system can be modified or created in accordance with the description in this disclosure.
[0024] The LAN / WAN / wireless adapter 112 is capable of connecting to a network 130 (not part of the computer system 100), which can be any public or private computer network or combination of networks known to those skilled in the art, including the Internet. The computer system 100 is capable of communicating with a server system 140 via the network 130, which is also not part of the computer system 100, but can be implemented, for example, as a separate computer system 100.
[0025] The disclosed implementation overcomes the technical shortcomings of current methods by implementing an "interpretability layer" capable of identifying and assembling elements of a causal chain for detected anomalies. The interpretability layer generates the causal chain, which is then displayed to the operator by the presentation layer.
[0026] Figure 2 An example of the layers of an AI / ML system 200 implemented in one or more computer systems 100 is shown to illustrate the interpretability architecture disclosed herein. The AI / ML system 200 includes an evidence layer 208, which can include elements such as model metric data 210, historical application domain data 212, and a causal chain graph 214. Each of the model metric data 210, historical application domain data 212, and causal chain graph 214, as well as other elements of the evidence layer 208, can communicate with the interpretability layer 204, which is described in more detail herein. The interpretability layer 204 can then send its output to a presentation layer 202 to present the output 208, such as detected anomalies and their causal chains and response actions.
[0027] The causal chain graph 214 can be implemented as a knowledge graph or other suitable knowledge base or data structure. For example, the causal chain graph 214 can be implemented as a directed graph that associates a specific anomaly or anomaly category with potential causes or other conditions that may produce the anomaly, or with indicators of potential problems that will identify the causes of the anomaly. The causal chain graph 214 can be a weighted graph, wherein the specific association between anomalies and potential causes is weighted based on, for example, the probability that a specific potential cause is an actual cause of the anomaly. The causal chain graph 214 can include response actions associated with a specific anomaly or anomaly category and a specific cause, wherein the response actions are those actions determined to address the corresponding cause to eliminate or prevent the anomaly.
[0028] Figure 3An example of a causal chain 300 according to the disclosed implementation is shown. The causal chain 300 includes an anomaly 302 (also referred to as a sign), a root cause 304, and a response action 306 that addresses the root cause and eliminates the anomaly 302. The assembled causal chain 300 can include one or more root causes 304 for an anomaly, and for each root cause 304, can include one or more response actions 306. Each root cause and response action is associated with a confidence level.
[0029] Figure 4 An example of an interpretability layer 204 according to the disclosed implementation is shown. The interpretability layer 400 includes multiple functional components implemented via executable instructions as part of an AI / ML system 200 implemented in one or more computer systems 100, which is based on... Figure 2 The information in evidence layer 206 is used for operation. Functional component 402 assembles the causal chain. Functional component 404 identifies the domain application context. Functional components 402 and 404 together can identify anomalies in AI applications, ML models, or the like. This can include, for example, identifying outliers, data points, or events in model metric data. In particular, this can include identifying whether a model metric exceeds a threshold within a given time period. In this context, "exceeding" a threshold should be understood as meaning a value greater than an upper threshold or less than a lower threshold.
[0030] Functional component 406 identifies the root cause of anomalies, which can be based on the identified anomaly, historical application data, model metric data, and causal chain diagrams. Functional component 408 identifies response actions. Functional component 410 constructs a causal chain model.
[0031] Figure 5 A flowchart of process 500 according to the disclosed implementation is shown, which can be executed, for example, by one or more computer systems 100 (hereinafter referred to as "systems") as disclosed herein, which implement interpretability layer 204 in ML / AI system 200. In process 500, the system assembles causal chains from detected anomalies and historical application domain data.
[0032] At point 502, the system receives evidence data. "Receiving," as used herein, can include loading from a storage device, receiving from another device or process, receiving via interaction from a user, or otherwise. Specifically, in this context, this can include the interpretability layer 204 receiving evidence data from the evidence layer 208.
[0033] Receiving evidence data can include receiving model measurement data 210, receiving historical application domain data 212, and receiving causal chain graphs 214. In particular, receiving evidence data does not necessarily have to be performed as a preparatory step for the rest of process 500, but can be performed concurrently with each of the other actions of process 500 or as needed.
[0034] At point 504, the system identifies anomalies in the AI or ML model. This can be performed, for example, based on model metric data 210.
[0035] In one exemplary process for identifying anomalies, the system frequently processes model metric data 210, such as every second, every 10 seconds, every 60 seconds, or every 3600 seconds. The system compares the model metric to thresholds (lower and / or upper thresholds) and determines whether the model metric exceeds one of these thresholds. If a component has identified that a model metric exceeds a threshold, that component can determine whether the model metric has exceeded or will exceed the model threshold within a defined minimum time period. If the model metric has exceeded or will exceed the threshold within at least the defined minimum time period, the model metric is identified as an anomaly.
[0036] Figure 6 An example of a metric threshold table 600 is shown, which identifies a specific model metric 602 and the corresponding lower threshold 604 (where appropriate), upper threshold 606 (where appropriate), and shortest duration 608.
[0037] A 504 output indicates the identified exception.
[0038] At point 506, the system identifies the application domain context corresponding to the identified anomaly. This can be performed using historical application domain data.
[0039] In one exemplary process for identifying application domain context, the system is able to identify a model ID (or an equivalent identifier for an AI / ML model) and read the application domain context of the identified model ID. This can be performed using an application domain context table, which can be part of model metric data and / or historical application domain data.
[0040] Figure 7 An example of an application domain context table 700 is shown, which identifies a specific AI / ML model 702 and the corresponding context for each model 702, such as production line 704 (where appropriate), workstation 706, material dependencies 708, and process dependencies 710. Of course, the nature and implementation of a particular AI / ML model will determine the corresponding context for its implementation.
[0041] The output of 506 is the identified application domain context.
[0042] At point 508, the system assembles one or more causal chains corresponding to the identified anomalies and the identified application domain context. This can be performed using historical application domain data 212 and the causal chain graph 214.
[0043] In one exemplary process for assembling a causal chain, the system can initiate a new causal chain instance CC.1. The system adds the identified anomaly to causal chain instance CC.1 and the identified application domain context to causal chain instance CC.1. The system identifies the root cause of the identified anomaly in the causal chain graph with a confidence level above a predetermined threshold. The number of root causes can be finite.
[0044] In this exemplary process, the system adds the identified root cause to the causal chain instance CC.1 and searches the causal chain graph for the response action of the identified root cause with a confidence level above a predetermined threshold. The number of response actions for each root cause can be finite.
[0045] In this exemplary process, the system adds the response action for each identified root cause to the causal chain CC.1.
[0046] The output of 508 is the assembled causal chain, such as causal chain 300 which includes exception 302, root cause 304, and response action 306.
[0047] At point 510, the system generates an output including a causal chain or components thereof. This output can be generated, for example, to a presentation layer that can display the detected anomaly, its causal chain, and the response action. The output can be stored in a tangible or non-transitory medium, and can be sent to another device process. In some cases, where the response action can be automated, the system can also automatically execute the response action at or after point 510 to address the root cause.
[0048] Figure 8 An exemplary output of an overview of reported events is shown in accordance with the disclosed implementation of the interpretability layer (such as that presented by the presentation layer).
[0049] Figure 9 Exemplary outputs showing indications and other features of an interpretability layer (such as that presented by a rendering layer) according to the disclosed implementation are shown.
[0050] Figure 10 An exemplary output of the application domain context (in this case, location) and other features of the interpretability layer (such as that presented by the presentation layer) according to the disclosed implementation is shown.
[0051] Figure 11Exemplary outputs of the root cause and other features of an interpretability layer (such as that presented by a rendering layer) according to the disclosed implementation are shown.
[0052] Figure 12 Exemplary outputs of immediate response actions and other features of an interpretability layer (such as one rendered by a presentation layer) according to the disclosed implementation are shown.
[0053] Figure 13 Exemplary outputs of responsive actions and other features of an interpretability layer (such as that rendered by a presentation layer) according to the disclosed implementation are shown.
[0054] Using the techniques and processes disclosed herein, computer systems can display anomalies, root causes, and response actions that are understandable to operators without a data science background. This enables operators to select and initiate response actions within minutes rather than days.
[0055] In various implementations, the identified response action is machine-readable and machine-executable.
[0056] The system disclosed herein can provide operators with or display specific content and controls for abnormal situations.
[0057] Of course, those skilled in the art will recognize that, unless the order of operations is specifically indicated or required, certain steps in the above process can be omitted, performed concurrently or sequentially, or performed in a different order.
[0058] Those skilled in the art will recognize that, for the sake of brevity and clarity, not all the complete structure and operation of a computer system suitable for use with this disclosure are depicted or described herein. Rather, only those parts of the computer system that are unique to this disclosure or necessary for understanding this disclosure are depicted and described. The remainder of the construction and operation of the computer system 100 can conform to various current implementations and practices known in the art.
[0059] It is important to note that although this disclosure is described in the context of a fully functional system, those skilled in the art will understand that at least part of the mechanisms of this disclosure can be distributed in any of a variety of forms as instructions contained in a machine-usable, computer-usable, or computer-readable medium, and this disclosure applies equally regardless of the specific type of instruction or signal-bearing medium or storage medium actually used to perform such distribution. Examples of machine-usable / readable or computer-usable / readable media include: non-volatile hard-coded type media, such as read-only memory (ROM) or erasable electrically programmable read-only memory (EEPROM), and user-recordable type media, such as floppy disks, hard disk drives, and optical disc read-only memory (CD-ROM) or digital versatile discs (DVDs).
[0060] Although exemplary embodiments of this disclosure have been described in detail, those skilled in the art will understand that various changes, substitutions, modifications and improvements can be made to the content disclosed herein without departing from the spirit and scope of the broadest form of this disclosure.
[0061] Nothing described in this application should be construed as implying that any particular element, step, or function is a necessary element that must be included within the scope of the claims: the scope of the patent subject matter is defined only by the permitted claims. Furthermore, none of these claims are intended to invoke 35 USC §112(f) unless the exact word “means for” is followed by a participle. The use of terms such as (but not limited to) “mechanism,” “module,” “device,” “unit,” “component,” “element,” “member,” “apparatus,” “machine,” “system,” “processor,” or “controller” in the claims should be understood as and intended to refer to structures known to a person skilled in the art, as further modified or enhanced by the features of the claims themselves, and is not intended to invoke 35 USC §112(f).
Claims
1. A method (500) executed by a computer system (100), comprising: Receive (502) evidence data (204), the evidence data (204) including model measurement data (210), historical application domain data (212) and causal chain diagram (214). The computer system (100) identifies (504) anomalies (302) in the model (152), wherein the model (152) is an artificial intelligence (AI) model or a machine learning (ML) model; The computer system (100) identifies (506) the application domain context (174) corresponding to the identified anomaly (302). The computer system (100) assembles (508) a causal chain (300) corresponding to the identified anomaly (302) and the identified application domain context (174); and The computer system (100) outputs (510) the causal chain (300), which includes a response action (306).
2. The method according to claim 1, further comprising the computer system (100) performing the response action (510).
3. The method according to claim 1, wherein, The anomaly (302) is identified based on the model measurement data (210).
4. The method according to claim 1, wherein, The anomaly (302) is identified by comparing the model metric data (210) with one or more thresholds to determine whether the model metric data (210) exceeds one or more thresholds.
5. The method according to claim 1, wherein, The application domain context (174) is identified based on the historical application domain data (212).
6. The method according to claim 1, wherein, The causal chain (300) includes the identified anomaly (302), the root cause (304) corresponding to the anomaly (302), and the response action (306).
7. The method according to claim 1, wherein, The output is generated to the presentation layer (202) of the computer system (100) for display to the user.
8. The method according to claim 1, wherein, The output is stored in a non-transitory medium (126).
9. The method according to claim 1, wherein, The method (500) is executed by an interpretability layer (204) implemented through the computer system (100).
10. A computer system (100), comprising: Processor (102); and Accessible memory (108), the computer system (100) is specifically configured to perform the method (500) according to any one of claims 1 to 9.
11. A non-transitory computer-readable medium (126) encoded with executable instructions, which, when executed by one or more computers (100), cause the one or more computers (100) to perform the method (500) according to any one of claims 1 to 9.