Providing alerts related to exception scores assigned to input data methods and systems
By receiving input data and using multiple anomaly detection models to automatically generate alarms, the problem of complex anomaly detection requiring manual intervention in existing technologies is solved, thereby improving the efficiency and accuracy of equipment analysis and monitoring.
Patent Information
- Application Number
- CN202180046563.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-30
- Filing Date
- 2021-06-30
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-06-30
AI Technical Summary
Existing computer software products, when using anomaly detection models to analyze, monitor, and control equipment, struggle to efficiently provide alerts related to anomaly scores in the input data, and require significant manual intervention and complex configuration processes.
By receiving input data, multiple anomaly detection models are used to determine anomaly scores, and alarms are provided when the differences exceed a threshold. This process is automated using computing units and interface systems.
It enables automated anomaly detection and alarm generation, reducing manual intervention and improving the efficiency and accuracy of equipment analysis, monitoring and control.
Smart Images

Figure CN115867873B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates generally to software management systems, and in particular to a system for providing alerts related to anomaly scores assigned to input data, such as detecting distribution drift of incoming data using anomaly detection models (collectively referred to herein as product systems). BACKGROUND
[0002] Recently, an increasing number of computer software products involving artificial intelligence, machine learning, etc. are used for performing various tasks. Such computer software products can for example be used for speech, image or pattern recognition purposes. Further, such computer software products can be used for analyzing, monitoring, operating and / or controlling devices, e.g. in an industrial environment, directly or indirectly, e.g. by embedding such computer software products in more complex computer software products. The present invention generally relates to computer software products providing alerts and to the management and e.g. updating of such computer software products.
[0003] Currently, there are product systems and solutions supporting the use of anomaly detection models for analyzing, monitoring, operating and / or controlling devices and supporting the management of such computer software products involving anomaly scores. Such product systems can benefit from improvements. SUMMARY
[0004] Various disclosed embodiments include methods and computer systems that can be used to help provide alerts related to anomaly scores assigned to input data and to manage computer software products.
[0005] According to a first aspect of the invention, a computer-implemented method can comprise:
[0006] receiving input data related to at least one device, wherein the input data comprises incoming data batches X related to at least N separable classes, where n e 1,..., N;
[0007] determining respective anomaly scores si,..., sn for respective incoming data batches X related to the at least N separable classes using N anomaly detection models Mn;
[0008] applying the (trained) anomaly detection models Mn to the input data to generate output data, which is suitable for analyzing, monitoring, operating and / or controlling the respective device;
[0009] determining a difference between, on the one hand, the determined respective anomaly scores si,..., sn for the at least N separable classes for respective incoming data batches X and, on the other hand, given respective anomaly scores Si,..., Sn of the N anomaly detection models Mn (130) for the at least N separable classes;
[0010] - if the respective determined difference therebetween is greater than the difference threshold, then: providing an alert related to the determined difference to a user, the respective device, and / or an IT system connected to the respective device.
[0011] As an example, the input data can be received with a first interface. Further, the respective anomaly detection model can be applied to the input data with a computing unit. In some examples, an alert related to the anomaly score assigned to the input data can be provided with a second interface.
[0012] According to a second aspect of the present application, a system (e.g. a computer system or an IT system) can be arranged and configured to perform the steps of the computer-implemented method. In particular, the system can comprise:
[0013] - a first interface configured for receiving input data related to at least one device, wherein the input data comprises an incoming data batch X related to at least N separable classes, wherein n e 1,..., N;
[0014] - a computing unit configured for
[0015] - determining, using the N anomaly detection models Mn, a respective anomaly score si,..., sn for the respective incoming data batch X related to the at least N separable classes;
[0016] - applying the anomaly detection model Mn to the input data to generate output data, the output data being suitable for analyzing, monitoring, operating, and / or controlling the respective device;
[0017] - determining, for the respective incoming data batch X, a difference between, on the one hand, the determined respective anomaly scores si,..., sn of the at least N separable classes and, on the other hand, the given respective anomaly scores Si,..., Sn of the N anomaly detection models Mn 130; and
[0018] - a second interface configured for providing an alert related to the determined difference to a user, the respective device, and / or an IT system connected to the respective device, if the respective determined difference therebetween is greater than the difference threshold.
[0019] According to a third aspect of the present application, a computer program can comprise instructions which, when executed by a system (e.g. an IT system), cause the system to perform the described method of providing an alert related to the anomaly score assigned to the input data.
[0020] According to a fourth aspect of the present application, a computer readable medium can comprise instructions which, when executed by a system (e.g. an IT system), cause the system to perform the described method of providing an alert related to an anomaly score assigned to input data. As an example, the described computer readable medium can be non-transitory and can also be a software component on a storage device.
[0021] The foregoing has outlined rather broadly the technical features of the present disclosure in order that the detailed description that follows can be better understood. Additional features and advantages of the present disclosure will be described hereinafter which form the subject of the claims. Those skilled in the art will appreciate that they can readily use the disclosed conception and specific embodiments as a basis for designing or modifying other structures for carrying out the same purposes of the present disclosure. Those skilled in the art will also realize that such equivalent constructions do not depart from the spirit and scope of the present disclosure.
[0022] Furthermore, before proceeding to the following detailed description, it should be understood that certain words and phrases have been used herein with respect to certain aspects and features thereof, which have been presented for purposes of interpretation of the claims. It is therefore intended that the definitions of certain terms and phrases be not limited to the specific embodiments presented herein, but be interpreted broadly under the ordinary and customary meaning of such terms and phrases. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 A functional block diagram of an exemplary system that facilitates providing alerts in a product system is shown.
[0024] Figure 2 Degradation of a trained model over time due to data distribution shift is shown.
[0025] Figure 3 Exemplary data distribution drift detection for a binary classification task is shown.
[0026] Figure 4 An exemplary box plot comparing two distributions of anomaly scores is shown.
[0027] Figure 5 A functional block diagram of an exemplary system that facilitates providing alerts in a product system and managing computer software products is shown.
[0028] Figure 6 A further flowchart of an exemplary method that facilitates providing alerts in a product system is shown.
[0029] Figure 7 An embodiment of an artificial neural network is shown.
[0030] Figure 8 An embodiment of a convolutional neural network is shown.
[0031] Figure 9 A block diagram of a data processing system capable of implementing embodiments is shown. DETAILED DESCRIPTION
[0032] Various techniques related to systems and methods for providing alerts and managing computer software products in a product system will now be described with reference to the drawings, where like reference numerals refer to like elements throughout. The drawings discussed below and the various embodiments used to describe the principles of the present disclosure in this patent document are by way of illustration only and should not be construed in any way to restrict the scope of the disclosure. Those skilled in the art will understand that the principles of the present disclosure can be implemented in any suitably arranged device. It should be understood that functionality described as being performed by certain system elements can be performed by multiple elements. Similarly, for example, an element can be configured to perform functionality described as being performed by multiple elements. Many innovative teachings are presented in this patent document that can be applied to a wide range of applications and industries.
[0033] With reference to Figure 1 , an example computer system or data processing system 100 is shown that facilitates providing alerts 150, in particular providing alerts 150 related to an anomaly score assigned to input data 140, such as detecting a distribution shift of incoming data 140 using an anomaly detection model 130. The processing system 100 can include at least one processor 102 configured to execute at least one application software component 106 from a memory 104 accessed by the processor 102. The application software component 106 can be configured (i.e., programmed) to cause the processor 102 to perform various actions and functions described herein. For example, the described application software component 106 can include and / or correspond to one or more components of an application configured to provide and store output data in a data store 108, such as a database.
[0034] It should be appreciated that providing alerts 150 in complex application and industrial environments can be difficult and time consuming. For example, advanced decoding knowledge of a user or IT specialist can be required or conscious selection of many options can be required, both of which involve many manual steps, which is a lengthy but inefficient process.
[0035] To enable the provision of the enhanced alert 150, the described product system or processing system 100 can comprise at least one input device 110 and, optionally, at least one display device 112, such as a display screen. The described processor 102 can be configured to generate a GUI 114 through the display device 112. Such a GUI 114 can comprise GUI elements, such as buttons, text boxes, images, scroll bars, which can be used by a user to provide inputs that can support the provision of the alert 150 through the input device 110.
[0036] In an exemplary embodiment, the application software component 106 and / or the processor 102 can be configured to receive input data 140 related to at least one device 142, wherein the input data 140 comprises a batch X of incoming data related to at least N separable classes, with n e 1,..., N. Further, the application software component 106 and / or the processor 102 can be configured to determine, using the N anomaly detection models Mn 130, a respective anomaly score si,..., sn for the respective batch X of incoming data related to the at least N separable classes. In some instances, the application software component 106 and / or the processor 102 can be further configured to apply the anomaly detection models Mn 130 to the input data 140 to generate output data 152 suitable for analyzing, monitoring, operating and / or controlling the respective device 142. The application software component 106 and / or the processor 102 can be further configured to determine, for the respective batch X of incoming data, a difference between the determined respective anomaly score si,..., sn of the at least N separable classes on the one hand and a given respective anomaly score Si,..., Sn of the N anomaly detection models Mn 130 on the other hand. Further, the application software component 106 and / or the processor 102 can be configured to provide, in case the respective determined difference therebetween is greater than a difference threshold, an alert 150 related to the determined difference to a user (e.g. via the GUI 114), the respective device 142 and / or an IT system connected to the respective device 142.
[0037] In some instances, the respective anomaly detection models Mn 130 are provided beforehand and stored in the data repository 108.
[0038] The input device 110 and the display device 112 of the processing system 100 can be considered optional. In other words, the subsystem or computing unit 124 comprised in the processing system 100 can correspond to a claimed system, e.g. an IT system, which can comprise one or more appropriately configured processors and memories.
[0039] As an example, the input data 140 can comprise incoming data batches X related to at least N separable categories, with n e 1,..., N. The data batches can for example comprise measured sensor data related to e.g. temperature, pressure, current and voltage, distance, speed or rate, acceleration, flow rate, electromagnetic radiation including visible light, or any other physical quantity. In some examples, the measured sensor data can also be related to chemical quantities such as acidity, concentration of a given substance in a mixture of substances, etc. The respective variables can for example characterize a respective device 142 or a state in which a respective device 142 is in. In some examples, the respective measured sensor data can characterize a processing or production step carried out or monitored by the respective device 142.
[0040] In some examples, the respective device 142 can be or comprise a sensor, an actuator such as a motor, a valve or a robot, and an inverter powering the motor, a gearbox, a programmable logic controller (PLC), a communication gateway, and / or other parts and components typically related to industrial automation products and industrial automation. The respective device 142 can be part of a complex production line or production plant such as a bottling machine, a conveyor, a welding machine, a welding robot, etc. In other examples, there can be input data messages 142 related to one or more variables of a plurality of such devices 142. Further, as an example, the IT system can be or comprise a manufacturing operations management (MOM) system, a manufacturing execution system (MES) and an enterprise resource planning (ERP) system, a supervisory control and data acquisition (SCADA) system, or any combination thereof.
[0041] The input data 140 can be used to generate output data 152 by applying the anomaly detection model Mn 130 to the input data 140. The anomaly detection model Mn 130 can for example interrelate the input data messages or respective variables with the output data 152. The output data 152 can be used to analyze or monitor the respective device 142, for example to indicate whether the respective device 142 is working properly or whether the respective device 142 is monitoring a production step that is working properly. In some instances, the output data 152 can indicate that the respective device 142 is damaged or that there can be a problem with the production step monitored by the respective device 142. In other instances, the output data 152 can be used to operate or control the respective device 142, for example to implement a feedback or control loop using the input data 140, by analyzing the input data messages 140 with the anomaly detection model Mn 130 and controlling or operating the respective device 142 based on the received input data 140. In some instances, the device 142 can be a valve in a process automation plant, wherein the input data messages comprise data about a flow rate as a physical variable, which is then analyzed with the anomaly detection model Mn 130 to generate the output data 152, wherein the output data 152 comprises one or more target parameters for the operation of the valve, for example a target flow rate or a target position of the valve.
[0042] The incoming data batches X of the input data 140 can be related to at least N separable classes. In a rather simple instance, there can be two classes: class 1 and class 2, class 1 indicating that the device 142 or the corresponding production plant is in a “normal” state and class 2 indicating that the device 142 or the corresponding production plant is in an “abnormal” state. For example, the device 142 can correspond to a bearing of a gearbox or a belt conveyor, wherein class 1 can indicate a correct operation of the device 142 and class 2 can indicate that the bearing does not have enough lubricant or that the belt of the belt conveyor will be lost. In general, different N classes can be related to typical scenarios of the monitored device, which in some instances can be a physical object. Thus, the N classes can correspond to a correct operation state of the physical device 142 and N-1 typical failure modes. In some instances, the domain model can separate the “normal” state from the “abnormal” state, wherein there can be sub-classes that specify in more detail in which “abnormal” state the device 142 is.
[0043] It will be appreciated that in some instances the anomaly detection models Mn 130 can be trained anomaly detection models Mn. The training of such trained anomaly detection models Mn can be done for example using a reference dataset or training dataset. The reference dataset can be provided in advance for example by identifying typical scenarios related to the typical variables or input data 140. Such typical scenarios can be for example scenarios when the respective device 142 is working properly, scenarios when the respective device 142 is monitoring a production step that is performed properly, scenarios when the respective device 142 is damaged, scenarios when the respective device 142 is monitoring a production step that is not performed properly, etc. As an example, the device 142 can be a bearing that becomes overheated and thus has increased friction during its operation. Such scenarios can be analyzed or recorded in advance to enable providing corresponding reference data. When receiving corresponding input data 140, the input data 140 can be compared to the reference dataset using the N anomaly detection models Mn 130 to determine a respective anomaly score sn for the respective batch of incoming data X related to at least N separable classes.
[0044] For each new and previously unseen batch X of input data, a descriptive statistic of the anomaly scores sn can be determined and compared to the corresponding descriptive statistic Sn obtained for each model Mn. Thus, the descriptive statistic of the respective anomaly scores sn or Sn can comprise a corresponding median, standard deviation and / or interquartile range of the respective anomaly scores sn or Sn. In this context, in the descriptive statistic, the interquartile range (IQR) (also known as the middle distribution, middle 50% or H distribution) is a measure of statistical dispersion equal to the difference between the 75th percentile and the 25th percentile, or between the upper quartile and the lower quartile, such that IQR = Q3 - Q1. In other words, the IQR is the difference between the third quartile and the first quartile; these quartiles can be clearly seen on a box plot about the data, examples of which are shown in Figure 4 This can be a trimmed estimator defined as a 25% trimming range and is a commonly used robust measure of scale.
[0045] The IQR can be considered as a measure of variability based on dividing a dataset into quartiles. Quartiles divide an ordered dataset into four equal parts. The values that separate these parts are called the first quartile, the second quartile, and the third quartile; and they are denoted by Q1, Q2, and Q3, respectively.
[0046] If the comparison of the abnormality score snwith the abnormality scores S1,...,Snof the N abnormality detection models Mn 130 (or the comparison of the corresponding descriptive statistics on snand Sn) reveals a significant difference, this can be the case if the determined difference is greater than a difference threshold, then a data distribution drift can be detected and a warning can be sent to the user, the respective device 142 and / or the IT system, which can indicate that a data drift has occurred and / or that the abnormality detection model can no longer be trustworthy.
[0047] In some instances, the respective abnormality scores S1,...,Snof the N abnormality detection models Mn 130 can be determined beforehand. As an example, typical scenarios of the monitored device 142 can be used to determine the respective abnormality scores S1,...,Sn, such typical scenarios including normal operating states and typical failure modes of the device 142. If the corresponding input data 140 is received, this can allow to identify such typical scenarios of the respective device 142.
[0048] It is to be understood that in some instances, the determined respective abnormality scores s1,...,snof the incoming data batch X can not fit well to the given respective abnormality scores S1,...,Sn, such that the respective abnormality scores differ from each other and the respective determined difference is greater than the difference threshold. Such a case can occur due to a distribution drift of the input data and can indicate that the used abnormality detection models Mn 130 can no longer be valid for the input data 140 of the respective device 142. In such a case, an alarm 150 is generated and provided to the user, the respective device 142 and / or the IT system connected to the respective device 140.
[0049] As an example, the input data 140 comprises data on several variables and there are n abnormality detection models Mn reflecting n different scenarios, with n > 1, for example one acceptable state scenario and n-1 different damage scenarios.
[0050] Further, the trained abnormality detection models Mn 130 with n > 1 can correspond to supervised learning (SL), i.e. a machine learning task of learning a function that maps inputs to outputs based on example input-output pairs. Such supervised learning infers a function from labeled training data consisting of training examples. In supervised learning, each example is a pair consisting of an input object (typically a vector) and a desired output value (also called the supervisory signal). A supervised learning algorithm analyzes the training data and produces an inferred function that can be used to map new examples. An optimal scenario would allow the algorithm to correctly determine the class label of unseen instances. This requires the learning algorithm to generalize from the training data to unseen situations in a “reasonable” way (see inductive bias).
[0051] The N anomaly detection models Mn 130 (trained or untrained) can then be used to determine respective anomaly scores sn for respective batches of incoming data X related to at least N separable classes. In addition, the N anomaly detection models Mn 130 (trained or untrained) can be applied to the input data 140 to generate output data 152 suitable for analyzing, monitoring, operating and / or controlling the respective device 142. Based on the determined respective anomaly scores sn, by comparing them with the given respective anomaly scores Sn of the anomaly detection models Mn 130, an alarm 150 can be generated and provided to a user, the respective device 142 and / or an IT system connected to the respective device 142. The alarm 150 related to the determined difference can be provided to a user (e.g. monitoring or supervising a production process involving the device 142) such that the user can trigger a further analysis of the device 142 or the related production step. In some instances, the alarm 150 can be provided to the respective device 142 or the IT system, e.g. and the respective device or the IT system can be or comprise a SCADA, MOM or MES system scenario.
[0052] It should also be understood that the determined anomaly scores sn of the anomaly detection models Mn 130 can be interpreted in terms of the trustworthiness of the anomaly detection models Mn 130. In other words, the determined anomaly scores sn can be indicative of whether the anomaly detection models Mn 130 are trustworthy. As an example, the generated alarm 150 can comprise the determined anomaly scores sn or information about the trustworthiness (level) of the anomaly detection models Mn 130.
[0053] In addition, in some instances, the anomaly values with respect to the input data 140 can be allowed such that not every input data 140 can trigger an alarm 150. For example, for a given number z of sequentially incoming data batches X, an alarm 150 can only be provided if the determined difference is greater than a given difference threshold.
[0054] As already mentioned above, Figure 1 The system 100 illustrated in Fig. 1 can correspond to or comprise the computing unit 124. In addition, the computing unit can comprise a first interface 170 for receiving the input data messages 140 related to at least one variable of at least one device 142 and a second interface 172 for providing an alarm 150 related to the determined difference to a user, the respective device 142 and / or an IT system connected to the respective device 142 in case the determined difference is greater than the difference threshold. Depending on to which device or system the alarm 150 is sent, the first interface 170 and the second interface 172 can be the same interface out of all different interfaces. In some instances, the computing unit 124 can comprise the first interface 170 and / or the second interface 172.
[0055] In some instances, the input data 140 undergoes an increased distribution shift involving a determined difference.
[0056] As an example, the input data 140 comprises a variable, wherein for a given time period, the values of the variable oscillate around a given average value. For some reason, at a later time, the values of the variable oscillate around a different average value, such that a distribution shift has occurred. In many instances, the distribution can involve an increase of a determined difference between the abnormality scores snand Sn. As an example, the distribution shift of a variable can occur due to wear, aging or other kinds of deterioration, e.g. for equipment subject to mechanical or stress. The concept of distribution shift leading to an increased difference is explained in more detail below in the context of Figure 2
[0057] In some instances, the proposed method can thus detect an increase of a difference due to a distribution shift of the input data 140.
[0058] It should also be understood that in some instances, the application software component 106 and / or the processor 102 can also be configured to determine a distribution shift of the input data 140 in case a second difference between the abnormality scores s1,..., snof an earlier batch of incoming data Xeand the abnormality scores s1,..., snof a later batch of incoming data X1is greater than a second threshold value; and provide a report related to the determined distribution shift to a user, to the respective equipment 142 and / or to an IT system connected to the respective equipment 142 in case the determined second difference is greater than the second threshold value.
[0059] In these instances, the trend of the input data 140 can be used to identify a distribution shift. To this end, a second difference is determined that takes into account the earlier incoming data batch Xeand the later incoming data batch X1of the input data 140. This second difference is compared to a second threshold to determine whether a report should be provided. For example, the respective anomaly scores s1,..., snof both the earlier incoming data batch Xeand the later incoming data batch X1involve a difference with respect to the given respective anomaly scores S1,..., Snthat is smaller than the difference threshold. However, the second difference can be larger than the second threshold, causing a report to be generated and provided to the user, the respective device 142, and / or the IT system connected to the respective device 142. In some instances, the second threshold can be equal to the difference threshold, and the respective anomaly scores of the earlier incoming data batch Xeand the later incoming data batch X1may constitute an acceptable deviation at the upper and lower bounds of the difference threshold, but the second difference can still be larger than the second threshold. In this case, this can occur when a dynamic change occurs at the respective device 142, such as a complete failure or breakage of a certain electrical or mechanical component of the respective device 142. As an example, a number of earlier incoming data batches Xeand a number of later incoming data batches X1may be considered, such that a single occurrence of an outlier can be categorized and does not result in the generation and provision of a report. In other instances, the report can correspond to the alert 150 mentioned above. Additionally, in other instances, the anomaly scores s1,..., snof the earlier incoming data batch Xemay correspond to the given anomaly scores S1,..., Sn, which can allow for a more dynamic process of generating the alert 150.
[0060] It should also be understood that, in some instances, the application software component 106 and / or the processor 102 can also be configured to assign the training data batch Xtto at least N separable classes of the anomaly detection model Mn 130 and determine the given anomaly scores S1,..., Snfor the at least N separable classes of the N anomaly detection models Mn 130.
[0061] In these instances, the anomaly detection model Mn 130 can be considered a training function, whereby the training can be done using artificial neural networks, machine learning techniques, etc. It should be understood that, in some instances, the anomaly detection model Mn 130 can be trained such that the N anomaly detection models Mn 130 can be used to determine whether a respective incoming data batch X belongs to the nth class or to any of the other N-1 classes. As an example, a (suitable) anomaly detection model can be trained that can distinguish between data distributions belonging to class 1 or any of the other N-1 classes. Then, a further anomaly detection model can be trained that can distinguish between data distributions belonging to class 2 and any of the other N-1 classes. This process can be repeated for the other N-2 classes.
[0062] With ground truth Y = {1, 2,..., N}, N anomaly detection models can be trained for each class belonging to Y. After step 1, M1, M2,..., Mn anomaly detection models can be obtained, which can predict whether the stream data batch X of input data 140 belongs to class 1 or any of the other N-1 classes, class 2 or any of the other N-1 classes, etc.
[0063] With trained anomaly detection models Mn, descriptive statistics can be obtained for anomaly scores si, s2,..., sn, each model can output this anomaly score for its class based on the training data set or training data batch Xt.
[0064] For example, there is a training data batch Xt that can be considered as ground truth Y. The input can include data points X (such as data points to be classified, e.g., a training data set or historical data), ground truth Y (e.g., labels of data points, e.g., products that data points originate from), and model M. The data batch X and ground truth Y can be related to each other via a function.
[0065] In some instances, N = 1. Thus, there is only one “separable” term and one anomaly detection model.
[0066] This case can correspond to instances of unsupervised learning (UL), which is an algorithm that learns patterns from unlabeled data. The hope is to force the machine to build a compact internal representation of its world by imitation, and then generate innovative content. In contrast to supervised learning (SL), where data is labeled by humans, e.g., as “car” or “fish”, etc., UL exhibits self-organization that captures patterns as neuron preferences or probability densities. Other levels in the supervision spectrum are reinforcement learning, where the machine is only given a numerical performance score as its guidance, and semi-supervised learning, where a small fraction of the data is labeled. Two broad approaches in UL are neural networks and probabilistic methods.
[0067] Thus, for an incoming data batch X of input data 140, only it is determined whether the monitored device 142 is in a “normal” state of normal operation or in an “abnormal” state of our normal operation. As an example, other features such as typical error or failure scenarios of the device 142 can not be identified or determined.
[0068] When the initial data set belongs to only one class, the unsupervised scenario with N = 1 can be considered as a boundary case of the supervised setting, so that there is only one anomaly detection model Mn 130. This unsupervised scenario with N = 1 typically means that there is no label available for an incoming batch X of input data 140.
[0069] In other examples, the application software component 106 and / or the processor 102 can also be configured to embed the respective N anomaly detection models Mn 130 into a software application for analyzing, monitoring, operating and / or controlling the at least one device 142 in case the determined difference is smaller than the difference threshold and to deploy the software application on the at least one device 142 or an IT system connected to the at least one device 142 such that the software application can be used for analyzing, monitoring, operating and / or controlling the at least one device 142.
[0070] The software application can be, for example, a monitoring application for analyzing the status of the respective device 142 or of a production step performed by the respective device 142 and / or for funding the status. In some examples, the software application can be an operating application or a control application for operating or controlling the respective device 142 or a production step performed by the respective device 142. The respective N anomaly detection models Mn 130 can be embedded into such a software application, for example, to derive status information of the respective device 142 or of a respective production step sequence and thereby to derive operating or control information for the respective device for the respective production step. The software application can then be deployed on the respective device 142 or on an IT system. The software application can then be provided with input data 140 which can be processed using the respective N anomaly detection models Mn 130 to determine output data 152.
[0071] In some instances, a software application can be understood to be deployed if, for example, activities required to make the software application available for use on a respective device 142 or IT system are performed by a user using the respective device 142 or a software application on the IT system. The deployment process of a software application can comprise several interrelated activities with possible transitions between them. These activities can take place on the producer side (e.g. by a developer of the software application) or on the customer side (e.g. by a user of the software application) or both. In some instances, the application deployment process can comprise at least the installation and activation of the software application, and optionally also the release of the software application. The release activity can follow the complete development process and is sometimes classified as part of the development process rather than the deployment process. The release activity can comprise the operations required to prepare a system (here: e.g. the processing system 100 or the computing unit 124) and transfer to a computer system (here: e.g. the respective device 142 or IT system) on which the software application will run in production, for a component. Thus, the release activity can sometimes involve determining the resources required for the system to operate with tolerable performance and planning and / or documenting the subsequent activities of the deployment process. For simple systems, the installation of a software application can involve some form of command, shortcut, script or service that establishes the software (manually or automatically) for executing the software application. For complex systems, the installation of the software application can involve the configuration of the system (possibly by asking the end user questions about the user’s intended use, or directly asking the end user how the user wants the system to be configured) and / or making all required subsystems ready for use. Activation can be the activity of starting the executable components of the software application for the first time (this is to be distinguished from the common use of the term activation that relates to software licensing, which is a function of digital rights management systems).
[0072] It should also be understood that in some instances, the application software component 106 and / or the processor 102 can also be configured to, in case the determined difference is greater than the difference threshold, correct the respective anomaly detection model Mn 130 such that the determined difference using the respective corrected anomaly detection model Mn 130 is smaller than the difference threshold, replace the respective anomaly detection model Mn 130 with the respective corrected anomaly detection model Mn 130 in the software application and deploy the corrected software application on the at least one device 142 or IT system.
[0073] If the determined difference is greater than the difference threshold, the respective anomaly detection model Mn 130 can be corrected, e.g. by introducing an offset or a factor with respect to the variable, such that the difference using the respective corrected anomaly detection model Mn 130 is smaller than the difference threshold. For determining the difference using the corrected trained function, the same procedure can be applied with respect to the respective anomaly detection model Mn 130, i.e. using the respective corrected N detection models Mn 130 for determining the respective anomaly scores si,..., sn for the respective incoming data batch X related to the at least N separable classes. As an example, the respective corrected N detection models Mn 130 can be found by varying the parameters of the respective N detection models Mn 130 and calculating the corresponding correction difference. If the correction difference for a given different parameter set is smaller than the difference threshold, the different parameters can be used in the corrected respective corrected N detection models Mn 130 that comply with the difference threshold.
[0074] In some examples, the correction of the respective N detection models Mn 130 can have been triggered at a slightly lower first difference threshold corresponding to a higher confidence. Thus, the respective N detection models Mn 130 can still produce an acceptable quality for analyzing, monitoring, operating and / or controlling the respective device 142, but with the better, respective corrected N detection models Mn 130 can be desirable. In this case, the correction of the respective N detection models Mn 130 can have been triggered to obtain an improved, corrected respective corrected N detection models Mn 130 resulting in a lower correction difference. This approach can allow to always have the respective N detection models Mn 130 with a high confidence, including scenarios with data distribution drifts related to wear, aging or other kinds of deterioration. Using a slightly lower first difference threshold can account for a certain delay between an increased difference of the respective N detection models Mn 130 and the determination of a corrected respective corrected N detection models Mn 130 with a lower difference and thus with a higher confidence. This scenario can correspond to an online retraining or a permanent retraining of the respective N detection models Mn 130.
[0075] In the software application, the respective N detection models Mn 130 can then be replaced by the respective corrected N detection models Mn 130, which can then be deployed at the respective device 142 or IT system.
[0076] In other instances, the application software component 106 and / or the processor 102 are further configurable to replace the deployed software application with a backup software application in case the correction of the anomaly detection model takes more time than the duration threshold and use the backup software application to analyze, monitor, operate and / or control the at least one device 142.
[0077] In some instances, it can take more time than the duration threshold to appropriately correct the respective N detection models Mn 130. This can for example occur in the online retraining scenario mentioned previously if suitable training data is lacking or if there is limited computational power. In this case, a backup software application can be used to analyze, monitor, operate and / or control the respective device 142. The backup software application can for example put the respective device 142 into a safe mode, for example to avoid damage or injury to personnel or the related production process. In some instances, the backup software application can shut down the respective device 142 or the related production process. In other instances, for example involving collaborating robots or other devices 142 that are intended to guide human-robot / device interaction within a shared space, or in case the human and the robot / device are in close proximity, the application can switch the corresponding device 142 to a slow mode, thereby also avoiding injury to personnel. Such scenarios can for example include a car manufacturing plant or other manufacturing facility with a production or assembly line, where machines and humans work in a shared space, and where the backup software application can switch the production or assembly line to such a slow mode.
[0078] It is further to be understood that in some instances, for a plurality of interconnected devices 142, the application software component 106 and / or the processor 102 are further configurable to embed the respective N detection models Mn 130 into a respective software application for analyzing, monitoring, operating and / or controlling the respective interconnected device 142, deploy the respective software application on the respective interconnected device 142 or an IT system connected to the plurality of interconnected devices 142, such that the respective software application can be used to analyze, monitor, operate and / or control the respective interconnected device (142), use the respective anomaly detection model Mn 130 to determine the respective difference, and provide an alert 150 related to the determined difference and the respective interconnected device 142 to a user, the respective device 142 and / or an automation system in case the determined difference is greater than the respective difference threshold, the corresponding respective software application being used to analyze, monitor, operate and / or control the respective interconnected device 142 for the alert.
[0079] As an example, the interconnected devices 142 can be part of a more complex production or assembly machine, or even constitute a complete production or assembly factory. In some examples, a plurality of respective anomaly detection models Mn 130 are embedded in respective software applications to analyze, monitor, operate and / or control one or more of the interconnected devices 142, wherein the respective anomaly detection models Mn 130 and the corresponding devices 142 can interact and cooperate. In such a scenario, it can be challenging to identify the origin of a problem that can occur during the operation of the interconnected devices 122. To overcome this difficulty, it is determined a respective discrepancy using the respective anomaly detection model Mn 130, and if the respective determined discrepancy is greater than a respective discrepancy threshold, an alert 152 related to the respective determined discrepancy and the respective interconnected device 142 can be provided. This approach allows a root cause analysis in a complex production environment involving a plurality of respective anomaly detection models Mn 130 embedded in corresponding software applications deployed on a plurality of interconnected devices 142. Thus, a particularly high transparency is achieved, allowing a fast and efficient identification and correction of errors. As an example, in such a complex production environment, it can be easy to identify a problematic device 142 among the plurality of interconnected devices 142, and by revising the respective anomaly detection model Mn 130 of the problematic device 142, the problem can be solved.
[0080] In the context of these examples, there can be scenarios with one respective set of anomaly detection models Mn 130 per device 142, with a plurality of respective anomaly detection models Mn 130 per device 142 or with a plurality of respective anomaly detection models Mn 130 for a plurality of devices 142. Thus, there can be a one-to-one correspondence, a one-to-many correspondence, a many-to-one correspondence or a many-to-many correspondence between the respective anomaly detection models Mn 130 and the devices 142.
[0081] It should also be understood that in other examples, the respective devices 142 are any one of a production machine, an automated device, a sensor, a production monitoring device, a vehicle or any combination thereof.
[0082] As already mentioned above, in some instances, the respective device 142 can be or comprise a sensor, an actuator such as a motor, a valve or a robot and an inverter powering the motor, a gearbox, a programmable logic controller (PLC), a communication gateway and / or other components typically associated with industrial automation products and industrial automation. The respective device 142 can be a part of a complex production line or production plant, e.g. a bottling machine, a conveyor, a welding machine, a welding robot, etc. Further, as an example, the respective device can be or comprise a manufacturing operation management (MOM) system, a manufacturing execution system (MES) and an enterprise resource planning (ERP) system, a supervisory control and data acquisition (SCADA) system or any combination thereof.
[0083] In an industrial embodiment, the proposed method and system can be implemented in the context of an industrial production facility for producing parts of a product device, e.g. printed circuit boards, semiconductors, electronic components, mechanical components, machines, devices, vehicles or parts of vehicles such as cars, bicycles, airplanes, ships, etc., or an energy generation or distribution facility, typically a power plant, a transformer, a switching device, etc. As an example, the proposed method and system can be applied to certain manufacturing steps such as milling, grinding, welding, forming, painting, cutting, etc. during the production of a product device, e.g. to monitor or even control a welding process during the production of a car. In particular, the proposed method and system can be applied to one or several plants performing the same task at different locations, whereby the input data can originate from one or several of these plants, which can allow for a particularly good database to further improve the quality of the analysis, monitoring, operation and / or control of the respective exception monitoring model Mn 130 and / or the device 142 or plant.
[0084] Here, the input data 140 can originate from a device 142 of such a facility, e.g. a sensor, a controller, etc., and the proposed method and system can be applied to improve the analysis, monitoring, operation and / or control of the device 142 or the related production or operation step. To this end, the respective exception detection model Mn 130 can be embedded into a suitable software application, which can then be deployed on the device 142 or a system, e.g. an IT system, such that the software application can be used for the mentioned purposes.
[0085] It should also be understood that in some instances the convergence of the mentioned training is not an issue, such that a stopping criterion can not be needed. This can be due to the respective anomaly detection model Mn 130 being a rather analytical function, thus can only require a limited number of iteration steps. With respect to artificial neural networks, the minimum number of nodes can typically depend on the details of the algorithm, whereby in some instances a random forest can be used for the present application. Further, the minimum number of nodes of the artificial neural network used can depend on the dimensionality of the input data 140, e.g. two dimensions (e.g. for two individual forces) or 20 dimensions (e.g. for 20 corresponding physical observable tabular data or time series data).
[0086] In exemplary embodiments of the present application one or more of the following steps can be used:
[0087] 1) receive input data comprising incoming data points X (data points to be classified - training set = historical data), optionally ground truth Y (= label of data points, e.g. product from which data points originate) and a model M; X and Y are related to each other via a function; optionally put input into some storage (e.g. buffer or low access time storage allowing for the desired sampling frequency); data points can comprise information about one or several variables, such as sensor data about current, voltage, temperature, noise, vibration, light signal, etc.
[0088] 2) optionally (if a trained anomaly detection model is not yet available): train a suitable anomaly detection model that can distinguish data distributions belonging to class 1 and not to any of the other N-1 classes
[0089] 3) optionally (if a trained anomaly detection model is not yet available): train a further model to distinguish data distributions belonging to class 2 and not to any of the other N-1 classes, etc.
[0090] 4) have ground truth Train N anomaly detection models for each class belonging to Y. After step 1 we have M1, M2,..., Mn anomaly detection models that can predict whether a stream data batch belongs to class 1 or to any of the other N-1 classes; belongs to class 2 or to any of the other N-1 classes, etc.
[0091] 5) with the trained anomaly detection models we obtain descriptive statistics for the anomaly scores si, s2,..., sn, each model outputs this anomaly score for its class based on the training data set.
[0092] 6) For each new and previously unseen incoming batch of data, we output the descriptive statistics of the anomaly score si and compare it to the corresponding descriptive statistics si obtained for each model M1, M2,..., Mn.
[0093] 7) The newly obtained anomaly score is compared to the reference anomaly score obtained based on the initial data. If significantly different, then: optionally data distribution shift is detected and a warning is sent that the trained AI model can no longer be trusted.
[0094] • The report can include an indication of "warning" if the determined difference is greater than a first threshold (accuracy < 98%; i.e. difference > 2%), then data can be started to be collected, the collected data can be data labeled (in a supervised case), and a use case machine learning model (e.g. a trained anomaly detection model) can be employed;
[0095] • If the determined difference is greater than a first threshold (accuracy < 95%, i.e. difference > 5%), the report can include an indication of "error" and the use case machine learning model (e.g. a trained anomaly detection model) can be replaced with a revised use case machine learning model (e.g. a revised trained anomaly detection model).
[0096] The embodiments have several advantages in conjunction with the invention, including:
[0097] • Full automatic detection of data distribution shift after deployment of an artificial intelligence (AI) model,
[0098] • No need for ground truth,
[0099] • Not focusing on computing possible contributors of data distribution shift, but rather employing all dimensions of the data set, and thus being more robust to multi-dimensional data sets. Moreover, the suggested solution is independent of the number of variables and can be utilized with appropriate constructed features.
[0100] In a more refined embodiment, the following considerations can be applied. To detect data distribution shift, machine learning techniques commonly used for anomaly detection are utilized. However, for the sake of generality, any other suitable method performing anomaly detection can be employed. The following settings can be covered:
[0101] 1) The AI task is solved in a supervised setting (initial training data is supplied by ground truth,
[0102] 2) The AI task is solved in an unsupervised setting (initial training data has no ground truth)
[0103] 1) Supervised setting
[0104] The Al task is formulated as follows: given data points X and ground truth Y = {1, 2,... N}, one needs to be able to build an analytical model that constructs a decision boundary that separates the streaming data between the different classes 1, 2,... N. To this end, one trains a machine learning model or uses any other analytical technique to obtain a model M. Here, the model M plays the role of a function that outputs from Y a predictor X that predicts belonging to one of the N classes. Thus, one obtains a general model M that is able to distinguish between different data distributions within the training dataset. However, this model can fail whenever the input is data that was not included in the initial training dataset. To detect this situation, one needs to determine that the incoming data distribution is different from all the data distributions that the model has seen before. To perform this detection, one trains any suitable anomaly detection model that is able to distinguish between data distributions that belong to class 1 and do not belong to any of the other N-1 classes. Then, one trains a further model to distinguish between data distributions that belong to class 2 and do not belong to any of the other N-1 classes, etc. The following workflow is established:
[0105] • Train N anomaly detection models for each class that belongs to Y, given the ground truth.
[0106] • After step 1, one has M1, M2,... Mn anomaly detection models that are able to predict whether a batch of streaming data belongs to class 1 or to any of the other N-1 classes; belongs to class 2 or to any of the other N-1 classes, etc.
[0107] • With the trained anomaly detection models, one is able to obtain descriptive statistics for the anomaly scores si, s2,... sn, each model outputting an anomaly score based on the training dataset over its class. This descriptive statistics can be: median, standard deviation, IQR, etc. of si, s2,... sn.
[0108] • For each new and previously unseen batch of incoming data, one outputs the descriptive statistics of the anomaly score si and compares it to the corresponding descriptive statistics si obtained for each model M1, M2,... Mn.
[0109] An example of this approach for a binary classification problem is shown in Figure 3 . The first model M1 (“first model”, square in Figure 3 ) has been trained with data that belongs only to class 1 of our initial dataset, model M2 (“second model”, circle in Figure 3The abnormal scores on the subsets belonging to class 1 and class 2 are obtained after training these two models. These abnormal scores are denoted as si and s2, and are distributed over timestamps 0 and 157. The median of si (0-156) and s2(0-156) together is 27.4. At timestamp 157, data belonging to other distributions starts streaming, and has been checked against the training models M1 and M2. These abnormal scores are distributed over timestamps 157 and 312, and can be denoted as si (157-312) s2(157-312). The median of the abnormal score distribution in this case is 5.4. For this instance, only one descriptive statistic is used, which is the median of s. As Figure 3 It can be seen that the data distribution drifts at timestamp 157.
[0110] To introduce robustness to the proposed method, the descriptive statistic is considered as the distribution itself for s, and for this, the distributions of s are plotted for comparison. These distributions are shown in Figure 4
[0111] It can be seen that the abnormal scores are mainly concentrated in the left box, which is a unimodal distribution with median 27.4 and with 4 outliers. The left box incorporates most of the data under the data distribution drift with median 5.4 and 2 outliers. These two distributions are completely separable, and the data distribution drift can be clearly seen. However, in cases where these distributions are not completely separable, a statistical test with the following hypotheses can be employed:
[0112] • H0: The distributions of s are identical.
[0113] • H1: H0 is not correct.
[0114] 2) Unsupervised case
[0115] When the initial dataset belongs to only one class, the proposed method can consider the unsupervised setting as a boundary case of the supervised setting. In this case, all the things described above are applicable and valid. The number of anomaly detection models drops to 1.
[0116] Sending an alert
[0117] After having trained the initial model and the model for data drift detection, the anomaly detection model can be deployed and monitor the abnormal scores in an automated way that cross-checks the newly obtained abnormal scores against the reference abnormal scores obtained based on the initial data. As previously described, if the newly obtained abnormal score distribution is significantly different from the reference abnormal score distribution, a data distribution drift is detected, and a warning can be sent that the AI model has been trained (e.g. the corresponding trained anomaly detection model Mn can no longer be trustworthy).
[0118] To avoid (unnecessary) false positives, the following workflow is suggested:
[0119] • The trained Al model and the anomaly detection model are deployed
[0120] • The data stream is started
[0121] • For each incoming batch of data, the above method is applied and a new anomaly score is obtained
[0122] • If the distribution of the newly obtained anomaly scores is significantly different from the reference distribution of the N incoming batches of the stream data in order, the user is warned that a data distribution shift has occurred.
[0123] If the difference in the anomaly score distribution does not occur in order or suddenly, the difference can be ignored and treated as an outlier.
[0124] Compared to other approaches, the suggested method provides the following advantages:
[0125] 1) To detect data distribution shifts after deploying an Al model, our solution does not require any ground truth and performs the detection in a fully automated way.
[0126] 2) The suggested solution does not focus on computing possible contributors of data distribution shifts, but rather takes all dimensions of the data set into account and is thus more robust for multi-dimensional data sets. Moreover, our solution is independent of the number of variables and can be utilized with appropriate constructed features.
[0127] 3) The approaches of other ways use one-dimensional distances and high-priced operations for their computations. To detect data distribution shifts, they suggest to manually set thresholds based on empirical knowledge, which has the disadvantage of having a lot of false positives or / and false negatives. An additional disadvantage is the difficulty to automate the manual steps.
[0128] 4) Most of the indicated competitors are operating in the e-commerce area or / and in the computer vision area and thus provide solutions that are mainly based on certain use cases. The suggested method is applicable on a large scale in a fully automated way to tabular and time series data and to supervised and unsupervised settings.
[0129] 5) Other ways usually rely on handcrafted thresholds, which are use case and data specific and can lead to a lot of false / missed positive detections.
[0130] 6) The suggested method is based on Al technology and performs the monitoring and decision making in a fully automated way. This approach replaces the manual threshold monitoring and provides room for scaling and generalization.
[0131] Generally, the proposed method provides:
[0132] • better performance and efficiency
[0133] • in addition, the robustness of the proposed method can be improved by performing statistical tests
[0134] • is more robust to the use of multidimensional data sets and can reduce the dimensionality
[0135] • the deployment is fully automated
[0136] • is computationally efficient and can run on any suitable edge device
[0137] • is suitable for various Siemens products and solutions.
[0138] In an industrial embodiment, the proposed method and system can be implemented in the context of an industrial production facility for parts of production equipment, e.g. printed circuit boards, semiconductors, electronic components, mechanical components, machines, devices, vehicles or parts of vehicles, such as cars, bicycles, airplanes, ships, etc., or energy generation or distribution facilities, typically power plants, transformers, switching devices, etc. As an example, the proposed method and system can be applied to certain manufacturing steps during the production of a device, such as milling, grinding, welding, forming, painting, cutting, etc., thereby monitoring or even controlling a welding process, for example, during the production of a car. In particular, the proposed method and system can be applied to one or several factories performing the same task at different locations, whereby the input data can originate from one or several of these factories, which can allow for a particularly good database to further improve the quality of the trained model and / or the analysis, monitoring, operation and / or control of the device or factory.
[0139] Here, the input data can originate from a device of such a facility, e.g. a sensor, a controller, etc., and the proposed method and system can be applied to improve the analysis, monitoring, operation and / or control of the device. To this end, the training functionality can be embedded into a suitable software application, which can then be deployed on the device or system, e.g. an IT system, such that the software application can be used for the mentioned purposes.
[0140] As an example, device input data can be used as input data, while device output data can be used as output data.
[0141] Figure 2 The degradation of a model over time due to a shift in the data distribution is shown. Herein, the model can correspond to a respective anomaly detection model Mn 130, wherein the (anomaly detection) model can be a trained model.
[0142] In an ideal situation, a model trained, e.g., based on collected data, must perform very well on an incoming data stream. However, an analytical model degrades over time and a model trained at time ti can perform worse at time t2.
[0143] For illustration purposes, consider a binary classification between classes A and B of a two-dimensional data set. At time ti, a data analyst trains a model capable of building a decision boundary 162 between data belonging to either class A (see data points 164) or class B (see data points 166). In this case, building the decision boundary 162 corresponds to the actual boundary 160 separating the two classes. At deployment, the model typically performs very well. However, at a later time t2>ti, the incoming data stream or input data messages 140 can experience a shift in the data distribution and, as a result, can have an impact on the performance of the model. At time t2, the data analyst can detect a performance degradation of the model and, as a result, can decide to retrain or update the model. Figure 2 This phenomenon can be seen on the right-hand side of the figure: data points 166 belonging to class B have shifted towards the lower right corner, while data points 164 belonging to class A have shifted in the opposite direction. Thus, the previously built decision boundary 162 does not correspond to the new data distribution of classes A and B, since the new actual boundary 160' separating the two classes has moved. As a result, the analytical model must be retrained or updated as quickly as possible.
[0144] One goal of the proposed approach can include developing a method for detecting a performance degradation or decrease of a trained model (e.g., a respective anomaly detection model Mn 130) as a function of a shift in the data distribution in a data stream, such as a sensor data stream or input data messages 140. It is to be noted that, in some instances, a high data shift alone does not imply a poor prediction accuracy of a trained model (e.g., a respective anomaly detection model Mn 130). It can finally be necessary to correlate this shift with the ability of the old model to handle data shifts (i.e., measure the current accuracy). In some instances, once a performance degradation or a greater discrepancy is detected, the data analyst can retrain the model based on new features of the incoming data.
[0145] Figure 3 An exemplary data distribution shift detection for a binary classification task is shown (see explanation above).
[0146] Figure 4 An exemplary boxplot comparing two distributions of anomaly scores is shown (see explanation above).
[0147] Figure 5 A functional block diagram of an exemplary system that facilitates providing alerts and managing computer software products in a product system is shown.
[0148] The overall architecture of the illustrated exemplary system can be divided into a development ("dev"), an operation ("ops"), and a big data architecture arranged between development and operation. In this context, dev and ops can be understood as in DevOps, a collection of practices that combines software development (Dev) and IT operations (Ops). DevOps aims to shorten the system development lifecycle and provide continuous delivery with high software quality. As an example, the above explained anomaly detection model can be developed or refined and then embedded into a software application in the "dev" area of the illustrated system, whereby the anomaly detection model of the software application is operated in the "ops" area of the illustrated system. The overall concept is to enable an adjustment or refinement of the anomaly detection model or the corresponding software solution based on operational data from "ops", which can be processed or worked on by the "big data architecture", whereby the adjustment or refinement is made in the "dev" area.
[0149] In the lower right, i.e. in the "ops" area, a deployment tool for an application, such as a software application, with various microservices, referred to as "productive pasture directory", is shown. This allows data import, data export, MQTT broker and data monitor. The productive pasture directory is part of a "productive cluster", which can belong to the operational side of the overall "digital service architecture". The productive pasture directory can provide a software application ("application program") that can be deployed as a cloud application in the cloud or as an edge application on an edge device, such as a device and machine used in an industrial production facility or an energy generation or distribution facility (as explained in detail above). The microservices can for example represent such an application or be included in such an application. The device on which the corresponding application is running (or the application running on the respective device) can deliver data, such as sensor data, control data, etc., for example as records or raw data (or for example input data), to the Figure 4 The cloud storage referred to as "big data architecture" in the middle.
[0150] The input data can be used on the development side ("dev") of the overall digital service architecture to check whether the anomaly detection model (see block "Your model" in block "Code orchestration framework" in section "Software and AI development") is still accurate or needs to be corrected (see determination of the difference and correction of the anomaly detection model in case the determined difference is above a certain threshold). In the "Software and A1 development" area, there can be templates and AI models and, optionally, training of new models can be performed. If correction is needed, the anomaly detection model is corrected accordingly and corrected during the "Automated CI / CD pipeline" (CI / CD = Continuous Integration / Continuous Delivery or Continuous Deployment) embedded in the application, which can be deployed as a cloud application in the cloud or as an edge application on an edge device, when transferred to the protection cluster (mentioned above) on the operational side of the overall digital service architecture.
[0151] The automated CI / CD pipeline can comprise:
[0152] • Build "base image" and "base application" -> build application image and application
[0153] • Unit tests: software tests, machine learning model tests
[0154] • Integration tests (containers technology on machines or clusters, e.g. Kubernetes cluster)
[0155] • HW (= hardware) integration tests (deployment on real edge device / edge box)
[0156] • New image can be obtained that is suitable for release / deployment in a productive cluster
[0157] For example, if a sensor or device is damaged, has a fault, usually needs to be replaced, the described update or correction of the anomaly detection model can be necessary. In addition, sensors and devices are aging, so that sometimes a new calibration can be necessary. Such events can result in anomaly detection models that are no longer trustworthy but need to be updated.
[0158] An advantage of the proposed method and system embedded in such a digital service architecture is that the update of the anomaly detection model can be performed as quickly as the replacement of a sensor or device, e.g. for a programming and deployment of a new anomaly detection model and the corresponding application including the new anomaly detection model, only a recovery time of 15 minutes is needed. A further advantage is that the update of the deployed anomaly detection model and the corresponding application can be performed completely automatically.
[0159] The described examples can provide an efficient way of providing alerts related to anomaly scores assigned to input data, such as detecting distribution drift of incoming data using an anomaly detection model, enabling the ability to drive digital transformations and impart machine learning applications impact and even shape processes. One important aspect of the invention contributes to helping ensure the trustworthiness of such applications in highly volatile environments with respect to a plant. The invention can support handling this challenge by providing a monitoring and alert system that helps react appropriately once a machine learning application is not working the way it was trained. Thus, the described examples can generally reduce the total cost of ownership of computer software products by improving the trustworthiness of computer software products and supporting keeping them up to date. This efficient provision of output data and management of computer software products can be leveraged in any industry, such as aerospace and defense, automotive and transportation, consumer goods and retail, electronics and semiconductors, energy and utilities, machinery and heavy equipment, marine, or medical devices and pharmaceuticals. This efficient provision of output data and management of computer software products can also be applicable to customers facing the need for trustworthy and up-to-date computer software products.
[0160] In particular, the above examples apply equally to the computer system 100, the corresponding computer program product, and the corresponding computer readable medium arranged and configured to perform the steps of the computer-implemented method of providing output data explained in this patent document, respectively.
[0161] Reference is now made to Figure 6 , showing a method 600 that helps provide alerts related to anomaly scores assigned to input data, such as detecting distribution drift of incoming data using an anomaly detection model. The method can start from 602, and the method can include several actions by operation of at least one processor.
[0162] The actions can comprise an action 604 of receiving input data related to at least one device, wherein the input data comprises incoming data batches X related to at least N separable classes, wherein n e 1,..., N; an action 606 of using N anomaly detection models Mn to determine respective anomaly scores si,..., sn for respective incoming data batches X related to the at least N separable classes; an action 608 of applying the (trained) anomaly detection models Mn to the input data to generate output data suitable for analyzing, monitoring, operating and / or controlling the respective device; an action 610 of determining a difference between, on the one hand, the determined respective anomaly scores si,..., sn for the at least N separable classes for the respective incoming data batches X and, on the other hand, the given respective anomaly scores Si,..., Sn of the N anomaly detection models Mn (130). And if the respective determined difference is greater than a difference threshold, then: an action 612 of providing an alert related to the determined difference to a user, the respective device and / or an IT system connected to the respective device. At 614, the method can end.
[0163] It will be appreciated that the method 600 can comprise other actions and features discussed previously in relation to the computer-implemented method of providing an alert related to an anomaly score assigned to input data, such as detecting a distribution shift of incoming data using anomaly detection models.
[0164] For example, the method can further comprise an action of determining a distribution shift of the input data in the event that a second difference between anomaly scores si,..., sn of an earlier incoming data batch Xe and anomaly scores si,..., sn of a later incoming data batch Xi is greater than a second threshold; and an action of providing a report related to the determined distribution shift to a user, the respective device and / or an IT system connected to the respective device in the event that the determined second difference is greater than the second threshold.
[0165] It will also be appreciated that, in some instances, the method can further comprise an action of assigning training data batches Xt to the at least N separable classes of the anomaly detection models Mn; and an action of determining the given anomaly scores Si,..., Sn for the at least N separable classes for the N anomaly detection models Mn.
[0166] In some instances, the method can further comprise, in the event that the determined precision value is equal to or greater than a precision threshold, an action of embedding the N anomaly detection models Mn in a software application for analyzing, monitoring, operating and / or controlling the at least one device in the event that the determined difference is less than the difference threshold; and an action of deploying the software application on the at least one device or an IT system connected to the at least one device such that the software application can be used for analyzing, monitoring, operating and / or controlling the at least one device.
[0167] In other instances, if the determined difference is greater than the difference threshold, the method can further include the acts of modifying the corresponding anomaly detection model Mn such that the determined difference using the corresponding modified anomaly detection model Mn is less than the difference threshold; replacing the corresponding anomaly detection model Mn with the corresponding modified anomaly detection model Mn in the software application; and deploying the modified software application on the at least one device or IT system.
[0168] It should also be appreciated that in some instances, the method can further include the acts of replacing the deployed software application with a backup software application and using the backup software application to analyze, monitor, operate, and / or control the at least one device in the event that the modification of the anomaly detection model takes more time than the duration threshold.
[0169] In some instances, for a plurality of interconnected devices, the method can further include the acts of embedding the corresponding N detection models Mn in corresponding software applications for analyzing, monitoring, operating, and / or controlling the corresponding interconnected devices; deploying the corresponding software applications on the corresponding interconnected devices or an IT system connected to the plurality of interconnected devices such that the corresponding software applications can be used to analyze, monitor, operate, and / or control the corresponding interconnected devices; determining the corresponding difference for the corresponding anomaly detection model; and further including the act of providing an alert to a user, the corresponding device, and / or an automated system related to the determined difference and the corresponding interconnected device in the event that the corresponding determined difference is greater than the corresponding difference threshold, the corresponding software application being used to analyze, monitor, operate, and / or control the corresponding interconnected device for the alert.
[0170] As previously discussed, the acts associated with these methods (other than any described manual acts, such as the act of manually making a selection by an input device) can be performed by one or more processors. Such processors can be included in one or more data processing systems, e.g., that execute software components operable to cause the acts to be performed by the one or more processors. In example embodiments, such software components can include computer-executable instructions corresponding to routines, subroutines, programs, applications, modules, libraries, execution threads, etc. Additionally, it should be appreciated that software components can be written and / or generated in a software environment / language / framework, such as Java, JavaScript, Python, C, C#, C++, or any other software tool that can produce components and graphical user interfaces configured to perform the acts and features described herein.
[0171] Figure 7An embodiment of an artificial neural network 2000 is shown, which can be used in the context of providing an alert related to an anomaly score assigned to input data, such as using an anomaly detection model to detect a distribution shift of incoming data. Alternative terms for "artificial neural network" are "neural network", "artificial neural net", or "neural net".
[0172] The artificial neural network 2000 comprises nodes 2020,..., 2032 and edges 2040,..., 2042, wherein each edge 2040,..., 2042 is a directed connection from a first node 2020,..., 2032 to a second node 2020,..., 2032. Typically, the first node 2020,..., 2032 and the second node 2020,..., 2032 are different nodes 2020,..., 2032, it is also possible that the first node 2020,..., 2032 and the second node 2020,..., 2032 are the same. For example, in the shown embodiment, the edge 2040 is a directed connection from the node 2020 to the node 2023, while the edge 2042 is a directed connection from the node 2030 to the node 2032. The edges 2040,..., 2042 from a first node 2020,..., 2032 to a second node 2020,..., 2032 are also denoted as "input edges" of the second node 2020,..., 2032 and "output edges" of the first node 2020,..., 2032. Figure 6
[0173] In this embodiment, the nodes 2020,..., 2032 of the artificial neural network 2000 can be arranged in layers 2010,..., 2013, wherein these layers can comprise an inherent order introduced by the edges 2040,..., 2042 between the nodes 2020,..., 2032. In particular, edges 2040,..., 2042 can only exist between adjacent layers of nodes. In the shown embodiment, there is an input layer 2010 comprising only nodes 2020,..., 2022 without incoming edges, an output layer 2013 comprising only nodes 2031, 2032 without outgoing edges, and hidden layers 2011, 2012 between the input layer 2010 and the output layer 2013. Typically, the number of hidden layers 2011, 2012 can be chosen arbitrarily. The number of nodes 2020,..., 2022 within the input layer 2010 is typically related to the number of input values of the neural network, while the number of nodes 2031, 2032 within the output layer 2013 is typically related to the number of output values of the neural network.
[0174] In particular, a (real) number can be assigned as a value to each node 2020,..., 2032 of the neural network 2000. Here, x(n)i denotes the value of the i-th node 2020,..., 2032 of the n-th layer 2010,..., 2013. The values of the nodes 2020,..., 2022 of the input layer 2010 are equivalent to the input values of the neural network 2000, and the values of the nodes 2031, 2032 of the output layer 2013 are equivalent to the output values of the neural network 2000. Furthermore, each edge 2040,..., 2042 can comprise a weight being a real number, in particular, the weight is a real number within the interval [-1, 2] or within the interval [0, 2]. Here, w (m,n) i,j denotes the weight of the edge between the i-th node 2020,..., 2032 of the m-th layer 2010,..., 2013 and the j-th node 2020,..., 2032 of the n-th layer 2010,..., 2013. Furthermore, the weight w (n,n+1) i,j The abbreviation w (n) i,j .
[0175] In particular, for calculating the output values of the neural network 2000, the input values are propagated through the neural network. In particular, the values of the nodes 2020,..., 2032 of the (n+1)-th layer 2010,..., 2013 can be calculated based on the values of the nodes 2020,..., 2032 of the n-th layer 2010,..., 2013 by
[0176]
[0177] In this context, the function f is a transfer function (further term: activation function). Known transfer functions are step functions, sigmoid functions (e.g. logistic function, generalized logistic function, hyperbolic tangent, arctangent function, error function, smooth step function) or rectifier functions. The transfer function is mainly used for normalization purposes.
[0178] In particular, the values are propagated through the neural network layer by layer, wherein the values of the input layer 2010 are given by the input of the neural network 2000, wherein the values of the first hidden layer 2011 can be calculated based on the values of the input layer 2010 of the neural network, wherein the values of the second hidden layer 2012 can be calculated based on the values of the first hidden layer 2011, etc.
[0179] For setting the values w (m,n) i,j The neural network 2000 has to be trained using training data. In particular, the training data comprises training input data and training output data (denoted as t i). For a training step, the neural network 2000 is applied to the training input data to generate computed output data. In particular, the training data and the computed output data comprise a number of values, the number being equal to the number of nodes of the output layer.
[0180] In particular, the comparison between the computed output data and the training data is used to recursively adjust the weights within the neural network 2000 (backpropagation algorithm). In particular, the weights are changed according to the following formula:
[0181]
[0182] where y is the learning rate and the number δ (n) j is recursively computed as:
[0183]
[0184] Based on δ(n+1)j, if the (n+1)th layer is not the output layer, and
[0185]
[0186] if the (n+1)th layer is the output layer 2013, where f' is the first derivative of the activation function, and y (n+1) j is the comparison training value of the jth node of the output layer 2013.
[0187] Figure 8 An embodiment of a convolutional neural network 3000 is shown, which can be used in the context of providing an alert related to an anomaly score assigned to input data, such as detecting a distribution drift of incoming data using an anomaly detection model.
[0188] In the shown embodiment, the convolutional neural network comprises 3000, an input layer 3010, a convolutional layer 3011, a pooling layer 3012, a fully connected layer 3013, and an output layer 3014. Alternatively, the convolutional neural network 3000 can comprise several convolutional layers 3011, several pooling layers 3012, and several fully connected layers 3013, as well as other types of layers. The order of the layers can be chosen arbitrarily, typically using the fully connected layer 3013 as the final layer before the output layer 3014.
[0189] In particular, within the convolutional neural network 3000, the nodes 3020,..., 3024 of one layer 3010,..., 3014 can be considered to be arranged as a d-dimensional matrix or d-dimensional image. In particular, in the two-dimensional case, the values of the nodes 3020,..., 3024 indexed with i and j in the nth layer 3010,..., 3014 can be represented as x (n)[i,j] However, the arrangement of the nodes 3020,..., 3024 of one layer 3010,..., 3014 has no influence on the computations performed within the convolutional neural network 3000, and thus, is given by the structure and weights of the edges only.
[0190] In particular, the convolutional layer 3011 is characterized by a structure and weights of the incoming edges based on a number of kernels forming a convolution operation. In particular, the structure and weights of the incoming edges are chosen such that the value x (n-1) of the node 3021 of the convolutional layer 3011 is formed based on the values x (n) k as a convolution x (n) k = K k * x (n-1) where the convolution * is defined in the two-dimensional case as:
[0191]
[0192] Here, the k-th kernel K k is a d-dimensional matrix (in this embodiment a two-dimensional matrix) which is typically small (e.g. a 3x3 matrix or a 5x5 matrix) compared to the number of nodes 3020,..., 3024. In particular, this means that the weights of the incoming edges are not independent but chosen such that they result in said convolution equation. In particular, for a kernel being a 3x3 matrix, there are only 9 independent weights (each entry of the kernel matrix corresponds to one independent weight) independent of the number of nodes 3020,..., 3024 in the respective layer 3010,..., 3014. In particular, for the convolutional layer 3011, the number of nodes 3021 in the convolutional layer is equivalent to the number of nodes 3020 in the preceding layer 3010 multiplied by the number of kernels.
[0193] If the nodes 3020 of the preceding layer 3010 are arranged as a d-dimensional matrix, the use of multiple kernels can be interpreted as adding an additional dimension (denoted as "depth" dimension) such that the nodes 3021 of the convolutional layer 3021 are arranged as a (d+1)-dimensional matrix. If the nodes 3020 of the preceding layer 3010 are already arranged as a (d+1)-dimensional matrix including the depth dimension, the use of multiple kernels can be interpreted as an extension along the depth dimension such that the nodes 3021 of the convolutional layer 3021 are also arranged as a (d+1)-dimensional matrix, wherein the size of the (d+1)-dimensional matrix with respect to the depth dimension is a factor larger than the number of kernels in the preceding layer 3010.
[0194] An advantage of using a convolutional layer 3011 is that the spatial local correlation of the input data can be exploited by implementing a local connectivity pattern between the nodes of adjacent layers, in particular by each node being connected only to a small area of the nodes of the previous layer.
[0195] In the shown embodiment, the input layer 3010 comprises 36 nodes 3020 arranged as a two-dimensional 6x6 matrix. The convolutional layer 3011 comprises 72 nodes 3021 arranged as two two-dimensional 6x6 matrices, each of which is the result of a convolution of the values of the input layer with a kernel. Equivalently, the nodes 3021 of the convolutional layer 3011 can be interpreted as being arranged as a three-dimensional 6x6x2 matrix, where the final dimension is the depth dimension.
[0196] The pooling layer 3012 can be characterized by the structure and weights of the incoming edges and the activation function of its nodes 3022, which form a pooling operation based on a non-linear pooling function f. For example, in the two-dimensional case, the value x (n - 1) of a node 3022 of the pooling layer 3012 can be computed as: (n)
[0197] x (n) [i,j] = f(x (n-1) [id1,jd2],...,x (n-1) [id1+d1-1,jd2+d2-1])
[0198] In other words, by using a pooling layer 3012, the number of nodes 3021, 3022 can be reduced by replacing a number d1-d2 of neighboring nodes 3021 in the previous layer 3011 with a single node 3022, which is computed as a function of the values of said number of neighboring nodes in the pooling layer. In particular, the pooling function f can be the max function, the average or the L2 norm. In particular, for the pooling layer 3012, the weights of the incoming edges are fixed and not modified by training.
[0199] An advantage of using a pooling layer 3012 is the reduction of the number of nodes 3021, 3022 and of the number of parameters. This results in a reduction of the computational effort in the network and in a control of overfitting.
[0200] In the shown embodiment, the pooling layer 3012 is max pooling, replacing four neighboring nodes with one node, the value being the maximum of the values of the four neighboring nodes. Max pooling is applied to each d-dimensional matrix of the previous layer; in this embodiment, max pooling is applied to each of the two two-dimensional matrices, reducing the number of nodes from 72 to 18.
[0201] The fully connected layer 3013 can be characterized by the fact that there are most of the edges, in particular all the edges, between the nodes 3022 of the previous layer 3012 and the nodes 3023 of the fully connected layer 3013, and wherein the weight of each of the edges can be adjusted individually.
[0202] In this embodiment, the nodes 3022 of the previous layer 3012 of the fully connected layer 3013 are shown as a two-dimensional matrix and are additionally shown as uncorrelated nodes (indicated as node lines, wherein the number of nodes is reduced for better presentability). In this embodiment, the number of nodes 3023 in the fully connected layer 3013 is equal to the number of nodes 3022 in the previous layer 3012. Alternatively, the number of nodes 3022, 3023 can be different.
[0203] Furthermore, in this embodiment, the values of the nodes 3024 of the output layer 3014 are determined by applying a Softmax function to the values of the nodes 3023 of the previous layer 3013. By applying the Softmax function, the sum of the values of all nodes 3024 of the output layer is 1 and all values of all nodes 3024 of the output layer are real numbers between 0 and 1. In particular, if the convolutional neural network 3000 is used to classify input data, the values of the output layer can be interpreted as probabilities that the input data falls into one of the different classes.
[0204] The convolutional neural network 3000 can also comprise a ReLU (acronym for "Rectified Linear Unit") layer. In particular, the number of nodes and the structure of the nodes comprised in the ReLU layer are equivalent to the number of nodes and the structure of the nodes comprised in the previous layer. In particular, the value of each node in the ReLU layer is calculated by applying a rectification function to the value of the corresponding node of the previous layer. Examples of rectification functions are f(x) = max(0, x), the hyperbolic tangent function or the sigmoid function.
[0205] In particular, the convolutional neural network 3000 can be trained based on a backpropagation algorithm. To prevent overfitting, regularization methods can be used, such as dropout of information of the nodes 3020,..., 3024, random pooling, use of artificial data, weight decay based on L1 or L2 norm or maximum norm constraint.
[0206] It is important to note that while the present disclosure includes descriptions of the embodiments in the context of fully functional systems and / or a series of acts, one skilled in the art will appreciate that the mechanisms and / or acts described herein can be distributed in the context of computer-executable instructions, which can be embodied in any of a variety of forms, including, but not limited to, program modules, applications, program modules, routines, libraries, objects, executables, components, data structures, etc. that are present in any of a variety of forms on a variety of non-transitory, machine or computer-readable media, and that can be executed by or to otherwise cooperate with one or more processors to effectuate aspects of the present disclosure. Non-transitory machine- readable or computer- readable media include any medium, or combination thereof, that is sufficient for a processor to read from or write to. Examples of non-transitory machine- readable or computer-readable media include, but are not limited to, RAM, EPROM, tape, floppy disks, hard disks, SSDs, flash memory, CD-ROMs, DVDs, and Blu-ray disks. Computer-executable instructions include, but are not limited to, routines, sub-routines, programs, applications, modules, libraries, executables, threads, etc. Still further, results of actions of the methods can be stored in computer-readable media, displayed on display devices, etc.
[0207] Figure 9 A block diagram of a data processing system 1000 (also referred to as a computer system) is shown in which embodiments can be implemented, for example, as part of a product system and / or other systems operably configured by software or otherwise configured to perform processes as described herein. The data processing system 1000 can include, for example, the above-mentioned computer or IT system or data processing system 100. The depicted data processing system includes at least one processor 1002 (e.g., CPU) that can be connected to one or more bridges / controllers / buses 1004 (e.g., northbridge, southbridge). For example, one of the buses 1004 can include one or more I / O buses, such as a PCI Express bus. Also connected to the various buses in the depicted example is a main memory 1006 (RAM) and a graphics controller 1008. The graphics controller 1008 can be connected to one or more display devices 1010. Note also that in some embodiments, one or more controllers (e.g., graphics, southbridge) can be integrated with the CPU (on the same chip or die). Examples of CPU architectures include IA-32, x86-64, and ARM processor architectures.
[0208] Other peripheral devices connected to one or more buses can include a communications controller 1012 (Ethernet controller, WiFi controller, cellular controller) operable to connect to a local area network (LAN), a wide area network (WAN), a cellular network, and / or other wired or wireless networks 1014 or communication equipment.
[0209] Other components connected to the various buses can include one or more I / O controllers 1016, such as a USB controller, a Bluetooth controller, and / or a dedicated audio controller (connected to speakers and / or a microphone). It should also be appreciated that various peripheral devices can be connected to the I / O controller (via various ports and connections), including input devices 1018 (e.g., keyboard, mouse, pen, touch screen, touch pad, drawing tablet, trackball, buttons, keypad, game controller, game pad, camera, microphone, scanner, motion sensing device that captures motion gestures), output devices 1020 (e.g., printer, speakers), or any other type of device operable to provide input to or receive output from a data processing system. Moreover, it should be appreciated that many devices referred to as input devices or output devices can provide both input and receive output to and from the data processing system. For example, a processor 1002 can be integrated into a housing (such as a tablet) that includes a touch screen that functions as both an input and a display device. Additionally, it should be appreciated that some input devices (such as a laptop computer) can include multiple different types of input devices (e.g., touch screen, touch pad, keyboard). Moreover, it should be appreciated that other peripheral devices 1022 connected to the I / O controller 1016 can include any type of device, machine, or component configured to communicate with the data processing system.
[0210] Additional components connected to the various buses can include one or more storage controllers 1024 (e.g., SATA). The storage controller can be connected to storage devices 1026 (such as one or more storage drives and / or any associated removable media), which can be any suitable non-transitory machine usable or machine readable storage media. Examples include nonvolatile memory, volatile memory, read only memory, writeable memory, ROM, EPROM, magnetic tape storage, floppy disk drive, hard disk drive, solid-state drive (SSD), flash memory, optical disk drive (CD, DVD, Bluray), and other known optical, electrical or magnetic storage device drives and / or computer media. Moreover, in some examples, a storage device (such as an SSD) can be connected directly to the I / O bus 1004, such as a PCI Express bus.
[0211] A data processing system according to an embodiment of the present disclosure can include an operating system 1028, software / firmware 1030, and data repositories 1032 (which can be stored on the storage device 1026 and / or the memory 1006). Such an operating system can employ a command line interface (CLI) shell and / or a graphical user interface (GUI) shell. A GUI shell permits multiple display windows to be presented in different graphical user interfaces, each one having different applications or different instances of the same application presented therein. A cursor or pointer in the graphical user interface can be manipulated by a user through a pointing device such as a mouse or touch screen. The position of the cursor / pointer can be changed and / or an event, such as a mouse click or touch screen tap, can be generated to actuate desired responses. Examples of operating systems that can be used in a data processing system include the Microsoft Windows, Linux, UNIX, iOS, and Android operating systems. Moreover, examples of data repositories include data files, data tables, relational databases (e.g., Oracle, Microsoft SQL Server), database servers, or any other structure and / or device that can store data retrievable by a processor.
[0212] The communication controller 1012 can be connected to the network 1014 (not part of the data processing system 1000) that can be any public or private data processing system network or combination of networks, known to those of skill in the art, including the Internet. The data processing system 1000 can be in communication with one or more other data processing systems, such as a server 1034 (also not part of the data processing system 1000), through the network 1014. However, alternative data processing systems can correspond to a plurality of data processing systems implemented as part of a distributed system in which the processors associated with the several data processing systems can communicate over one or more network connections and can collectively perform tasks described as being performed by a single data processing system. Thus, it should be understood that when referring to a data processing system, such a system can be implemented across several data processing systems organized in a distributed system in communication with each other over a network.
[0213] Furthermore, the term "controller" means any device, system or part thereof that controls at least one operation, whether the device is implemented in hardware, firmware, software or some combination of at least two of the same. It should be noted that the functionality associated with any particular controller can be centralized or distributed, whether locally or remotely.
[0214] Further, it should be appreciated that the data processing system can be implemented as a virtual machine architecture or a virtual machine in a cloud environment. For example, the processor 1002 and associated components can correspond to a virtual machine executing in a virtual machine environment of one or more servers. Examples of virtual machine architectures include VMware ESCi, Microsoft hyper-V, Xen, and KVM.
[0215] Those of ordinary skill in the art will appreciate that the hardware depicted for the data processing system can vary depending on the particular implementation. For example, the data processing system 1000 in this example can correspond to a computer, workstation, server, PC, notebook, tablet, mobile phone, and / or any other type of device / system operable to process data and perform the functionality and features described herein in association with the operation of the data processing systems, computers, processors, and / or controllers discussed herein. The depicted examples are provided for the purpose of explanation only, and are not meant to imply architectural limitations for the disclosure.
[0216] Further, it should be noted that the processors described herein can be located in a server remote from the displays and input devices described herein. In such examples, the described display devices and input devices can be included in a client device in communication with the server (and / or a virtual machine executing on the server) over a wired or wireless network, which can include the Internet. In some embodiments, such a client device may, for example, execute a remote desktop application, or can correspond to a portal device with a server for a remote desktop protocol, to send input from the input device to the server and receive visual information from the server for display by the display device. Examples of such remote desktop protocols include Teradici's PCoIP, Microsoft's RDP, and RFB protocol. In such examples, the processors described herein can correspond to virtual processors of a virtual machine executing in a physical processor of the server.
[0217] As used herein, the terms "component" and "system" are intended to encompass hardware, software, or a combination of hardware and software. Thus, for example, a system or component can be a process, a process executing on a processor, or a processor. Additionally, a component or system can be located on a single device or distributed across several devices.
[0218] Further, as used herein, a processor corresponds to any electronic device configured to process data via hardware circuitry, software, and / or firmware. For example, the processors described herein can correspond to one or more of (or a combination of) a microprocessor, a CPU, an FPGA, an ASIC, or any other integrated circuit (IC) or other type of circuit that can process data, which can take the form of a controller board, a computer, a server, a mobile phone, and / or any other type of electronic device.
[0219] Those skilled in the art will recognize that, for simplicity and clarity, the full structure and operation of all data processing systems suitable for use with the present disclosure is not being depicted or described herein. Instead, only so much of a data processing system as is unique to the present disclosure or necessary for an understanding of the present disclosure is depicted and described. The remainder of the construction and operation of the data processing system 1000 can conform to any of the various current implementations and practices known in the art.
[0220] Further, it should be understood that the words or phrases used herein in reference to particular embodiments should not be construed as limiting, unless explicitly so limited in some instances. For example, the terms “comprising” and / or “comprise,” and variations thereof, mean including, but not limited to, and the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Additionally, the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The term “or” is inclusive, meaning and / or, unless the context clearly indicates otherwise. The phrases “associated with” and “associated therewith,” as well as derivatives thereof, can mean to include, be included within, interconnect with, contain, be contained within, connect to or with, couple to or with, be communicable with, cooperate with, interleave, be proximate to, be bound to or with, have a property of, have, have a property of, or the like.
[0221] Further, although the terms “first,” “second,” “third,” etc. can be used herein to describe various elements, functions or acts, such elements, functions or acts should not be limited by these terms. Rather, these terms are used merely to distinguish one element, function, or act from another element, function, or act. For example, a first element, function, or act could later be referred to as a second element, function, or act, and similarly, a second element, function, or act could be later referred to as a first element, function, or act, without departing from the scope of the present disclosure.
[0222] Further, phrases such as "processor configured to," variations thereof and / or "configured to" can refer to a processor being operable or operably configured to perform a function or process via software, firmware and / or hardware. For example, a processor configured to perform a function / process can correspond to a processor that is executing software / firmware programmed to cause the processor to perform the function / process and / or a processor that has software / firmware in memory or storage that the processor can execute to perform the function / process. It should also be noted that a processor that is "configured to" perform one or more functions or processes can correspond to a processor that is specially designed or "hardwired" to perform the functions or processes (e.g., an ASIC or FPGA design). Additionally, the phrase "at least one" preceding the recitation of one or more elements of an element (e.g., processor) configured to perform more than one function can correspond to one or more elements (e.g., processors) each performing a function and can also correspond to two or more of the elements (e.g., processors) each performing a different one of the one or more different functions.
[0223] Further, unless otherwise indicated, the term "adjacent" can mean that one element is in relatively close proximity to, but not in contact with, another element; or that the element is in contact with another part.
[0224] While exemplary embodiments of the disclosure have been described in detail, those skilled in the art should understand that various changes, substitutions, variations and improvements can be made to what is disclosed herein without departing from the spirit and scope of the disclosure in its broadest form.
[0225] The description in the present patent document should not be interpreted as implying any particular element, step, action or function is essential to the practice of the claims: the scope of the patent subject matter is defined only by the claims as allowed.
Claims
1. A computer-implemented method, the method comprising: - receiving input data (140) related to at least one device (142), wherein the input data (140) comprises incoming data batches X related to at least N separable classes, wherein n e 1,..., N, wherein the at least N separable classes are related to typical scenarios of the respective device, and wherein the at least N separable classes correspond to a correct operating state of the device and N-1 typical failure modes; - determining respective anomaly scores si,..., sn for the respective incoming data batches X related to the at least N separable classes using N anomaly detection models Mn (130), wherein the anomaly detection models are trained anomaly detection models, wherein the trained anomaly detection models are accomplished using a reference data set pre-provided by identifying typical scenarios related to typical variables or input data, wherein the typical scenarios comprise scenarios in which the respective device is working properly and scenarios in which the respective device is damaged; - applying the anomaly detection models Mn (130) to the input data (140) to generate output data (152) suitable for analyzing, monitoring, operating and / or controlling the respective device (142); - determining, for the respective incoming data batches X, a difference between the determined respective anomaly scores si,..., sn for the at least N separable classes on the one hand and the given respective anomaly scores Si,..., Sn of the N anomaly detection models Mn (130) on the other hand; and - if the respective determined difference between the determined respective anomaly scores for the at least N separable classes and the given respective anomaly scores of the N anomaly detection models is greater than a difference threshold: providing an alert (150) related to the determined difference to a user, the respective device (142) and / or an IT system connected to the respective device (142), wherein the method further comprises, if the determined difference is smaller than the difference threshold: embedding the N anomaly detection models Mn in a software application for analyzing, monitoring, operating and / or controlling the at least one device (142); and deploying the software application on the at least one device (142) or at an IT system connected to the at least one device (142) such that the software application can be used for analyzing, monitoring, operating and / or controlling the at least one device (142); if the determined difference is greater than the difference threshold: revising the respective anomaly detection model Mn (130) such that the determined difference using the respective revised anomaly detection model Mn (130) is smaller than the difference threshold; replacing the respective anomaly detection model Mn (130) with the respective revised anomaly detection model Mn (130) in the software application; and deploying the revised software application on the at least one device (142) or the IT system; if the anomaly detection model's correction takes more time than a duration threshold: replacing the deployed software application with a backup software application; and using the backup software application to analyze, monitor, operate and / or control the at least one device (142).
2. The computer-implemented method of claim 1, wherein, The input data (140) undergoes a distribution shift, which involves an increase of a determined difference.
3. The computer-implemented method according to any of the preceding claims, further comprising: - determining a distribution shift of the input data (140) if a second difference between the anomaly scores s1,...,sn of an earlier incoming data batch Xe and the anomaly scores s1,...,sn of a later incoming data batch X1 is greater than a second threshold; and - providing a report related to the determined distribution shift to a user, the respective device (142) and / or an IT system connected to the respective device (142) if the determined second difference is greater than the second threshold.
4. The computer-implemented method according to claim 1 or 2, further comprising: - assigning a training data batch Xt to the at least N separable classes of the anomaly detection model Mn (130); and - determining a given anomaly score S1,...,Sn for the at least N separable classes for N of the anomaly detection models Mn (130).
6. The computer-implemented method according to claim 1 or 2, further comprising for a plurality of interconnected devices (142):
5. The computer-implemented method of claim 1 or 2, wherein, N=1。 - embedding a respective N detection model Mn (130) into a respective software application for analyzing, monitoring, operating and / or controlling a respective interconnected device (142); - deploying the respective software application on the respective interconnected device (142) or on an IT system connected to the plurality of interconnected devices (142) such that the respective software application can be used for analyzing, monitoring, operating and / or controlling the respective interconnected device (142); - determining a respective difference of the respective anomaly detection model; and - if the respective, determined difference is greater than a respective difference threshold: providing an alert (150) related to the determined difference and the respective interconnected device (142) to a user, the respective device (142) and / or an automation system, the corresponding respective software application being used for analyzing, monitoring, operating and / or controlling the respective interconnected device (142) for the alert. The respective device (142) is a production machine, an automation device, a sensor, a production monitoring device, a vehicle or any combination thereof.
7. The computer-implemented method of claim 1 or 2, wherein, 8. An IT system, the system comprising: - A first interface (170) configured to receive input data (140) associated with at least one device (142), wherein the input data (140) includes incoming data batches X associated with at least N separable categories, wherein n∈1, ...,N, wherein the at least N separable categories are associated with typical scenarios of the corresponding devices, and wherein the at least N separable categories correspond to the correct operating state of the device and N-1 typical fault modes; - Computation unit (124), the computation unit being configured for - Use N anomaly detection models Mn (130) to determine corresponding anomaly scores s1, ..., sn for the corresponding incoming data batches X associated with the at least N separable categories, wherein the anomaly detection models are trained anomaly detection models, wherein the trained anomaly detection models are performed using a reference dataset provided in advance by identifying typical scenarios associated with typical variables or input data, wherein the typical scenarios include scenarios where the corresponding device is working normally and scenarios where the corresponding device is damaged; - The anomaly detection model Mn (130) is applied to the input data (140) to generate output data (152), which is suitable for analyzing, monitoring, operating and / or controlling the corresponding device (142). - For the corresponding incoming data batch X, determine the difference between the corresponding anomaly scores s1, ..., sn for the determination of the at least N separable categories and the corresponding anomaly scores S1, ..., sn of the N anomaly detection models Mn (130) given on the other hand. If the determined difference is less than the difference threshold: the N anomaly detection models Mn are embedded in a software application for analysis, monitoring, operation and / or control of the at least one device (142); and the software application is deployed on the at least one device (142) or on an IT system connected to the at least one device (142), so that the software application can be used to analyze, monitor, operate and / or control the at least one device (142). If the determined difference is greater than the difference threshold: modify the corresponding anomaly detection model Mn (130) such that the determined difference using the corresponding modified anomaly detection model Mn (130) is less than the difference threshold; replace the corresponding anomaly detection model Mn (130) with the corresponding modified anomaly detection model Mn (130) in the software application; and deploy the modified software application on the at least one device (142) or the IT system; If the correction of the anomaly detection model takes longer than the duration threshold: replace the deployed software application with a backup software application; and use the backup software application to analyze, monitor, operate, and / or control the at least one device (142); and - a second interface (172) configured for providing an alert related to the determined difference to a user, to the respective device (142) and / or to an IT system connected to the respective device (142) in case the respective determined difference between the determined respective anomaly score for the at least N separable categories and the given respective anomaly score of the N anomaly detection models is larger than a difference threshold.
9. A computer program product comprising computer program code which, when executed by a system (100), causes the system (100) to perform the method according to any one of claims 1 to 7.
10. The computer program product of claim 9, wherein, The system is an IT system.
11. A computer readable medium comprising computer program code which, when executed by a system (100), causes the system (100) to perform the method according to any one of claims 1 to 7.
12. The computer readable medium of claim 11, wherein, The system is an IT system.
Citation Information
Patent Citations
Fleet anomaly detection method
CN101354316A
Computer system and method for monitoring the technical state of industrial process systems
CN110431503A