Event Detection in Data Streams
The challenges in high-dimensional data stream processing and analysis are solved by using automatic encoders and reinforcement learning algorithms in the IoT environment, achieving efficient and adaptive event detection capabilities.
Patent Information
- Application Number
- CN201980101296.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-10-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2039-10-09
AI Technical Summary
In the Internet of Things (IoT) environment, prior art is difficult to efficiently process and analyze high-dimensional data streams from multiple devices, especially in the presence of drift in the data distribution, resulting in limited accuracy and efficiency of event detection.
The automatic encoder is used to condense the information in the data stream, and the hyperparameters of the automatic encoder are improved through reinforcement learning algorithms, combined with logic verification to drive reinforcement learning, and adaptive optimization of event detection is achieved.
Improves the accuracy and efficiency of detection of events in the data flow, and can adapt to different use cases in the absence of verification data, and realizes autonomous event detection and management.
Smart Images

Figure CN114556359B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method and a system for performing event detection on a data stream, and to a method and a node for managing an event detection process performed on a data stream. The present disclosure also relates to a computer program and a computer program product configured to perform, when run on a computer, a method for performing event detection and managing an event detection process. Background Art
[0002] The "Internet of Things" (IoT) refers to devices that can be connected to a communication network so that these devices can be remotely managed, and data collected or required by the devices can be exchanged between individual devices and between the devices and an application server. Such devices (examples of which can include sensors and actuators) are typically but not necessarily subject to strict limitations imposed by their operating environment or situation on processing power, storage capacity, energy supply, device complexity, and / or network connectivity, and can thus be referred to as constrained devices. Constrained devices typically connect to a core network via a gateway using short-range radio technology. The information collected from the constrained devices can then be used to create value in a cloud environment.
[0003] The IoT is widely regarded as an enabler of the digital transformation of business and industry. The IoT's ability to assist in monitoring and managing devices, environments, and industrial processes is a key component of achieving such digital transformation. For example, by deploying a large number of sensors to monitor a range of physical conditions and device states, substantially continuous monitoring can be achieved. The data collected by such sensors typically needs to be processed in real time and converted into information about the monitored environment representing available intelligence, and can be trigger actions performed within the monitored system. Data from individual IoT sensors can highlight specific, individual problems. However, even when evaluated by a person with expertise, the concurrent processing of data from many sensors (referred to herein as high-dimensional data) can highlight system behaviors that may not be apparent in a single reading.
[0004] The ability to highlight system behaviors can be particularly relevant in fields such as intelligent transportation and intelligent manufacturing, and the communication networks (including radio access networks) that serve them. In such fields, the large number of sensors and the large amount of data generated mean that expertise-based methods can quickly become cumbersome.
[0005] In the automotive and transportation domains, sensors are deployed to monitor the status of vehicles and their environment as well as the status of the passengers or cargo being transported. Status monitoring systems can improve the management of vehicles and their cargo by enabling predictive maintenance, rerouting, and expediting the delivery of perishable goods and optimizing transportation routes based on contractual requirements. Similarly, in the field of smart manufacturing, large amounts of data collected by industrial IoT devices can be used by status monitoring systems for equipment predictive maintenance, thereby reducing facility and equipment downtime and increasing production. In a radio access network (RAN), data collected from devices and sensors can be used to compute specific key performance indicators (KPIs) that reflect the current state and performance of the network. Rapid processing of data originating from RAN sensors helps identify issues that affect latency, throughput, and cause packet loss.
[0006] The domains discussed above represent examples of industrial and commercial activities where processes are used for data monitoring, which need to be as "hands-free" as possible so that these processes can run continuously and adapt to changes in the monitored environment and thus to changes in the monitored data. Such changes can include drifts in the data distribution. It will be understood that the requirements of any one domain may be very different from the requirements of another domain. Therefore, intelligence may be required to drive the monitoring process to meet the very different needs of different applications. IoT data analysis is thus not suitable for the design of machine learning (ML) models and preloading into monitoring nodes. The scope of use cases and application scenarios for IoT data analysis is wide, and providing a process that can provide IoT data analysis independently adaptable to different use cases is an ongoing challenge. Summary of the Invention
[0007] The object of the present disclosure is to provide methods, systems, nodes, and computer-readable media that at least partially address one or more of the challenges discussed above.
[0008] According to a first aspect of the present disclosure, there is provided a method for performing event detection on a data stream that includes data from a plurality of devices connected via a communication network. The method includes using an autoencoder to condense information in the data stream, where the autoencoder is configured according to at least one hyperparameter and events are detected from the condensed information. The method further includes generating an evaluation of the detected events based on the logical compatibility between the detected events and a knowledge base, and using a reinforcement learning (RL) algorithm to improve at least one hyperparameter of the autoencoder, where a reward function of the RL algorithm is computed based on the generated evaluation.
[0009] The above aspects of the present disclosure thus incorporate features of event detection from condensed data, using reinforcement learning to improve hyperparameters for condensed data, and using logical validation to drive the reinforcement learning. Well-known methods for improving model hyperparameters rely on validation data to trigger and drive learning. However, in many IoT and other systems, such validation data is simply not available. The above aspects of the present disclosure use an assessment of logical compatibility with a knowledge base to drive reinforcement learning to improve model hyperparameters. Contrary to data-based validation approaches, this use of logical validation means that the above methods can be applied to a wide range of use cases and deployments, including those where validation data is not available. Additionally, it will be understood that the resulting assessment of detected events is used to improve hyperparameters of an autoencoder for information condensation, rather than hyperparameters of an ML model that can be used for event detection itself. In this way, the process of data condensation is adjusted based on the quality of event detection that can be performed on the condensed data.
[0010] According to another aspect of the present disclosure, there is provided a system for performing event detection on a data stream that includes data from a plurality of devices connected via a communication network. The system is configured to use an autoencoder to condense information in the data stream, where the autoencoder is configured to detect events from the condensed information according to at least one hyperparameter, generate an assessment of the detected events based on the logical compatibility between the detected events and a knowledge base, and use a reinforcement learning (RL) algorithm to improve at least one hyperparameter of the autoencoder, where a reward function of the RL algorithm is calculated based on the generated assessment.
[0011] According to another aspect of the present disclosure, there is provided a method for managing an event detection process performed on a data stream that includes data from a plurality of devices connected via a communication network. The method includes receiving a notification of a detected event, where the event has been detected from information condensed from the data stream using an autoencoder configured according to at least one hyperparameter. The method further includes: receiving an assessment of the detected event, where the assessment has been generated based on the logical compatibility between the detected event and a knowledge base; and using a reinforcement learning (RL) algorithm to improve at least one hyperparameter of the autoencoder, where a reward function of the RL algorithm is calculated based on the generated assessment.
[0012] According to another aspect of the present disclosure, there is provided a node for managing an event detection process performed on a data stream, the data stream including data from a plurality of devices connected via a communication network. The node includes a processing circuit and a memory containing instructions executable by the processing circuit, whereby the node can be used to receive a notification of a detected event, wherein the event has been detected from information condensed from the data stream using an autoencoder configured according to at least one hyperparameter. The node can also be used to: receive an evaluation of the detected event, wherein the evaluation has been generated based on the logical compatibility between the detected event and a knowledge base; and use a reinforcement learning (RL) algorithm to improve at least one hyperparameter of the autoencoder, wherein a reward function of the RL algorithm is calculated based on the generated evaluation.
[0013] According to another aspect of the present disclosure, there is provided a computer program product including a computer-readable medium containing computer-readable code configured to cause, when executed by a suitable computer or processor, the computer or processor to perform a method according to any aspect or example of the present disclosure.
[0014] According to an example of the present disclosure, the above-mentioned knowledge base may include at least one of rules and / or facts, and their logical compatibility can be evaluated. At least one rule and / or fact can be generated from at least one of the following: the operating environment of at least some of the plurality of devices, the operating domain of at least some of the plurality of devices, the service protocols applied to at least some of the plurality of devices, and / or the deployment specifications applied to at least some of the plurality of devices. According to such an example, the knowledge base can be populated based on any one or more of the following: the physical environment in which the device is operating, the operating domain of the device (communication network operator, third-party domain, etc. and applicable rules), and / or the service level agreement (SLA) and / or the system and / or deployment configuration determined by the administrator of the device. Information about the above factors related to the device may be available even when the validation dataset of the device is not available.
[0015] According to an example of the present disclosure, the plurality of devices connected via the communication network may include a plurality of constrained devices. For the purposes of the present disclosure, a constrained device includes a device that complies with the definition specified for a "constrained node" in Section 2.1 of RFC 7228.
[0016] According to the definition in RFC 7228, a constrained device is a device in which "some of the characteristics of Internet nodes that are almost taken for granted at the time of writing cannot be implemented, usually due to cost constraints and / or physical limitations on characteristics such as size, weight, and available power and energy. The strict limitations on power, memory, and processing resources result in hard limits on state, code space, and processing cycles, making the optimization of energy and network bandwidth usage a major consideration in all design requirements. In addition, some Layer 2 services such as full connectivity and broadcast / multicast may be missing". Constrained devices are thus clearly distinguishable from server systems, desktops, laptops or tablets, and powerful mobile devices such as smartphones. Constrained devices can include, for example, machine-type communication devices, battery-powered devices, or any other device with the limitations discussed above. Examples of constrained devices can include, for example, sensors that measure temperature, humidity, and gas content in a room or when transporting and storing goods, motion sensors for controlling light bulbs, sensors that measure light that can be used to control blinds, heart rate monitors (continuously monitoring blood pressure, etc.) and other sensors for personal health, actuators, and connected electronic door locks. A constrained network correspondingly includes "a network in which some of the features of the link layer that are almost taken for granted in the Internet at the time of writing cannot be implemented", and more generally can include: a network including one or more constrained devices as defined above. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To better understand the present disclosure and to more clearly show how the present disclosure can be implemented, reference will now be made, by way of example, to the following drawings, in which:
[0018] Figure 1 is a flowchart showing a method for performing event detection on a data stream;
[0019] Figure 2a 、 Figure 2b and Figure 2c is a flowchart showing another example of a method for performing event detection on a data stream;
[0020] Figure 3 shows an autoencoder;
[0021] Figure 4 shows a stacked autoencoder;
[0022] Figure 5 illustrates event detection according to an example method;
[0023] Figure 6 shows a graphical user interface;
[0024] Figure 7aand Figure 7b shows an adaptive loop;
[0025] Figure 8 shows adaptive knowledge retrieval;
[0026] Figure 9 shows functions in a system for performing event detection on a data stream;
[0027] Figure 10 is a block diagram showing an exemplary embodiment of a method according to the present disclosure;
[0028] Figure 11 is a flowchart showing process steps in a method for managing an event detection process;
[0029] Figure 12 is a block diagram showing functional units in a node;
[0030] Figure 13a 、 Figure 13b and Figure 13c show information transformation when passing through an intelligent pipeline;
[0031] Figure 14 is a conceptual representation of an intelligent pipeline;
[0032] Figure 15 shows the composition of an exemplary IoT device;
[0033] Figure 16 shows the functional composition of an intelligent execution unit;
[0034] Figure 17 is a functional representation of an intelligent pipeline;
[0035] Figure 18 shows an IoT scenario; and
[0036] Figure 19 shows the orchestration of a method for performing event detection on a data stream within an IoT scenario. DETAILED DESCRIPTION
[0037] Artificial intelligence (AI), and specifically machine learning (ML), is widely regarded as the essence of autonomous solutions to meet industrial and commercial needs. However, in many AI systems, the deployment of ML models and the adjustment of model hyperparameters still highly rely on the input and expertise of human engineers. It will be understood that the "hyperparameters" of a model are parameters external to the model, and their values cannot be estimated from the data processed by the model, but still determine how the model learns its internal parameters. Model hyperparameters can be adjusted for a given problem or use case.
[0038] In the IoT ecosystem, it is very important that machines can continuously learn and retrieve knowledge from data streams to support industrial automation (also known as Industry 4.0). A high-level autonomous intelligent system can minimize the need for human engineer input and insights. However, the IoT deployment environment is constantly changing, and data drift can occur at any time, rendering existing artificial intelligence models ineffective. Currently, this problem is almost always manually solved by engineer intervention to readjust the model. Different from many other AI scenarios in highly specified fields including, for example, machine vision and natural language processing, due to the broadness of the IoT application domain and the heterogeneity of the IoT environment, it is difficult to find a single learning model suitable for all IoT data. Therefore, the adaptive ability to learn and retrieve knowledge from IoT data is highly anticipated to address these challenges. End-to-end automation is also expected to minimize the need for human intervention.
[0039] Existing technologies related to retrieving intelligence from IoT data cannot provide such automation and adaptability.
[0040] When considered in a deployment environment with automation requirements, traditional machine learning-based knowledge retrieval solutions have several disadvantages:
[0041] - The model is pre-built before adding relevant hardware for deployment. In many cases, such a model still highly depends on human engineer intervention to update the model to process data streams in real time.
[0042] - The deployment of AI models, especially hyperparameter tuning, is not automated but relies on human intervention
[0043] - The IoT environment generates a wide variety and highly dynamic data. Therefore, pre-loaded static models can easily lose accuracy.
[0044] - The above limitations make the development of a single model for extracting data from IoT deployments highly challenging.
[0045] The following criteria thus represent the desired features of methods and systems that can help retrieve intelligence from IoT data:
[0046] - The data processing algorithm should be dynamic so that the input size and model shape can be adjusted according to specific requirements;
[0047] - The algorithm itself should scale based on the number of data processing nodes, the number of data sources, and the amount of data.
[0048] - The analysis should be performed online to enable fast event detection and fast prediction;
[0049] - Dependence on prior domain knowledge, including training labels, validation data, and useful data models, should be minimized as such knowledge is often unavailable in IoT systems.
[0050] When considering the above-expected criteria, recent attempts at automated knowledge retrieval have shown significant limitations.
[0051] For example, highly automated solutions that can be implemented close to where data is generated are extremely rare. Most automated solutions send all data to the cloud, where knowledge is extracted before downloading conclusions for appropriate actions. Additionally, most existing solutions still require a large amount of manual intervention, cannot handle dynamically evolving data, and lack flexibility and scalability in the core algorithms. International Patent Application PCT / EP2019 / 066395 discloses a method and system seeking to overcome some of the above challenges. This disclosure seeks to enhance various aspects of the solution proposed in PCT / EP2019 / 066395, specifically the autonomous ability to adapt to data, monitored systems, and environmental changes.
[0052] Aspects of this disclosure thus provide an automated solution that implements adaptability in methods and systems available for retrieving intelligence from real-time data streams. Examples of this disclosure provide the possibility of automating the operating loop to adjust hyperparameters of a model such as a neural network according to changes in a dynamic environment (without data labels for training). Examples of this disclosure are thus adaptive and can be deployed in a wide variety of use cases. Examples of this disclosure minimize dependence on domain expertise.
[0053] The adaptability of examples of this disclosure is based on an iterative loop built on a reinforcement learning agent and a logic validator. Feature extraction allows reducing dependence on domain expertise. Based on a knowledge base that can be populated without specific domain knowledge, examples of this disclosure apply logical validation of results. Such a knowledge base can be constructed from data including environmental, physical, and business data and can thus be regarded as a "common sense" check on whether the results are consistent with the known monitored system and / or environment and the business requirements of a specific deployment. Such requirements can be specified, for example, in a service level agreement (SLA) when applied to a communication network. Examples of this disclosure provide a solution without any specific model; model hyperparameters are adjusted through a reinforcement learning loop.
[0054] Figure 1 and FIG. 2 is a flowchart showing methods 100, 200 for performing event detection on a data stream according to an example of this disclosure, the data stream including data from multiple devices connected via a communication network. Figure 1 and FIG. 2 provides an overview of these methods, showing how the functions discussed above are implemented. Then refer toFigures 3 to 8 Discuss each method step in detail, including implementation details.
[0055] First, refer to Figure 1 , method 100 includes: using an autoencoder in a first step 110 to condense information in a data stream, where the autoencoder is configured according to at least one hyperparameter. The method then includes: detecting an event from the condensed information in step 120; and generating an evaluation of the detected event based on the logical compatibility between the detected event and a knowledge base in step 130. Finally, in step 140, the method includes using a reinforcement learning RL algorithm to improve at least one hyperparameter of the autoencoder, where a reward function of the RL algorithm is calculated based on the generated evaluation.
[0056] The data stream can include data from multiple devices. Such devices can include devices for environmental monitoring, devices for facilitating intelligent manufacturing, devices for facilitating intelligent vehicles, and / or devices in a communication network such as a radio access network (RAN) or devices connected to the communication network. In some embodiments, the data stream can include data from network nodes, or can include data collected by software and / or hardware. Examples of specific devices can include temperature sensors, audiovisual devices such as cameras, video devices or microphones, proximity sensors, and device monitoring sensors.
[0057] The data in the data stream can include real-time or near-real-time data. In some examples, method 100 can be executed in real time such that there is minimal or little latency between the collection and processing of the data. For example, in some examples, method 100 can be executed at a rate comparable to the rate at which the data in the data stream is generated such that no significant backlog of data starts to accumulate. Multiple devices are connected via a communication network. Examples of communication networks can include radio access network (RAN), wireless local area network (WLAN or WIFI), and wired networks. In some examples, the devices can form part of a communication network, such as part of a RAN, part of a WLAN, or part of a WIFI network. In some examples, such as in intelligent manufacturing or intelligent vehicle deployments, the devices can communicate via the communication network.
[0058] Referring to step 110, it will be understood that an autoencoder is a machine learning algorithm that can be used to condense data. The autoencoder is trained to take a set of input features and reduce the dimensionality of the input features while minimizing information loss. Training the autoencoder is generally an unsupervised process, and the autoencoder is divided into two parts: an encoding part and a decoding part. The encoder and decoder can include, for example, a deep neural network including neuron layers. If the decoder can recover the original data stream with tolerable data loss, the encoder has successfully encoded or compressed the data. Training can include minimizing a loss function that describes the difference between the input (original) and output (decoded) data. Training the encoder part thus involves optimizing the data loss of the encoder process. An autoencoder can be considered for condensing data (e.g., as opposed to simply reducing dimensionality) because the essential or salient features in the data are not lost. It will be understood that the autoencoder used in accordance with method 100 can actually include multiple autoencoders that can be configured to form a distributed stacked autoencoder, as discussed in further detail below. A stacked autoencoder includes two or more individual autoencoders that are arranged such that the output of one autoencoder is provided as input to another autoencoder. In this way, the autoencoder can be used to sequentially condense the data stream, reducing the dimensionality of the data stream in each autoencoder operation. A distributed stacked autoencoder includes a stacked autoencoder implemented across multiple nodes or processing units. Thus, the distributed stacked autoencoder provides an extended way to condense information along the intelligent data pipeline. Additionally, since each autoencoder residing in each node (or processing unit) is linked to each other, the distributed stacked autoencoder can be used to grow according to the information complexity of the input data size.
[0059] Referring to step 120, the detected events can include any data readings of interest, including, for example, statistically anomalous data points. In some examples, the events can be related to anomalies. In other examples, the events can be related to system performance metrics (e.g., key performance indicators (KPIs)) and can indicate anomalous or undesired system behavior. Examples of events can vary significantly depending on the specific use case or domain in which the examples of the present disclosure can be implemented. In the field of smart manufacturing, examples of events can include temperature, humidity, or pressure readings that exceed the operating window for such readings, which can be configured manually or established based on historical readings of such parameters. Anomalous readings from temperature, pressure, humidity, or other sensors can indicate a failure of a particular device, or that the process is no longer operating within the optimal parameter range, etc. In the field of communication networks, example events can include KPI readings that are outside the window of expected system behavior or KPI readings that fail to meet the targets specified in a service level agreement or other business protocol. Examples of such KPIs for a radio access network can include average and maximum cell throughput in downloads, average and maximum cell throughput in uploads, cell availability, total upload traffic, etc.
[0060] Referring to step 140, reinforcement learning is a technique for developing self-learning software agents that can learn and optimize policies for controlling a system or environment (e.g., the autoencoder of method 100) based on observed system states and a reward system tailored to achieve a specific goal. In method 100, the goal can include improving the assessment of detected events and thus increasing the accuracy of event detection. When executing a reinforcement learning algorithm, the software agent establishes the state St of the system. Based on the state of the system, the software agent selects an action to perform on the system, and once the action has been executed, it receives a reward rt resulting from the action. The software agent selects actions based on the system state with the aim of maximizing the expected future reward. The reward function can be defined such that a greater reward is received for actions that cause the system to enter a state close to the target end state of the system, which is consistent with the overall goal of the entity managing the system. In the case of method 100, the target end state of the autoencoder can be a state in which the hyperparameters are such that event detection in the condensed data stream has reached a desired accuracy threshold, as indicated by the resulting assessment of the detected events.
[0061] Figures 2a to 2cA flowchart is shown which depicts the process steps in another example of a method 200 for event detection performed on a data stream that includes data from multiple devices connected via a communication network. The steps of method 200 illustrate one example way in which the steps of method 100 can be implemented and supplemented to achieve the functions discussed above and additional functions. Method 200 can be performed by multiple devices that cooperate to implement the different steps of the method. The method can be managed by a management function or node that can orchestrate and coordinate certain method steps and can facilitate the scaling of the method to accommodate changes in the number of devices generating data, the amount of data generated, the number of nodes, the functions or processes available to perform different method steps, etc.
[0062] First, referring to Figure 2a , in a first step 202, the method includes collecting one or more data streams from multiple devices. As Figure 2a shown, in example method 200, the devices are constrained devices or IoT devices, although it will be understood that method 200 can be used for event detection in data streams generated by devices other than constrained devices. The devices are connected via a communication network which can include any type of communication network as described above. In step 204, method 200 includes transforming and aggregating the collected data before accumulating the aggregated data, and in step 206, dividing the accumulated data stream into multiple consecutive windows, each window corresponding to a different time interval.
[0063] In step 210, method 200 includes using a distributed stacked autoencoder to condense the information in the data stream, where the autoencoder is configured according to at least one hyperparameter. The at least one hyperparameter can include a time interval associated with a time window, a scaling factor, and / or a layer reduction rate. The distributed stacked autoencoder can be used to condense the information in the windowed data according to the time windows generated in step 212. This step is also referred to as feature extraction as the data is condensed such that the most relevant features are retained. As shown in step 210, using a distributed stacked autoencoder can include using an unsupervised learning (UL) algorithm to determine the number of layers in the autoencoder and the number of neurons in each layer of the autoencoder based on at least one of the parameters associated with the data stream and / or at least one of the at least one hyperparameters. The parameters associated with the data stream can include, for example, at least one of the following: the data transmission frequency associated with the data stream and / or the dimensionality associated with the data stream. A complete discussion of different equations for calculating the number of layers and the number of neurons per layer is provided below. The UL process can implement the training discussed above where the encoding loss is minimized by comparing the original input data with the decoded output data.
[0064] As Figure 2aAs shown, the process of using a distributed stacked autoencoder may include: dividing a data stream into one or more data sub-streams in step 210a, using different autoencoders of the distributed stacked autoencoder to condense the information in each corresponding sub-stream in step 210b, and providing the condensed sub-streams to another autoencoder in another layer of the hierarchical structure of the stacked autoencoder in step 210c.
[0065] In step 212, method 200 includes: accumulating the condensed data in the data stream over time before detecting an event (now referring to Figure 2b ) from the condensed information in step 220. As shown in step 220, this may include comparing different portions of the accumulated condensed data. In some examples, the cosine difference may be used to compare different portions of the accumulated condensed data. In some examples, as Figure 2b shown, detecting an event may further include: generating labels for a training data set using at least one event detected by comparing different portions of the accumulated condensed data in step 220a, where the training data set includes the condensed information from the data stream. Detecting an event may then further include training a supervised learning (SL) model using the training data set in step 220b, and detecting an event from the condensed information using the SL model in step 220c. In some examples, only those detected events with a suitable evaluation score (e.g., a score above a threshold) may be used to generate labels for the training data set, as discussed in further detail below.
[0066] In step 230, method 200 includes generating an evaluation of the detected event based on the logical compatibility between the detected event and a knowledge base. In some examples, the evaluation score may also be generated based on error values generated during at least one of condensing the information in the data stream or detecting an event from the condensed information. Further discussion of the machine learning components for evaluating the detected event is provided below.
[0067] As Figure 2b shown, generating an evaluation of the detected event based on the logical compatibility between the detected event and a knowledge base may include: converting the parameter values corresponding to the detected event into a logical assertion in step 230a, and evaluating the compatibility of the assertion with the content of the knowledge base in step 230b, where the content of the knowledge base includes at least one of rules and / or facts. The knowledge base may contain one or more rules and / or facts, which may be generated from at least one of the following:
[0068] the operating environment of at least some of the plurality of devices;
[0069] the operating domain of at least some of the plurality of devices;
[0070] A service agreement applicable to at least some of a plurality of devices; and / or
[0071] A deployment specification applicable to at least some of a plurality of devices.
[0072] Thus, a knowledge base can be populated based on the physical environment in which the device is operating, the device's operational domain (network operator, third-party domain, etc. and applicable rules), and / or a service agreement such as an SLA, and / or system / deployment configurations determined by a device administrator. As discussed above, such information may be available in the case of an IoT deployment event when a complete validation data set is not available.
[0073] Step 230b of evaluating the compatibility of the assertion with the content of the knowledge base can include performing at least one of incrementing or decrementing an evaluation score for each logical conflict between the assertion and a fact or rule in the knowledge base. A detection event indicating multiple logical conflicts with the knowledge base is unlikely to be a correctly detected event. Evaluating events in this way and using the evaluation to refine the model hyperparameters for condensing data streams can thus result in condensing the data in a way that maximizes the potential for accurate event detection.
[0074] Now referring to Figure 2c , method 200 further includes using a reinforcement learning RL algorithm to refine at least one hyperparameter of the autoencoder, wherein a reward function of the RL algorithm is calculated based on the generated evaluation. As shown, this can include using the RL algorithm to experiment with different values of at least one hyperparameter and determining the value of at least one hyperparameter associated with the maximum value of the reward function. Steps 240a to 240d illustrate how this can be achieved. In step 240a, the RL algorithm can establish a state of the autoencoder, where the state of the autoencoder is represented by the value of at least one hyperparameter. In step 240b, the RL algorithm selects an action to perform on the autoencoder based on the established state, where the action is selected from a set of actions including incrementing and decrementing the value of at least one hyperparameter. In step 240c, the RL algorithm causes the selected action to be performed on the autoencoder, while in step 240d, the RL algorithm calculates the value of the reward function after performing the selected action. Action selection can be driven by a policy seeking to maximize the value of the reward function. Since the reward function is based on the generated evaluation of the detected events, maximizing the value of the reward function will seek to maximize the evaluation score of the detected events and thus maximize the accuracy of detecting events.
[0075] In step 242, method 200 includes updating the knowledge base to include detected events that are logically compatible with the knowledge base. This can include adding assertions corresponding to the detected events as rules to the knowledge base. In this way, correctly detected events can contribute to the knowledge used to evaluate future detected events. Thus, conflicts with previously correctly detected events can reduce the evaluation score of future detected events. Finally, in step 244, method 200 includes exposing the detected events to the user. This can be implemented in any practical way suitable for a particular deployment or use case. The detected events can be used to trigger actions in one or more devices and / or the system or environment in which the devices are deployed.
[0076] The above methods 100 and 200 provide an overview of how aspects of the present disclosure can implement adaptive and autonomous event detection, which can be used to obtain actionable intelligence from one or more data streams. These methods can be implemented in a range of different systems and deployments, and aspects of these systems and deployments are now introduced. Then, how to implement the steps of the above methods will be discussed in detail.
[0077] Systems or deployments in which the methods discussed above can be run can include the following elements:
[0078] 1) One or more devices, which can be constrained devices such as IoT devices. Each device can include sensors and sensor units for collecting information. This information can relate to the physical environment, the operating state of the device, physical, electrical, and / or chemical processes, etc. Examples of sensors include: environmental sensors, including temperature, humidity, air pollution, acoustics, sound, vibration, etc.; sensors for navigation, such as altimeters, gyroscopes, internal navigators, and magnetic compasses; optical devices, including light sensors, thermal imagers, photodetectors, etc.; and many other sensor types. Each device can also include a processing unit for processing sensor data and sending the results via a communication unit. In some examples, the processing unit of the device can assist in performing some or all of the method steps discussed above. In other examples, the device can simply provide the data of the data stream, while the method steps are performed in a distributed manner in other functions, nodes, and elements. Each device can also include a communication unit for sending the sensor data provided by the sensor unit. In some examples, the device can send sensor data from the processing component unit.
[0079] 2) One or more computing units, which can be implemented in any suitable device such as a gateway or other node in a communication network. The computing units can additionally or alternatively be implemented in a cloud environment. Each computing unit can include a processing unit and a communication unit, where the processing unit is used to implement one or more of the above method steps and appropriately manage communication with other computing units. The communication unit can receive data from heterogeneous radio nodes and (IoT) devices via different protocols, exchange information between intelligent processing units, and expose data and / or insights, detected events, conclusions, etc. to other external systems or other internal modules.
[0080] 3) A communication agent that facilitates the collection of device sensor data and the exchange of information between entities. The communication agent can include, for example, a message bus, a persistent storage unit, a peer-to-peer communication module, etc.
[0081] 4) A repository for the knowledge base.
[0082] The steps of methods 100, 200 can be implemented via the collaborative elements that are the separate intelligent execution units discussed above, and the separate intelligent execution units include:
[0083] "Data input" (data source): Defines how data is retrieved. Depending on the specific method step, the data can vary, and this data includes sensor data, monitoring data, aggregated data, feature matrices, reduced features, distance matrices, etc.
[0084] "Data output" (data sink): Defines how data is sent. Depending on the specific method step, the data can vary as discussed above for the input data.
[0085] "Mapping function": Specifies how data should be accumulated and preprocessed. This can include complex event processing (CEP) functions such as accumulate (acc), window, average, last, first, standard deviation, sum, minimum, maximum, etc.
[0086] "Transformation": Refers to any type of execution code required to perform operations in the method step. Depending on the specific operation of the method step, the transformation operation can be a simple protocol conversion function, an aggregation function, or an advanced algorithm.
[0087] "Interval": Depending on the protocol, it can be appropriate to define the window size to perform the requested calculations.
[0088] The intelligent execution units discussed above can be connected together to form an intelligent (data) pipeline. It will be understood that in the proposed method, each step can be considered as a computational task for a certain independent intelligent execution unit, and the automated composition of the integrated set of intelligent execution units constitutes the intelligent pipeline. The intelligent execution units can be deployed in a "click and run" manner via software, with a configuration file for initialization. The configuration of the data processing model can be adapted after initialization according to the method described herein. The intelligent incentive units can be distributed across multiple nodes for resource orchestration, maximizing the utilization and performance of the nodes. Such nodes can include devices, edge nodes, fog nodes, network infrastructure, cloud, etc. In this way, the existence of a central failure point is also avoided. Using a participant-based architecture, intelligent execution units can be easily created in batches using an initial configuration file. Implementing in this way contributes to the scalability of the method proposed herein and its deployment automation.
[0089] It will be understood that the deployment of a distributed cluster can be automated in the sense of providing an initial configuration file and then "clicking to run". The configuration file provides general configuration information of the software architecture. This file can be provided to a single node (the root participant) once to create the entire system. Then the computational model can be adjusted based on the shape of the data input. It will also be understood that some steps of the method can be combined with other steps to be deployed as a single intelligent execution unit. The steps of methods 100, 200 together form an interactive autonomous loop. The loop can continuously adjust the configuration of the algorithm and the model in response to a dynamic environment.
[0090] Some steps of methods 100, 200 introduced above are now discussed in more detail. It will be understood that the following details relate to different examples and embodiments of the present disclosure.
[0091] Step 202: Collect data stream (collect relevant sensor data or any relevant data in the stream)
[0092] IoT is a data-driven system. The purpose of step 202 is to retrieve and collect the raw data from which actionable intelligence is to be extracted. This step can include collecting available data from all devices that provide data to the data stream. Step 202 can integrate multiple heterogeneous devices, which may have multiple different communication protocols, multiple different data models, and multiple different serialization mechanisms. This step may therefore require the system integrator to understand the data payload (data model and serialization) in order to collect and unify the data format.
[0093] There are various combinations of protocols, data models, and serialization of data that can provide a unified way to seamlessly collect sensor data from multiple sources in the same system. In this step, the "sources" and "transformation functions" discussed earlier with reference to the intelligent execution unit that performs these steps can be particularly useful for generating consistent and homogeneous data. For example, in a system where IoT device "X" sends raw data in a specific format and IoT device "Y" sends JSON data, step 202 will allow the data from device "X" to be transformed into JSON. Subsequent processing units can manage the data seamlessly without additional data transformation. The way a certain unit forwards data to the next unit is defined in the "sink". In the above example, the "sink" can be specified to ensure that the output data is provided in JSON format.
[0094] Step 204: Transform and aggregate the data in the stream;
[0095] It will be understood that step 202 can be performed in a distributed and parallel manner, and step 204 can thus provide a central aggregation to collect all sensor data, based on which a data frame in high-dimensional data can be created.
[0096] The data can be aggregated based on system requirements. For example, if environmental analysis or analysis of a certain business process requires data collected from specific multiple distributed sensors / data sources, then all the data collected from these sensors and sources should be aggregated. In many cases, the data collected will be sparse; as the number of categories within the collected data increases, the output can ultimately become a high-dimensional sparse data frame.
[0097] It will be understood that the number of intelligent execution units, and in some examples, the number of physical and / or virtual nodes that perform the processing of steps 202 and 24 can vary according to the number of devices from which data is to be collected and aggregated and the amount of data generated by these devices.
[0098] Step 206: Accumulate high-dimensional data and generate a window
[0099] To trigger the condensation of data and establish / improve a suitable model, a time window of a specific size should be defined within which the data can be accumulated. This step groups small batches of data according to the window size. The size of the window can be specific to a particular use case and can be configured according to different requirements. Depending on the requirements, the data can be accumulated in memory or a persistent storage device. According to the description of the intelligent execution unit through which examples of the present disclosure can be implemented, this step uses a "map function" to implement. In some examples, the operation of this step can accumulate the data into an array only using the "map function".
[0100] Step 110 / 210: (Using hyperparameters optimized from previous iterations of the method), establish a deep autoencoder-based model and perform feature extraction / information condensation on the data for each time window
[0101] Once data has been accumulated over a sufficient number of windows, feature extraction / data condensation can be triggered. Initially, a deep autoencoder is constructed based on the accumulated data. Feature extraction is then performed by applying the autoencoder model to the accumulated data, with the data being processed by the stacked autoencoders of the deep encoder. Figure 3 A single autoencoder 300 is shown. As described above, the autoencoder includes an encoder section 310 and a decoder section 320. High-dimensional input data 330 is input into the encoder section, while condensed data 340 from the autoencoder section 310 is output from the autoencoder 300. The condensed data is fed into the decoder section 320 that reconstructs the high-dimensional data 350. The comparison between the input high-dimensional data 330 and the reconstructed high-dimensional data 340 is used to learn the parameters of the autoencoder model.
[0102] Figure 4 A stacked autoencoder 400 is shown. The stacked autoencoder 400 includes a plurality of individual autoencoders, each of which outputs its condensed data for input into another autoencoder, thereby forming a hierarchical arrangement based on which the data is continuously condensed.
[0103] Each sliding window defined in the previous step outputs data frames in chronological order at a defined interval. In step 110 / 210, feature extraction can be performed in two dimensions:
[0104] (a) Compress the information carried by the data in the time dimension. For example, deployed sensors may send data every 10 milliseconds. A sliding window with a duration of 10 seconds will thus accumulate 1000 data items. Feature extraction can be achieved by reducing the time length of the data frames to summarize the information of the data frames and provide comprehensive information.
[0105] (b) Condense the information carried by the data in the feature dimension. Due to the high-dimensional features of sensor data, the sensor data collected from IoT sensors may be very complex. In many cases, it is almost impossible for even domain experts to understand the operating state of the entire IoT system by looking at the collected sensor data sets. Therefore, high-dimensional data can be processed to extract the most important features and reduce any unnecessary information complexity. The features extracted from one or a group of deep autoencoders can be the input to another deep autoencoder.
[0106] When the output of one or several deep autoencoders is used as the input to another deep autoencoder, this will form as discussed above and as Figure 4The stacked deep autoencoder shown. The stacked deep autoencoder provides an extended way to concentrate information along the intelligent data pipeline. The stacked deep autoencoder can also grow according to the information complexity of the input data dimension and can be fully distributed to avoid computational bottlenecks. As output, step 110 / 210 sends the concentrated and extracted information carried by the collected data. A large amount of high-dimensional data thus becomes more manageable for subsequent calculations.
[0107] As discussed above, each encoder part and decoder part of the autoencoder can be implemented using a neural network. Having too many layers in the neural network will introduce unnecessary computational burden and cause computational latency, while having too few layers risks weakening the expressive power of the model and may affect performance. The optimal number of layers for the encoder can be obtained using the following formula:
[0108] Number_of_hidden_layers
[0109] =int(Sliding_windows_interval
[0110] *data_transmitting_frequency / (2*scaling_factor
[0111] *number_of_dimensions))
[0112] where the scaling factor is a configurable hyperparameter that generally describes how the shape of the model will change from short and wide to long and narrow.
[0113] The deep autoencoder can also introduce a hyperparameter layer_number_decreasing_rate (e.g., 0.25) to create the output size of each layer. For each layer in the encoder:
[0114] Encoder_Number_of_layer_(N+1)=
[0115] int(Encoder_Number_of_layer_N*(1-
[0116] Layer_Number_Decreasing_Rate));
[0117] The scale factor, time interval, and layer number decreasing rate are all examples of hyperparameters that may have been optimized during previous iterations of method 100 and / or 200.
[0118] The number of each layer in the decoder corresponds to the number of each layer in the encoder. The purpose of the unsupervised learning process for constructing a stacked deep autoencoder is to concentrate information by extracting features from high-dimensional data. The computational accuracy / loss of the autoencoder can be performed using K-fold cross-validation, where, unless otherwise configured, the number K defaults to 5. The validation calculation can be performed inside each mapping function. For a stacked deep autoencoder, then the validation is performed on each individual run of the autoencoder.
[0119] Step 212: Accumulate the extracted features (based on the optimized hyperparameters from the earlier iterations of the method)
[0120] Step 212 can be implemented using the "mapping function" of the intelligent execution unit as described above. The mapping function accumulates the reduced features over a certain period of time or a certain sample size. In this step, each mapping function is a Figure 4 deep autoencoder as shown. As discussed in the previous steps, feature extraction is chained and can quickly iteratively approach the data source. The extracted features are accumulated as the sliding window moves. For example, if the feature extraction described in the previous step is performed every 10 seconds, it can be envisioned that the system needs to perform anomaly detection over a one-hour time period during this accumulation step. The time range for buffering samples can be set to 60 seconds, and 360 data will be accumulated through monitoring and features will be extracted from the data generated every 10 milliseconds. Such accumulation lays the foundation for concentrating information in the time dimension. As can be seen from the example, after the processing of the stacked deep autoencoder, the raw data generated every 10 milliseconds for each data block is concentrated into data generated every 1 second for every 6 data blocks.
[0121] Steps 120 / 220: Perform event detection - Retrieve insights from the concentrated data / extracted features (based on the optimized hyperparameters from the earlier iterations of the method)
[0122] Steps 120 / 220 analyze the accumulated reduced feature values and compare them in order to evaluate in which time windows anomalies occur. For example, time slots with a large distance from other time slots can be suggested as anomalies in the cumulative time. This step is based on the accumulation of the previously extracted features. Event detection can be carried out in two stages: The first stage detects events based on distance calculation and comparison. These events are logically verified in steps 130 / 230, and then the labels of the training data set are assembled using those events that pass the logical verification. The training data set is used to train a supervised learning model, which performs event detection on the concentrated and accumulated data in the second stage of event detection. The two stages of event detection are as Figure 5 shown.
[0123] In a first stage, for the results from step 110 / 210 (and 212, if performed), distance measurements can be used to calculate pairwise distances between elements in the cumulative output of feature extraction. Step 110 / 210 is represented by a single deep autoencoder 510, although it will be understood that in many embodiments, step 110 / 210 can be performed by a stacked deep autoencoder as discussed above. The output of the autoencoder 510 is input to a distance calculator 520. The distances calculated in the distance calculator 520 can be cosine distances, and the pairwise distances can accordingly form a distance matrix. Since the Markov Chain Monte Carlo (MCMC) method is used to reconstruct the extracted features, the output is randomly generated, which maintains the same probability distribution characteristics. The distribution is unknown, which means that measuring distances (e.g., Euclidean distances) may not be meaningful in many cases. The angle between vectors is of most interest, and thus cosine distance may be the most effective measurement.
[0124] If in an example, the previous steps have accumulated N outputs, labeled {F 0 , F 1 , F 2 , … F n-1}; then the distance matrix can be calculated as follows:
[0125] Table 1: Forming the distance matrix
[0126]
[0127]
[0128] In the field named "Distance Avg", for each extracted feature, the average distance of it from the remaining extracted features in the same buffered time window is calculated. The calculated result is written from the cache to the storage device, which can facilitate visualization.
[0129] The events detected through distance comparison in the distance calculator are then passed to a logic validator 530 to evaluate compatibility with the content of the knowledge base 540. This step will be discussed in more detail below. Then the verified events are used to generate labels for the training dataset. Thus, the condensed data corresponding to the events detected through distance calculation and comparison is labeled as corresponding to the events. In the second stage of event detection, this labeled data in the form of a training dataset is input to a supervised learning model, implemented as, for example, a neural network 550. Thus, the training data is used to train the neural network 550 to detect events in the condensed data stream. It will be understood that the training of the neural network 550 can be delayed until a training dataset of appropriate size has been generated through event detection using distance calculation.
[0130] Step 130 / 230: Generate an evaluation - logical verification of the detected events to detect conflicts with the knowledge base and provide verification results for reinforcement learning
[0131] This step performs a logical verification on the events detected in the first stage of event detection to rule out events that are detected to have a logical conflict with the common sense that is already assembled into the knowledge base within the domain. The evaluation score of an event can reflect the number of logical conflicts with the content of the knowledge base. A penalty table can be created by linking the logical verification results to the current configuration and can drive a reinforcement learning loop through which the hyperparameters of the model used for data condensation and optional event detection can be improved.
[0132] Based on the available information about the device generating the data, the physical environment of the device, the operation domain of the device, the business protocols associated with the device, the deployment priorities or rules on how the deployment should operate, etc., the logical verification performs a logical conflict check on the detected events against the facts and / or rules that have already been assembled in the knowledge base. This information can be filled without the need for domain - specific expertise. For example, the outdoor temperature in Stockholm in June should be above 0 degrees. It will be understood that the knowledge base is thus very different from a validation dataset, which is typically used to check the event detection algorithms in existing event detection solutions. A validation dataset can compare the detected events with the ground truth events and can only be assembled with a large amount of input from domain experts. Also, in many IoT deployments, the data for the validation dataset is simply not available. In contrast to the validation dataset, the present disclosure includes an evaluation of the logical conflict between the detected events and the knowledge base, which includes facts and / or rules generated from information that is easily obtainable even for those without expertise in the relevant domain. The logical verification can be used to filter out false detected events as well as populate a penalty table that describes the number of logical conflicts for a given detected event. The penalty table can be used as a handle for adaptive model improvement by driving reinforcement learning, as described below.
[0133] The logical verifier can include an inference engine that verifies, based on second - order logic, whether the detected anomalies have any logical conflict with the existing knowledge that is relevant to a given deployment and populated into the knowledge base. The detection events from the previous step are streamed to the processor implementing the logical evaluation in the form of assertions.
[0134] For example: Consider a system defined by a set of parameters labeled S = {a, b, c, d... k}, and suspicious behavior can be detected by satisfying {a = XY, b = XXY, c = YY, d = YXY... k = Y}. This statement can be easily converted into the following assertion: Assertion = ({a = XY, b = XXY, c = YY, d = YXY... k = Y} = 0). This assertion can be verified by running a logic-based engine based on an existing knowledge base. The knowledge base can include two parts: a fact base and a rule base. Both of these parts can provide a graphical user interface (GUI) for interacting with the user. In the present disclosure, the proposed pattern for the knowledge base is based on the assumption of a closed space, which means that logical verification is performed using the available knowledge that can be called "common sense", and thus can be easily obtained from the known characteristics of the device, the environment of the device, and / or any business or operational principles or protocols. The population of the knowledge base can be guided by the GUI as shown in Figure 6 shown.
[0135] The semantic patterns "Subject + Verb + Object" and "Subject + Copula + Predicative" can be mapped to the fact space, as shown in Figure 6 shown. For example, in an intelligent manufacturing use case, the production line includes a robotic arm designed to rotate and whose temperature cannot exceed 100 degrees Celsius. This "common sense" knowledge about how the production line operates can be transformed into a fact space for performing the following logical verification: "The robotic arm that is rotating is healthy" can be semantically mapped to (Robotics rotate == True)=>(Robotics health == True), while "The robotic arm above 100 degrees Celsius is unhealthy" can be semantically mapped to (Robotics temp ≥ 100)=>(Robotics health == False). In this case, if event detection generates an anomaly represented as the following assertion: (Robotics rotate == True) V (Robotic temp ≥ 100)=>(Robotics health == False), the logical verification will provide an "error" judgment because this assertion represents a logical conflict with the facts that the rotating robotic arm is healthy and that a healthy robotic arm cannot exceed 100 degrees Celsius. For adaptive training performed via reinforcement learning, in subsequent steps, each "error" judgment indicates a logical conflict, and the number of logical conflicts provides the number of facts and / or rules that have a direct or indirect conflict with the provided assertion.
[0136] Logical verification can be used to generate a JSON document that describes the current hyperparameters of the stacked autoencoder and the number of detected logical conflicts, as follows. For example, the following JSON can define the input configuration for subsequent adaptive learning steps.
[0137]
[0138] Step 242: Update the verified insights into the knowledge base to support supervised learning;
[0139] In step 212, the recommended assertions after verification can be updated to the existing knowledge base. The assertions can be updated using, for example, the JSON format. The data repository can include within it a knowledge base and data labels that are generated from the verified detection events and are used to train a supervised learning model for second-stage event detection. The assertions updated to the knowledge base can adopt the format described in the semantic pattern "Subject + Verb + Object".
[0140]
[0141] Step 140 / 240: Improve hyperparameters using an RL algorithm - Adaptive model improvement
[0142] This step adaptively improves the model for data condensation / feature extraction through reinforcement learning. The reinforcement learning is driven by the information in the penalty table, which is populated based on logical verification and in some examples also based on ML errors according to the evaluation of the detection events. The reinforcement learning improves the hyperparameters of the autoencoder model to optimize the effectiveness and accuracy of event detection in the condensed data. In some examples, the RL algorithm can be Q-learning, although other RL algorithms can also be envisioned.
[0143] The adaptive loop is formed by using the state-action-reward-state process of reinforcement learning. As described above, RL can be performed using various different algorithms. For the purposes of methods 100, 200, the RL algorithm is expected to meet the following conditions:
[0144] (1) Support model-free reinforcement learning, which can be applied to various highly dynamic environments
[0145] (2) Support value-based reinforcement learning, which is particularly aimed at improving the evaluation results by optimizing hyperparameters
[0146] (3) Support updating in a time-difference manner, such that the configuration of the computational model can be continuously improved during training until the best solution is found
[0147] (4) Low cost and easy for rapid iteration, which means reduced resource and capacity requirements for the device on which the algorithm is to be run.
[0148] The Q-learning algorithm is an option that basically meets the above conditions and can be integrated with the data condensation / feature extraction model of the currently proposed method through the optimization of the hyperparameters of these models. The adaptive model improvement process strengthens the hyperparameter adjustment based on the evaluation results. Then the evaluation results are mapped to the table shown below.
[0149] Table 2: Evaluation Results
[0150]
[0151] lndr: layer_number_decreasing_rate;
[0152] scf: scaling_factor;
[0153] swi: sliding_windows_interval;
[0154] mle: ml error;
[0155] lge: logic_error
[0156] Each column of the table represents an adjustment operation for the hyperparameters of the data condensation / feature extraction autoencoder model. Each row represents the ml error and the logic error after the corresponding operation. The reward γ for each state is:
[0157] reward_for_mle = 0 - mle × 100, and reward_for_lge = 0 – lge
[0158] The quality score Q of the action in a given state can be calculated according to the Bellman equation; each adjustment action on the configuration is labeled as d; the error state is labeled as e; and each iteration is labeled as t. Therefore, the Q-function value of each action in the current state is expressed as Q(e t ,d t ).
[0159] The initial state Q(e 0 ,d 0 ) = 0 and Q is updated as:
[0160]
[0161] The adaptive process is an iterative loop in which Q(e t ,d t ) is updated until Q updated (et ,d t ) == Q(e t ,d t ). Once this condition is met, the condensed features with the current adjusted hyperparameter values are output, and the process continues to the next sliding window.
[0162] The adaptive loop is as Figure 7a and Figure 7b shown. Initially referring to Figure 7a , the deep stacked autoencoder (represented by autoencoder 710) is configured according to certain hyperparameters. The autoencoder outputs the condensed data 720 to event detection via distance calculation and comparison (not shown). Then the detected event is verified in the logic validator 730, where an evaluation of the detected event is generated. The result of this evaluation 740 is used to drive reinforcement learning to update the Q - table 740, which is used to evaluate the adjustment actions performed on the hyperparameters according to which the autoencoder 710 is configured. Once the further increase in the Q - score is negligible, the optimal values of the model hyperparameters have been found, and the resulting condensed data from the current time window 760 can be output to event detection, e.g., via supervised learning, as discussed above with reference to steps 120 / 220.
[0163] Now referring to Figure 7b , the data condensation performed by the autoencoder 710 is shown in more detail, where the temporal accumulation of the condensed data is also shown. In addition to using the verification results of the detected events to drive the hyperparameter optimization of the autoencoder 710 in the adaptive learner 760, Figure 7b it is also shown that the detected verified events are used to generate a training dataset for supervised learning 770 to detect events in the condensed data.
[0164] The adaptive knowledge retrieval is as Figure 8 shown. Referring to Figure 8, presents the input data 802 of the stacked autoencoder 810. The condensed data / extracted features are forwarded to the distance calculator 820 that detects events. The detected events are verified within the above-mentioned logic framework 830, and the output of this verification is updated to the knowledge base 840. The knowledge base 840 can be used to populate the penalty table 850, which stores the configuration of the autoencoder 810 (i.e., hyperparameter values), the assertions corresponding to the detected events, and the errors corresponding to those assertions after logical evaluation. The penalty table drives the adaptive learner 860, on which an RL algorithm runs to improve the hyperparameters of the autoencoder 810 in order to maximize the evaluation score of the detected events. In some examples, a lightweight operation implementation is provided, and updates to the described configuration file can be performed only when new data is collected. The update is part of a closed iterative loop, according to which the hyperparameter values in the configuration file are updated based on the evaluation results, as described above. The iteration is time-frozen, as Figure 8 shown, which means that the same data frames extracted from the same time window will continue to iterate until the best configuration for the current data set is obtained, at which point the computation will move on to consider the data in the next time window.
[0165] Step 244: Expose the obtained insights;
[0166] This step can be implemented and executed using external storage. In this step, the system exposes the data to other components, thus facilitating the integration of the proposed method with the IoT ecosystem. For example, the results can be exposed to a database for visualization by users. In this way, in addition to supporting automated analysis, the method proposed herein can enrich human knowledge to support decision-making. By integrating with process components, these methods can additionally enhance business intelligence. The detected events and associated data can be exposed to other systems including management or enterprise resource planning (ERP), or the actuation or visualization systems (including LEDs, LCDs, or any other device that provides feedback to the user) within the same computing unit can be made available.
[0167] As described above, examples of the present disclosure also provide a system for performing event detection on a data stream that includes data from multiple devices connected via a communication network. An example of such a system 900 is in Figure 9is shown and configured to use an autoencoder (which may be a stacked distributed autoencoder) to condense information in a data stream, where the autoencoder is configured according to at least one hyperparameter. System 900 is further configured to detect an event from the condensed information and generate an evaluation of the detected event based on the logical compatibility between the detected event and a knowledge base. System 900 is further configured to use a reinforcement learning RL algorithm to improve at least one hyperparameter of the autoencoder, where a reward function of the RL algorithm is calculated based on the generated evaluation.
[0168] As Figure 9 shown, the system may include: a data processing function 910, configured to use an autoencoder to condense information in a data stream, where the autoencoder is configured according to at least one hyperparameter; and an event detection function 920, configured to detect an event from the condensed information. The system may further include: an evaluation function 930, configured to generate an evaluation of the detected event based on the logical compatibility between the detected event and a knowledge base; and a learning function 940, configured to use an RL algorithm to improve at least one hyperparameter of the autoencoder, where a reward function of the RL algorithm is calculated based on the generated evaluation.
[0169] One or more of functions 910, 920, 930, and / or 940 may include virtualized functions running in the cloud and / or may be distributed across different physical nodes.
[0170] The evaluation function 930 may be configured to generate an evaluation of the detected event based on the logical compatibility between the detected event and a knowledge base by: converting parameter values corresponding to the detected event into a logical assertion and evaluating the compatibility of the assertion with the content of the knowledge base, where the content of the knowledge base includes at least one of rules and / or facts. The evaluation function may further be configured to: generate an evaluation of the detected event based on the logical compatibility between the detected event and the knowledge base by performing at least one of incrementing or decrementing an evaluation score for each logical conflict between the assertion and a fact or rule in the knowledge base.
[0171] Figure 10 is a block diagram showing an example implementation of a method according to the present disclosure. Figure 10 An example implementation of may, for example, correspond to an intelligent manufacturing use case, which may be considered a possible application of the method proposed herein. Intelligent manufacturing is an example of a use case involving multiple heterogeneous and / or geographically distributed devices, and where it is necessary to understand the performance of the automated production line followed by the deployed devices and detect abnormal behavior. Further, a solution that expects to retrieve insights / knowledge from an automated industrial manufacturing system should be automated, and the data processing model should be self-updating.
[0172] As described above, domain experts often cannot provide manual evaluations of data and validation datasets, especially when newly introducing or assembling production lines. Additionally, in modern factories, the categories of deployed devices / sensor devices can be very complex and are often updated or changed, and these changes can cause a decline in the performance of pre-existing data processing models. The devices / sensor devices can serve different production units in different geographical locations (heterogeneity and distributed topology of devices), and production can be scaled up and / or down for different environments or business priorities. It is desirable for data processing to be real-time and produce results quickly.
[0173] Reference Figure 10 , each block can correspond to an intelligent execution unit, which, as described above, can be virtualized, distributed across multiple physical nodes, etc. Figure 10 The processing flow of
[0174] It will be understood that methods 100, 200 can also be implemented by the execution of a single node or virtualized function of a node-specific method. Figure 11 is a flowchart showing the process steps in one such method 1100. Reference Figure 11 , a method 1100 for managing an event detection process performed on a data stream includes, in a first step 1110, receiving a notification of a detected event, the data stream including data from a plurality of devices connected via a communication network, wherein the event has been detected from information condensed from the data stream using an autoencoder configured according to at least one hyperparameter. In step 1120, method 1100 includes receiving an evaluation of the detected event, wherein the evaluation has been generated based on the logical compatibility between the detected event and a knowledge base. In step 1130, the method includes using a reinforcement learning RL algorithm to improve at least one hyperparameter of the autoencoder, wherein a reward function of the RL algorithm is calculated based on the generated evaluation.
[0175] Figure 12is a block diagram showing example node 1200 which, for example, when receiving appropriate instructions from computer program 1206, can implement method 1100 according to an example of the present disclosure. Refer to Figure 12 , node 1200 includes a processor or processing circuitry 1202 and may include a memory 1204 and an interface 1208. The processing circuitry 1202 can be used to perform some or all of the steps of method 1100 as discussed above with reference to Figure 11 . The memory 1204 may contain instructions executable by the processing circuitry 1202 such that node 1200 can be used to perform some or all of the steps of method 1100. The instructions may also include instructions for performing one or more telecommunication and / or data communication protocols. The instructions may be stored in the form of a computer program 1206. In some examples, the processor or processing circuitry 1002 may include one or more microprocessors or microcontrollers, and other digital hardware which may include a digital signal processor (DSP), application specific digital logic, etc. The processor or processing circuitry 1202 may be implemented by any type of integrated circuit, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc. The memory 1204 may include one or several types of memory suitable for the processor, such as read only memory (ROM), random access memory, cache memory, flash memory devices, optical storage devices, solid state disks, hard disk drives, etc.
[0176] Figure 13a , Figure 13b and Figure 13c shows how information is transformed when passing through an intelligent pipeline formed by intelligent execution units connected to implement the method according to the present disclosure, as shown in an example implementation of Figure 10 . The intelligent pipeline receives data from sensors which may be geographically distributed (e.g., across Figure 10 the intelligent manufacturing sites of the implementation). The data collected represents information from different production components and their environments. After data processing such that the data is transformed into an appropriate format and normalized, a high-dimensional data set is obtained, as shown in Figure 13a . Then feature extraction / data condensation is performed, resulting in the extracted features, as shown in Figure 13b . After accumulating features from a time window, the distances between the final extracted features are calculated pairwise, and then the average distance of each feature to other features in a given time slot is calculated, as shown in Figure 13c . These results can be exposed to an external visualization tool such as Kibana.
[0177] Figure 14Provides a conceptual representation of an intelligent pipeline 1400 formed by one or more computing devices 1402, 1404, on which intelligent execution units 1408 perform the steps of methods 100, 200 on data generated by IoT devices 1406.
[0178] Figure 15 Illustrates the composition of an example IoT device that can generate data for a data stream. The IoT device 1500 includes a processing unit 1502, a sensor unit 1504, a storage device / memory unit 1506, and a communication unit 1508.
[0179] Figure 16 Illustrates the functional composition of an intelligent execution unit 1600, which includes a data source 1602, a mapping function 1604, a transformation function 1606, and a data sink 1608.
[0180] Figure 17 Is a functional representation of an intelligent pipeline, which shows point-to-point communication between processors and communication using an external agent.
[0181] As described above, each step in the methods disclosed herein can be implemented at different locations and / or different computing units. Examples of the methods disclosed herein can thus be implemented within an IoT environment consisting of devices, edge gateways, base stations, network infrastructure, fog nodes, and / or clouds, as Figure 18 shown. Figure 19 Illustrates an example of how an intelligent pipeline of intelligent execution units implementing the methods according to the present disclosure can be orchestrated in an IoT environment.
[0182] Examples of the present disclosure provide a technical solution to address the challenges of performing event detection in data streams, which can independently adapt to changes in data streams and different types, quantities, and complexities of data, minimize the requirements for domain expertise, and are fully scalable, reusable, and replicable. The proposed solution can be used to provide online anomaly analysis for data by implementing an automated intelligent data pipeline that accepts raw data and generates actionable intelligence with minimal input from human engineers or domain experts.
[0183] Examples of the present disclosure can exhibit one or more of the following advantages:
[0184] · Adaptive semi-supervised learning with rich capabilities for integration with data-intensive autonomous systems. The ability to adapt to changes through self-revision of the model.
[0185] · Can provide data labels for model creation or training without domain expertise; relies only on easily available "common sense" information for logical verification.
[0186] · Provide batch-based online machine learning for IoT data streams in real time, not limited to using stream data models: Continuously apply semi-machine learning algorithms to real-time data streams to obtain insights by using sliding windows and condensing data in time and feature dimensions.
[0187] · A highly scalable solution that is easy to deploy. A cluster can be created by providing an initial configuration file to the root node. Then, the cluster, including the model itself and the underlying computing resource orchestration, can be scaled up / down according to the number of devices generating data and the quantity and complexity of the data. The machine learning model is automatically improved and updated.
[0188] · A highly reusable and replicable solution.
[0189] Examples of the present disclosure apply deep learning in semi-supervised methods for retrieving knowledge from raw data and checking insights via logical verification. Then, the model configuration for data condensation is adjusted by optimizing model hyperparameters using a reinforcement agent, ensuring that these methods can adapt to changing environments and a wide variety of deployment scenarios and use cases.
[0190] The example methods proposed herein use stacked deep autoencoders to provide batch-based online machine learning for obtaining insights from high-dimensional data. The proposed solution is dynamically configurable, scalable in model and system architecture, independent of domain expertise, and thus highly replicable and reusable in different deployments and use cases.
[0191] The example methods proposed herein first apply unsupervised learning (stacked deep autoencoders) to extract and condense features from the raw data set. This unsupervised learning does not require any pre-existing labels to train the model. The example methods then apply common sense logical verification to exclude detected events that have logical conflicts with common sense in the relevant domain and form a Q-table. Based on the Q-table, Q-learning can be performed to obtain the optimal configuration of the unsupervised learning model (autoencoder), and then the existing model can be updated. In addition, the verified detected events can be used as labels for supervised learning to perform event detection in the condensed data. The example methods disclosed herein can thus be used in use cases where labeled training data is not available.
[0192] The methods of the present disclosure may be implemented in hardware or as software modules running on one or more processors. The methods may also be performed in accordance with the instructions of a computer program, and the present disclosure also provides a computer-readable medium having stored thereon a program for performing any of the methods described herein. The computer programs embodying the present disclosure may be stored on a computer-readable medium, or it may, for example, be in the form of a signal (such as a downloadable data signal provided from an Internet website), or it may be in any other form.
[0193] It should be noted that the above examples illustrate rather than limit the present disclosure, and those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. The word "comprising" does not exclude the presence of elements or steps other than those listed in the claims, "a" or "an" does not exclude a plurality, and a single processor or other unit may perform the functions of several units recited in the claims. Any reference signs in the claims should not be construed as limiting their scope.
Claims
1. A method for performing event detection on a data stream, the data stream including data from multiple devices connected via a communication network, the method comprises: using an autoencoder to condense information in the data stream, wherein the autoencoder is configured according to at least one hyperparameter; detecting events from the condensed information; generating an evaluation of the detected events based on the logical compatibility between the detected events and a knowledge base; and using a reinforcement learning RL algorithm to improve the at least one hyperparameter of the autoencoder, wherein the reward function of the RL algorithm is calculated based on the generated evaluation, wherein generating an evaluation of the detected events based on the logical compatibility between the detected events and a knowledge base includes: converting the parameter values corresponding to the detected events into logical assertions; and evaluating the compatibility of the assertions with the content of the knowledge base, wherein the content of the knowledge base includes at least one of rules and / or facts.
2. The method according to claim 1, wherein using an autoencoder to condense information in the data stream, wherein the autoencoder is configured according to at least one hyperparameter, comprises: using an unsupervised learning UL algorithm to determine the number of layers in the autoencoder and the number of neurons in each layer of the autoencoder based on at least one of the following: parameters associated with the data stream; or the at least one hyperparameter.
3. The method according to claim 2, wherein, parameters associated with the data stream include at least one of the following: the data transmission frequency associated with the data stream; and / or the dimensionality associated with the data stream.
4. The method according to any one of claims 1-3, wherein, the at least one hyperparameter includes: a time interval associated with a window; a scaling factor; a layer decrement rate.
5. The method according to any one of claims 1-3, wherein, the autoencoder includes a distributed stacked autoencoder, and wherein using the distributed stacked autoencoder includes: dividing the data stream into one or more data sub-streams; using an autoencoder different from the distributed stacked autoencoder to condense information in each corresponding sub-stream; and providing the condensed sub-streams to another autoencoder in another layer in the hierarchical structure of the stacked autoencoder.
6. The method according to any one of claims 1-3, further comprises: accumulating data in the data stream; and dividing the accumulated data stream into multiple consecutive windows, each window corresponding to a different time interval; and wherein using the autoencoder includes: condensing information in the windowed data.
7. The method according to any one of claims 1-3, wherein, detecting events from the condensed information includes: accumulating the condensed information over time; and comparing different parts of the accumulated condensed data.
8. The method according to claim 7, wherein, detecting events from the condensed information further includes: using cosine difference to compare different parts of the accumulated condensed data.
9. The method according to claim 7, wherein, detecting an event from the condensed information further comprises: generating labels for a training data set using at least one event detected by comparing different parts of the accumulated condensed data, the training data set comprising condensed information from the data stream; using the training data set to train a supervised learning SL model; and using the SL model to detect events from the condensed information.
10. The method according to any one of claims 1 - 3, wherein, the knowledge base contains at least one of rules and / or facts, and wherein the at least one rule and / or fact is generated from at least one of the following: the operating environment of at least some of the plurality of devices; the operating domain of at least some of the plurality of devices; service protocols applicable to at least some of the plurality of devices; deployment specifications applicable to at least some of the plurality of devices.
11. The method according to any one of claims 1 - 3, wherein, generating an evaluation of the detected event based on the logical compatibility between the detected event and the knowledge base further comprises: for each logical conflict between the assertion and a fact or rule in the knowledge base, performing at least one of the following: incrementing or decrementing an evaluation score.
12. The method according to any one of claims 1 - 3, further comprises: updating the knowledge base to include the detected events that are logically compatible with the knowledge base.
13. The method according to any one of claims 1 - 3, wherein, generating an evaluation of the detected event further comprises: generating the evaluation based on the logical compatibility between the detected event and the knowledge base and based on error values generated during at least one of the following: during condensing the information in the data stream, or during detecting events from the condensed information.
14. The method according to any one of claims 1 - 3, wherein, using a reinforcement learning RL algorithm to improve at least one hyperparameter of the autoencoder, wherein a reward function of the RL algorithm is calculated based on the generated evaluation, the method comprising: using the RL algorithm to experiment with different values of the at least one hyperparameter and determining the value of the at least one hyperparameter associated with the maximum value of the reward function.
15. The method according to any one of claims 1 - 3, wherein, using a reinforcement learning RL algorithm to improve at least one hyperparameter of the autoencoder, wherein a reward function of the RL algorithm is calculated based on the generated evaluation, the method comprising: establishing a state of the autoencoder, wherein the state of the autoencoder is represented by the value of the at least one hyperparameter; selecting an action to be performed on the autoencoder according to the established state; causing the selected action to be performed on the autoencoder; and calculating the value of the reward function after performing the selected action; Wherein, selecting an action to be performed on the autoencoder according to the established state includes: selecting an action from a set of actions, the set of actions including increasing and decreasing the value of the at least one hyperparameter.
16. The method according to any one of claims 1-3, wherein, the plurality of devices connected via the communication network includes a plurality of restricted devices.
17. A system for performing event detection on a data stream, the data stream including data from a plurality of devices connected via a communication network, the system being configured to: use an autoencoder to condense information in the data stream, wherein the autoencoder is configured according to at least one hyperparameter; detect events from the condensed information; generate an evaluation of the detected events based on the logical compatibility between the detected events and a knowledge base; and use a reinforcement learning RL algorithm to improve the at least one hyperparameter of the autoencoder, wherein the reward function of the RL algorithm is calculated based on the generated evaluation, wherein, generating an evaluation of the detected events based on the logical compatibility between the detected events and a knowledge base includes: converting the parameter values corresponding to the detected events into logical assertions; and evaluating the compatibility of the assertions with the content of the knowledge base, wherein the content of the knowledge base includes at least one of rules and / or facts.
18. The system according to claim 17, wherein, the system includes: a data processing function configured to use an autoencoder to condense information in the data stream, wherein the autoencoder is configured according to at least one hyperparameter; an event detection function configured to detect events from the condensed information; an evaluation function configured to generate an evaluation of the detected events based on the logical compatibility between the detected events and a knowledge base; and a learning function configured to use a reinforcement learning RL algorithm to improve the at least one hyperparameter of the autoencoder, wherein the reward function of the RL algorithm is calculated based on the generated evaluation.
19. The system according to claim 18, wherein, at least one of the functions includes a virtualized function.
20. The system according to claim 18 or 19, wherein, the functions are distributed on different physical nodes.
21. The system according to claim 17, wherein, the evaluation function is further configured to: generate an evaluation of the detected events based on the logical compatibility between the detected events and the knowledge base by at least performing one of increasing or decreasing an evaluation score for each logical conflict between the assertions and the facts or rules in the knowledge base.
22. A method for managing an event detection process performed on a data stream, the data stream including data from a plurality of devices connected via a communication network, the method comprises: receiving a notification of a detected event, wherein the event has been detected from information condensed from the data stream using an autoencoder configured according to at least one hyperparameter; Receive an evaluation of the detected event, where the evaluation has been generated based on the logical compatibility between the detected event and a knowledge base; and Use a reinforcement learning RL algorithm to improve the at least one hyperparameter of the autoencoder, where the reward function of the RL algorithm is calculated based on the generated evaluation, wherein, the evaluation being generated based on the logical compatibility between the detected event and a knowledge base includes: converting a parameter value corresponding to the detected event into a logical assertion; and evaluating the compatibility of the assertion with the content of the knowledge base, where the content of the knowledge base includes at least one of rules and / or facts.
23. A node for managing an event detection process performed on a data stream, the data stream including data from a plurality of devices connected via a communication network, the node including a processing circuit and a memory, the memory containing instructions executable by the processing circuit, whereby the node operates to: Receive a notification of a detected event, wherein, the event has been detected from information condensed from the data stream using an autoencoder configured according to at least one hyperparameter; Receive an evaluation of the detected event, where the evaluation has been generated based on the logical compatibility between the detected event and a knowledge base; and Use a reinforcement learning RL algorithm to improve the at least one hyperparameter of the autoencoder, where the reward function of the RL algorithm is calculated based on the generated evaluation, wherein, the evaluation being generated based on the logical compatibility between the detected event and a knowledge base includes: converting a parameter value corresponding to the detected event into a logical assertion; and evaluating the compatibility of the assertion with the content of the knowledge base, where the content of the knowledge base includes at least one of rules and / or facts.
24. A computer program product including a computer-readable medium having computer-readable code embodied therein, the computer-readable code being configured such that, when executed by a suitable computer or processor, causes the computer or processor to perform the method according to any one of claims 1 to 16 or 22.
Citation Information
Patent Citations
Mobile robot path planning method with combination of depth automatic encoder and Q-learning algorithm
CN105137967A
Stacking velocity spectrum pickup method based on deep reinforcement learning and processing terminal
CN109031421A