Context-Aware Feature Embedding Using Deep Recurrent Neural Networks and Anomaly Detection of Sequential Log Data

The context-aware feature embedding through deep recursive neural networks solves the difficulties in network traffic and operation log abnormal detection in the prior art, and achieves efficient and accurate abnormal detection effect.

CN113190843BActive Publication Date: 2025-07-01ORACLE INT CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202110516260.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-09-05
Filing Date
2019-07-23
Publication Date
2025-07-01
Estimated Expiration
2039-09-11

AI Technical Summary

Technical Problem

The prior art has difficulties in detecting abnormalities of network traffic and operation logs, including the inability to effectively detect unknown malicious activities, poor generalization ability, high dependence on human experts, and the problems of redundancy in information and loss of contextual information in log analysis.

Method used

Deep recursive neural network (RNN) is used for context-aware feature embedding, and dense feature vectors are generated for anomaly detection by sequence prediction of network packet flow and log messages. This technology combines sparse feature encoding and intensive feature transcoding, and uses graph embedding and unsupervised training to improve feature embedding for log tracking.

Benefits of technology

It realizes efficient abnormal detection of network traffic and operation logs, can effectively identify unknown malicious activities, reduce dependence on human experts, retain context information, and improve detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113190843B_ABST
    Figure CN113190843B_ABST
Patent Text Reader

Abstract

The present invention relates to context-aware feature embedding using deep recursive neural networks and anomaly detection of sequential log data. Techniques are provided herein for contextually embedding features of network traffic or operation logs for anomaly detection based on sequential prediction. In an embodiment, a computer has a predictive recursive neural network (RNN) for detecting anomalous network flows. In an embodiment, the RNN transcodes a sparse feature vector representing a log message into a dense feature vector in context, which may be predictive or used to generate a predictive vector. In an embodiment, graph embedding improves the feature embedding of log traces. In an embodiment, the computer detects and feature-encodes independent traces from related log messages. These techniques can detect malicious activities by performing anomaly analysis on context-aware feature embedding of network packet flows, log messages, and / or log traces.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a PCT application that entered the Chinese national phase, with an international filing date of July 23, 2019, a national application number of 201980067519.2, and an invention title of "Context-Aware Feature Embedding Using Deep Recurrent Neural Networks and Anomaly Detection of Sequential Log Data".

[0002] Cross-reference to related applications

[0003] The entire contents of the following related references are incorporated by reference:

[0004] · U.S. Patent Application No. 16 / 122,398, titled "MALICIOUS ACTIVITY DETECTION BY CROSS-TRACE ANALYSIS AND DEEP LEARNING", filed by Juan Fernandez Peinador et al. on September 5, 2018.

[0005] · U.S. Patent Application No. 16 / 122,664, titled "MALICIOUS NETWORK TRAFFIC FLOW DETECTION USING DEEP LEARNING", filed by Zhou Guang-Tong Zhou et al. on September 5, 2018.

[0006] · W.I.P.O. Patent Application No. PCT / US2017 / 033698, titled "MEMORY-EFFICIENT BACKPROPAGATING THROUGH TIME", filed by Marc Lanctot et al. on May 19, 2017;

[0007] · U.S. Patent Application No. 15 / 347,501, titled "MEMORY CELL UNIT AND RECURRENT NEURAL NETWORK INCLUDING MULTIPLE MEMORY CELL UNITS", filed by Daniel Neil et al. on November 9, 2016;

[0008] · U.S. Patent Application No. 14 / 558,700, titled "AUTO-ENCODER ENHANCED SELF-DIAGNOSTIC COMPONENTS FOR MODEL MONITORING", filed by Jun Zhang et al. on December 2, 2014; and

[0009] · "EXACT CALCULATION OF THE HESSIAN MATRIX FOR THE MULTI-LAYER PERCEPTRON" by Christopher M. Bishop, published in Neural Computation 4, No. 4 (1992), pp. 494-501. Technical Field

[0010] The present disclosure relates to sequence anomaly detection. Techniques are presented herein for contextually embedding features of network traffic or operational logs for anomaly detection based on sequence prediction. Background Art

[0011] For network-based systems such as enterprises and cloud data centers, network security is a major challenge. These systems are complex and dynamic and operate in an ever-evolving network environment. Although increasingly important from a network security perspective, analyzing the large amounts of data flowing between hosts and the distributed processing accompanying the traffic has exceeded the workload of human security experts. In some respects, traffic and activity analysis are more or less unmanageable with some techniques.

[0012] The basic representation of network data is the raw network traffic carried by network packets. Most malicious activities occur in the application layer of the TCP / IP network model, where applications pass flows of network packets between hosts. Evidence of malicious activity can be more or less hidden within the network flows of the packets.

[0013] Most existing industrial solutions make full use of rule- or signature-based techniques to detect malicious activities in network flows. Some techniques require security experts to comprehensively examine known malicious flows to extract rules or signatures therefrom. If a new flow matches any existing rule or signature, then it is detected as a malicious flow. Rule- or signature-based techniques have three distinct drawbacks: (i) they can only detect known malicious activities; (ii) patterns and rules are often difficult to generalize and thus often miss malicious activities that have changed slightly; and (iii) there is a significant requirement for the involvement of human security experts.

[0014] Fortunately, most network devices (such as servers, routers, and firewalls) summarize the activities and events that occur on the device in the form of text log messages. For example, on a Linux server, the operating system writes auditd (audit demon) logs for security-related activities such as logins, logouts, file accesses, etc. Evidence of malicious activity is more or less hidden in these operational logs.

[0015] Log analysis involves a large amount of log data from various sources, even for small companies and especially for large enterprises with a potentially confusing mix of multiple domain silos, multiple external interfaces, multiple middleware layers, and scheduling and self-organization activities. Therefore, manual log analysis can be futile. Most entries in the log data are uninteresting, making it like looking for a needle in a haystack to read through them. Additionally, manual log analysis depends on the expertise of the human operator to conduct the analysis.

[0016] Rule-based log filtering and analysis require manually crafted rules that are prone to human error for each known type of malicious behavior. The rule set is limited to known attacks and may be difficult to maintain manually over time to adapt to new types of malicious behavior.

[0017] Even with the assistance of machine learning, manual work can remain. Selecting informative, discriminative, and independent features is a key step in building an effective classification machine learning model. Features are often extracted from individual log messages, thus ignoring all possible interrelationships between them and losing context information. As a result, the effectiveness of the machine learning model may be severely hindered. Training the model on individual log messages may mislead the model into trying to detect abnormal log messages independently rather than abnormal activities across multiple log messages or network packets.

[0018] Tools such as Splunk can provide search and exploration capabilities for Windows and Linux logs (such as audit logs). Splunk can predict numerical fields (linear regression), predict categorical fields (logistic regression), detect numerical outliers (distribution statistics), detect categorical outliers (probability measurement), predict time series, and cluster numerical events. Effective use of such a tool requires in-depth knowledge of the context, provided that the user can only select a limited number of log message fields for each search query. However, log message parsing is static and tailored for dashboard tools rather than security tools. Therefore, Splunk cannot discover relationships between different log message fields or even between log messages.

[0019] The Elastic (ELK) stack is a log collection, transformation, normalization, and visualization framework that includes time series analysis of a set of log message fields selected by the user. Effective use of ELK requires in-depth knowledge of the problem and application context, as the task of ELK users is to set up the time series analysis pipeline and analyze the results. Unfortunately, tools such as ELK are prone to false positives, which must be filtered out by domain expert users.

[0020] Generally speaking, both Splunk and ELK come with detection tools that are static and have little or no learning ability. Therefore, the applicability of Splunk and ELK is limited because any fresh data will trigger a re - analysis of the entire data set (Splunk) or the last analysis window (ELK). As a result, neither of these two tools can associate log messages into meaningful groups, thus losing context and the opportunity to detect malicious events.

[0021] Structured logs mainly consist of key - value pairs, such as for category fields. In typical machine learning, categorical / status variables are usually vectorized via one - hot encoding. Therefore, depending on the total number of category fields in the log message and their associated values, the resulting feature vectors can be very large but sparse at the same time. On the other hand, there is not much semantic or context information encoded in the resulting vectors. For example, in a one - hot encoding vector space, two states (e.g., field values) that have more in common are equidistant from two completely independent states.

[0022] An embedding model is needed that not only provides dense and reduced feature vectors but also provides optimized vectors with more semantics. New technologies are needed that are different from the basic models that act on individual log messages or simply aggregate information into log messages by summing or averaging functions. On the one hand, for many network attack scenarios that affect the inter - relationships between log messages, individual log message analysis is very weak. On the other hand, heuristic methods such as averaging ignore some important information in sequence data, such as the ordering of log messages. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In the drawings:

[0024] Figure 1 is a block diagram depicting an example computer in an embodiment, the example computer having a predictive recurrent neural network (RNN) for detecting abnormal network flows;

[0025] Figure 2 is a flowchart depicting an example process in an embodiment for using a predictive RNN to detect abnormal network flows;

[0026] Figure 3 is a block diagram depicting an example computer in an embodiment for generating dense feature vectors by transcoding sparse unprocessed feature vectors;

[0027] Figure 4 is a flowchart depicting an example process in an embodiment for generating dense feature vectors by transcoding sparse unprocessed feature vectors;

[0028] Figure 5 is a block diagram of an example computer depicting a warning anomaly network flow in an embodiment;

[0029] Figure 6 is a flowchart of an example process for warning anomaly network flow in an embodiment;

[0030] Figure 7 is a block diagram of an example computer depicting the use of a recursive topology of an RNN to generate a packet anomaly score from which a flow anomaly score can be synthesized in an embodiment;

[0031] Figure 8 is a flowchart of an example process for using a recursive topology of an RNN to generate a packet anomaly score from which a flow anomaly score can be synthesized in an embodiment;

[0032] Figure 9 is a block diagram of an example computer depicting sparse encoding of features based on the anatomy of a generic packet of a given communication protocol in an embodiment;

[0033] Figure 10 is a block diagram of an example RNN configured and trained by a computer in an embodiment;

[0034] Figure 11 is a block diagram of an example computer with an RNN that transcodes a sparse feature vector representing a log message contextually into a dense feature vector that can be used to predict or generate a predictive vector in an embodiment;

[0035] Figure 12 is a flowchart of an example process for using an RNN to transcode a sparse feature vector representing a log message contextually into a dense feature vector for generating a predictive vector in an embodiment;

[0036] Figure 13 is a flowchart of an example process for densifying the same sparse feature vector differently based on context in an embodiment;

[0037] Figure 14 is a block diagram of an example computer with a training harness for improving densification in an embodiment;

[0038] Figure 15 is a flowchart of an example process for a training harness for improving densification in an embodiment;

[0039] Figure 16is a schematic diagram depicting an example activity diagram of computer system activities occurring on one or more interoperating computers in an embodiment;

[0040] Figure 17 is a block diagram of an example computer depicting feature embeddings using graph embeddings to improve log tracing in an embodiment;

[0041] Figure 18 is a flowchart of an example process for feature embeddings using graph embeddings to improve log tracing in an embodiment;

[0042] Figure 19 is a block diagram of an example log containing related traces from which a temporally pruned sub - graph can be created in an embodiment;

[0043] Figure 20 is a block diagram of an example computer in an embodiment having a trainable graph embedder and a trainable anomaly detector, where the trainable graph embedder generates context feature vectors and the trainable anomaly detector consumes the context feature vectors;

[0044] Figure 21 is a block diagram of an example computer that detects and feature - encodes independent traces from related log messages in an embodiment;

[0045] Figure 22 is a flowchart of an example process for detecting and feature - encoding independent traces from related log messages in an embodiment;

[0046] Figure 23 is a tabular diagram of an example log in an embodiment that contains semi - structured operational (e.g., diagnostic) data from which log messages and their features can be parsed and extracted;

[0047] Figure 24 is a block diagram illustrating a computer system on which embodiments of the present invention can be implemented;

[0048] Figure 25 is a block diagram illustrating a basic software system that can be used to control the operation of a computing system. Detailed Description

[0049] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be evident, however, that the present invention may be practiced without these specific details. In other instances, well - known structures and devices are shown in block diagram form to avoid unnecessarily obscuring the present invention.

[0050] Embodiments are described herein according to the following overview:

[0051] 1.0 Overall Overview

[0052] 2.0 Example Computer

[0053] 2.1 Network Packet

[0054] 2.2 Feature Extraction and Encoding

[0055] 2.3 Recurrent Neural Network

[0056] 3.0 Anomaly Detection Process

[0057] 4.0 Context Encoding

[0058] 4.1 Benefits of Feature Embedding

[0059] 4.2 Unsupervised Training

[0060] 5.0 Context Encoding Process

[0061] 6.0 Network Flow Alert

[0062] 6.1 Anomaly Score

[0063] 6.2 Future Proofing

[0064] 7.0 Alert Process

[0065] 8.0 Sequence Prediction

[0066] 8.1 Packet Anomaly Score

[0067] 9.0 Sequence Prediction Process

[0068] 10.0 Packet Decomposition

[0069] 10.1 Network Flow Demultiplexing

[0070] 11.0 Training

[0071] 12.0 Embedding Recorded Features

[0072] 13.0 Log Feature Embedding Process

[0073] 14.0 Context Encoding Meaning

[0074] 15.0 Training Bundle

[0075] 15.1 Decoding for Reconstruction

[0076] 15.2 Measured Error

[0077] 16.0 Training Using Reconstruction

[0078] 17.0 Activity Diagram

[0079] 18.0 Graph Embedding

[0080] 18.1 Pruning

[0081] 18.2 Log Trace

[0082] 19.0 Graph Embedding Process

[0083] 20.0 Trace Aggregation

[0084] 21.0 Trainable Anomaly Detection

[0085] 22.0 Trace Composition

[0086] 22.1 Composition Criteria

[0087] 23.0 Trace Detection Process

[0088] 24.0 Declarative Trace Detection

[0089] 24.1 Declarative Rules

[0090] 24.2 Example Operations

[0091] 25.0 Machine Learning Model

[0092] 25.1 Artificial Neural Network

[0093] 25.2 Illustrative Data Structure for Neural Network

[0094] 25.3 Backpropagation

[0095] 25.4 Deep Context Overview

[0096] 26.0 Hardware Overview

[0097] 27.0 Software Overview

[0098] 28.0 Cloud Computing

[0099] 1.0 General Overview

[0100] This document provides techniques for contextually embedding features of network traffic or operational logs for anomaly detection based on sequence prediction. These techniques can detect malicious activities by performing anomaly analysis on context-aware feature embeddings of network packet flows, log messages, and / or log traces.

[0101] In an embodiment, a computer has a predictive recurrent neural network (RNN) that detects anomalous network flows. The computer generates a sequence of actual dense feature vectors corresponding to a sequence of network packets. Each feature vector in the sequence of actual dense feature vectors represents a corresponding network packet in the sequence of network packets. The RNN generates a sequence of predicted dense feature vectors representing the sequence of network packets based on the sequence of actual dense feature vectors. The sequence of network packets is processed based on the sequence of predicted dense feature vectors.

[0102] In an embodiment, the RNN transcodes sparse feature vectors representing log messages into dense feature vectors in context, which can be predictive or used to generate predictive vectors. The computer processes each log message in a sequence of related log messages as follows. Features are extracted from the log message to generate a sparse feature vector representing the features. The sparse feature vector is applied as an excitation input to a corresponding step of an encoder RNN. A corresponding embedded feature vector is output from the encoder RNN, which is based on the features and one or more log messages that occurred earlier in the sequence of related log messages. One or more of the embedded feature vectors output from the encoder RNN are processed to determine a predicted next related log message that should occur in the sequence of related log messages.

[0103] In an embodiment, graph embedding improves the feature embedding of log traces. The computer receives independent feature vectors. Each independent feature vector occurs in context as follows. The independent feature vector represents a corresponding log trace. The corresponding log trace represents a corresponding single action. The corresponding log trace is based on one or more log messages generated by the corresponding single action. The corresponding log trace indicates one or more network identities.

[0104] Graph embedding requires generating one or more edges of a connected graph that includes a particular vertex generated from a particular log trace represented by a particular independent feature vector. Each edge connects two vertices of the connected graph, where the two vertices are generated from two corresponding log traces represented by two corresponding independent feature vectors. The two corresponding log traces indicate the same network identity. Embedded feature vectors are generated based on the independent feature vectors that represent the log traces from which the vertices of the connected graph are generated. Based on the embedded feature vectors, a particular log trace is indicated as anomalous.

[0105] In an embodiment, a computer detects individual traces from relevant log messages and encodes their characteristics. Key-value pairs are extracted from each log message. A trace representing a single action is detected based on a subset of log messages whose key-value pairs satisfy grouping criteria. A suspicious feature vector representing the trace is generated based on the key-value pairs in the subset of log messages. Based on one or more feature vectors including the suspicious feature vector, an anomaly detector indicates that the suspicious feature vector is anomalous.

[0106] 2.0 Example Computer

[0107] Figure 1 is a block diagram depicting example computer 100 in an embodiment. Computer 100 has a predictive recurrent neural network (RNN) that detects anomalous network flows. Computer 100 can be one or more computers, such as an embedded computer, a personal computer, a rack server (e.g., blade, mainframe, virtual machine), or any computing device capable of executing an artificial neural network (ANN) (such as having hyperbolic functions, differential equations, and matrix operations such as multiplication). Example ANN implementations and techniques will be discussed in the "Overview of Artificial Neural Networks" section below.

[0108] Figure 1 Data flowing from left to right is shown, and data transformation occurs at each stage of the workflow. Sequences 110, 130, and 150 can reside in the random access memory (RAM) of computer 100 as data structures. The original sequence 110 contains network packets 121 - 123 that occur in sequence in the network flow. The network flow can form a conversation or other flow between two computers (not shown).

[0109] 2.1 Network Packets

[0110] A network packet is the smallest unit of transmission, such as a protocol data unit (PDU) of the network layer, such as an Internet Protocol (IP) datagram or a data frame of the data link layer (such as an Ethernet frame). Each of packets 121 - 123 is or was a live packet that traversed a network link (not shown) monitored by a network element. In an embodiment, computer 100 is a network element within a live network route through which packets 121 - 123 typically flow. For example, computer 100 can be a switch, router, bridge, repeater, or proxy.

[0111] In an embodiment, computer 100 is a live network element outside the routing of packets 121 - 123. For example, a network flow can be split / forked / sent to feed copies of packets 121 - 123 to computer 100 more or less in real - time while the original packets are further transmitted under the opposite fork of the split. In an embodiment, computer 100 is offline such that the live network flow is not available and the original sequence 110 has been previously recorded for delayed (e.g., scheduled) analysis. For example, the original sequence 110 can be spooled persistently to a file or database.

[0112] In an embodiment, the network flow occurs within network traffic that simultaneously includes multiple flows. For example, a data link can be multiplexed (e.g., in time, such as with time slots). For example, packets 121 - 123 appear consecutive within the original sequence 110, but during live transmission they may have been interleaved with other packets (not shown) of other flows.

[0113] In an embodiment, computer 100 receives the mixed flows and demultiplexes them (i.e., untangles) into individual flows for recording and / or analysis. For example, each packet can carry identifiers of the sender and the receiver that can together identify the flow. Each packet can carry additional identifiers (such as for a session, application, account, or principal (e.g., end - user)) that may also be required to identify the flow. Thus, individual packets can be associated with the same or different flows. In an embodiment, computer 100 receives the demultiplexed (i.e., separated) flows or only receives one flow. In any case, the original sequence 110 and packets 121 - 123 are part of the same (e.g., non - tangled) flow.

[0114] 2.2 Feature Extraction and Encoding

[0115] In operation, computer 100 directly or indirectly converts the original sequence 110 into an observed sequence 130 that represents packets 121 - 123 in a format that is accepted as an input stimulus by a recurrent neural network (RNN) 170. Each of packets 121 - 123 can be individually converted into a corresponding feature vector 141 - 143. In the illustrated embodiment, the feature vectors 141 - 143 are densely encoded such that most or all bits of the feature vectors represent meaningful attributes of the network packets. However, the dense encoding need not be discrete such that a particular bit always represents a particular feature.

[0116] In an embodiment, feature extraction and encoding are used to directly generate the observed sequence 130 from the original sequence 110, such as using a predefined semantic mapping. In an embodiment, intermediate encoding (not shown) occurs between the original sequence 110 and the observed sequence 130, such as sparse feature encoding. Sparse encoding and subsequent dense transcoding are discussed later in this document.

[0117] 2.3 Recurrent Neural Network

[0118] In operation, the observed sequence 130 that is specifically encoded for network flow is consumed by the RNN 170. Different from a conventional ANN, the RNN is stateful. The RNN is naturally suitable for analyzing sequences of related items, including identifying interesting orderings of items within the sequence. Therefore, the RNN naturally implements context analysis, which can be crucial for identifying anomalies. For example, even if all packets of a flow individually appear normal, the network flow may be abnormal. For example, one ordering of packets can be abnormal while another ordering of the same packets can be normal, which can be the content for training the RNN 170 to facilitate detection.

[0119] In operation, the RNN 170 outputs a predicted next feature vector of a predicted sequence 150 that is expected to match the feature vector of the next packet in the original sequence 110 for the observed sequence, such as 130. For the observed sequence, such as 130, the RNN 170 outputs the corresponding predicted sequence 150. The computer 100 can compare the feature vector sequences 130 and 150 with each other to detect whether they match (e.g., bitwise), thereby achieving context sensitivity, such as incorporating information about surrounding (i.e., temporally or semantically related) packets rather than just the current packet in isolation.

[0120] For example, when activated with the actual feature vector 141 of packet 121, the RNN 170 can generate a predicted feature vector 162 as the predicted next feature vector of the next packet 122 represented by the actual feature vector 142. Thus, packet 122 can be predicted based on the previous packet sequence (e.g., 121). The computer 100 can detect that the feature vector 142 does not match 162 and thereby identify that the network flow is abnormal because a packet different from the expected packet has occurred. The operation of the RNN will be further discussed later in this document. Example RNN implementations and techniques will be discussed in the "Deep Context Overview" section below.

[0121] The RNN 170 can have various internal architectures based on neuron hierarchies and repeating units (i.e., cells), where each cell consists of a number of specialized and specially arranged neurons, such as a long short-term memory (LSTM) network. In Python embodiments, a third-party library such as Keras can provide an RNN with or without LSTM. In Java embodiments, deeplearning4j can be used as a third-party library. In C++ embodiments, TensorFlow can be used. In embodiments, a graphics processing unit (GPU) or other single instruction multiple data (SIMD) infrastructure, such as a vector processor, can provide hardware acceleration for the training and / or production use of the RNN 170.

[0122] 3.0 Anomaly Detection Process

[0123] Figure 2 is a flowchart depicting the computer 100 in an embodiment using a predictive RNN to detect anomalous network flows. Refer to Figure 1 for discussion Figure 2 .

[0124] Step 202 generates a sequence of actual dense feature vectors corresponding to the sequence of network packets. For example, the computer 100 initially receives the original sequence 110 of network packets 121 - 123 (or obtains a copy thereof). The network packets 121 - 123 can be obtained in batches, or naturally, the packets can arrive individually at various time delays. The original sequence 110 is directly or indirectly converted into the observed sequence 130 of actual dense feature vectors 141 - 143, which represent the packets 121 - 123 in a format that the RNN 170 accepts as input stimuli. In an embodiment, each network packet is individually converted into a corresponding actual dense feature vector.

[0125] In step 204, the RNN generates a sequence of predicted dense feature vectors representing the sequence of network packets based on the sequence of actual dense feature vectors. For example, the RNN 170 outputs the predicted sequence 150 of predicted dense feature vectors 161 - 163. The advantages of densely encoded feature vectors include the acceleration and size reduction of ANNs, which will be discussed in the "Benefits of Feature Embedding" section below.

[0126] In step 206, a sequence of network packets is processed based on the predicted sequence of dense feature vectors. For example, the predicted sequence 150 can be compared with the observed sequence 130. Each predicted dense feature vector 161 - 163 can be individually compared with the corresponding actual dense feature vector 141 - 143. Any individual mismatch or multiple mismatches of dense feature vectors within the sequence can indicate an anomaly. Thus, an anomaly can be detected based on an initial subsequence of dense feature vectors within the sequence. For example, an anomalous mismatch between dense feature vectors 142 and 162 can be detected before receiving network packet 123.

[0127] The original sequence 110 can be further processed based on whether the predicted sequence 150 is anomalous. For example, computer 100 can issue an alert, perform further analysis and / or monitoring on the original sequence 110, isolate the original sequence 110, terminate the network connection, or lock the account. If the predicted sequence 150 is not anomalous, then the original sequence 110 can be relayed to the original intended recipient.

[0128] 4.0 Context Encoding

[0129] Figure 3 is a block diagram depicting an example computer 300 in an embodiment. Computer 300 generates dense feature vectors by transcoding sparse unprocessed feature vectors. Computer 300 can be an implementation of computer 100.

[0130] Feature encoding can affect the accuracy of regression, such as anomaly detection. Naive, simple, or straightforward encodings (such as sparse feature encoding) tend to treat all features as equally important, which cannot eliminate noise from unprocessed data (such as network packets and flows). For example, some features may be merely distracting, which can mean the so - called "curse of dimensionality", which can cause overfitting.

[0131] The unprocessed feature vectors 361 - 363 of the intermediate sequence 350 are sparse direct encodings of the corresponding network packets 321 - 323. The advantage of sparse encoding is that it is justified to start with sparse encoding (such as through an RNN (not shown)) before transcoding to dense encoding for downstream deep analysis. The main advantage of sparse encoding is that it can be transcoded and does not require a large amount of (e.g., skilled) design. For example, sparse encoding does not require knowledge of the natural (and possibly counterintuitive) relationships between features. In fact, advanced techniques such as feature selection can be more or less avoided during sparse encoding. The techniques and mechanisms of sparse encoding will be discussed below for Figure 9 will be discussed.

[0132] 4.1 Benefits of Feature Embedding

[0133] Encoding into a reduced (i.e., denser) space can be lossy, which can actually be beneficial, such as dimensionality reduction. If noisy features are not emphasized (e.g., lost) and important features are emphasized (e.g., amplified or at least retained), then lossy dense encoding can actually improve the accuracy of anomaly detection. For example, the raw feature vector 361 can contain thousands of features. While the dense feature vector 341 can contain only hundreds of the most interesting features. Thus, transcoding from the readily available sparse intermediate sequence 350 to the optimized dense observed sequence 330 can improve the accuracy of an RNN (not shown) that consumes the observed sequence 330. This transcoding is performed by the encoder 390.

[0134] The encoder 390 transcodes each raw feature vector 361 - 363 individually (i.e., one at a time) into the corresponding dense feature vector 341 - 343. Since the transcoding does not need to consider the temporal context (i.e., the relationship between multiple packets of the stream), the encoder 390 can be stateless. In an embodiment, the encoder 390 includes an artificial neural network (ANN), such as a multi - layer perceptron (MLP) for deep learning. In the illustrated embodiment, the encoder 390 has a stack of at least neural layers 371 - 372, with neurons of adjacent layers being fully connected for deep learning. More layers and connections help the MLP identify more patterns and subtle differences between patterns. In this sense, deep means many layers (not shown).

[0135] Compared to the sparse encoding of the intermediate sequence 350, the dense encoding of the observed sequence 330 can achieve multiple performance improvements. Dense encoding can emphasize semantic context (such as the relationships between features), which can be as important or more important than discrete (i.e., separate) features. Dense encoding can emphasize feature selection, such as de - emphasizing (e.g., discarding) noisy features.

[0136] Context and feature selection can intersect, such as when a particular feature depends on whether a context such as a combination of features is relevant or not. Thus, dense encoding can have an optimal vocabulary of labels (i.e., symbols representing combinations of features and values), which, while dense (i.e., reduced), is actually based on complex and / or numerous heuristics that imply the static connection weights of the ANN.

[0137] 4.2 Unsupervised Training

[0138] During supervised training, a predefined label can be given to a (neural or non-neural) trainable encoder 390. Unfortunately, the labels are domain- and application-specific. Thus, defining the labels can be an expensive design effort in terms of research and development (R&D) time and expertise. In theory, there can be an infinite variety and / or type of possible anomalies, including anomalies unknown during training. It may be impossible to label all possible anomalies. Fortunately, by performing unsupervised training on the encoder 390, a priori labeling can be avoided. All machine learning models discussed herein are suitable for unsupervised training.

[0139] In an embodiment, the encoder 390 is an autoencoder, which is an unsupervised deep learning MLP (or a stack of dedicated MLPs) trained as an encoder / decoder (codec). The codec stacks an MLP decoder downstream of the MLP encoder. In an embodiment, the encoder 390 is only the MLP encoder layer of the autoencoder. For example, an autoencoder can be trained or retrained as a complete MLP stack and then decomposed into a subset of layers to isolate the encoder MLP for production deployment of the encoder 390. In an embodiment discussed later herein, the autoencoder is split into the encoder 390 and a decoder, and then a predictive RNN is stacked (i.e., inserted) downstream of the encoder 390 and upstream of the decoder.

[0140] During unsupervised training of the autoencoder, a vocabulary for dense coding spontaneously emerges and is learned. Attributed to the feedback loop that the codec can be configured for (especially for training), the autoencoder learns a dense vocabulary (i.e., more or less noise-free and filled-in), but not necessarily the densest. For an autoencoder, optimality is based on accuracy, where density can be an important factor, but not the only one. Dense coding, transcoding, and autoencoding will be further discussed later in this document. Example implementations and techniques (including training) of the autoencoder will be discussed in the "Deep Training Overview" section below.

[0141] 5.0 Context Encoding Process

[0142] Figure 4 is a flowchart depicting a computer 300 in an embodiment generating a dense feature vector by transcoding a sparse, unprocessed feature vector. Refer to Figure 3 for discussion Figure 4 .

[0143] Before or during step 402, an original sequence 310 of network packets 321 - 323 is received. Step 402 generates a sequence of actual unprocessed feature vectors corresponding to the sequence of network packets. For example, network packets 321 - 323 are directly encoded as corresponding sparse unprocessed feature vectors 361 - 363 of an intermediate sequence 350. The techniques and mechanisms of sparse coding will be discussed hereinafter for Figure 9 discussion.

[0144] Step 404 encodes the sequence of actual unprocessed feature vectors into a sequence of actual dense feature vectors. For example, encoder 390 can be an autoencoder (as discussed later herein) or other MLP that transcodes each unprocessed feature vector 361 - 363 individually (i.e., one at a time) into a corresponding dense feature vector 341 - 343.

[0145] Step 406 is exemplary. Other embodiments can alternatively perform alternative actions in addition to the actions depicted in step 406. Step 406 applies the sequence of actual dense feature vectors as an input stimulus to a predictive RNN, such as for anomaly detection. For example, the observed sequence 330 can also be the observed sequence 130 applied to Figure 1 the RNN 170 in to generate a predicted sequence 150, which can be compared with the observed sequence 130 / 330 for anomaly detection.

[0146] 6.0 Network Flow Alerts

[0147] Figure 5 is a block diagram depicting an example computer 500 in an embodiment. Computer 500 alerts of an abnormal network flow. Computer 500 can be an implementation of computer 100.

[0148] Computer 500 includes a complete stack 505 of dedicated MLPs. Although the complete stack 505 does not actually detect anomalies, it does perform sequence predictions that can facilitate downstream anomaly detection, as shown below.

[0149] The complete stack 505 consists essentially of transcoders 561 - 562 and an RNN 570 therebetween. The transcoders 561 - 562 are used together in series as an encoder - decoder to mediate between the internally dense - coded complete stack 505 and the rest of computer 500 that provides and expects unprocessed (i.e., sparse) coding. Data flows into and out of the complete stack 505 as follows.

[0150] Computer 500 sparsely encodes the original sequence 510 of network packets 511-512 into an intermediate sequence 520 of actual unprocessed feature vectors 521-522, which is transcoded by encoder 561 into an observed sequence 530 of actual dense feature vectors 331-332. RNN 570 consumes the observed sequence 530, and the RNN 570 responsively generates a predicted sequence 540 of predicted dense feature vectors 541-542, which is reverse transcoded by decoder 562 back into an expected sequence 550 of predicted unprocessed (i.e., sparse) feature vectors 551-552.

[0151] In an embodiment, the original sequence 510 is buffered or otherwise recorded such that data can be transformed as a batch of vectors from the original sequence 510 and transmitted along the illustrated loop path to the expected sequence 550. In (e.g., streaming) embodiments, each individual network packet or corresponding feature vector is transformed and transmitted individually along the illustrated data path. For example, packet 511 can traverse the illustrated data path to output the predicted feature vector of the expected next packet into the expected sequence 550 before the next packet actually exists (i.e., is generated and output by the source network element). For each packet that arrives at computer 500 in real time, the complete stack 505 can have made a prediction (i.e., unprocessed feature vector) for this packet, which can be used to evaluate the suspiciousness of the actually received packet by comparison.

[0152] 6.1 Anomaly Score

[0153] Anomaly detection can occur as follows. Computer 500 can compare the predicted expected sequence 550 with the actually received intermediate sequence 520. The content difference between sequences 520 and 550 can be measured as an anomaly score 580. Later in this document are embodiments that compare each predicted unprocessed feature vector with the corresponding actual unprocessed feature vector individually and use techniques such as mean squared error to calculate the anomaly score 580.

[0154] If the anomaly score 580 exceeds a threshold 590, then an anomaly is detected and an alert 595 can be issued. The alert 595 can be recorded and / or issued interactively to a human network administrator. The alert 595 can analyze security issues and risks immediately or with a delay. The alert 595 can automatically or indirectly cause a security intervention, which can impose various countermeasures on a party, application, or computer involved in a violating network flow, such as auditing, increased monitoring, and / or suspension of permissions or capabilities.

[0155] In a streaming transmission embodiment of analyzing network traffic step by step (i.e., as each packet arrives), the threshold 590 can be crossed before receiving or analyzing the entire packet stream. For example, a subsequence of the packets of the stream may be sufficient to detect an anomaly. Thus, live countermeasures can occur in real time, such as triggering deeper analysis or monitoring, such as deep packet inspection, or terminating a suspicious network connection.

[0156] The alert 595 can contain details about the anomaly, the problematic packet, the original sequence 510, or metadata about the suspicious network traffic. The alert 595 at least indicates the anomaly, which typically represents a network attack (e.g., an intrusion). Depending on its complexity (e.g., depth of the layer), the predictions made by the RNN 570 can vary (i.e., be anomalous) for various types or styles of attacks. In an embodiment, the RNN 570 makes different predictions for one, some, or all kinds of attacks, including hypertext transfer protocol (HTTP) denial of service (DoS), brute force secure shell (SSH), or simple network management protocol (SNMP) reflection amplification.

[0157] 6.2 Future proof

[0158] Due to the unsupervised training for predicting more or less general packet sequences, the RNN 570 can help reveal anomalies that occur during new types of attacks that may not have been present during the training of the RNN 570. Thus, based on supervised training or a rule base, the RNN 570 can achieve versatility beyond competing technologies. Thus, the RNN 570 can be more or less future proof because of its natural tendency to show unfamiliar patterns, which is inherently surprising.

[0159] Because of the versatility of the RNN 570, the RNN 570 can be used in a production environment to analyze the network traffic of somewhat opaque applications. For example, even if the application emitting the network traffic is an application lacking source code and / or documentation or an unfamiliar black box (i.e., opaque) application, the RNN 570 remains effective.

[0160] Moreover, because of the versatility of the RNN 570, the RNN 570 can be placed in various corners of the network topology. For example, the RNN 570 can reside and / or analyze the traffic that occurs outside the firewall or inside the demilitarized zone (DMZ). The RNN 570 can analyze the traffic between third parties in real time, such as for man-in-the-middle inspection by an operating company, possibly through port mirroring or cable splitting.

[0161] Tools such as packet capture (pcap) and / or terminal-based Wireshark (TShark) can be used to obtain packets. Packet capture files can be emitted by tools such as Transmission Control Protocol dump (tcpdump). Files containing logged packets from any of these tools can be analyzed according to the techniques herein, such as for anomaly detection.

[0162] 7.0 Alert Process

[0163] Figure 6 is a flowchart depicting a computer 500 in an embodiment warning of an anomalous network flow. Refer to Figure 5 for discussion Figure 6 。

[0164] Before or during step 602, the RNN 570 generates a predicted sequence 540 of predicted dense feature vectors 541 - 542, which represents an original sequence 510 of predicted network packets 511 - 512. Step 602 generates a sequence of predicted unprocessed feature vectors by decoding the sequence of predicted dense feature vectors. For example, the decoder 562 can be an auto-decoder (each auto-encoder, as discussed later herein) or other MLP that transcodes each dense feature vector 541 - 542 individually (i.e., one at a time) into a corresponding predicted unprocessed feature vector 551 - 552.

[0165] Based on the sequence of actual unprocessed feature vectors and the sequence of predicted unprocessed feature vectors, step 604 generates an anomaly score for the sequence of network packets. For example, the computer 500 can compare the predicted expected sequence 550 with the actual received intermediate sequence 520. The content difference between sequences 520 and 550 can be measured as the anomaly score 580. Later herein are embodiments that compare each predicted unprocessed feature vector with the corresponding actual unprocessed feature vector individually and use techniques such as mean squared error to calculate the anomaly score 580.

[0166] Step 606 also processes the sequence of network packets based on the anomaly score. For example, when the anomaly score 580 exceeds a threshold 590, an anomaly is detected, which can cause an alert 595 to be raised. If the original sequence 510 of network packets 511 - 512 is not anomalous, then the original sequence 510 can be processed normally, such as by relaying the sequence 510 to a downstream consumer.

[0167] 8.0 Sequence Prediction

[0168] Figure 7is a block diagram depicting an example computer 700 in an embodiment. The computer 700 uses a recursive topology of an RNN to generate a packet anomaly score, and a flow anomaly score can be synthesized based on the packet anomaly score. The computer 700 may be an implementation of the computer 100.

[0169] The RNN 720 includes a plurality of recursive steps, such as 721 - 723. Each recursive step may include an MLP (not shown) having a certain number of gated states, which, according to control gates, are stateful neural arrangements that can be latched or erased as needed. For example, the gated states may be based on Long Short - Term Memory (LTSM). Each recursive step 721 - 723 corresponds to sequential time steps, such as one for each packet of a network flow. In an embodiment, the RNN 720 has as many recursive steps as there are packets in the longest expected network flow, which typically has fewer than 100 packets.

[0170] The RNN 720 perceives context (i.e., sequence) in two ways (not shown). First, the gated states facilitate stateful (i.e., history - sensitive) processing. Second, a previous recursive step notifies (i.e., activates) its successive next recursive step, thereby cumulatively and propagating history through logical time. For example, the recursive step 721 may have a connection (not shown) that activates the recursive step 722. Thus, each recursive step (except the first step 721, as it has no previous step) receives inputs from two sources: a) the corresponding observed features, such as 712, and b) cross - activation through the previous recursive step.

[0171] Cross - activation facilitates potentially cumulative and propagating history across all recursive steps (e.g., the entire flow), such that the processing of the current (e.g., last) packet may be affected by some or all of the previous packets in the flow. The RNN 720 is applicable to various internal implementation topologies for implementing recursion. In an embodiment, each recursive step has its own MLP. In various embodiments, the MLP for each step has: a) its own personalized connection weights, or b) a copy of the weights shared by all steps. In an embodiment, all steps share the same MLP, such that cross - activation between steps requires a cyclic backward (i.e., reverse) edge to be present within the shared MLP.

[0172] Regardless of the internal topology of the RNN 720, the input and output of the RNN 720 as a black box operate as follows. Each observed feature 711-713 of the actual sequence 710 includes a dense feature vector, which is applied as a stimulus input to the corresponding recursive step 721-723. For example, the observed feature 711 may represent the first packet (not shown) of the network flow. Similarly, the recursive step 721 may be the first recursive step of the RNN 720. Thus, the observed feature 711 is applied to the recursive step 721.

[0173] The actual sequence 710 may be a pair (not shown) of differently encoded parallel sequences, such as a sparse / raw sequence and a corresponding transcoded dense sequence. The dense feature vectors of the dense sequence are applied to the recursive steps 721-723.

[0174] Similarly, the predicted sequence 730 may be a pair (not shown) of parallel sequences, such as a predicted dense sequence and a corresponding decoded sparse sequence. The RNN 720 outputs the dense sequence. For example, the recursive step 721 outputs the predicted feature 731. The horizontal line correlates the observed feature 712 with the predicted feature 731, which implies a skew.

[0175] For example, the predicted feature 731 is the first vector in the predicted sequence 730, and the associated observed feature 712 is the second vector in the actual sequence 710. This is because the first prediction (generated from the first packet) actually predicts the next (i.e., the second) packet. This skew is Figure 7 explicitly depicted in, but may be implied (i.e., not shown) in other figures in this document showing parallel sequences of input and output vectors without skew (e.g., Figure 1 ).

[0176] The horizontal line correlating the observed feature 712 with the predicted feature 731 implies that the vectors 712 and 731 can be compared and should more or less match without anomalies. The meaning of the skew is as follows. One vector of each sequence 710 and 730 is not relevant and cannot be compared to the corresponding vector. This includes the first vector 711 of the input sequence 710 and the last vector 733 of the output sequence 730. This is because the RNN 720: a) does not predict the first packet, and b) outputs a completely speculative (i.e., fictional) last prediction because, for example, when the network flow ends, there is no actual next packet to receive.

[0177] 8.1 Packet Anomaly Score

[0178] As shown by the horizontal line, all other vectors are correlated and compared in pairs. Each individual comparison requires a pair of opposite vectors, such as 712 compared to 731. Based on a prediction error calculation (such as mean squared), the fitness (i.e., closeness) or lack of fitness of each comparison can be measured as a grouped anomaly score, such as 741 - 742. For example, if vector 712 is most similar to 731 and vector 713 is divergent (i.e., different) from 732, then the grouped anomaly score 742 will exceed the grouped anomaly score 741.

[0179] Even though the individual grouped anomaly scores can clearly indicate which individual grouping is more surprising, the individual grouped anomaly scores need not have an actionable security association at the operational level. For actionable security, what matters is the flow anomaly score 750, which integrates the grouped scores 741 - 742 into a total score that reflects whether the overall network flow is anomalous. Different embodiments can use various mathematical integrations of the grouped anomaly scores 741 - 742 to arrive at the flow anomaly score 750. For example, the score 750 can be the maximum (or minimum) or mean of the scores 741 - 742.

[0180] 9.0 Sequence Prediction Process

[0181] Figure 8 is a flowchart depicting the computer 700 in an embodiment using the recursive topology of an RNN to generate grouped anomaly scores from which a flow anomaly score can be synthesized. Refer to Figure 7 for discussion Figure 8 .

[0182] For each feature vector of a sequence of actual dense feature vectors (such as actual sequence 710), steps 801 - 804 are repeated. Step 801 applies the current feature vector to the corresponding recursive step in the sequence of recursive steps. For example, the observed feature 711 is applied as a stimulus input to the recursive step 721 of the RNN 720. Since the RNN 720 is stateful, the recursive steps 721 - 723 can receive their corresponding inputs at different times, such as sequentially and possibly separated by various time delays.

[0183] In step 802, the recursive step outputs the next predicted dense feature vector, which approximates the next actual dense feature vector that appears in the sequence of actual dense feature vectors after the current feature vector. In other words, the RNN 720 may predict the next vector in the sequence before actually receiving the next vector. For example, the recursive step 721 generates the predicted feature 731 as an approximation of the expected observed feature 712, regardless of whether the observed feature 712 has actually been received.

[0184] An embodiment that compares the actual dense sequence with the predicted dense sequence may skip step 803, such as when both sequences 710 and 730 are dense. An embodiment that compares the actual unprocessed / sparse sequence with the predicted unprocessed / sparse sequence (as Figure 8 shown) should not skip step 803 when sequences 710 and 730 are both unprocessed / sparse. For example, the content shown as sequences 710 and 730 may actually each be a pair of sequences (not shown), where one is unprocessed / sparse and the other is dense, and one sequence in the pair is transcoded from the other sequence, as discussed elsewhere herein.

[0185] For example, RNN 720 may be the only component that accepts a dense vector as input and generates a dense vector as output. However, the rest of computer 700 may expect unprocessed / sparse vectors. Thus, step 803 generates the next predicted unprocessed feature vector in the sequence of predicted unprocessed feature vectors by decoding the next predicted dense feature vector. For example, the decoding may occur between recursive step 721 and predicted feature 731. Similarly, the encoding may occur between observed feature 711 and recursive step 721.

[0186] Step 804 compares the next predicted unprocessed feature vector with the next actual unprocessed feature vector to generate an individual anomaly score for the next actual packet. An embodiment may need to wait for the next actual unprocessed feature vector to arrive before performing step 804. For example, predicted feature 731 is compared with observed feature 711 to generate packet anomaly score 741.

[0187] After step 804 is performed on each feature vector in the sequence, all packet anomaly scores 741 - 742 for a given network flow have been calculated. If the individual scores of packet anomaly scores 741 - 742 are too high (i.e., clearly indicate an anomaly even if some actual sequences 710 have not been received and / or some packet anomaly scores have not been calculated), then the embodiment may stop repeating steps 801 - 804 and skip step 805.

[0188] Step 805 calculates the overall anomaly score of the network flow based on the individual packet anomaly scores. For example, flow anomaly score 750 may be the arithmetic sum of packet anomaly scores 741 - 742. Although only two packet anomaly scores are shown, the flow may have many packets and packet scores.

[0189] If the flow anomaly score 750 is incrementally computed, such as when computing each packet anomaly score at different times, such as when packets arrive as a flow in real-time and are separated by various delays, then embodiments can maintain / update the flow anomaly score 750 as a running total during a particular network flow. Embodiments can detect whether the flow anomaly score 750 exceeds a threshold (not shown) during each update of the flow anomaly score 750. Thus, in embodiments, the actual sequence 710 can be detected as an anomaly before all of the actual sequence 710 has arrived.

[0190] 10.0 Packet Decomposition

[0191] Figure 9 is a block diagram depicting an example computer 900 in an embodiment. The computer 900 sparsely encodes features based on the decomposition of a general packet of a given communication protocol. The computer 900 may be an implementation of the computer system 100.

[0192] The unprocessed feature vector 940 is more or less a direct encoding of the data from the packet 910. Although shown as not conforming to any particular protocol, it is generally expected that the packet 910 conforms to a given communication protocol. Various embodiments expect the packet 910 to be a datagram of the Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), or Simple Network Management Protocol (SNMP).

[0193] In an embodiment not shown, the unprocessed feature vector 940 is filled by copying a single bit string from the packet 910. In an embodiment, the single string of copied bits starts from the first bit of the packet 910. In an embodiment, the unprocessed feature vector 940 has a fixed size (i.e., amount of bits). In an embodiment, the single string of copied bits is padded or truncated to match the fixed size of the unprocessed feature vector 940.

[0194] In embodiments expecting packets of a given protocol, the packet 910 contains predefined fields such as the payload 929, and protocol metadata such as the header fields 920 - 928. In embodiments having a semantic mapping of features, specific fields of the packet 910 are encoded as specific semantic features at specific offsets within the unprocessed feature vector 940.

[0195] Typically, a specific fixed subset of positions is reserved for each encoded feature in the unprocessed / sparse feature vector. For example, packet 910 has a destination address 921 as a packet header field, and the destination address 921 is directly copied into bytes 5 - 8 of the sparse feature vector 940. In an embodiment, each legend 1950 ignores irrelevant fields (i.e., not encoded into the unprocessed feature vector 940), such as the checksum 922. Since computer 900 can extract content from all packets 910 in the form of discrete fields or the entire unprocessed byte string, the techniques presented herein can be implemented, supplemented, or otherwise combined with deep packet inspection, possibly accompanied by port mirroring or cable splitting.

[0196] In an embodiment, the fields of packet 910 are encoded into the unprocessed feature vector 940 according to data transformation or data conversion rather than a direct copy such as having a numerical range normalization (such as a unit (i.e., 0 - 1) range) or compressing a year (such as encoding 1970 as the zero year). In an embodiment, packet 910 may contain a category field having enumerated values, such as 928. For example, category 928 may have a range of string literal values that can be encoded into the unprocessed feature vector 940 as a dense integer or a sparse bitmap (such as using one - hot encoding, which reserves one bit for each literal within the value range).

[0197] One - hot encoding is illustrated by bitmap 960, which encodes category 928 into bytes 14 - 15 of the unprocessed feature vector 940. For example, if category 928 is month and the value is April (i.e., the fourth month of the year), then bit 4 of bitmap 960 is set and bits 1 - 3 and 5 - 12 are cleared. Naturally sorted (i.e., sortable) categories (such as months on a calendar) should be encoded as integers to preserve the sort order. For example, one of the twelve months can be encoded as a nibble (i.e., a 4 - bit integer with 16 possible values). Naturally unsorted categories (such as colors in a palette) should be one - hot encoded.

[0198] 10.1 Network Flow Demultiplexing

[0199] In embodiments of untying multiple intertwined network flows, a combination of header fields can identify a flow. For example, a flow can be identified by its pair of endpoints. For example, the combination of destination address 921 and destination port 924 can identify the destination endpoint, while other (e.g., similar) fields of packet 910 identify the source endpoint. In an embodiment, the endpoint identifier is used by the transport layer of the network stack such that the endpoint identifier is unique within a network or internetwork. In an embodiment, packet header fields 920 - 921 and 923 - 924 identify link layer (i.e., current hop) endpoints. In an embodiment, packet header fields 920 - 921 and 923 - 924 identify transport layer endpoints that are the original and final endpoints of a multi-hop (i.e., store-and-forward) route.

[0200] In an embodiment, a network flow is bidirectional between two endpoints such that the sequence of packets (not shown) consumed or emitted by the RNN (not shown) for prediction and the corresponding sequence of feature vectors include packets flowing in opposite directions, such as having a network round-trip, such as for request / response, such as for client / server, such as for HTTP, SSH, or telnet. As a stateful RNN, it can well tolerate the network latency inherent in the round-trip. Protocol metadata fields of packets that can facilitate (e.g., bidirectional) flow demultiplexing include source IP address, destination IP address, source port, destination port, protocol flag, and / or timestamp.

[0201] 11.0 Training

[0202] Figure 10 is a block diagram depicting an example RNN 1000 in an embodiment. RNN 1000 is configured and trained by a computer (not shown). RNN 1000 can be Figure 1 an implementation of RNN 170.

[0203] Before production deployment, RNN 1000 should be trained. RNN 1000 implements deep learning with a stack of multiple neural layers (such as 1011 - 1012). Each neural layer contains many neurons. For example, neural layer 1011 contains neuron 1021. Although not shown, each neuron has an activation value that measures the degree to which the neuron is fired (i.e., activated).

[0204] Adjacent layers are interconnected by many neural connections (such as 1030). Each connection has a learned (i.e., during training) weight that indicates the importance of the connection. For example, the weight of connection 1030 is 1040. Each connection unidirectionally passes a numerical value from a neuron in one layer to a neuron in the next layer based on the activation value of the source neuron and the weight of the connection.

[0205] For example, when calculating the activation value of neuron 1021, connection 1030 scales the activation value according to weight 1040 and then delivers the scaled value to neuron 1022. Although not shown, adjacent layers are typically richly interconnected (e.g., fully connected) such that each neuron has many connections in and out, and the activation value of a neuron is based on the sum of the scaled values delivered to the neuron through connections from neurons in the previous layer.

[0206] The configuration of RNN 1000 has two sets of attributes, namely weights (such as 1040) and hyperparameters (such as 1050). The weights are configured (i.e., learned) during training. The hyperparameters are configured before training. Examples of hyperparameters include the count of neural layers in RNN 1000, and the count of hidden units (i.e., neurons) in each hidden (i.e., internal) layer. The size of RNN 1000 is proportional to the number of layers or neurons in RNN 1000. The accuracy and training time are proportional to the size of RNN 1000.

[0207] RNN 1000 can have dozens of hyperparameters, where each hyperparameter has few or millions of possible values, such that the combination of hyperparameters and their values can be difficult to handle. For example, optimizing hyperparameters for RNN 1000 can be non-deterministic polynomial (NP) hard, which is important because sub-optimal hyperparameter values can significantly increase the training time or significantly reduce the accuracy, and manually tuning hyperparameters is a very slow process. Various third-party hyperparameter optimization libraries can be used for hyperparameter auto-tuning, such as Bergstra's hyperopt for Python, and other libraries available elsewhere that can be used for C++, Matlab, or Java.

[0208] During training, RNN 1000 will be activated with example network flows. The training does not need to be supervised, such as with some flows known to be anomalies and other flows known not to be anomalies. Instead, the training can be unsupervised such that the comparison of the predicted next packet with the actual next packet of the flow is sufficient to calculate the training error, even when it is not known whether the flow is actually an anomaly. The error during training causes the connection weights (such as 1040) to be adjusted. Techniques for optimal adjustment include backpropagation and / or gradient descent, as discussed in the following "Overview of Deep Training" section. The stateful (i.e., recursive) topology of RNN 1000 may impede typical backpropagation, and specialized techniques (such as backpropagation through time) are more suitable for training RNN 1000.

[0209] Training can end when the measured error drops below a threshold. After training, the RNN 1000 can more or less substantially be well represented as a two-dimensional matrix of learned connection weights. After training, the RNN 1000 can be deployed into a production environment as part of a network anomaly detector (such as for intrusion detection).

[0210] 12.0 Embedding Recorded Features

[0211] As discussed above, network packet analysis provides a way to monitor computer system usage. However, network traffic does not cover all computer system activities that may occur. For example, a computer virus can perform suspicious activities that are all local to the infected host and do not emit network traffic. However, a computer can record log messages as various operation messages, such as semantically rich diagnostic messages, for forensic purposes such as debugging and / or auditing of functions and / or security.

[0212] Each log message more or less details fine-grained computer activities among multiple activities that occur to implement a coarse-grained action. For example, a shell command can perform high-level work, such as listing a file system directory. The list command can retrieve metadata about each of multiple files in the directory and output the metadata for each file on a separate text line, which can be displayed in the display terminal of the command.

[0213] Each line of text can be recorded as a log message. Log messages can be spooled (i.e., recorded sequentially) to a text file, or stored as records in a relational database table. Log messages can be streamed, or buffered and flushed to standard output (stdout), and thus are suitable for I / O redirection and / or inter-process pipelines.

[0214] Figure 11 is a block diagram depicting an example computer 1100 in an embodiment. The computer 1100 has an RNN that transcodes a sparse feature vector representing a log message into a dense feature vector that can be predictive or used to generate a predictive vector. The computer 1100 can be an implementation of the computer 100.

[0215] Figure 11 Two additional concepts are explored. First, is the insight that deep learning for dense feature encoding and prediction is applicable to console log messages rather than network packets. Second, is that the transcoding from a sparse vector to a dense vector can include context states. Thus, dense encoding can be affected by vector ordering and data dependencies between vectors, which have been introduced earlier in this document and will be further discussed below.

[0216] The computer 1100 processes log messages (such as 1111 - 1113) rather than network packets. Each log message more or less details the fine-grained computer activities among the multiple activities that occur to achieve a coarse-grained action. For example, a shell command can perform high-level work, such as listing a file system directory. The list command can retrieve metadata for each of the multiple files in the directory and output the metadata for each file on a separate text line, which can be displayed in the display terminal of the command.

[0217] In some embodiments, each text line can be recorded as a log message. In other embodiments, only the list command is recorded as a log message, and the text lines are not recorded. In an embodiment, the log messages of the original log sequence 1110 are spooled (i.e., recorded sequentially) to a text file. In an embodiment, the log messages of the original log sequence 1110 are stored as records in a relational database table. In an embodiment, the original log sequence 1110 is flushed (i.e., streamed) to the standard output (stdout) and is suitable for I / O redirection and / or pipelining.

[0218] Each horizontal dashed arrow depicts the data flow for one log message as its data is transmitted and transformed to achieve the conversion from the original log sequence 1110 to the dense sequence 1140. The data of the log message can be decomposed into features. For example, the log message can contain an identifier for a session or transaction, which can be a feature. As shown, the first log message 1111 is encoded into the sparse feature vector 1121, which is applied to the recursive step 1131 of the encoder RNN 1130, and the encoder RNN 1130 serves as a transcoder to generate the dense feature vector 1141.

[0219] The log messages 1111 - 1112 may be the same. Since sparse coding is stateless (e.g., rule-based), the fact that log message 1111 precedes log message 1112 has no effect. Stateless sparse coding results in the same message being sparsely encoded identically, in which case the sparse feature vectors 1121 - 1122 will also be the same. For example, the log message 1111 can be encoded into the sparse feature vector 1121 by a non-recursive neural encoder (not shown) (such as a feed-forward MLP). In an embodiment, the adjacent neural layers of the encoder MLP are fully connected.

[0220] However, due to the encoder RNN 1130, the transcoding from the sparse sequence 1120 to the dense sequence 1140 is stateful. Thus, the transcoding here is context-sensitive because the encoder RNN 1130 has a stateful neural circuitry (e.g., LSTM), and because each recurrent step of the encoder RNN 1130 cross-activates the adjacent next recurrent step. For example, the recurrent step 1132 receives activations from both the sparse feature vector 1122 and the recurrent step 1131. In an embodiment, the RNN 1130 has as many recurrent steps as there are log messages in the longest expected original log sequence 1110.

[0221] The result of context (i.e., stateful) transcoding can be counterintuitive. For example, even if the sparse feature vectors 1121 - 1122 can be the same, they can be transcoded differently. For example, their corresponding dense feature vectors 1141 - 1142 can be different.

[0222] Even if the log messages 1111 - 1112 are almost the same but not absolutely identical, the dense sequence 1140 can still have context encoding that cannot be achieved by other encoding techniques. For example, the log message 1111 can have a default value (e.g., null (empty), zero, negative one, empty string) with a few more features than 1112. Important environmental (i.e., context) information can be encoded into the dense sequence 1140, such as whether the log message is before or after an almost identical message with more default values, or whether these two almost identical messages are adjacent or separated by (one or more) other messages.

[0223] Although not shown, the dense sequence 1140 can be used as an input to a predictive additional RNN (not shown), such as has been discussed for the previous figures herein. This predictive additional RNN can achieve more accurate predictions than the examples before herein because the transcoded vectors of the sequence 1140 are not only dense but also context-related.

[0224] In an embodiment not shown, the RNN 1130 is both an encoder RNN and a predictive RNN, and there is no additional RNN. For example, in the case of skewness, the dense vector 1141 can instead be a prediction of the next log message 1112.

[0225] The original sequence 1110 is shown to consist of log messages. In an embodiment, the original sequence 1110 instead contains network packets. The encoding, prediction, and anomaly detection discussed herein apply equally to log messages and network packets. For example, in the separate practice areas of applying logging and network communication, sequence anomaly detection can be an important security technique.

[0226] 13.0 Log Feature Embedding Process

[0227] Figure 12 is a flowchart depicting the computer 1100 in the embodiment using an RNN to transcode a sparse feature vector representing a log message into a dense feature vector for generating a predictive vector in context. Refer to Figure 11 for discussion Figure 12 .

[0228] Steps 1202, 1204, and 1206 are repeated for each log message in a sequence of related log messages (such as the raw log sequence 1110). Step 1202 extracts each feature from the log message to generate a sparse feature vector representing the feature. For example, a third-party log parser (not shown) (such as Splunk, Scalyr, Loggly, or Microsoft Log Parser) can extract the raw values from each of the log messages 1111 - 1113.

[0229] As discussed elsewhere herein, the extracted features of the log message can be sparsely encoded as a sparse feature vector. For example, the features of log message 1111 are encoded as sparse feature vector 1121. The extracted categorical features can be one-hot encoded, for example, as discussed for Figure 9 discussion.

[0230] Based on the sparse feature vector, step 1204 activates the corresponding step of the encoder RNN. For example, the sparse feature vector 1121 is applied as an excitation input to the recurrent step 1131 of the encoder RNN 1130.

[0231] In step 1206, the encoder RNN outputs a corresponding embedded feature vector based on the features of the current log message and the log messages that occurred earlier in the sequence of related log messages. For example, the recurrent step 1133 generates the dense feature vector 1143 based on the sparse feature vector 1123 as a direct input to the recurrent step 1133, and also generates the dense feature vector 1143 based on the cross-activation of the internal states of the previous recurrent steps 1131 - 1132 based on the previous sparse feature vectors 1121 - 1122. Thus, in context, the features are embedded into the dense feature vector 1143 based on multiple log messages of the raw log sequence 1110.

[0232] After performing step 1206 on each log message in the original log sequence 1110, some dense sequences 1140 have been generated. As has been described elsewhere in this document, step 1208 can process the dense / embedded feature vectors either batchwise or incrementally to determine the predicted next relevant log message that may occur in the sequence of relevant log messages. For example, the dense sequence 1140 can be applied as an excitation input to a predictive RNN (not shown) to predict the next log message in the sequence. For example, the dense feature vector 1141 can be applied to the first recurrent step of the predictive RNN to generate a next predicted dense feature vector (not shown) expected to match / approximate the next dense feature vector 1142, which can be used to detect whether the dense feature vector 1142 is contextually anomalous.

[0233] 14.0 Context Encoding Meaning

[0234] Figure 13 is a flowchart depicting the same sparse feature vector that is densely encoded differently according to context in the embodiment. Refer Figure 11 - 12 to discuss Figure 13 .

[0235] Figure 13 is an example of a specific scenario, which is a specialization of the general scenario of Figure 12 . For example, for the same log message sequence, the same process can perform Figure 12 - 13 processing steps.

[0236] Steps 1302 and 1304 process the first log message of the sequence. Steps 1306 and 1308 perform more or less the same processing on the second log message of the sequence. Thus, in Figure 13 , the steps 1202, 1204, and 1206 (e.g., a control loop) that are repeated for each log message in Figure 12 are instead shown as a loop that unfolds into separate steps for each of two adjacent log messages.

[0237] Specifically, step 1302 can be step 1202 for the first log message. Step 1304 can be steps 1204 and 1206 for the first log message. Step 1306 can again be step 1202 for the second log message. Step 1308 can again be steps 1204 and 1206 for the second log message.

[0238] The purpose of depicting the unfolding loop is to illustrate that more or less identical log messages can have different embeddings (i.e., context encodings) because they appear in different message sequences or at different positions within the same sequence, perhaps adjacent to each other. Step 1302 extracts the features of the first log message and encodes them into a first sparse feature vector. For example, the parser can extract feature values from each log message in log messages 1111. The features of log messages 1111 are encoded as sparse feature vector 1121.

[0239] Step 1304 embeds the first sparse feature vector into a first dense feature vector. The feature embedding is context-sensitive such that the message sequence and position within the sequence of the sparse feature vector affect the feature embedding. For example, how encoder RNN 1130 embeds sparse feature vector 1121 into dense feature vector 1141 depends on the fact that sparse feature vector 1121 is in sparse sequence 1120 and is the first vector in that sequence, in which case encoder RNN 1130 has most recently been reset and has not yet accumulated any internal state.

[0240] Step 1306 extracts each feature from the second log message to generate a second sparse feature vector. Since sparse coding is not context-sensitive, each log message can be sparsely encoded individually (i.e., without regard to other log messages). For example, the sparse coding of log messages 1111 - 1112 occurs independently. Since log message 1112 is identical to log message 1111 in this example, their sparse codings are the same even though log messages 1111 - 1112 are transcoded independently without knowledge of each other. Thus, sparse feature vectors 1121 - 1122 are the same when log messages 1111 - 1112 are the same.

[0241] However, context encoding (i.e., feature embedding) does not preserve identity. In step 1308, the encoder RNN outputs a second embedded feature vector that is based on the feature and one or more log messages that occurred earlier in the sequence of relevant log messages. For example, unlike log message 1111, log message 1112 is not the first in the sequence and encoder RNN 1130 has since accumulated some internal state that affects the feature embedding. Thus, even though log messages 1111 - 1112 are identical and sparse feature vectors 1121 - 1122 are the same, dense feature vectors 1141 - 1142 are different.

[0242] 15.0 Training harness

[0243] Figure 14 is a block diagram depicting an example computer 1400 in an embodiment. Computer 1400 has a training harness to improve dense coding. Computer 1400 can be an implementation of computer 100.

[0244] When deployed in production, computer 1400 has a prediction pipeline (i.e., the white boxes shown according to legend 1490) that flows more or less from left to right. This pipeline contains layers trained with the depth (i.e., many layers) of stacked RNNs 1410 and 1430, which perform dense encoding and prediction respectively to calculate the anomaly score 1450. The encoder RNN 1410 can transcode a sparse feature vector (not shown) of a log message (not shown) into a dense vector (not shown) of a dense sequence 1420, and this dense sequence 1420 is applied to the predictor RNN 1430.

[0245] 15.1 Decoding for reconstruction

[0246] A potential obstacle to training the encoder RNN 1410 may be that backpropagation through time requires too many layers (RNNs 1410 + 1430) to be properly tuned. Instead, a training harness (i.e., the shaded blocks) containing the decoder RNN 1460 can be used to improve the training accuracy and acceleration, especially for the encoder RNN 1410. This training harness converts the dense vectors of the dense sequence 1420 back into sparse vectors. The training harness does not need to be deployed into production.

[0247] The reconstructed unprocessed (i.e., sparse) vectors (such as 1465) generated by the decoder RNN 1460 can be compared with the original unprocessed vectors (e.g., 1451 - 1452) and consumed by the encoder RNN 1410 to calculate the reconstruction loss 1470 that measures the accuracy of the RNNs 1410 (and 1460). For example, in the illustrated embodiment, the original unprocessed vector 1452 contains the session type and transaction type as one-hot encoded categorical features. For example, the session type is a categorical feature with three mutually exclusive possible values. For example, in the original unprocessed vector 1452, the middle value of the three values for the session type is set (circled).

[0248] 15.2 Measured error

[0249] During training, two pipelines (i.e., RNNs 1430 and 1460) consume the dense sequence 1420. Each pipeline has its own error function such that the reconstruction loss 1470 measures the encoding accuracy of the current log message, and the prediction loss 1440 measures the prediction accuracy of the next log message. The reconstruction loss 1470 can be calculated based on comparing the original and reconstructed raw vectors of the dense sequence 1420. For example, the reconstructed raw vector 1465 can be compared to the original raw vector 1452. In an embodiment, the values within the reconstructed raw vector 1465 are probabilistic such that an original value of "1" is approximated by a probability close to "1", such as "0.81" as shown in the circle. Similarly, the prediction loss 1440 can be calculated based on comparing the original vector of the dense sequence 1420 to the predicted vector. Either or both of the losses 1440 and 1470 may require calculating and integrating (e.g., summing) the individual losses for each vector in the sequence. For example, the individual losses for the raw / sparse or dense vectors can be measured by comparing the corresponding vectors (not shown) of the original sequence to the predicted sequence (such as for Figure 7 discussed).

[0250] Both the losses 1440 and 1470 can provide relevant feedback to the RNNs of the pipelines. In an embodiment, the losses 1440 and 1470 are integrated into a combined training loss 1480 that is used for backpropagation through the full stack of RNNs 1410, 1430, and 1460, which should increase the fitness of the dense encoding of sequences such as 1420. The losses 1440 and 1470 can be summed or weighted averaged to calculate the training loss 1480.

[0251] The anomaly score 1450 can indicate how strange (i.e., unexpected) a given log message sequence (such as the log message sequence from which the dense sequence 1420 is generated) is. In production, there may be no training harness. Thus, the anomaly score 1450 is based on the prediction loss 1440. In an embodiment not shown, an individual message anomaly score is calculated for each actual log message of the sequence, and these scores are integrated (e.g., averaged) to calculate the anomaly score 1450.

[0252] When the anomaly score 1450 exceeds a threshold (not shown), an alert can be generated. For example, the anomaly score 1450 can be used to detect and warn of various potential problems, such as online security intrusions or Internet of Things (IoT) malfunctions. In the case of detecting an IoT malfunction, the computer 1400 analyzes a central log of remote telemetry messages from potentially many IoT devices. The central logging may need to be unwound (i.e., demultiplexed) such that the dense sequence 1420 contains only the feature vectors for a particular IoT device. The dense sequence 1420 can also be restricted to a particular subsequence of log messages, such as log messages related to the same (sub) activity (e.g., trace), as discussed later in this document.

[0253] 16.0 Training with Reconstruction

[0254] Figure 15 is a flowchart depicting a training harness for improving dense coding in an embodiment. Refer Figure 14 to discuss Figure 15 . Figure 15 Illustrates the behavior of the training, some of which may also occur in production or may not occur in production. These behaviors include decoding, measuring the reconstruction loss, and backpropagation.

[0255] Based on the embedded feature vector, step 1502 activates the decoder RNN to decode the embedded feature vector into a reconstructed sparse feature vector that approximates the original sparse feature vector representing the features of the log message. For example, the dense sequence 1420 can have dense feature vectors as context embeddings of the log messages 1451 - 1452. The decoder RNN 1460 can decode the dense vectors of the dense sequence 1420 back into a reconstructed unprocessed vector (such as 1465) representing the original unprocessed vector 1452.

[0256] Step 1504 compares the original sparse feature vector with the reconstructed sparse feature vector to calculate the reconstruction loss. Ideally, the original and the reconstructed sparse feature vectors are the same, in which case the reconstruction loss 1479 is zero. For example, the reconstruction loss 1470 can be the sum of the absolute values of the differences between the corresponding numerical fields in the original and the reconstructed sparse feature vectors.

[0257] Step 1506 sums the prediction error and the reconstruction error to calculate the training error. For example, the predictor RNN 1430 can output a predicted sequence that is different from the actual dense sequence 1420, which is measured as the prediction loss 1440. Similarly, the reconstruction loss 1470 is a comparable measure of the loss observed elsewhere within the neural topology for the same sequence of log messages. The losses 1440 and 1470 are arithmetically summed to calculate the total training loss 1480 that indicates the overall accuracy.

[0258] Step 1508 backpropagates the training error. Backpropagation over time can be used to penetrate deep neural layers to affect deep learning. In an embodiment, convergence is achieved and training stops when the training loss 1480 drops below a threshold or fails to improve (i.e., decrease) by at least a threshold amount.

[0259] 17.0 Activity Diagram

[0260] Figure 16 is a diagram depicting the example activity diagram 1600 in an embodiment. The activity diagram 1600 represents computer system activities occurring on one or more interoperating computers.

[0261] Each vertex in the activity diagram 1600 represents a log trace. A log trace is a closely related subsequence of log messages generated during and about the same action. For example, during an observed action, a sequence of log messages can be appended to a message log, such as a console log file. The message log can be generated directly by a software application or delegated by the application to a structured logging framework (such as syslog for Unix or auditd).

[0262] For example, auditd can sometimes issue ten log messages for a single user login, which can be one trace. Trace detection (i.e., forensically associating multiple log messages into a single trace) is discussed and can be an orthogonal issue to the activity diagrams later in this document. For example, some embodiments can analyze activities but not traces, and vice versa. In this example, each vertex in the activity diagram 1600 can represent multiple log messages but only one log trace.

[0263] The logging framework can aggregate logs from multiple processes, applications, and / or computers. The logging framework can insert semantic features into the log messages and suppress recording or adjust the format of some messages according to tunable verbosity, message characteristics, or additional criteria. Some telemetry frameworks (such as Graphite from Orbitz) can be used as remote log spooling (i.e., telemetry and archival) frameworks. The logging framework can provide more or less support for trace detection, such as using some form of explicit (e.g., text adornment) delimiters.

[0264] Each log message can contain values for multiple semantic features. Similarly, each log trace can contain values for multiple semantic features, which can be extracted from or otherwise derived from the characteristics of the log messages of the trace. Traces can be correlated with each other based on the consistency of their (one or more) feature values.

[0265] For example, traces can be correlated, such as in the case where a sequence of suspicious script execution actions (i.e., traces) is carried out to achieve a malicious goal. Traces can be related to a computer, a network address (such as an IP address), a session, a transaction, a connection, an application, a web browser, a shell, a login, an account, a principal, an end user, or other abstractions of an agent in a computer or computer network. Such an agent can be generalized as a network identity, and each network identity has a network identifier, which can be used to merge related log messages into a log trace, as discussed later in this article.

[0266] As shown, the connection between two vertices (i.e., traces) indicates that the two traces are related, such as through the consistency of eigenvalue. As shown, the connection can be marked to reveal which (which) consistent features and which value are consistent under the connection. For example, the login, file access, and logout vertices of the activity diagram 1600 are interconnected by edges (i.e., connections) that are consistent on the same user John (i.e., through their association), as the same value of the same feature. However, the file access and data transfer vertices are related by different types of features.

[0267] In fact, with log messages with many features and long decorations, traces can become intricately correlated with each other, so that the traces together form a more or less connected activity diagram 1600. The edges between traces provide rich topological context, and this topological context can be used to implement the context feature encoding of each trace from the activity diagram 1600. For example, the login itself may not be suspicious, and in this case, any feature vector encoding the login will not be regarded as suspicious in isolation. However, the login may actually be suspicious (i.e., abnormal) in the context of a certain pattern of related traces.

[0268] The potential problem is that although there is topological connectivity, many edges may be more or less contextually irrelevant. For example, given a specific vertex (such as the file access trace (shown in the center of the activity diagram 1600)), many or most of the activity diagram 1600 (i.e., unlabeled edges and vertices) are more or less irrelevant (i.e., noise). What is actually relevant (relative to the file access trace) is only the subgraph of the labeled edges and vertices. Therefore, the file access trace can be embedded in a highly relevant subgraph, which is the essence of graph embedding, as a preface to the context encoding of features (such as traces).

[0269] Defining the scope of the relevant subgraph may require thoroughly pruning the activity diagram 1600 by selecting edges that meet the relevance criteria (such as those with specific features and / or values). Another relevance criterion may require traversing the scope (i.e., the maximum radius of the traversal path of multiple edges starting from the file access vertex).

[0270] 18.0 Graph Embedding

[0271] Figure 17 is a block diagram depicting an example computer 1700 in an embodiment. The computer 1700 uses graph embedding to improve the feature embedding of log traces. The computer 1700 may be an implementation of the computer 100.

[0272] A log trace is a closely related subsequence of log messages generated during the same action. For example, during the observed action 1713, a series of log messages as shown in the log trace 1733 are appended to the log 1720, and the log message may be a console log file.

[0273] Each log trace in the log 1720 can be encoded sparsely or densely as a raw trace feature vector, such as 1751 - 1757. Each raw vector 1751 - 1757 contains encoded values of features, such as network identities. The trace 1733 represents the log messages for one activity. The horizontal arrow passing through the log trace 1733 depicts the data flow from the trace 1733 to the raw vector 1753 to the embedded trace vector 1763, including transformation and transmission, all of which represent the recording or encoding of the observed action 1713.

[0274] Parallel to the horizontal arrow passing through the trace 1733 are other horizontal arrows, which represent the data flows of other log traces that are somewhat similar. Even if not shown, each horizontal arrow has its own observed action that causes its own log trace in the log 1720. Thus, the log 1720 has many traces, and each trace may have multiple log messages. Depending on the embodiment, and although the log 1720 is shown as containing discrete traces (e.g., 1733), the log 1720 may alternatively contain only log messages from which traces must be inferred. Trace detection will be discussed further later in this document.

[0275] A suspicious process, script, or interactive session may perform many more or less related activities (such as 1713). Related traces can be identified when the values of the same feature or set of features match. For example, the raw vectors 1752 - 1753 and 1755 - 1756 are related because they have 55 as the same network identity. For example, the network identity may include one or more fields such as: username, computer hostname, Internet Protocol (IP) address, or an identifier for a session or transaction.

[0276] Each of the original vectors 1751 - 1757 can be a vertex in the exhaustive logic graph 1740, which may or may not be actually constructed (e.g., in the RAM of computer 1700). Associated original vectors are shown connected by edges. For example, edge 1771 connects original vectors 1752 - 1753 associated with the same network identity 55. Original vectors can be associated with other kinds of features. For example, edge 1774 connects original vectors 1755 and 1757 associated with features that are not network identities (not shown).

[0277] 18.1 Pruning

[0278] Graph embedding is the construction of subgraph 1740 as a context for a given trace (such as represented by original vector 1753). Subgraph 1740 is connected such that all vertices 1752 - 1753 and 1755 - 1756 can reach each other directly or indirectly (i.e., via multi - edge paths). While not necessarily fully connected (i.e., all vertices of subgraph 1740 can be directly adjacent to each other), the connection means that all vertices of subgraph 1740 have some traversal path to the vertex of the current focus (i.e., the vertex being context - encoded currently) (such as 1753). Graph embedding and trace detection can be orthogonal problems. Some embodiments can perform trace detection without graph embedding and vice versa.

[0279] Prune subgraph 1740 such that edges that do not meet the trace detection criteria are excluded from subgraph 1740. For example, original vectors 1755 and 1757 share the same trace feature value that does not meet the trace detection criteria, such that edge 1774 is excluded (i.e., pruned) from subgraph 1740. Identifiers or other values may be too common to be useful. For example, the same trace feature value can be a very common value that appears in many (e.g., most) traces, which would define a subgraph that is (e.g., almost) as large as the exhaustive graph 1740, which is too large to be relevant for context encoding. In an embodiment, common identifiers are excluded from forming edges.

[0280] 18.2 Log Traces

[0281] Although not shown, a trainable algorithm (e.g., a deep - learning algorithm such as an RNN or an unsupervised artificial neural network) can also context - encode the original trace vector 1753 as an embedded trace vector 1763 (i.e., a trace associated with trace 1733) based on the original vectors 1752 and 1755 - 1756 of the traces that occur in subgraph 1740. For example, each of the original vectors 1752 - 1753 and 1755 - 1756 can be applied to a corresponding step of the RNN to cause the RNN to generate the embedded trace vector 1763.

[0282] Traces within log 1720 can overlap (such as in the case where the log messages that compose it are interleaved. Interleaving is typically caused by concurrency, such as due to multithreading, symmetric or other multiprocessing (SMP) (such as multi-core), or distributed computers that stream log messages to the same consolidated (i.e., aggregated) log 1720 stream. Thus, trace detection may require de-interleaving (i.e., demultiplexing) the log messages into separate traces.

[0283] As a result of trace overlap, multiple traces (e.g., live) can be in progress at the same point within log 1720. Thus, the logic for trace detection should detect multiple traces, or have multiple concurrent instances of trace detection logic (e.g., compute threads), each instance detecting a different trace. For example, subgraph 1740 can be one of multiple subgraphs constructed and / or processed simultaneously from graph 1740.

[0284] In an embodiment, a log message can be only part of one trace, and subgraphs do not overlap (i.e., share some vertices / traces). In such an embodiment, each step (i.e., fold) of the context encoder RNN (not shown) can generate a separate embedded trace vector for each corresponding raw vector in the subgraph. For example, the RNN generates not only the embedded trace vector 1763 from subgraph 1740, but also separate embedded trace vectors for the corresponding raw vectors 1752 and 1755 - 1756 from the corresponding steps of the RNN.

[0285] In an embodiment, computer 1700 uses third-party Python tools for graph embedding (i.e., subgraph definition and context feature encoding), such as GraphSAGE, DeepWalk, or Node2Vec for TensorFlow. DeepWalk scales horizontally. TensorFlow can be accelerated with GPUs and is also available in JavaScript for deployment in web pages.

[0286] 19.0 Graph Embedding Process

[0287] Figure 18 is a flowchart depicting graph embedding in an embodiment to improve the feature embedding of log traces. Refer to Figure 17 for discussion Figure 18 .

[0288] Step 1802 receives independent feature vectors, each representing a corresponding log trace. For example, log 1720 has traces such as 1733, which are encoded as raw feature vectors 1751 - 1757. In an embodiment, the raw feature vectors 1751 - 1757 are sparsely encoded, which maintains discrete fields (i.e., features) such that the raw vectors can be correlated based on matching values of specific fields.

[0289] Step 1804 generates the edges of a connected graph that includes a particular vertex generated from a particular log trace represented by a particular independent feature vector. For example, the original vector 1753 may be the feature vector that is the current focus of the graph embedding. Based on the matching values of the (one or more) particular fields, other original vectors may be directly or indirectly related to the original vector 1753.

[0290] The filtering criteria may specify which fields should match to include an original vector in the pruned connected subgraph 1740. For example, edges 1771 - 1773 that define the pruned connected subgraph 1740 may be created based on the matching according to the filtering criteria. Thus, the pruned connected subgraph 1740 is the graph embedding of the original vector 1753.

[0291] Step 1806 generates an embedded feature vector based on the independent feature vector that represents the log trace from which the vertex of the connected graph is generated. For example, the original vector 1753 may be transcoded into an embedded trace vector 1763 based on the pruned connected subgraph 1740. Thus, the embedded trace vector 1763 is the graph embedding (i.e., context encoding) and dense transcoding of the original vector 1753.

[0292] Thus, the embedded trace vector 1763 may contain more information than the original vector 1753, and of course more relevant information. For example, the additional information (i.e., context) in the embedded trace vector 1763 that is not in the original vector 1753 may result in the difference between an accurate alert and a false negative. The context relevance can be maximized by embedding performed by a trained graph embedder (not shown) such as the MLP or RNN discussed herein.

[0293] Step 1808 depicts possible further processing after the graph embedding. Embodiments not shown may instead perform alternative processing at step 1808. Based on the embedded feature vector, step 1808 indicates that a particular log trace is anomalous.

[0294] For example, the computer 1700 may generate other pruned connected subgraphs in addition to 1740 as the graph embeddings of other subsets of traces from the log 1720. The sequence of these graph embeddings may activate a predictive RNN (not shown) whose output may be compared to the original sequence to detect an anomalous predicted sequence. Multiple pruned connected subgraphs of the log 1720, such as the subgraph 1740, may overlap such that they share vertices (original vectors).

[0295] In an embodiment, each raw vector within pruned connected subgraph 1740 is a central vertex within a corresponding separated subgraph. For example, raw vector 1752 may be a central vertex of a different pruned connected subgraph (not shown) that slightly overlaps subgraph 1740.

[0296] In an embodiment, the pruned connected subgraphs do not overlap, and each original vector appears in exactly one pruned connected subgraph. For example, each of the original vectors 1752-1753 and 1755-1756 appears only in the pruned connected subgraph 1740. The pruned connected subgraph 1740 can be operated as a graph embedding for all the original vectors 1752-1753 and 1755-1756, in which case an additional embedded trace vector (not shown) is generated from the pruned connected subgraph 1740 in addition to the embedded trace vector 1763.

[0297] 20.0 Trace Aggregation

[0298] Figure 19 1 is a block diagram depicting an example log 1900 in an embodiment. A computer creates a temporally pruned subgraph from related traces in log 1900. Although log 1900 is shown as storing traces 1-9, this document is specific to Figure 19 The described time pruning techniques can be applied to sequences of network packets or log messages, rather than traces.

[0299] Graph pruning aims to improve feature relevance by reducing the context of traces (i.e., related traces) to include only semantically close traces (i.e., subgraphs). In addition to pruning content (i.e., semantics), subgraphs can also be pruned in time. For example, traces 1-9 can be time series events, so that the contents of log 1900 are time-ordered (e.g., naturally spooled), so that the contents of log 1900 can be sorted along the timeline according to the shown order of traces 1-9.

[0300] For example, trace 1 appears first, and trace 9 appears last. In an embodiment, traces may have timestamps, such as times 1-9, such as the timestamp of the first or last log message (not shown) of each trace. As shown, traces 1-9 are ordered, but not necessarily equidistant in time. For example, the duration between times 1-2 may be different from the duration between times 2-3. For example, times 1-2 may occur simultaneously.

[0301] A fixed amount of time or a time window of traces may be imposed during subgraph pruning. For example, according to the example 1950, only traces 3-7 are currently considered to be contained in the same subgraph. Each subgraph may have its own time window.

[0302] In the illustrated embodiment, if trace 3 is recognized as starting a new subgraph (e.g., when trace 3 is received), and the size of all time windows is set to include five time units or traces, then a window for the subgraph is created to span time 3 - 7 or traces 3 - 7. In an embodiment, the time window spans at least an entire day, which can be resource intensive as the time window can also operate more or less as a retention window for buffering real-time traces and / or log messages.

[0303] In an embodiment with time windows measured in terms of traces instead of time, the maximum number of traces in a subgraph does not exceed the size of the window. Regardless of how the window is measured, trace feature semantics can provide additional subgraph inclusion criteria such that only some of traces 3 - 7 in the window of the subgraph are eligible to actually be included in the subgraph (not shown).

[0304] In a different embodiment (also shown), the time window can stretch (i.e., time or trace expansion) to accommodate a larger subgraph. According to legend 1950, so far the window can be a small peephole at the end of the subgraph such that whenever another trace is added to the subgraph, the peephole slides (i.e., shifts), which is similar to max_gap described later in this document. For example, as shown, trace 5 can be received, and the peephole covering two time units or trace size should cover traces 6 - 7 that should appear after trace 5 and may not have been received yet.

[0305] For example (not shown), if trace 6 is received and the semantic criteria to be included in the same subgraph as trace 5 are met, then the 2-unit window can slide from time / trace 6 - 7 to time / trace 7 - 8. If the sliding window passes (i.e., none of the covered traces meet the eligibility for subgraph inclusion), then the filling of the current subgraph closes (i.e., is completed), and the traces currently covered by the window will be included in one or more other subgraphs.

[0306] 21.0 Trainable Anomaly Detection

[0307] Figure 20 is a block diagram depicting an example computer 2000 in an embodiment. Computer 2000 has a trainable graph embedder that generates context feature vectors, and a trainable anomaly detector that consumes the context feature vectors. Computer 2000 can be an implementation of computer 1700. Although time window 2010 shows traces 2021 - 2025, the anomaly detection techniques described herein for Figure 20 can be applied to sequences of log messages or network packets instead of traces.

[0308] A subset of the traces 2021 - 2025 in the time window 2010 satisfies the semantic relevance criterion, including traces 2021, 2023, and 2025, which were initially encoded as the corresponding raw vectors 2031, 2033, and 2035, and the trainable graph embedder 2060 transcodes them into the corresponding embedded trace vectors 2041, 2043, and 2045. The trainable graph embedder 2060 can include deep learning algorithms such as RNN or unsupervised artificial neural networks such as for Figure 17 discussed.

[0309] The trainable anomaly detector 2050 learns how to decide (yes or no as shown in the figure) whether the sequence of the embedded trace vectors 2041, 2043, and 2045 is anomalous. Depending on the embodiment, the embedded trace vectors 2041, 2043, and 2045 are applied to the trainable anomaly detector 2050 incrementally (i.e., sequentially) or more or less simultaneously. For example, the embedded trace vectors 2041, 2043, and 2045 can be concatenated to create a combined feature vector, which will be applied to the trainable anomaly detector 2050.

[0310] In an embodiment, the trainable anomaly detector 2050 includes at least one of the following non - neural algorithms: principal component analysis (PCA), isolation forest (iForest), or one - class support vector machine (OCSVM). PCA is popular due to its ease of configuration and use as well as its predictable (i.e., stable) performance. iForest operates efficiently through optimized sub - sampling - based calculations and avoids computationally expensive mathematical operations such as distances or densities. OCSVM is fast, tolerant of noise, and has unsupervised training. Example neural and non - neural implementations based on trainable artificial intelligence (AI) are discussed in the "Machine Learning Overview" section below.

[0311] The trainable anomaly detector 2050 can include deep learning algorithms such as RNN or unsupervised artificial neural networks such as for Figure 17 discussed. In an embodiment, the trainable anomaly detector 2050 is an auto - encoder with unsupervised training, as for Figure 3 discussed.

[0312] 22.0 Trace Composition

[0313] Figure 21 is a block diagram depicting an example computer 2100 in an embodiment. The computer 2100 detects independent traces from relevant log messages and encodes their features. Earlier in this document, event aggregation based on feature consistency and time windows was discussed to merge multiple log traces into Figure 19 - 20 a sub - graph. Figure 21 - 23Have a finer-grained unprocessed input, enabling multiple log messages to be merged into a log trace, and the criteria for time and semantic merging will be discussed in more detail.

[0314] In an embodiment, the computer 2100 can generate a trace vector consumed by the computer 2000 as input for graph embedding. For example, Figure 21 The suspicious feature vector 2140 of can subsequently be used as Figure 20 The original vector 2031 of.

[0315] The log (not shown) can contain many log messages, such as 2131 - 2133. By parsing each message into key-value pairs, features are extracted from the log messages. For example, the log message 2131 has a key-value pair where the session ID is the key and XYZ is the value.

[0316] The computer 2100 can use a third-party log parser (not shown) such as Splunk to parse the log messages 2131 - 2133 and extract the key-value pairs. In an embodiment, the log parser encodes and outputs the key-value pairs in a serialized format such as Extensible Markup Language (XML), comma-separated values (CSV), JavaScript Object Notation (JSON), YAML, or Java object serialization.

[0317] 22.1 Composition criteria

[0318] The filtering criteria 2110 specify one or more scenarios in which a subset of the log messages 2131 - 2133 can be merged into a single trace (such as 2120). Each specified scenario identifies one or more features whose actual values should match all the log messages in the trace. For example, the log messages 2131 - 2132 can be merged into the trace 2120 because the filtering criteria 2110 specify a scenario that identifies the session ID and time zone as features, and all the log messages in the same trace should share the same values for that feature. In this case, although not shown, the log message 2132 has the same session ID and time zone values as the log message 2131. While the session ID and / or time zone of the log message 2133 can have different values, or one or both of these keys can be missing, such that the log message 2133 is not merged into the trace 2120.

[0319] The filtering criteria 2110 can specify other scenarios, such as merging log messages with the same color. Even if the log messages 2131 - 2132 do not share the same color, the log messages 2131 - 2132 can be merged into the trace 2120 as long as the log messages 2131 - 2132 meet the different scenarios of the filtering criteria 2110.

[0320] Based on the techniques herein, trace 2120 can be encoded densely or sparsely as a suspicious feature vector 2140. A trainable anomaly detector 2150 can analyze the suspicious feature vector 2140 to determine whether the trace 2120 is anomalous. In an embodiment, the trainable anomaly detector 2150 analyzes feature vectors of additional traces (not shown) to decide whether the suspicious feature vector 2140 is anomalous, such as according to the context techniques herein. Regarding Figure 22 - 23 Trace detection according to the filtering criteria 2110 is discussed in more detail.

[0321] 23.0 Trace Detection Process

[0322] Figure 22 is a flow chart depicting the detection and feature encoding of individual traces from related log messages in an embodiment. Refer to Figure 21 Discuss Figure 22 。

[0323] Step 1802 extracts key-value pairs from each log message. For example, each of the log messages 2131 - 2133 can be independently parsed into various key-value pairs. The key-value pairs can be stored in memory in one or more of various formats, such as a hash table or other associative data structure and / or sparse feature vector.

[0324] Step 1804 detects log traces representing a single action based on a subset of log messages whose key-value pairs satisfy the filtering criteria. In other words, a set (i.e., trace) of (not necessarily adjacent) log messages is identified as related to the execution of a single observed action (not shown). Thus, log traces (such as 2120) have a coarser granularity than individual log messages, which facilitates context encoding. Thus, the logs that originate as sequences of log messages can be abstracted into more meaningful multiple traces, where each trace consists of one or more log messages.

[0325] Traces do not overlap and do not share log messages. Thus, a log message belongs to exactly one trace. Thus, there is a rather strict inclusion (i.e., nesting) hierarchy for log messages within a trace and traces within a log (not shown).

[0326] Step 1806 generates a suspicious feature vector representing the log trace based on the key-value pairs from the subset of log messages of the trace. For example, the log messages 2131 - 3132 of trace 2120 can be embedded (i.e., contextually encoded) as a dense suspicious feature vector 2140.

[0327] Based on the feature vectors including the suspicious feature vector, the anomaly detector indicates in step 2208 that the suspicious feature vector is an anomaly. For example, the trainable anomaly detector 2150 can contextually analyze a sequence of embedded / dense feature vectors of relevant traces including the trace 2120 to detect whether the suspicious feature vector 2140 is anomalous. For example, a predictive RNN (not shown) can predict a sequence of feature vectors, including corresponding predicted feature vectors that do or do not match the suspicious feature vector 2140. The trainable anomaly detector 2150 can decide (yes or no) whether the suspicious feature vector 2140 has or has not been approximately predicted with a threshold accuracy.

[0328] 24.0 Declarative Trace Detection

[0329] Figure 23 is a block diagram depicting an example log 2300 in an embodiment. The log 2300 contains semi-structured operational (e.g., diagnostic) data from which log messages and their features can be parsed and extracted. The log 2300 can be a live data stream, a log file, a spreadsheet, or a database table containing a (e.g., time) sequence of log messages that are initially output to one or more consoles of one or more software applications residing on one or more host computers.

[0330] The log 2300 is depicted as a table where each row represents a log message and each column represents a field (i.e., a feature). The trace and offset columns to the left of the log 2300 are declarative and do not actually exist in the log 2300. The offset column uniquely identifies each log message and can be a text line number in a file or stream, a database record / row number, or a byte offset in a file or stream (in which case the offset would increase monotonically but could be discontinuous (different from that shown)).

[0331] The features (i.e., columns) can have application-specific semantics (e.g., price) or can represent metadata explicitly shown in the (one or more) log messages. For example, the host field can indicate which server computer generated the log message or which client computer caused the server computer to generate the log message. The log 2300 can integrate log messages generated by many server computers.

[0332] The trace column indicates which log messages are merged into the same trace. For example, trace A consists of log messages 0 - 2. Traces can overlap (i.e., have interleaved log messages). For example, even though adjacent log messages 11 and 13 belong to trace C, log message 12 belongs to trace D.

[0333] 24.1 Declarative Rules

[0334] Apply various logical rules to determine which log messages to merge into which traces, such as the following. If a rule applies to an interesting merge scenario discussed herein, then the relevant feature values of the relevant log messages that satisfy the rule are depicted on the black box. Log messages 0 - 31 are merged into traces A - I according to the following rules.

[0335] Each feature (i.e., the same key for key - value pairs) presents a dimension (i.e., degree of freedom), which can be used to merge log messages into traces. Rules (e.g., Figure 21 the filtering criterion 2110 on

[0336] node_filter: “node =”

[0337] entity_filter: “*pid =”

[0338] time_window_filter: (min_duration: 10s, max_duration: 60s, max_gap: 10s)

[0339] exclude_filter: “exclude =”

[0340] nomatch_heuristic: “closest active trace”

[0341] Each type of filter in the above example has predefined semantics and configurable property(ies). For example, a predefined node filter to identify that multiple log messages relate to the same server (or client) machine. The node filter can be configured to use the node feature of log 2300 (which happens to be in the host field) to detect the host machine of the log message. For example, log messages 10 - 11 relate to different host machines and thus will not be merged into the same trace.

[0342] A predefined entity filter to identify that multiple log messages relate to the same operating entity, such as the same user, principal, customer, account, or session. The feature name can have wildcards, such as an asterisk or a regular expression. Additionally, the same feature can appear in multiple fields. For example, *pid can match the pid in field X of log message 4 and also match the spid in field Z of log message 8.

[0343] A predefined time_window_filter is used to identify that multiple log messages occur simultaneously. The time_window_filter has various configurable attributes. The min_duration attribute can be configured such that, in cases where the current trace being merged does not already contain the first and last log messages with a time interval of at least min_duration, it generally prevents adjacent log messages from being isolated into separate traces. Thus, the length of a normal trace is at least min_duration.

[0344] The max_duration attribute limits how much time can elapse between the first and last log messages of a trace. Thus, a normal trace does not exceed max_duration.

[0345] Even within the min_duration and max_duration, time constraints can cause log messages to be isolated into separate traces. For example, the max_gap attribute specifies the longest time between any two adjacent log messages within a trace. Thus, max_gap is the longest time that the trace detector can wait for another log message to be appended to the current trace.

[0346] Filtering rules can have priorities such that even if a log message can satisfy multiple rules, in fact only one rule is applied. A predefined exclude_filter is used to preempt all other rules when satisfied. Based on the above example set of filtering rules, log 2300 can be interpreted as follows, including the exclude_filter with the following operations.

[0347] 24.2 Example Operations

[0348] Log messages 0 - 2 are marked as excluded in field X (which is a column of log 2300). Field X can naturally appear in the log message in which it was originally generated. In an embodiment, field X is synthetic such that the log parser inserts field X into some or all of the log messages that did not originally contain field X. Log messages can be selected for adornment (i.e., synthetic field insertion) according to various criteria. For example, a keep-alive heartbeat can be logged and then marked as excluded. In the illustrated embodiment, adjacent excluded log messages are merged into the same isolated trace (such as A).

[0349] In an embodiment (not shown), adjacent excluded log messages are isolated (i.e., excluded from adjacent traces), and are merged into the same trace only when their excluded values match. For example, each of log messages 0-2 has a different excluded value, and thus, each of log messages 0-2 will have its own trace (i.e., three separate traces for the three excluded log messages).

[0350] Log message 3 does not match the entity filter or any filter (i.e., filtering rule). In an embodiment, there is a bias for traces having only one message. Since there are no active traces for host 1 (except for the excluded trace A), log message 3 is stored in a temporary trace (not shown) for unmatched messages until a matching log message appears or the time_window_filter is satisfied. In an embodiment, unmatched messages in the temporary trace can eventually be reassigned to a previous or subsequent adjacent trace, which can be configured with rules for unmatched predefined rules, as shown in the above example rules. For example, log message 4 arrives in time to create trace B for message 4. Log message 3 is moved from the temporary trace to trace B.

[0351] In this example, almost any trace can be temporary. For example, due to the entity filter, messages 5-6 are initially placed in a different trace (not shown) from message 4, which has a different pid in field Y. However, message 8 arrives soon after, and its fields Y-Z indicate that pids 10 and 20 belong to the same trace. Thus, messages 5-6 are moved from their temporary trace to trace B. Entity matching can be transitive. For example, the combination of messages 8-9 indicates that pids 10 and 30 match, even though no individual log message suggests so. Thus, messages 3-10 belong to trace B.

[0352] When generating trace C-1 from log messages 11-31, the correlations between pids 10, 20, and 30 are remembered and used. In an embodiment, each correlation of a pair of pids can expire after a duration. For example, a pid can eventually be reused for an unrelated entity. Similarly, the max_gap criterion can prevent some adjacent messages from being merged into the same trace. For example, although there are adjacent messages 23-24, traces G-H are more than one second apart, as indicated by the milliseconds in the timestamp field, and this interval is too long to be merged into the same trace according to max_gap.

[0353] 25.0 Machine learning model

[0354] A machine learning model is trained using a specific machine learning algorithm. Once trained, an input is applied to the machine learning model for prediction, which can also be referred to as a prediction output or output herein.

[0355] A machine learning model includes a model data representation or model artifact. The model artifact includes parameter values, which may be referred to herein as theta values, and which are applied by a machine learning algorithm to an input to generate a predicted output. Training a machine learning model requires determining the theta values of the model artifact. The structure and organization of the theta values depend on the machine learning algorithm.

[0356] In supervised training, training data is used by a supervised training algorithm to train a machine learning model. The training data includes an input and a “known” output. In an embodiment, the supervised training algorithm is an iterative process. In each iteration, the machine learning algorithm applies the model artifact and the input to generate a predicted output. An error or variance between the predicted output and the known output is calculated using an objective function. In effect, the output of the objective function indicates the accuracy of the machine learning model based on a particular state of the model artifact in the iteration. By applying an optimization algorithm based on the objective function, the theta values of the model artifact can be adjusted. An example of the optimization algorithm is gradient descent. The iteration can be repeated until a desired accuracy is achieved or some other criterion is met.

[0357] In a software implementation, when a machine learning model is said to receive an input, perform and / or generate an output or prediction, a computer system executing the machine learning algorithm processes applying the model artifact to the input to generate a predicted output. The computer system processes execute the machine learning algorithm by executing software configured to cause the algorithm to execute.

[0358] Problem categories that machine learning (ML) excels at include clustering, classification, regression, anomaly detection, prediction, and dimensionality reduction (i.e., simplification). Examples of machine learning algorithms include decision trees, support vector machines (SVMs), Bayesian networks, stochastic algorithms such as genetic algorithms (GAs), and connectionist topologies such as artificial neural networks (ANNs). Implementations of machine learning can rely on matrices, symbolic models, and hierarchical and / or associative data structures. Parametric (i.e., configurable) implementations of the best-of-breed machine learning algorithms can be found in open-source libraries such as Google's TensorFlow for Python and C++, or the Georgia Institute of Technology's MLPack for C++. Shogun is an open-source C++ ML library with adapters for several programming languages including C#, Ruby, Lua, Java, Matlab, R, and Python.

[0359] 25.1 Artificial Neural Networks

[0360] An artificial neural network (ANN) is a machine learning model that models a system of neurons interconnected by directed edges at a high level. An overview of the neural network is described in the context of a layered feedforward neural network. Other types of neural networks share the characteristics of the neural network described below.

[0361] In a layered feedforward network such as a multi-layer perceptron (MLP), each layer includes a set of neurons. A layered neural network includes an input layer, an output layer, and one or more intermediate layers called hidden layers.

[0362] The neurons in the input layer and the output layer are called input neurons and output neurons, respectively. The neurons in the hidden layer or the output layer may be referred to as activation neurons in this article. Activation neurons are associated with an activation function. The input layer does not contain any activation neurons.

[0363] From each neuron in the input layer and the hidden layer, there is one or more directed edges pointing to an activation neuron in a subsequent hidden layer or output layer. Each edge is associated with a weight. The edge from a neuron to an activation neuron represents the input from the neuron to the activation neuron, as regulated by the weight.

[0364] For a given input to the neural network, each neuron in the neural network has an activation value. For an input node, the activation value is simply the input value of the input. For an activation neuron, the activation value is the output of the corresponding activation function of the activation neuron.

[0365] Each edge from a particular node to an activation neuron represents that the activation value of the particular neuron is the input to the activation neuron, i.e., the input to the activation function of the activation neuron, as regulated by the weight of the edge. Thus, the activation neurons in the subsequent layer represent that the activation value of the particular neuron is the input to the activation function of the activation neuron, as regulated by the weight of the edge. An activation neuron may have multiple edges pointing to an activation neuron, and each edge represents that the activation value from the source originating neuron (as regulated by the weight of the edge) is the input to the activation function of the activation neuron.

[0366] Each activation neuron is associated with a bias. To generate the activation value of the activation node, the activation function of the neuron is applied to the weighted activation value and the bias.

[0367] 25.2 Illustrative Data Structures for Neural Networks

[0368] The artifacts of a neural network may include matrices of weights and biases. Training a neural network can iteratively adjust the matrices of weights and biases.

[0369] For a layered feedforward network and other types of neural networks, the artifact can include one or more matrices of edges W. The matrix W represents the edges from layer L-1 to layer L. Assuming the number of nodes in layers L-1 and L are N[L-1] and N[L], respectively, then the matrix W has dimensions of N[L-1] columns and N[L-1] rows.

[0370] The bias for a particular layer L can also be stored in a matrix B of one column with N[L] rows.

[0371] The matrices W and B can be stored in the RAM memory as vectors or arrays, or as a comma-separated set of values in the memory. When the artifact is persistently stored in a persistent storage device, the matrices W and B can be stored as comma-separated values in a compressed and / or serialized form or other suitable persistent form.

[0372] A particular input applied to the neural network includes the values of each input node. The particular input can be stored as a vector. The training data includes multiple inputs, each called a sample in the set of samples. Each sample includes the values of each input node. The samples can be stored as vectors of input values, and multiple samples can be stored as a matrix, where each row in the matrix is a sample.

[0373] When an input is applied to the neural network, activation values are generated for the hidden layers and the output layer. For each layer, the activation values can be stored in one column of a matrix A, which has one row for each node in the layer. In the vectorized method for training, the activation values can be stored in a matrix that has one column for each sample in the training data.

[0374] Training the neural network requires storing and processing additional matrices. The optimization algorithm generates a matrix of derivative values for the matrices that regulate the weights W and the bias B. Generating the derivative values can use and requires storing matrices of intermediate values generated when computing the activation values for each layer.

[0375] The number of nodes and / or edges determines the size of the matrices required to implement the neural network. The fewer the number of nodes and edges in the neural network, the smaller the matrices and the memory required to store the matrices. In addition, a smaller number of nodes and edges reduces the amount of computation required to apply or train the neural network. Fewer nodes mean fewer activation values need to be computed during training and / or fewer derivatives need to be computed.

[0376] The characteristics of the matrices used to implement a neural network correspond to neurons and edges. A cell in matrix W represents a specific edge from a node in layer L-1 to a node in layer L. An activated neuron represents the activation function of that layer including the activation function. The activated neurons in layer L correspond to the rows of the weights of the edges between layer L and layer L-1 in matrix W and the columns of the weights of the edges between layer L and layer L+1 in matrix W. During the execution of the neural network, the neurons also correspond to one or more activation values stored in matrix A for that layer and generated by the activation function.

[0377] ANNs are subject to vectorization for data parallelization, which can utilize vector hardware such as single instruction multiple data (SIMD), such as by using a graphics processing unit (GPU). Matrix partitioning can achieve horizontal scaling, for example by using symmetric multiprocessing (SMP), such as by using a multi-core central processing unit (CPU) and / or multiple coprocessors (such as GPUs). Feed-forward computations within an ANN can occur with only one step per neural layer. The activation values in a layer are calculated based on the weighted propagation of the activation values of the previous layer, such that the values are calculated sequentially for each subsequent layer (such as with the corresponding iteration of a for loop). The layering imposes a non-parallelizable computational order. Thus, the network depth (i.e., the number of layers) can cause computational latency. Deep learning requires a multi-layer perceptron (MLP) with multiple layers. Each layer implements data abstraction, and complex (i.e., multi-dimensional with several inputs) abstractions require multiple layers that implement cascaded processing. Implementations of ANNs based on reusable matrices and matrix operations for feed-forward processing are readily available and parallelizable in neural network libraries such as Google's TensorFlow for Python and C++, OpenNN for C++, and the Fast Artificial Neural Network (FANN) from the University of Copenhagen. These libraries also provide model training algorithms such as backpropagation.

[0378] 25.3 Backpropagation

[0379] The output of an ANN can be more or less correct. For example, an ANN that recognizes letters might mistake an I for an L because these letters have similar characteristics. The correct output can have a (one or more) specific value, while the actual output can have a slightly different value. The arithmetic or geometric difference between the correct output and the actual output can be measured as an error according to a loss function, such that zero represents error-free (i.e., completely accurate) behavior. For any edge in any layer, the difference between the correct output and the actual output is an incremental value.

[0380] Backpropagation requires the error to be distributed backward in different amounts to all the connection edges within the ANN through the layers of the ANN. The propagation of the error causes the adjustment of the edge weights, which depends on the gradient of the error on each edge. The gradient of an edge is calculated by multiplying the error increment of the edge by the activation value of the upstream neuron. When the gradient is negative, the greater the magnitude of the error contributed by the edge to the network, the more the weight of the edge should be decreased, which is negative reinforcement. When the gradient is positive, positive reinforcement requires increasing the weight of the edge whose activation would decrease the error. The edge weights are adjusted according to a percentage of the gradient of the edge. The steeper the gradient, the greater the adjustment. Not all edge weights are adjusted by the same amount. As the model training continues with additional input samples, the error of the ANN should decrease. When the error stabilizes (i.e., stops decreasing) or falls below a threshold (i.e., approaches zero), the training can stop. Christopher M. Bishop taught example mathematical formulas and techniques for feedforward multi-layer perceptrons (MLPs) in the related reference "EXACT CALCULATION OF THE HESSIAN MATRIX FOR THE MULTI-LAYER PERCEPTRON", including matrix operations and backpropagation.

[0381] Model training can be supervised or unsupervised. For supervised training, the desired (i.e., correct) output is already known for each example in the training set. The training set is configured by pre-assigning class labels to each example by (e.g., human experts). For example, a training set for optical character recognition can have blurred photos of individual letters, and the experts can pre-label each photo according to which letter is shown. As explained above, error calculation and backpropagation occur.

[0382] More unsupervised model training is involved because the desired output needs to be discovered during training. Unsupervised training can be more easily adopted because it does not require human experts to pre-label the training examples. Thus, unsupervised training saves manpower. The natural way to implement unsupervised training is to utilize an autoencoder, which is a type of ANN. The autoencoder serves as an encoder / decoder (codec) with two sets of layers. The first set of layers encodes the input example into a condensed code that needs to be learned during model training. The second set of layers decodes the condensed code to regenerate the original input example. The two sets of layers are trained together as a combined ANN. The error is defined as the difference between the original input and the input regenerated upon decoding. After sufficient training, the decoder outputs more or less exactly the original input.

[0383] For each input example, the autoencoder relies on the condensed code as an intermediate format. The intermediate condensed code does not initially exist but rather emerges only through model training, which may be counterintuitive. Unsupervised training can achieve the vocabulary of the intermediate encoding based on features and distinctions of unintended correlations. For example, which examples and which labels are used during supervised training can be somewhat unscientific (e.g., anecdotal) or incomplete depending on the human expert's understanding of the problem space. However, unsupervised training discovers the appropriate intermediate vocabulary more or less entirely based on statistical trends that reliably converge to optimality with sufficient training due to internal feedback generated by the re-generated decoding. The implementation and integration techniques of the autoencoder are taught in the related U.S. Patent Application No. 14 / 558,700 entitled "AUTO-ENCODER ENHANCED SELF-DIAGNOSTIC COMPONENTS FOR MODEL MONITORING". This patent application elevates supervised or unsupervised ANN models to a first class of objects that are subject to management techniques such as monitoring and governance during model development (such as during training).

[0384] 25.4 Deep Context Overview

[0385] As described above, an ANN can be stateless such that the timing of activation is more or less irrelevant to the ANN behavior. For example, the recognition of a particular letter can occur in isolation without context. More complex classifications can depend more or less on additional context information. For example, the information content (i.e., complexity) of an instantaneous input can be less than the information content of the surrounding context. Thus, semantics can occur based on context, such as across a time series of inputs or an extended pattern within an input example (e.g., a composite geometry). Various techniques have emerged to make deep learning context-aware. One general strategy is context encoding, which packs the stimulus input and its context (i.e., surrounding / relevant details) into the same (e.g., dense) encoding unit that can be applied to the ANN for analysis. One form of context encoding is graph embedding, which constructs and prunes (e.g., limits its scope) a logical graph of related events or records (e.g., temporally or semantically). Graph embedding can be used as context encoding and as an input stimulus to the ANN.

[0386] The hidden state (i.e., the memory) is a powerful ANN enhancement for (especially temporal) sequence processing. Sorting can facilitate prediction and operational anomaly detection, which can be important techniques. A Recurrent Neural Network (RNN) is a stateful MLP that is arranged in topological steps and can operate more or less as stages of a processing pipeline. In the folded / rolled-up embodiment, all steps have the same connection weights, and a single one-dimensional weight vector can be shared for all steps. In the recursive embodiment, only one step recycles some of its output back to that one step where the recursion implements sorting. In the unfolded / unrolled embodiment, each step can have different connection weights. For example, the weights for each step can appear in the corresponding columns of a two-dimensional weight matrix.

[0387] The input sequence can be applied to the corresponding steps of the RNN either simultaneously or sequentially to effect an analysis of the entire sequence. For each input in the sequence, the RNN predicts the next sequential input based on all the previous inputs in the sequence. The RNN can predict or otherwise output almost all of the input sequence that has been received as well as the next sequential input that has not yet been received. Predicting the next input by itself can be valuable. Comparing the predicted sequence with the sequence that was actually received (and applied) can facilitate anomaly detection. For example, an RNN-based spelling model can predict that U follows Q when reading a word letter by letter. If the letter that actually follows Q is not the expected U, then an anomaly is detected.

[0388] Unlike a neural layer composed of individual neurons, each recurrent step of an RNN can be an MLP composed of cells, each cell containing some specially arranged neurons. RNN cells operate as units of memory. RNN cells can be implemented by Long Short-Term Memory (LSTM) cells. The way LSTM arranges column neurons is different from how transistors are arranged in a flip-flop, but the common goal of LSTM and digital logic is to specially arrange several stateful control gates. For example, a neural memory cell can have an input gate, an output gate, and a forget (i.e., reset) gate. Different from binary circuits, the input gate and the output gate can conduct numerical values (e.g., unit-normalized) retained by the cell, also as numerical values.

[0389] Compared with other MLPs, RNNs have two main internal enhancement functions. The first is the localized memory cell, such as the LSTM, which involves microscopic details. The other is the cross-activation of the recursive step, which is macroscopic (i.e., overall topology). Each step receives two inputs and outputs two outputs. One input is the external activation from the item in the input sequence. The other input is the output of the adjacent previous step that can embed details from some or all of the previous steps, which realizes the sequential history (i.e., temporal context). The other output is the predicted next item in the sequence. Example mathematical formulas and techniques for RNNs and LSTMs are taught in the related U.S. Patent Application No. 15 / 347,501 titled "MEMORY CELL UNIT AND RECURRENT NEURAL NETWORK INCLUDING MULTIPLE MEMORY CELL UNITS".

[0390] Complex analysis can be achieved by a so-called MLP stack. An example stack can sandwich an RNN between an upstream encoder ANN and a downstream decoder ANN, and either or both of the upstream encoder ANN and the downstream decoder ANN can be an autoencoder. The stack can have fan-in and / or fan-out between the MLPs. For example, the RNN can directly activate two downstream ANNs, such as an anomaly detector and an auto-decoder. The auto-decoder may only exist during model training, such as for monitoring the visibility of training or for use in a feedback loop for unsupervised training. RNN model training can use backpropagation through time, which is a technique that can achieve higher accuracy for RNN models compared to ordinary backpropagation. Example mathematical formulas, pseudocode, and techniques for training RNN models using backpropagation through time are taught in the related W.I.P.O. Patent Application No. PCT / US2017 / 033698 titled "MEMORY-EFFICIENT BACKPROGATING THROUGH TIME".

[0391] 26.0 Hardware Overview

[0392] According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing device can be hard-wired to perform the techniques, or can include digital electronic devices such as one or more application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) persistently programmed to perform the techniques, or can include one or more general hardware processors programmed to perform the techniques according to program instructions in firmware, memory, other storage, or a combination. Such special-purpose computing devices can also combine custom hard-wired logic, ASICs, or FPGAs with custom programming to implement the techniques. The special-purpose computing device can be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that combines hard-wired and / or program logic to implement the techniques.

[0393] For example, Figure 24 is a block diagram of a computer system 2400 on which embodiments of the present invention can be implemented. The computer system 2400 includes a bus 2402 or other communication mechanism for conveying information, and a hardware processor 2404 coupled to the bus 2402 for processing information. The hardware processor 2404 can be, for example, a general-purpose microprocessor.

[0394] The computer system 2400 also includes a main memory 2406 coupled to the bus 2402 for storing information and instructions to be executed by the processor 2404, such as random access memory (RAM) or other dynamic storage device. The main memory 2406 can also be used to store temporary variables or other intermediate information during execution of instructions by the processor 2404. When stored in a non-transitory storage medium accessible to the processor 2404, these instructions cause the computer system 2400 to become a special-purpose machine customized to perform the operations specified in the instructions.

[0395] The computer system 2400 also includes a read only memory (ROM) 2408 or other static storage device coupled to the bus 2402 for storing static information and instructions for the processor 2404. A storage device 246, such as a magnetic disk or optical disk, is provided and coupled to the bus 2402 for storing information and instructions.

[0396] The computer system 2400 can be coupled via a bus 2402 to a display 2412 (such as a cathode ray tube (CRT)) for displaying information to a computer user. An input device 2414 including alphanumeric keys and other keys is coupled to the bus 2402 for transmitting information and command selections to the processor 2404. Another type of user input device is a cursor control 2416 (such as a mouse, trackball, or cursor direction keys) for transmitting direction information and command selections to the processor 2404 and for controlling cursor movement on the display 2412. Such input devices typically have two degrees of freedom in two axes (a first axis (e.g., x) and a second axis (e.g., y)), which allows the device to specify a position in a plane.

[0397] The computer system 2400 can implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that in combination with the computer system causes the computer system 2400 to be a special-purpose machine or programs the computer system 2400 to be a special-purpose machine. According to one embodiment, the computer system 2400 performs the techniques herein in response to execution by the processor 2404 of one or more sequences of one or more instructions contained in the main memory 2406. These instructions can be read from another storage medium (such as the storage device 246) into the main memory 2406. Execution of the instruction sequence contained in the main memory 2406 causes the processor 2404 to perform the processing steps described herein. In an alternative embodiment, hardwired circuitry may be used in place of or in combination with software instructions.

[0398] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as the storage device 246. Volatile media includes dynamic memory, such as the main memory 2406. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge.

[0399] Storage media is different from transmission media but can be used in combination with transmission media. Transmission media participates in transferring information between storage media. For example, transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that comprise the bus 2402. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.

[0400] Various forms of media can participate in carrying one or more sequences of one or more instructions to the processor 2404 for execution. For example, the instructions can initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to the computer system 2400 can receive the data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector can receive the data carried in the infrared signal, and appropriate circuitry can place the data on the bus 2402. The bus 2402 carries the data to the main memory 2406, from which the processor 2404 retrieves and executes the instructions. The instructions received by the main memory 2406 can optionally be stored on the storage device 246 before or after being executed by the processor 2404.

[0401] The computer system 2400 also includes a communication interface 2418 coupled to the bus 2402. The communication interface 2418 provides two-way data communication coupled to a network link 2420, where the network link 2420 is connected to a local network 2422. For example, the communication interface 2418 can be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem that provides a data communication connection to a corresponding type of telephone line. As another example, the communication interface 2418 can be a Local Area Network (LAN) card to provide a data communication connection to a compatible LAN. A wireless link can also be implemented. In any such implementation, the communication interface 2418 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.

[0402] The network link 2420 typically provides data communication through one or more networks to other data devices. For example, the network link 2420 can provide a connection through the local network 2422 to a main computer 2424 or to a data device operated by an Internet Service Provider (ISP) 2426. The ISP 2426 in turn provides data communication services through the global packet data communication network (now commonly referred to as the “Internet” 2428). Both the local network 2422 and the Internet 2428 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through the various networks and signals on the network link 2420 and through the communication interface 2418, which carry digital data to and from the computer system 2400, are example forms of transmission media.

[0403] The computer system 2400 can send messages and receive data including program code via one or more networks, network link 2420, and communication interface 2418. In an Internet example, the server 2430 can send the requested code for an application program via the Internet 2428, ISP 2426, local network 2422, and communication interface 2418.

[0404] The received code can be executed by the processor 2404 when received, and / or stored in the storage device 246 or other non-volatile memory for later execution.

[0405] 27.0 Software Overview

[0406] Figure 25 is a block diagram of a basic software system 2500 that can be used to control the operation of the computing system 2400. The software system 2500 and its components (including their connections, relationships, and functions) are merely exemplary and are not meant to limit the implementation of one or more example embodiments. Other software systems suitable for implementing one or more example embodiments may have different components, including components with different connections, relationships, and functions.

[0407] The software system 2500 is provided to direct the operation of the computing system 2400. The software system 2500, which can be stored on the system memory (RAM) 2406 and fixed storage device (e.g., hard disk or flash memory) 246, includes a kernel or operating system (OS) 2510.

[0408] The OS 2510 manages the low-level aspects of computer operation, including managing the execution of processes, memory allocation, file input and output (I / O), and device I / O. One or more applications represented as 2502A, 2502B, 2502C... 2502N can be "loaded" (e.g., transferred from the fixed storage device 246 into the memory 2406) for execution by the system 2500. Applications or other software intended to be used on the computer system 2400 can also be stored as downloadable computer-executable instruction sets, e.g., for downloading and installation from an Internet location (e.g., a web server, app store, or other online service).

[0409] The software system 2500 includes a graphical user interface (GUI) 2515 for receiving user commands and data in a graphical manner (e.g., "click" or "touch gesture"). In turn, these inputs can be operated on by the system 2500 according to instructions from the operating system 2510 and / or one or more applications 2502. The GUI 2515 is also used to display the results of operations from the OS 2510 and one or more applications 2502, from which the user can provide additional input or terminate the session (e.g., log off).

[0410] OS 2510 can be executed directly on the bare hardware 2520 of the computer system 2400 (e.g., the (one or more) processors 2404). Alternatively, a hypervisor or virtual machine monitor (VMM) 2530 can be inserted between the bare hardware 820 and the OS 2510. In this configuration, the VMM 2530 acts as a software “buffer” or virtualization layer between the OS 2510 and the bare hardware 2520 of the computer system 2400.

[0411] The VMM 2530 instantiates and runs one or more virtual machine instances (“guest machines”). Each guest machine includes a “guest” operating system (such as OS 2510), and one or more applications (such as the (one or more) applications 2502) designed to execute on the guest operating system. The VMM 2530 presents a virtual operating platform to the guest operating system and manages the execution of the guest operating system.

[0412] In some instances, the VMM 2530 can allow a guest operating system (OS) to run as if it were running directly on the bare hardware 2520 of the computer system 2400. In these instances, the same version of the guest operating system configured to execute directly on the bare hardware 2520 can also execute on the VMM 2530 without modification or reconfiguration. In other words, the VMM 2530 can provide full hardware and CPU virtualization to the guest operating system in some cases.

[0413] In other instances, the guest operating system can be specifically designed or configured to execute on the VMM 2530 for increased efficiency. In these instances, the guest operating system “is aware” that it is executing on the virtual machine monitor. In other words, the VMM 2530 can provide para-virtualization to the guest operating system in some cases.

[0414] A computer system process includes the allocation of hardware processor time, as well as the allocation of memory (physical and / or virtual), where the allocation of memory is for storing instructions executed by the hardware processor, for storing data generated by the execution of instructions by the hardware processor, and / or for storing the hardware processor state (e.g., the contents of registers) between allocations of hardware processor time when the computer system process is not running. A computer system process runs under the control of an operating system and can also run under the control of other programs executable on the computer system.

[0415] 28.0 Cloud Computing

[0416] This document generally uses the term "cloud computing" to describe a computing model that enables on-demand access to a shared pool of computing resources, such as computer networks, servers, software applications, and services, and allows for the rapid provisioning and release of resources with minimal administrative effort or service provider interaction.

[0417] Cloud computing environments (sometimes referred to as cloud environments or the cloud) can be implemented in a variety of different ways to best suit different requirements. For example, in a public cloud environment, the underlying computing infrastructure is owned by an organization that makes its cloud services available to other organizations or the public. In contrast, a private cloud environment is generally only used by or within a single organization. A community cloud is intended to be shared by several organizations within a community; while a hybrid cloud includes two or more types of clouds (e.g., private, community, or public) bound together through data and application portability.

[0418] Generally speaking, the cloud computing model enables some of those responsibilities that might previously have been provided by an organization's own information technology department to instead be delivered as service layers within the cloud environment for use by consumers (either within or outside the organization depending on the public / private nature of the cloud). Depending on the specific implementation, the precise definition of the components or features provided by or within each cloud service layer may vary, but common examples include: Software as a Service (SaaS), where the consumer uses software applications running on the cloud infrastructure while the SaaS provider manages or controls the underlying cloud infrastructure and applications. Platform as a Service (PaaS), where the consumer can use software programming languages and development tools supported by the PaaS provider to develop, deploy, and otherwise control their own applications while the PaaS provider manages or controls other aspects of the cloud environment (i.e., everything under the runtime execution environment). Infrastructure as a Service (IaaS), where the consumer can deploy and run any software applications, and / or provision processing, storage devices, networks, and other basic computing resources while the IaaS provider manages or controls the underlying physical cloud infrastructure (i.e., everything below the operating system layer). Database as a Service (DBaaS), where the consumer uses a database server or database management system running on the cloud infrastructure while the DbaaS provider manages or controls the underlying cloud infrastructure and applications.

[0419] The foregoing basic computer hardware and software, as well as cloud computing environments, are provided to illustrate the basic underlying computer components that can be used to implement one or more example embodiments. However, one or more example embodiments need not be limited to any particular computing environment or computing device configuration. Instead, one or more example embodiments can be implemented in any type of system architecture or processing environment that those skilled in the art will understand, in view of this disclosure, to be capable of supporting the features and functionality presented herein for one or more example embodiments.

[0420] In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. The sole and exclusive indication of the scope of the invention and what the applicant intends to be the scope of the invention is the literal and equivalent scope of the resulting claims in the specific form of a set of claims that issue from this application, including any subsequent corrections.

Claims

1. A method for sequence anomaly detection, comprising: For each console log message in a sequence of relevant console log messages generated by an application as console output: Parse each of one or more features from the console log message to generate a sparse feature vector representing the one or more features; Based on the sparse feature vector, activate the corresponding step of an encoder recurrent neural network (RNN); and Generate, by the encoder RNN, a corresponding embedded feature vector based on: the one or more features and one or more console log messages that occurred earlier in the sequence of relevant console log messages; Process the one or more embedded feature vectors generated by the encoder RNN to determine a predicted next relevant console log message that may occur in the sequence of relevant console log messages; Calculate an anomaly score based on the predicted next relevant console log message and without decoding the one or more embedded feature vectors; Wherein, the method further comprises: Extract each of the one or more features from a second log message to generate a second sparse feature vector identical to the sparse feature vector; Generate, by the encoder RNN, a second embedded feature vector based on: the one or more features and a second one or more log messages that occurred earlier in the sequence of relevant console log messages; Wherein: The one or more embedded feature vectors include the second embedded feature vector and the corresponding embedded feature vector; The second embedded feature vector is different from the corresponding embedded feature vector; Wherein the method is executed by one or more computers.

2. The method according to claim 1, wherein: The one or more features include at least one categorical feature; The console log message contains a corresponding value for each categorical feature of the at least one categorical feature; Extracting each feature from the console log message to generate a sparse feature vector includes one-hot encoding each corresponding value as a corresponding part of the sparse feature vector.

3. The method according to claim 1, wherein the corresponding step of activating the encoder RNN based on the sparse feature vector comprises: Based on the sparse feature vector, activate a corresponding non-recurrent neural network that activates the corresponding step of the encoder RNN.

4. The method according to claim 3, wherein: The corresponding non-recurrent neural network includes at least a first neural layer and a second neural layer; Each neuron of the first neural layer is connected to each neuron of the second neural layer.

5. The method according to claim 1, wherein the encoder RNN includes long short-term memory (LSTM).

6. The method according to claim 1, wherein determining the predicted next relevant console log message includes predicting an embedded feature vector of a feature representing the predicted next relevant console log message.

7. The method according to claim 1, wherein processing the one or more embedded feature vectors generated by the encoder RNN includes activating a predictor RNN based on the one or more embedded feature vectors to predict the predicted next relevant console log message.

8. The method according to claim 7, wherein the predictor RNN includes an encoder RNN.

9. The method according to claim 1, further comprising comparing the actual relevant log messages with the predicted relevant log messages to calculate a prediction error.

10. The method according to claim 1, wherein calculating the anomaly score includes averaging the prediction errors of each log message in the sequence of relevant console log messages.

11. The method according to claim 10, further comprising indicating that the anomaly score exceeds a threshold.

12. The method according to claim 11, wherein indicating that the anomaly score exceeds a threshold includes indicating an Internet of Things (IoT) failure.

13. The method according to claim 1, wherein said calculating the anomaly score comprises: The mean squared error is used to calculate the anomaly score.

14. The method according to claim 1, wherein processing the one or more embedded feature vectors generated by the encoder RNN to determine the predicted next relevant console log message does not include decoding the one or more embedded feature vectors.

15. The method according to claim 1, wherein generating the respective embedded feature vectors based on the one or more features by the encoder RNN comprises: Based on the corresponding embedded feature vectors, activate a decoder RNN to decode the corresponding embedded feature vectors into reconstructed sparse feature vectors, which approximate the sparse feature vectors representing the one or more features.

16. The method according to claim 15, further comprising comparing the sparse feature vectors with the reconstructed sparse feature vectors to calculate a reconstruction error.

17. The method according to claim 16, further comprising: Summing the prediction error and the reconstruction error to calculate a training error; Backpropagating the training error.

18. The method according to claim 15, wherein the reconstructed sparse feature vectors approximating the sparse feature vectors include the reconstruction loss of each of the one or more features.

19. The method according to claim 1, wherein the encoder RNN has as many recurrent steps as the longest expected sequence of relevant log messages.

20. One or more non-transitory computer-readable media storing one or more sequences of instructions, which when executed by one or more processors, cause the execution of the method according to any one of claims 1-19.

21. A device for sequence anomaly detection, comprising: One or more processors; And A memory coupled to the one or more processors and including instructions stored thereon, which when executed by the one or more processors, cause the execution of the method according to any one of claims 1-19.

22. A computer program product including instructions, which when executed by one or more processors of a computer, cause the computer to execute the method according to any one of claims 1-19.

Citation Information

Patent Citations

  • Memory cell unit and recurrent neural network including multiple memory cell units

    US10032498B2

  • Malicious activity detection by cross-trace analysis and deep learning

    US11451565B2

  • Auto-encoder enhanced self-diagnostic components for model monitoring

    US11836746B2

  • Flyback converter with secondary side regulation

    US20170033698A1

  • Malicious activity detection by cross-trace analysis and deep learning

    US20200076840A1