Machine learning for network traffic analysis, attack detection, and traffic filtering
Patent Information
- Application Number
- PCT/US2026/020278
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2026-01-09
- Filing Date
- 2026-03-23
- Publication Date
- 2026-10-01
Smart Images

Figure US2026020278_01102026_PF_FP_ABST
Abstract
Description
Docket No. AKAM-446-PCTTITLEMACHINE LEARNING FOR NETWORK TRAFFIC ANALYSIS, ATTACK DETECTION, AND TRAFFIC FILTERINGDocket No. AKAM-446-PCTBACKGROUNDTechnical Field
[0001] This application generally relates to the use of machine learning in internet infrastructure and distributed computing systems.Brief Description of the Related Art
[0002] Trained machine learning (ML) models are used to evaluate an input (such as network traffic) for a desired purpose, such as to classify the input. Models are typically trained in an offline, non-production environment, e.g., by curating data for training and feeding that data into a carefully designed ML model in a lab setting. Once trained, models are deployed to live production systems, where the model is used for inferencing, meaning that new, live data is applied to the ML model. The ML model produces an output, such as the classification with a confidence score. The output is commonly referred to as the ML model’s ‘prediction’. Over time, however, ML model’s performance may degrade, which is a problem sometimes referred to as ‘drift’.
[0003] Drift is a problem, for example, in the realm of network security, in which a trained ML model might be used to classify traffic as benign or an attack. Drift can occur for many reasons, but the underlying issue can be that the training data has become stale. As a remedy, the ML model can be retrained on updated or more accurate data, and then the production ML model can be upgraded to the newly trained ML model. This approach however is cumbersome.
[0004] The teachings hereof present new and improved approaches in the field of ML models and the application of such models to analyze network traffic. The teachings hereof also present new and improved approaches for integrating ML models into the operation of network infrastructure, including in particular with security platforms.Docket No. AKAM-446-PCTBRIEF DESCRIPTION OF THE DRAWINGS
[0005] The invention will be more fully understood from the following detailed description taken in conjunction with the accompanying drawings, in which:
[0006] FIG. 1 is a diagram illustrating a ML-based system in accordance with an embodiment of the teachings hereof;
[0007] FIG. 2 A is a diagram illustrating aspects of the MMRC 104 component shown in FIG. 1, in accordance with an embodiment of the teachings hereof;
[0008] FIG. 2B is a diagram illustrating aspects of the encoder module 202 component shown in FIG. 2A, in accordance with an embodiment of the teachings hereof;
[0009] FIG. 2C is a diagram illustrating aspects of the MMRC 104 component shown in FIG. 1, in accordance with an embodiment of the teachings hereof;
[0010] FIG. 2D is a diagram illustrating aspects of the MMRC 104 component shown in FIG. 1, in accordance with an embodiment of the teachings hereof; and,
[0011] FIG. 3 is a flowchart illustrating operation of the ML-based system, in accord with an embodiment of the teachings hereof;
[0012] FIG. 4 is a flowchart illustrating operation of the ML-based system, in accord with an embodiment of the teachings hereof; and,
[0013] FIG. 5 is a diagram illustrating computer hardware that may be used to implement the teachings hereof, in accordance with an embodiment of the teachings hereof.
[0014] Numerical labels are provided in some FIGURES solely to assist in identifying elements being described in the text; no significance should be attributed to the numbering unless explicitly stated otherwise.Docket No. AKAM-446-PCTBRIEF SUMMARY
[0001] This section describes some pertinent aspects of this invention. Those aspects are illustrative, not exhaustive, and they are not a definition of the invention.
[0002] An ML-based traffic analyzer, attack detector, and rule generator is described. An ML model is trained on semi -structured data from network traffic and learns to predict sequences in that data that correspond to any one of several target attributes. The ML model tags the location of such sequences using a parts of speech technique. The target attribute can be an attack of a given type, for example, with the tagged sequences identifying the location of a malicious substring associated with the attack vector. Such information can be used to develop rulesets for use in a network filter to block such traffic. In addition to making predictions about traffic, the system may also automatically determine the need for, and generate, updated rulesets for network filters. Such updated configurations can include new or improved filtering rules for the network filter that operate to find and block predicted attack substrings, as well as automatically optimized rulesets.
[0003] The claims are incorporated by reference into this section, in their entirety.Docket No. AKAM-446-PCTDETAILED DESCRIPTION
[0015] The following description sets forth embodiments of the invention to provide an overall understanding of the principles of the structure, function, manufacture, and use of the methods and apparatus disclosed herein. All of the subject matter described in this application and illustrated in the accompanying drawings are non-limiting examples; the claims alone define the scope of protection that is sought. The features described or illustrated in connection with one exemplary embodiment may be combined with the features of other embodiments. Such modifications and variations are intended to be included within the scope of the present invention. All patents, patent application publications, other publications, and references cited anywhere in this document are expressly incorporated herein by reference in their entirety, and for all purposes. The term “e.g ” used throughout is used as an abbreviation for the non-limiting phrase “for example.”
[0016] The teachings hereof may be realized in a variety of systems, methods, apparatus, and non-transitory computer-readable media. It should also be noted that the allocation of functions to particular machines is not limiting, as the functions recited herein may be combined or split amongst different hosts in a variety of ways.
[0017] Any reference to advantages or benefits refer to potential advantages and benefits that may be obtained through practice of the teachings hereof. It is not necessary to obtain such advantages and benefits in order to practice the teachings hereof.
[0018] All references to HTTP should be interpreted to include an embodiment using encryption (HTTP / S), such as when TLS secured connections are established. While context may indicate the hardware or the software exclusively, should such distinction be appropriate, the teachings hereof can be implemented in any combination of hardware and software. Hardware may be actual or virtualized.
[0019] The acronym ‘ML’ refers to “machine learning”.
[0020] Overview
[0021] An ML-based traffic analyzer, attack detector, and rule generator for a network filter is described. An ML model is trained on semi -structured data from network messages and learns to predict sequences in that data that correspond to any one of severalDocket No. AKAM-446-PCTtarget attributes. The ML model tags the location of such sequences. The target attribute can be an attack of a given type, for example, with the tagged sequences identifying the location of a malicious substring that evidence the attack. Such information can be used to configure the operation of the network filter to block such traffic. Beyond making predictions about traffic, the ML model can also learn to generate actionable configurations for the network infrastructure, such as new or improved filtering rules for the network filter. The teachings hereof can be applied not just to one network filter but also to a platform of many distributed network filters.
[0022] As one example, assume the network filter is an application layer firewall. The target attributes can be attack vectors — such as SQL injection, cross site scripting, and remote code execution — related to attacks that may be directed to a web application, API, or other application layer endpoint. In this case, the sequence predicted by the ML model is a sequence of characters in an HTTP request that represents or is indicative of the target attack. For a given message, multiple contiguous sequences corresponding to multiple attacks are predicted The ML model tags the location of the attacks using an OB IE tagging scheme, where ‘O’ is not part of the attack (“Outside”), ‘B’ is the beginning of the attack sequence, ‘E’ is the end of the attack sequence, and ‘I’ is the interior of the attack sequence. This scheme is referred to in the art as “Part of Speech” tagging.
[0023] Continuing the example, to train the ML model, a set of training data can be derived from the operations of an existing application layer firewall. The existing application layer firewall executes a filtering ruleset on network traffic, looking for attacks. These rulesets are typically expressed in the form of regular expressions, or ‘regexes’ for short. When the firewall detects an attack, the existing application layer firewall captures the input network message containing the detected attack and stores it with a label indicating the attack vector. The training data can be created from these captured messages and these labels. The training data can be processed by an encoder that performs normalization, tokenization, and other encodings, as well as complex embedding tasks. Hidden layers perform deep learning on the labeled data. A cross entropy loss function provides feedback; preferably this function is customized to penalize the model when it misses attacks found by the existing application layer firewall, but not when the model finds attacks that were not detected by the existing application layer firewall.Docket No. AKAM-446-PCT
[0024] Extending the example above, the predicted sequences can be fed into a second stage of the ML model that is trained to generate regular expressions that will detect the tagged attack sequences. The generated regular expressions can be used in a number of ways, e.g., to supplement, improve, replace or be combined with existing regular expressions installed in the existing application layer firewall.
[0025] Multiple ways to implement the second stage are disclosed. One way is to use a neural net based decoder (“NN decoder” for short) and another way is to use a decoder implemented with a clustering approach (“clustering decoder” for short).
[0026] For the NN decoder, the input is a network message and a predicted and tagged sequence of characters in the network message. The sequence is a substring within the larger network message. With HTTP, the substring is a substring within the string of semi-structured data. The NN decoder model is trained to predict a cluster for the attack substring and, for each cluster, one or more regular expressions for finding the constituent substrings.
[0027] To train the NN model, data collected from the existing application layer firewall can be used. As mentioned, the existing application layer firewall can detect an attack based on regular expression being matched. Hence the existing application layer firewall can produce a set of training data including: the input network message, the detected attack vector label, and an indication of the regular expression that detected the attack.
[0028] The clustering decoder operates in a different way. The clustering decoder builds a set of clusters of attack substrings, e.g., a clustering algorithm such as k-means clustering or DBSCAN, for example, on known attack substrings. Hence, each cluster is a group of substrings with similar features. Each cluster is also associated with one more regular expressions that were developed using regex generation libraries or otherwise, which will be described in more detail later.
[0029] In operation, the cluster decoder associates the newly detected attack substring to one of the clusters based on similarity scoring. A substring falling within a given cluster means that the regular expression(s) for that cluster are likely to find it. The level of coverage can be assessed over time and new regular expressions generated asDocket No. AKAM-446-PCTneeded to cover the clusters. Newly formed clusters will also result in new regular expressions associated therewith.
[0030] With additional processing, and from time to time, the system can analyze the set of predicted regular expressions from either of the NN approach and / or the clustering approach, and attempt to optimize them, e.g., by merging them or simplifying them.Similarly, the set of predicted regular expressions can be analyzed in view of the existing regular expressions in the existing application layer firewall, attempting to merge, supplement, or otherwise optimize them. In some cases, it may be necessary to generate a new cluster and new regular expression, if the attack substring is sufficiently different that it cannot be classified into one of the existing clusters.
[0031] Thus, in one embodiment, the ML model has a first stage (encoder) and a second stage (NN based decoder or clustering decoder). The encoder learns Part of Speech tagging on the semi structured messages, and the decoder produces regular expressions to install in a network filter to find them.
[0032] Figure 1 & System Diagram
[0033] FIG. 1 illustrates an embodiment of the overall system at a high level.
[0034] Network filter 100 is deployed on a network and it filters traffic according to a configuration of regular expressions (the filtering rules). In this example, the network filter functions as an application layer firewall that receives HTTP messages (typically HTTP requests, but also responses) and performs deep packet inspection. The network filter’s purpose is to detect attacks on an upstream server that the firewall is protecting (not shown). The teachings hereof can be extended to other use cases, including filtering for other purposes and different types of traffic. HTTP messages are semi -structured data and the teachings hereof apply without limitation to such semi -structured data.
[0035] The deep packet inspection performed by the network filter 100 involves applying the regexes to the HTTP messages, determining whether a regex is triggered indicating an attack (sometimes referred to as ‘matching’ on the message). Each regex looks for a particular type of attack. If the regex is triggered, the network filter 100 identifies the attack vector and performs an action (e.g., block the traffic, allow the traffic but generate anDocket No. AKAM-446-PCTalert) per an internal configuration stored at the network filter 100. More information about basic operation of a network filter in the role of an application layer firewall can be found in US Patent No. 8,458,769, titled “Cloud based firewall system and service”, issued June 4 2013, and in US Patent No. 11,012,416, titled “Symbolic Execution For Web Application Firewall Performance” issued May 182021, the teachings of which are hereby incorporated by reference in their entirety.
[0036] Messages that do not trigger any rules are passed to the destination.Messages that trigger one or more rules and hence result in an action cause the network filter 100 to generate a receipt 101. The receipt 101 contains the message or at least a portion of the message that triggered the regex. The receipt 101 also contains the regex that was triggered, and the attack vector. These receipts may be aggregated or compressed in conventional ways, and they are stored in data repository 102.
[0037] Though not shown explicitly in FIG. 1, the threat research 103 team may add known bad traffic or good traffic, or otherwise to data repository 102. For example, OpenAppSec provides examples of traffic with known good or bad characteristics.
[0038] The module labeled ML Model & Related Components 104 (MMRC 104) represents the core of the system and will be described in more detail below. In FIG. 1, the MMRC 104 is shown as having two stages. The first stage is sometimes referred to as the encoder stage. In this implementation, the first stage is a neural network based ML architecture and more specifically a transformer architecture. The neural network may use recurrence and some may consider it a recurrent neural network for that reason. The input to the first stage is the aforementioned receipt data 101 and other data stored in the data repository 102. The goal of the first stage is to take in a network message, or metadata about that message, and to generate a prediction about whether the message has one or more attributes. The attributes it is looking for are the attack vectors. Put another way, an attack vector is an attribute.
[0039] The receipts 101 and / or other data are labeled training data (label being the attribute of an attack vector). Once trained, the output of the first stage is a prediction of an attack vector in an input network message. The first stage uses a sequence tagging approach to express its prediction. Characters in the message are tagged, or labeled, according to an OB IE system. In the OB IE system, O stands for outside of an attack, B stands for beginningDocket No. AKAM-446-PCTof an attack sequence,! stands for the interior of the attack sequence, and E stands for the end or final character of the attack. If an attack is found, the substring indicating the attack is labeled BIE, where the number of I labels varies with the length of the substring. If no attack is found, then every character in the message would be labeled as ‘O’.
[0040] When the MMRC 104 identifies an attack that the network filter 100 missed, that prediction can be fed back to the threat research 103 function, as shown in FIG. 1. Once threat research 103 validates the finding (e.g., that there was an attack and hence the network filter had a false negative or FN), that network message can be labeled with the attack vector and stored into the data repository 102. That represents a reinforcement learning loop for the MMRC 104. Similarly, it is possible that the network filter 100 finds one attack in a network message, but misses others also present therein. This is because network filter 100 operation is expensive, so generally a network filter 100 will stop processing once a first attack is found and take the configured action to block or otherwise mitigate. If the MMRC 104 finds additional attacks in a network message, this signal is fed back to the threat research 103 function for reinforcement learning, too.
[0041] The foregoing is a generalized summary. MMRC 104 stage 1 is shown and described in more detail in FIG. 2 A and 2B.
[0042] As mentioned earlier, the second stage of the MMRC 104 is a decoder that takes the predicted attack substrings as input and generates one or more regexes for use by the network filter to find the attack sequence in network messages. There are two versions of the decoder are presented herein. The NN decoder uses a neural network trained on the receipt 101 data plus the tagged sequences, which collectively contain the network message, the substring constituting the attack, and the regex in the network filter 100 that found the attack. The NN decoder is presented and described in more detail with respect to FIG. 2C. The clustering decoder clusters predicted attack sequences using a clustering algorithm such as DBSCAN or k-means and for each cluster, one or more regexes are developed. The clustering decoder is presented and described in more detail with respect to FIG. 2D.
[0043] Generated regexes are sent from the MMRC 104 to the regex optimization component 105. The function of the regex optimization 105 is to take a set of two or more regexes and optimize them. This function may include combining, merging, simplifying and consolidating the set of regexes to provide a regex list that is easier for the network filter 100Docket No. AKAM-446-PCTto process. The regex optimization 105 may operate in at least two areas: optimizing predicted regexes from stage 2 of the MMRC 104 from across clusters, and optimizing the predicted regexes with the regexes that are already installed in the network filter 100. The result is an updated regex list that is pushed to the network filter 100 as an update.
[0044] Figure 2A & Detailed Description On MMRC 104
[0045] FIG. 2 A illustrates an embodiment of the MMRC 104 in more detail.
[0046] At block 200 the MMRC 104 pulls and pre-processes the data from the data repository 102. Pre-processing the data means that the network messages, the attack vectors, and other data must be generalized. Literals and variables may be replaced with constant filler value. Certain things such as defined elements (e.g., in HTTP elements, SQL expressions) are identified and tokenized so they remain intact.
[0047] Many conventional parsing and normalization tools are available that can perform the necessary functions. For example, Antlr is a community supported tool with available parser languages. There is support for various dialects of SQL, Javascript, User-Agent (separately) and other programming languages. Additionally, it allows for easier extension of parsing expressions and literals.
[0048] Tokenization 201 can be performed in a variety of ways. Tokenization 201 maps normalized expressions to tokens in a system dictionary. Hence certain things like HTTP header fields, method types, or the like are mapped to defined token values. To expand the vocabulary with new expressions, the system can select the top recurring expressions and add these to the custom encoder. Then it can re-tokenize the receipts 100 and other data from the data repository 102 with the new expressions. This information is used throughout the MMRC 104.
[0049] The tokenization process 201 can include intent tokenization 205. With intent tokenization, an ML model learns to interpret a corpus of data such as a programming language or text, and to produce a prediction of what it is trying to do — the “intent”. In the MMRC 104, an intent model is utilized to analyze data in the receipts 101 — particularly the network messages. So for example the intent model produces an encoded prediction of intent of what the HTTP request is doing, and / or what the payload such as a SQL statementDocket No. AKAM-446-PCTis doing. A suitable intent model is the CodeT5 pretrained model, which was developed for programming languages but which can be repurposed for the teachings hereof. The predicted intent is tokenized and attached to the receipts 101 data. More specifically, the steps are:First decode substrings using tokenizers.Encoding substrings using CodeT5 tokenizer.Using CodeT5 to get a summarization of substrings.Use encoded, summarization text for hidden layer generation.
[0050] The encoder module 202 in FIG. 2A contains many subcomponents to support a transformer based ML model. It is shown in more detail in FIG. 2B.
[0051] Figure 2B & Detailed Description of Encoder Module 202
[0052] The encoder module 202 performs several embeddings as shown in FIG. 2B.
[0053] Word Embedding 205 produces a representation of input tokens in higher dimensional space.
[0054] Position Embedding 206 produces positional information of each token in higher dimensional space.
[0055] Segment embedding 207 produces segment information (e.g., identifying parts of HTTP request) represented in higher dimensional space. More specifically, segment embedding 207 involves identifying locations in the network message that may be of significance in detecting attacks. Segment embedding 207 expresses locations to the model so that the model can learn where attacks tend to occur in the context of the network message’s structure. For example, attacks of certain types often are found (and the network filter 100 may have found an attack) in a particular message header field, or a message body. The information is added to the segment embedding layer appended to the word and positional embedding layer. For example, assuming the network message is an HTTP request, a simplified example with three user agent words and four cookie words, this mayDocket No. AKAM-446-PCTlook like an array = [<user-agent-id>,<user-agent-id><user-agent-id>,<cookie-id>,<cookie-id>,<cookie-id>,<cookie-id>]
[0056] The above example means that the first three words of the HTTP message string are the user agent header field, and the second three words are the cookie header field.
[0057] The segment embedding information can be appended to the word and positional embedding layer.
[0058] Attribute Embedding 208 produces attack vector and group information embedded, as a question, represented in higher dimensional space. More specifically, attribute embedding is used as a question to the language model. For example, adding attribute X to the embedding layer means the model is expected to produce results corresponding to X. All the attributes that the model will be trained on can be defined in this embedding. These are attack vectors, sometimes referred to as attack groups. The training dataset will repeat the same receipt 101 for each attack vector, but with one attack vector having non-null values, and the rest being null values. On testing, each request to the model will query for each attribute (each attack vector).
[0059] Input embedding 209 takes the preceding embeddings and produces a merge for input to the encoder layers 210. A variety of merge approaches can be used, including sum, concatenation or hybrid merge of word, positional, segment and attribute embedding.
[0060] The encoder layers 210 are neural network (NN) layers that employ a selfattention mechanism to produce an output using a part of speech tagging technique. Part of speech tagging involves tagging portions of an input sequence with tags from a limited tag set. In this embodiment, the input network message (e.g., the HTTP request) is converted into a sequence of tags from the set {O,B,I,E}. As mentioned earlier, ‘O’ is not part of the attack (“Outside”), ‘B’ is the beginning of the attack sequence, ‘E’ is the end of the attack sequence, and ‘I’ is the interior of the attack sequence. The input sequence may be tagged on a character by character basis.
[0061] The encoder layers 210 produce a sequence of tags, a collection of characters ({B, 1,0, E}) of the same length as the input sequence. It represents one or more subsequences of beginning and end indexes within the input token.Docket No. AKAM-446-PCT
[0062] The contextualized embedding 211 will merge the network message in the receipt 101 which is encoded, with the pre-processing mentioned before: normalized substring, identified expressions, segment embedding, intent embedding from CodeT5 model, and so on. The recommended merge approach is concatenation, so as to preserve the data.
[0063] The contextualized embedding 211 produced by the encoder layer 210 are processed by a linear condition random fields (CRF) class. CRF is a probabilistic function that is applied on the output logits of an encoder / decoder. This function adds a layer of probability on the token being predicted, the probability is dependent on the input and on the previous tokens also predicted. Essentially, this function adds the inherent dependencies between output tokens.
[0064] CRF Tensor Conversion. To convert the CRF output in 212 to classification probabilities, the approach can include: get probabilities by softmax / or argmax, convert tokens to string.
[0065] The CRF output 212 produced are linearly transformed to a string containing four classes of sequence tagging as required for this application, namely the OBIE tagging. The CRF tensor conversion can be implemented using known programming libraries such as PyTorch. Generally, the process is:Import the module torchcrfAdd the layer to the encoder’s self - torchcrf. CRFLayer(num labels)In forward, the output of the transformer will be logits.The logits will then be passed to the linear fc layer for classification.The fc layer output will be passed to CRF.loss = -self.crf(logits, tags, mask) (here tags is actual sequence tagging from the training dataset).best_tags = self.crf. decode(logits, mask)Docket No. AKAM-446-PCT
[0066] The CRF output from the encoder module 202 are tagged sequences. They can be used to generate masking information. The masking information is a mask that leaves only the substring identified by the encoder as an attack. For each of the substring identified in the network message, the NN decoder will be provided with masking information. This will be described in more detail in FIG. 2C.
[0067] A custom cross entropy loss function 213 is applied to the output of the encoder 202 to assess loss during training. As known in the art, the purpose of the loss function is to quantify the difference between the label and the prediction of the model. Loss functions are known in the art. In this implementation, however, a custom function is used because the ML model might detect attacks that the network filter 100 missed (false negative). Detections of these false negatives are differences but the model should not be penalized for them but rather incentivized to find them. It also should be incentivized to find the attacks that the network filter did (true positives), while not flagging the benign traffic (true negatives). Put another way, the ML model should be at least as good as the network filter 100 in finding attacks, and hopefully better.
[0068] Figure 2C & Detailed Description of NN Type Decoder
[0069] FIG. 2C is a continuation of FIG. 2 A and it illustrates details of stage 2 implemented using the NN decoder.
[0070] The custom masking 204 component is the interface between the encoder (stage 1) and the NN decoder (stage 2) in the overall MMRC 104 architecture. At box 214, custom masking 204 involves producing an encoded version of the input data from the receipt 101 (the HTTP request or other network message, the regex, etc.) using the same input embedding scheme described earlier in connection with FIG. 2B, 205 through 209. Custom masking 204 then masks these inputs using the information from the CRF encoder tokens 212 from the encoder (as shown and mentioned above for FIG. 2B).
[0071] The masking hides aspects of the network messages other than the sequences (the substrings) in the network messages that the MMRC 104 predicted to be attacks.
[0072] For training purposes, the decoder layers 214 ingest this masked input, along with the regex from the receipt 101 that detected the attack, and the cluster identification.Docket No. AKAM-446-PCTThe decoder layers 214 are neural network (NN) layers. The decoder layers 214 learn to predict a regex (or multiple regexes) that will find the unmasked data in the network messages, and also predicts a cluster identification. A confidence score may be included.
[0073] One or more regexes is generated for each of the clusters.
[0074] The purpose of producing a cluster assignment for a given input attack substring is to provide interoperability with the clustering decoder in FIG. 2D, and to help merge groups of predicted regexes (e.g., FIG. 1, box 105).
[0075] Two loss functions are used. Simple Loss - This uses regex library output as the prediction output compared with expected regex using a cross-entropy loss function, which is a loss function that is common and known in the art. The idea is to let the model in the right path. Optimization Loss - uses a genetic algorithm to optimize the regex and then calculate the fitness of regex to optimize it.
[0076] Figure 2D & Detailed Description of Clustering Decoder
[0077] FIG. 2D is a continuation of FIG. 2A and it illustrates details of stage 2 implemented using the alternative of the clustering decoder.
[0078] This embodiment involves clustering similar attack substrings, with each cluster being associated with one or more regular expressions that detect the cluster’s attack substrings in network traffic.
[0079] Initially, a corpus of known attack substrings (e.g., through operation of the network filter 100 and / or threat research 103) can be clustered using an algorithm such as DBSCAN or k-means. This is shown in FIG. 2D as the pre-trained clustering operation 215. Put another way, each cluster is a set of attack substrings or tagged attack sequences, that have similar attributes.
[0080] When new attack substrings are identified by the system (tagged attack substrings #1, #2, etc. in the diagram), they are evaluated for similarity to an existing cluster. Similarity scoring and cluster assignment can be accomplished using a variety of conventional techniques. There are three possibilities as a result of this evaluation:Docket No. AKAM-446-PCT1. The new attack substring is sufficiently similar to an existing cluster. (E.g., similarity score above a configured threshold.) The attack substring is placed into that cluster and the existing regular expressions are tested to see if they will find the new attack substring. To conduct this test, a known tool such as PCRE (Perl Compatible Regular Expressions library) can be used. If the existing regular expressions are sufficient, then the process is complete.2. The new attack substring is sufficiently similar to an existing cluster, however the existing regular expressions for that cluster do not cover the new attack substring. In this case, a new regular expression is generated and added to the regular expression set for the cluster. Then, an optimization procedure is executed to optimize the set of regular expressions. Optimization can be performed at regular intervals, also.3. The new attack substring does not fit into any existing cluster. A new cluster can be created for this attack substring, with a newly generated regular expression.Preferably, before creating a new cluster the system monitors network traffic to see if this new attack substring occurs with sufficient frequency and across a sufficient breadth of traffic to warrant the creation of a new cluster.
[0081] Regular Expression Generation
[0082] To generate a new regular expression for an attack substring, known tools for regex generation such as ‘DEAP’ (Distributed Evolutionary Algorithms in Python), ‘rstr’ (in Python) or ‘regexgen’ (in Python) can be used. An iterative process can be used in which, first, anchor points in a given attack substring (e.g., keywords) are identified and reflected as root nodes in a tree structure. Then the algorithm walks both forward and backwards from the anchor points, adding regular expression logic to find the necessary sequence of characters to reach the anchor point (root node). Walking stops when the algorithm reaches an adjacent anchor point, or a boundary between anchor points, or an end of the attack string.
[0083] Regular Expression Optimization
[0084] The goal of generating a set of regular expressions is to install them into the network filter and use them to detect incoming attacks. Executing regular expressions against network traffic comes with a performance impact and cost. Therefore it is desirableDocket No. AKAM-446-PCTto minimize the number of regular expressions per cluster, as well as the complexity (which is often indicated by the length) of each regular expression.
[0085] According to this disclosure, a genetic algorithm can be used to find an optimized set of regular expressions that meet the above needs. As known in the art, a genetic algorithm is an algorithm that starts with an initial population of individuals, where the individuals represent solutions to a problem. The algorithm iterates across N generations to find optimized individuals and / or a population, given a fitness function that describes desirable and perhaps undesirable properties. With each generation, operations such as mutation or crossover are performed on the individuals. Features of the individuals (such as elements of a regular expression, as relevant here) are thus changed over generations to arrive at an “evolved” solution.
[0086] In one embodiment, a genetic algorithm is used to find a set of regular expressions that covers the attack substrings in the cluster while excluding substrings outside of the cluster. The genetic algorithm is also incentivized to find shorter regular expressions to accomplish these goals and can be configured with a length limit.
[0087] The initial population can be set as the currently installed regular expressions for the attack, plus a newly generated but naive regular expression for the new attack substring. A naive regular expression can be produced using anchor point and tree structure approach, as described in the prior section. A naive regular expression might be produced, for example, as an ordered search for the language-specific characters or elements in a substring (e.g., the SQL characters / elements) with wildcards for other positions.
[0088] The fitness function is preferably a non-dominant algorithm tuned for precision, such that coverage of the attack substrings in the cluster are weighted positively, while coverage of substrings outside the cluster (e.g., which would cause false positives) are weighted negatively. The overall fitness is thus dependent both on maximizing coverage of true positives and minimizing false positives. Furthermore, a constraint on the length of a regular expression is applied, and a negative weight on length (if less than the length constraint) can be used.
[0089] The nature of genetic algorithmic solutions are to find a best fit individual(s), which put another way is a solution that performs best according to the fitness function.Docket No. AKAM-446-PCTHowever, in accord with the teachings hereof, preferably the fitness function also includes other guardrails that prohibit certain solutions, which are sometimes referred to as negative examples. For example, a regular expression that matches on all substrings is prohibited, as is one that matches on all punctuation, or common database operators or data formats, or one that matches on impossible or overbroad application layer messages, and so on. The negative examples can be developed from the following principles:(i) A regex should not find a substring that violates constraints imposed by the targeted application. Put another way, the genetic algorithm should not waste processing to find (and does not need to find) a regex that covers an input that the application cannot interpret, i.e., will produce an error. For example, if the substrings are from a SQL based attack, then a negative example is a substring that contains no SQL keywords, or that contains no spaces, or that places a FROM before SELECT rather than vice versa.Similar negative examples can be developed for an attack aimed at exploiting applications, including HTTP itself.(ii) A regex cannot be a solution if it does not contain a keyword or order of keywords necessary for an application targeted by the attack to function. Hence, a negative example would be a substring that reverses the order of SQL keywords (reverse nested SELECT statements or other keywords).
[0090] To expedite and improve the genetic algorithm’s execution, nearby clusters are also identified, and the attack substrings in those clusters are applied as constraints, meaning that the solution should not cover those attack substrings. This approach helps to bound the solution with the nearest neighbor clusters, and force the algorithm to more quickly explore the appropriate solution space. It also tends to limit the length of the regex solution.
[0091] After running through a configured number of generations, the genetic algorithm produces its solutions, which are one or more regular expressions that cover the attack substrings in the cluster. This can be tested and then used as an updated regular expression set, representing a subset of the overall set of regular expressions installed in the network filter.Docket No. AKAM-446-PCT
[0092] ‘Early stop’ triggers can be implemented in the genetic algorithm such that the algorithm terminates searching upon reaching a minimum number N of generations and the current population has a solution that meets certain ‘good enough’ criteria (e.g., detect at least X percent of substrings in the cluster and detect at most Y percent of substrings out of the cluster). This approach can be used to reduce processing time and is a variant of a known approach referred to as ‘hall of fame’ in which the top N performing solutions are saved as the algorithm continues to run. The use of these kinds of improvements is often necessary or desirable because given its nature, a genetic algorithm may not stay in an optimal solution space but rather over time move away from the optimal solution in its search.
[0093] The approach described above has focused on the generation of an optimized set of regular expressions for a cluster, but the approach can be used to generate a new regular expression for the newly identified attack substring, and then this regular expression can be added to the cluster’s regular expressions. Also, the approach outlined above can be adapted to other types of network filter rules that may not be considered classic regular expressions.
[0094] FIG. 3 & System Flowchart
[0095] FIG. 3 illustrates an example of system operation in a flowchart form, including the stages described earlier. This example uses the clustering approach for the decoder.
[0096] First, unfiltered data (“UF data”) is processed by the network filter using existing regular expressions. Unfiltered data here includes network traffic that contains both benign and attack traffic. This part of the diagram thus refers to deployed network components, e.g., network firewalls.
[0097] If the network filter detects an attack, then the flow in FIG. 3 ends. This means that the network filter simply operates as currently configured, e.g., blocking or alerting on the attack in accord with the current configuration. If the network filter does not detect an attack (NO), then it passes the traffic as filtered UF data to the system for analysis.Docket No. AKAM-446-PCT
[0098] The remainder of this workflow is dedicated to finding attacks in the traffic that the network filter thought was benign (false negatives) and then to developing a suitable regular expression for updating the network filter. It should be understood that in other embodiments, all traffic can be sent to the inference stage, because even if the network filter finds one attack, it may have missed others— this was mentioned earlier in the document. Also, in other embodiments, the system could operate on unfiltered traffic that has not passed through a network filter, as it could be used as a replacement for the network filter after tuning.
[0099] In the Inference stage, the input traffic is applied to the trained ML model that was described earlier in connection with FIGS. 1, 2A and 2B, and MMRC 104. Attack detection and sequence tagging is performed. If no attacks are found, the flow ends. If yes, then the tagged attack substring undergoes similarity analysis to determine whether it belongs to an existing cluster. (“Is Attack Similar (cluster colocation)?” step)
[0100] If the attack is similar to an existing cluster, then the regex coverage of the cluster is tested to see if it still covers the newly detected attack substring (“Existing Regex Coverage?”). If not, then the regex optimization routine is executed to adjust and improve the regex set. This updated set is sent for in-system testing (“Benchmark Efficacy Evaluations”), keeping in mind that this process is entirely automated.
[0101] The efficacy benchmarking process involves applying the new regex rules to a corpus of unfiltered traffic with known attacks embedded therein. A key performance indicator (‘KPI’) is established for true positive and false positive rates (e.g., the ruleset must find at least X percent of all attacks with less than Y percent false positives). If the efficacy standards are not met, then the ruleset must be improved and retested. Once the ruleset meets efficacy standards, it can be exported to the network-deployed network filters and installed to use against live traffic. (“Regex update to network”)
[0102] Returning to the “Is Attack Similar (cluster colocation)?” step, if the newly detected attack substring does not fit into any existing cluster, then the flow moves to the “Frequency Counting of attack” step. This means that the system monitors to see if this attack is detected with sufficient frequency over time. There could also be a requirement that the attack show up across a breadth of traffic (e.g., in many geographies, across a certain number of tenant / customer domains, and so on). This gating function helps ensure thatDocket No. AKAM-446-PCTspurious or immaterial attacks do not cause the generation of new clusters and new regex sets for the network filter to execute.
[0103] FIG. 4 and System Inference Stages
[0104] FIG. 4 provides another view of an example system, a view that focuses on the inference stages in the pipeline.
[0105] The first stage is ‘Attack Detection and Tagging’. Processing begins with the transformation and standardization being applied to the filtered UF data. This stage filters network traffic for configuration-enabled customers of the system. Preferably this stage executes periodically, e.g., every 15 minutes. It filters requests for attack groups enabled for detection by that customer, detecting attack substrings found in traffic the network filter considered benign, and performs attack substring tagging. The tagged results are post processed. If an attack was identified then the results are passed to the next steps. If no attack was found then the evaluation process ends. The ‘attack tagging tokenizer artifact’ and the ‘attack tagging model artifact’ refer to the trained ML models described earlier (with respect to the MMRC 104). The data output for this stage includes, for a detected attack, a vectorized and tagged piece of network traffic (the tagged portion of an HTTP request).
[0106] The key performance indicators output for this stage include the execution volumes and performance (e.g., of inputs and outputs for each execution cycle), as well as efficacy metrics such as total positives and negatives, and false positives and negatives.
[0107] The next stage is ‘Attack Similarity Analysis’. This stage processes only the attack substring(s) identified in the prior stage. It clusters similar attacks together, evaluating similarity with pre-defined clusters. This stage can be executed periodically, e.g., every 15 minutes after the first stage. The Cluster Model Artifact refers to the trained ML clustering model that was mentioned earlier. Key performance indicators for this stage can include execution volumes and performance, as well as metrics such as Attacks Covered / day; Attacks Requiring regex Update / day; Cluster distance average / day; Number of system-vulnerable attacks Identified (to date); Number of system-vulnerable attack similarity counts (to date).Docket No. AKAM-446-PCT
[0108] The next stage is ‘Protection Optimization’. This stage processes the similarity scores for the attack substring(s) that were produced in the prior stage. It identifies whether existing regex suffice to protect against new attacks, and it optimizes regexes to protect against recurrent new attacks. This stage can execute periodically, e.g., once a day. The Cluster Sampled Dataset refers to the clusters with constituent attack substrings and the regexes for each cluster. Generally, this stage determines one of the following:a. Attack substrings in a cluster have a high similarity score, then regex coverage validation is not required.b. Attack substrings in a cluster have medium or lower similarity score, then regex coverage validation is requiredRegex coverage validation checks if the regex associated to the closest cluster for a given attack substring, is able to 'match' (or block) the request. If coverage validation fails, then optimization is required. If a new attack substring is identified multiple times, and is not covered, then the regex needs to be updated. This is done by capturing system-vulnerable attacks in operational storage, and calculating similarity of new attacks with them. Key performance indicators for this stage can include execution volumes and performance, as well as Attacks Covered / day; Attacks Requiring regex Update / day; Cluster distance average / day; Number of attacks Identified (to date); Number of attack similarity counts (to date).
[0109] Computer Based Implementation
[0110] The teachings hereof may be implemented using conventional computer systems, but modified by the teachings hereof, with the components and / or functional characteristics described above realized in special-purpose hardware, general-purpose hardware configured by software stored therein for special purposes, or a combination thereof, as modified by the teachings hereof.
[0111] Software may include one or several discrete programs. Any given function may comprise part of any given module, process, execution thread, or other such programming construct. Generalizing, each function described above may be implementedDocket No. AKAM-446-PCTas computer code, namely, as a set of computer instructions, executable in one or more microprocessors to provide a special purpose machine. The code may be executed using an apparatus - such as a microprocessor in a computer, digital data processing device, or other computing apparatus - as modified by the teachings hereof. In one embodiment, such software may be implemented in a programming language that runs in conjunction with a proxy on a standard Intel hardware platform running an operating system such as Linux. The functionality may be built into the proxy code, or it may be executed as an adjunct to that code.
[0112] While in some cases above a particular order of operations performed by certain embodiments is set forth, it should be understood that such order is exemplary and that they may be performed in a different order, combined, or the like. Moreover, some of the functions may be combined or shared in given instructions, program sequences, code portions, and the like. References in the specification to a given embodiment indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic.
[0113] FIG. 5 is a block diagram that illustrates hardware in a computer system 500 upon which such software may run in order to implement embodiments of the invention. The computer system 500 may be embodied in a client device, server, personal computer, workstation, tablet computer, mobile or wireless device such as a smartphone, network device, router, hub, gateway, or other device. Representative machines on which the subject matter herein is provided may be a computer running a Linux or Linux-variant operating system and one or more applications to carry out the described functionality.
[0114] Computer system 500 includes a microprocessor 504 coupled to bus 501. In some systems, multiple processor and / or processor cores may be employed. Computer system 500 further includes a main memory 510, such as a random access memory (RAM) or other storage device, coupled to the bus 501 for storing information and instructions to be executed by processor 504. A read only memory (ROM) 508 is coupled to the bus 501 for storing information and instructions for processor 504. A non-volatile storage device 506, such as a magnetic disk, solid state memory (e.g., flash memory), or optical disk, is provided and coupled to bus 501 for storing information and instructions. Other application-specificDocket No. AKAM-446-PCTintegrated circuits (ASICs), field programmable gate arrays (FPGAs) or circuitry may be included in the computer system 500 to perform functions described herein.
[0115] A peripheral interface 512 may be provided to communicatively couple computer system 500 to a user display 514 that displays the output of software executing on the computer system, and an input device 515 (e.g., a keyboard, mouse, trackpad, touchscreen) that communicates user input and instructions to the computer system 500. However, in many embodiments, a computer system 500 may not have a user interface beyond a network port, e.g., in the case of a server in a rack. The peripheral interface 512 may include interface circuitry, control and / or level-shifting logic for local buses such as RS-485, Universal Serial Bus (USB), IEEE 1394, or other communication links.
[0116] Computer system 500 is coupled to a communication interface 516 that provides a link (e.g., at a physical layer, data link layer,) between the system bus 501 and an external communication link. The communication interface 516 provides a network link 518. The communication interface 516 may represent an Ethernet or other network interface card (NIC), a wireless interface, modem, an optical interface, or other kind of input / output interface.
[0117] Network link 518 provides data communication through one or more networks to other devices. Such devices include other computer systems that are part of a local area network (LAN) 526. Furthermore, the network link 518 provides a link, via an internet service provider (ISP) 520, to the Internet 522. In turn, the Internet 522 may provide a link to other computing systems such as a remote server 530 and / or a remote client 531. Network link 518 and such networks may transmit data using packet-switched, circuit-switched, or other data-transmission approaches.
[0118] In operation, the computer system 500 may implement the functionality described herein as a result of the processor executing code. Such code may be read from or stored on a non-transitory computer-readable medium, such as memory 510, ROM 508, or storage device 506. Other forms of non-transitory computer-readable media include disks, tapes, magnetic media, SSD, CD-ROMs, optical media, RAM, PROM, EPROM, and EEPROM, flash memory. Any other non-transitory computer-readable medium may be employed. Executing code may also be read from network link 518 (e.g., following storage in an interface buffer, local memory, or other circuitry).Docket No. AKAM-446-PCT
[0119] It should be understood that the foregoing has presented certain embodiments of the invention but they should not be construed as limiting. For example, certain language, syntax, and instructions have been presented above for illustrative purposes, and they should not be construed as limiting. It is contemplated that those skilled in the art will recognize other possible implementations in view of this disclosure and in accordance with its scope and spirit. The appended claims define the subject matter for which protection is sought.
[0120] It is noted that any trademarks appearing herein are the property of their respective owners and used for identification and descriptive purposes only, and not to imply endorsement or affiliation in any way.
Claims
Docket No. AKAM-446-PCTCLAIMS1. A method of detecting attributes in application layer network messages, comprising:generating a trained neural network model, at least by:receiving a first set of data describing events at one or more network filters that occurred in response to receiving first application layer network messages;receiving at least a portion of each of the application layer network messages related to the events;performing input embedding on the first set of data to encode elements in the first application layer network messages;training a neural network with the input embeddings to predict a sequence of characters in the first application layer network messages for an attribute;applying the trained neural network model to a second set of data describing events at the one or more network filters that occurred in response to receiving second application layer network messages, wherein the trained neural network model generates predicted sequences of characters in the second application layer messages;based on the trained neural network model’s predicted sequence of characters, obtaining a rule to install in the one or more network filters, the rule being executable to identify the attribute in subsequent application layer network messages.
2. The method of claim 1, wherein the first and second application layer messages comprise semi-structured language elements of a network protocol, and said input embedding comprises encoding semi -structured language elements.
3. The method of claim 2, wherein the network protocol comprises HTTP.
4. The method of claim 1, wherein the first set of data describing events comprises: a rule in the one or more network filters that detected an attack associated with an attack vector.
5. The method of claim 4, wherein the attack vector is used as a label for training the neural network model.Docket No. AKAM-446-PCT6. The method of claim 4, wherein the attack vector comprises any of SQL injection, cross site scripting attack, remote code injection, an OWASP top web application attack.
7. The method of claim 1, wherein the attribute comprises an attack vector.
8. The method of claim 1, wherein the training further comprises training the neural network model to tag predicted sequences of characters as an attack substring.
9. The method of claim 8, wherein said tagging comprises applying parts of speech tagging.
10. The method of claim 1, further comprising, during said training, applying a loss function that (i) penalizes the neural network model for missing attributes found by the one or more network filters and (ii) does not penalize the neural network model for finding attributes missed by the one or more network filters.
11. The method of claim 1, wherein the one or more network filters comprises one or more application layer firewalls.
12. The method of claim 1, wherein the first set of data describing events comprises, for a given event: a regular expression the one or more network filters applied to the application layer network message to trigger the event.
13. The method of claim 1, wherein the neural network comprises a transformer model having a self attention mechanism for processing of encoded language elements.
14. A system comprising:at least one first computer having circuitry forming memory for storing program instructions and forming a processor for executing the program instructions to:generate a trained neural network model, at least by:receiving a first set of data describing events at one or more network filters that occurred in response to receiving first application layer network messages;receiving at least a portion of each of the application layer network messages related to the events;Docket No. AKAM-446-PCTperforming input embedding on the first set of data to encode elements in the first application layer network messages;training a neural network with the input embeddings to predict a sequence of characters in the first application layer network messages for an attribute;at least one second computer having circuitry forming memory for storing program instructions and forming a processor for executing the program instructions to:apply the trained neural network model to a second set of data describing events at the one or more network filters that occurred in response to receiving second application layer network messages, wherein the trained neural network model generates predicted sequences of characters in the second application layer messages;at least one third computer having circuitry forming memory for storing program instructions and forming a processor for executing the program instructions to:obtain a rule to install in the one or more network filters, the rule being executable to identify the attribute in subsequent application layer network messages.
15. A non-transitory computer readable medium storing computer program instructions for execution on one or more hardware processors in one or more computers, the computer program instructions when executed causing the one or more computers to operate at least by:generating a trained neural network model, at least by:receiving a first set of data describing events at one or more network filters that occurred in response to receiving first application layer network messages;receiving at least a portion of each of the application layer network messages related to the events;performing input embedding on the first set of data to encode elements in the first application layer network messages;training a neural network with the input embeddings to predict a sequence of characters in the first application layer network messages for an attribute;Docket No. AKAM-446-PCTapplying the trained neural network model to a second set of data describing events at the one or more network filters that occurred in response to receiving second application layer network messages, wherein the trained neural network model generates predicted sequences of characters in the second application layer messages;based on the trained neural network model’s predicted sequence of characters, obtaining a rule to install in the one or more network filters, the rule being executable to identify the attribute in subsequent application layer network messages.