Machine learning architecture for detecting malicious files using data streams
By performing machine learning classification on the data blocks of streaming files at the edge device, the problems of low efficiency and latency in streaming file classification in the existing technology are solved, early classification and efficient detection are achieved, and it is suitable for edge device processing of malicious files and other files.
Patent Information
- Application Number
- CN202380092715.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-31
- Filing Date
- 2023-12-22
- Publication Date
- 2025-09-16
AI Technical Summary
Existing systems are inefficient and have latency when classifying streaming files at edge devices, and are unable to effectively classify files before they are received or processed, especially under memory constraints.
A machine learning model is used to classify the blocks of streaming files at the edge device. Feature extraction and classification are performed by aligning a predetermined number of data blocks. The classifier is trained using deep learning technology, and state information is saved for iterative classification to ensure that the same set of bytes is aligned for analysis.
It enables early classification of streaming files at edge devices, reduces latency and improves processing efficiency. It can take action before files are received or processed, making it suitable for malicious file detection and other file classification tasks.
Smart Images

Figure CN120660086A_ABST
Abstract
Description
Background Art
[0001] Malicious individuals attempt to compromise computer systems in a variety of ways. As an example, such an individual may embed or otherwise include a malicious file in an email attachment and transmit (or cause the malicious file to be transmitted) to an unsuspecting user. When executed, the malicious file damages the victim's computer. Some types of malicious files instruct the compromised computer to communicate with a remote host. For example, a malicious file can turn a compromised computer into a "zombie" in a "botnet," thereby receiving instructions from and / or reporting data to a command and control (C&C) server under the control of the malicious individual. One way to mitigate the damage caused by malicious files is for security companies (or other appropriate entities) to attempt to identify the malicious file and prevent it from reaching / executing on the end-user computer. Another approach is to try to prevent the compromised computer from communicating with the C&C server. Unfortunately, the authors of malicious files are using increasingly sophisticated techniques to obfuscate the operation of their software. Therefore, there is a continuous need for improved technologies to detect malware and prevent its harm. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Various embodiments of the invention are disclosed in the following detailed description and accompanying drawings.
[0003] Figure 1 is a block diagram of an environment in which malicious traffic is detected or suspected, according to various embodiments.
[0004] Figure 2 is a block diagram of a system for classifying files according to various embodiments.
[0005] Figure 3 is a block diagram of the method used to classify the model.
[0006] Figure 4 Illustrated is a system for classifying streaming files based on subsets of their chunks, according to various embodiments.
[0007] Figure 5 Illustrated is a system for classifying streaming files based on subsets of their blocks according to various embodiments.
[0008] Figure 6 Graph illustrating performance of file classification using a subset of chunks of a streaming file, according to various embodiments.
[0009] Figure 7 is a flow chart of a method for classifying a streaming file before processing the entire content of the streaming file, according to various embodiments.
[0010] Figure 8is a flow chart of a method for classifying a streaming file before processing the entire content of the streaming file, according to various embodiments.
[0011] Figure 9 is a flow chart of a method for classifying a streaming file before processing the entire content of the streaming file, according to various embodiments.
[0012] Figure 10 is a flowchart of a method for training a classification model according to various embodiments.
[0013] Figure 11 is a diagram of a set of chunks associated with a streaming file, according to various embodiments.
[0014] Figure 12 is a flow chart of a method for detecting malicious files according to various embodiments.
[0015] Figure 13 is a block diagram of a system for classifying streaming files based on block data, according to various embodiments.
[0016] Figure 14 is a block diagram illustrating a classification of a set of chunks obtained in streaming data of a file, according to various embodiments.
[0017] Figure 15 is a block diagram illustrating a classification of a set of chunks obtained in streaming data of a file, according to various embodiments.
[0018] Figure 16 is a flow chart of a method for classifying streaming data of a file according to various embodiments.
[0019] Figure 17 is a flow chart of a method for detecting malicious files according to various embodiments.
[0020] Figure 18 is a flow chart of a method for detecting malicious files according to various embodiments.
[0021] Figure 19 is a flow chart of a method for detecting malicious files according to various embodiments. DETAILED DESCRIPTION
[0022] The present invention can be implemented in a variety of ways, including as: a process; an apparatus; a system; a composition of matter; a computer program product embodied on a computer-readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on or provided by a memory coupled to the processor. In this specification, these embodiments, or any other form that the invention may take, may be referred to as techniques. In general, the order of the steps of the disclosed processes may be changed within the scope of the present invention. Unless otherwise stated, a component described as being configured to perform a task, such as a processor or memory, may be implemented as a general component that is temporarily configured to perform a task at a given time, or as a specific component that is manufactured to perform a task. As used herein, the term "processor" refers to one or more devices, circuits, and / or processing cores that are configured to process data, such as computer program instructions.
[0023] A detailed description of one or more embodiments of the present invention is provided below together with the accompanying drawings that illustrate the principles of the present invention. The present invention is described in conjunction with these embodiments, but the present invention is not limited to any embodiment. The scope of the present invention is limited only by the claims and the present invention encompasses many alternatives, modifications and equivalents. In order to provide a thorough understanding of the present invention, many specific details are set forth in the following description. These details are provided for illustrative purposes, and the present invention can be practiced according to the claims without some or all of these specific details. For the purpose of clarity, technical material known in the technical field related to the present invention is not described in detail so as not to unnecessarily obscure the present invention.
[0024] As used herein, an edge device may include a device (e.g., a hardware system) that controls the flow of data at the boundary between two networks. As an example, an edge device is a device that provides an entry point into an enterprise or service core network. Examples of edge devices include inline security entities such as firewalls. Other examples of edge devices include routers, routing switches, integrated access devices, multiplexers, and wide area network access devices.
[0025] As used herein, an inline security entity may include a network node (e.g., a device) that enforces one or more security policies on information such as network traffic, files, and the like. As an example, the security entity may be a firewall. As another example, the inline security entity may be implemented as a router, a switch, a DNS resolver, a computer, a tablet, a notebook, a smartphone, and the like. Various other devices may be implemented as security entities. As another example, the inline security entity may be implemented as an application running on a device, such as an anti-malware application. As another example, the inline security entity may be implemented as an application running on a container or a virtual machine.
[0026] Various embodiments include systems, methods, and devices for classifying streaming files. In some embodiments, the classification of streaming files includes security processing at an inline security entity. The method includes obtaining streaming data for a file at an edge device, processing a set of chunks associated with the streaming data for the file using a machine learning model, and classifying the file at the edge device before processing the entire contents of the file.
[0027] Various embodiments include systems, methods, and devices for classifying streaming files. In some embodiments, the classification of streaming files includes security processing at an inline security entity. The method includes obtaining streaming data for a file at an edge device, aligning a predetermined amount of data in blocks associated with the streaming data for the file, processing a plurality of aligned blocks associated with the streaming data for the file using a machine learning model, and classifying the file at the edge device based at least in part on the classification of the plurality of aligned blocks.
[0028] Prior art systems for classifying files (including streamed files) perform classification after receiving the entire contents of the file. For example, prior art systems classify files by using all (or substantially all) of the file to predict the classification of the file. Classification of files by prior art systems can include performing feature extraction across the entire file (or substantially the entire contents of the file) and querying a model, such as a machine learning model, to obtain a prediction of the file classification (e.g., the likelihood that the file is malicious, etc.). As an example, prior art systems use an XGBoost machine learning model to perform classification of non-streamed files at an edge device.
[0029] Prior art systems are generally not viable techniques for classifying streaming files because such prior art systems need to wait for the entire file to complete a transaction (e.g., be downloaded) in order for the system to perform feature extraction on the streaming file. Because prior art systems wait for the entire file to be received before performing classification (e.g., using a model for feature extraction and classification), prior art systems are inefficient and create delays in the consumption of streaming data in the streaming file. Additionally, due to memory constraints, it is not feasible to use prior art systems at edge devices. Edge devices are generally unable to store chunks (e.g., packets) of data locally at the edge device, and therefore, portions of the streaming file are forwarded to the connected device before the prior art system can discern the classification of the streaming file (such as whether the streaming file is malicious).
[0030] Various embodiments disclose systems, methods, and devices for performing classification for a streaming file at an edge device (e.g., a firewall) and before the entire corresponding streaming file has been processed at the edge device (e.g., before the entire streaming file has been received). The system can perform classification of the streaming file based at least in part on one or more blocks of the streaming file. As an example, a block can be a predefined number of bytes of data (e.g., 1500 bytes of data). In some embodiments, the system analyzes each block sequentially (e.g., concurrently with the receipt of the block) and performs a prediction of the classification of the streaming file before the entire streaming file has been received / processed. The system can perform proactive measures for the streaming file in response to a particular classification of a block of the streaming file (e.g., if the prediction that the file corresponds to a particular classification exceeds a predefined classification threshold).
[0031] In the case where the classification is in the context of detecting malicious files, if the block does not indicate a malicious file (e.g., the file is not classified as malicious based on the block), the system processes the block sequentially and allows the block to pass through the system (e.g., performed by the device), and if the block indicates that the streamed file is malicious (e.g., the file is classified as malicious based on the block), active measures are performed for additional blocks. An example of an active measure may be preventing the remaining blocks of the streamed file from passing through the edge device (or being processed by the edge device).
[0032] In some embodiments, the system uses a machine learning model to facilitate the classification of streaming files at the edge device, the machine learning model being trained using streamlined deep learning techniques. The machine learning model is trained to classify files one block at a time (e.g., sequentially classifying a predefined number of bytes of data). Due to strict memory constraints at the edge device, it is impractical to store the entire file. However, various embodiments save some state information indicating the state of the streaming file. The information indicating the state is used to classify the current block, and the system then iteratively saves the state information and uses this information to classify the next block. In some embodiments, the state information corresponds to the result of a max pooling operation performed on a subset of the streaming file (e.g., one or more blocks of the streaming file).
[0033] In some cases, the profile of the file received at the edge device is non-linear. For example, some file types have header information included in the first block (e.g., the first packet). However, in order for the classification of streaming files based on block-level classification (e.g., using a single block at a time until the analysis / prediction is completed) to be deterministic, the classifier (e.g., a machine learning model) needs to always analyze the same type of bytes (e.g., bytes including non-header information). Various embodiments implement alignment of the blocks of the streaming file and ensure that classification is performed on the same set of bytes.
[0034] Various embodiments improve upon prior art systems because streaming files can be classified at the edge device, and the classification can be performed before the entire streaming file has been received or processed. Thus, various embodiments enable the system to take action earlier with respect to streaming files based on the classification of the streaming files, before the entire streaming file has been received or processed.
[0035] Despite the combination Figure 1-19 The embodiments described by the examples illustrated in are primarily described in the context of detection of malicious files / traffic (e.g., classifying a file as malicious / non-malicious based on analysis of a subset of the file's blocks), but the various embodiments may be implemented in other contexts for classifying streaming files. Examples of other contexts include, but are not limited to, classifying files as including / associated with: financial information, HIPPA information, personally identifiable information (PII), copyrighted material, General Data Protection Regulation (GDPR) data, etc. As an example, the various embodiments classify (or predict) whether a streaming file includes copyrighted material based on analysis of a subset of the streaming file's blocks (e.g., before the entire contents of the streaming file are received or processed).
[0036] Figure 1is a block diagram of an environment in which malicious traffic is detected or suspected, according to various embodiments. In the example shown, client devices 104-108 (respectively) are a laptop, desktop computer, and tablet computer present in an enterprise network 110 (belonging to "ACME Corporation"). Data device 102 (e.g., an edge device) is configured to enforce policies (e.g., security policies) regarding communications between client devices (such as client devices 104 and 106) and nodes external to enterprise network 110 (e.g., reachable via external network 118). Examples of such policies include policies that manage traffic shaping, quality of service, and traffic routing. Other examples of policies include security policies, such as policies that require scanning for threats in the following: incoming (and / or outgoing) email attachments, website content, input to an application portal (e.g., a web interface), files exchanged via an instant messaging program, and / or other file transfers. In some embodiments, data device 102 is also configured to enforce policies with respect to traffic that remains within (or does not enter) enterprise network 110. For example, the data device 102 may enforce policies regarding the leakage or improper transmission of certain data (such as GDPR data, PII, etc.).
[0037] In the example shown, the data device 102 is an inline security entity. However, various other embodiments may include a data device that is another type of edge device (e.g., a device that does not specifically provide inline security processing). The data device 102 performs low-latency processing / analysis of incoming data (e.g., traffic data) and determines whether to offload any processing of the incoming data to a cloud system (such as the security platform 140). As an example, the data device 102 processes streaming files and classifies the streaming files locally. In some embodiments, the data device 102 classifies the streaming files based on a subset of the streaming data before the entire content of the corresponding streaming file is received / processed. For example, the data device 102 can perform classification using individual blocks (e.g., packets or a predefined number of bytes). In connection with performing classification using individual chunks, the data device sequentially performs feature extraction on the chunks and classifies the streaming file based at least in part on the feature extraction, and then continues to iteratively perform such analysis on a chunk-by-chunk basis (e.g., in the order in which the chunks are received) until the earlier of: (i) the streaming file is classified (e.g., a prediction obtained based on the classification exceeds a predefined threshold, such as a predefined maliciousness threshold), and (ii) the streaming file has been fully received or processed. For example, the data device 102 queries a classifier or model (e.g., a machine learning model) stored locally at the data device 102 based at least in part on the feature extraction for a particular chunk to obtain a prediction of the classification of the streaming file using the chunk.
[0038] The techniques described herein can be used in conjunction with various platforms (e.g., desktops, mobile devices, gaming platforms, embedded systems, etc.) and / or various types of applications (e.g., Android .apk files, iOS applications, Windows PE files, Adobe Acrobat PDF files, Microsoft Windows PE installers, etc.). Figure 1 In the example environment shown, client devices 104-108 are a laptop, a desktop computer, and a tablet (respectively) that reside within enterprise network 110. Client device 120 is a laptop that resides outside enterprise network 110.
[0039] Data device 102 can be configured to work in conjunction with a remote security platform 140. Security platform 140 can be a cloud system such as a cloud service security entity. Security platform 140 can provide various services, including: performing static and dynamic analysis on malware samples; providing a list of signatures of known exploits (e.g., malicious input strings, malicious files, etc.) to a data device (such as data device 102) as part of a subscription; detecting exploits such as malicious input strings or malicious files (e.g., on-demand detection, or periodic updates of a mapping of an input string or file to an indication of whether the input string or file is malicious or benign); providing a probability that an input string or file is malicious or benign; providing / updating a whitelist of input strings or files that are considered benign; providing / updating input strings or files that are considered malicious; identifying malicious input strings; detecting malicious input strings; detecting malicious files; predicting whether an input string or file is malicious; and providing an indication that an input string or file is malicious (or benign). In various embodiments, the analysis results (as well as additional information related to applications, domains, etc.) are stored in a database 160. In various embodiments, the security platform 140 comprises one or more dedicated commercially available hardware servers (e.g., with multi-core processor(s), 32GB+ of RAM, gigabit network interface adapter(s), and hard drive(s)) running a typical server-class operating system (e.g., Linux). The security platform 140 can be implemented across a scalable infrastructure comprising multiple such servers, solid-state drives, and / or other suitable high-performance hardware. The security platform 140 can include several distributed components, including components provided by one or more third parties. For example, part or all of the security platform 140 can be implemented using Amazon Elastic Compute Cloud (EC2) and / or Amazon Simple Storage Service (S3). Additionally, as with the data device 102, whenever the security platform 140 is referred to as performing a task (such as storing data or processing data), it should be understood that one or more subcomponents of the security platform 140 (whether acting alone or in collaboration with third-party components) can collaborate to perform the task. As an example, the security platform 140 can optionally collaborate with one or more virtual machine (VM) servers to perform static / dynamic analysis. An example of a VM server is a physical machine that includes commercially available server-grade hardware (e.g., a multi-core processor, 32+ gigabytes of RAM, and one or more gigabit network interface adapters) running commercially available virtualization software (such as VMware ESXi, Citrix XenServer, or Microsoft Hyper-V). In some embodiments, the VM server is omitted.Additionally, the virtual machine server may be under the control of the same entity that manages security platform 140, but may also be provided by a third party. As one example, the virtual machine server may rely on EC2, with the remainder of security platform 140 provided by dedicated hardware that is owned and under the control of the operator of security platform 140.
[0040] In some embodiments, system 100 uses security platform 140 to perform processing, such as processing involving heavy computations, on traffic data offloaded from data device 102. Security platform 140 provides one or more services to data device 102, client device 120, and the like. Examples of services provided by security platform 140 (e.g., a cloud service entity) include data loss prevention (DLP) services, application cloud engine (ACE) services (e.g., services that identify application types based on traffic patterns or fingerprints), machine learning command control (MLC2) services, advanced URL filtering (AUF) services, threat detection services, enterprise data leakage services (e.g., detecting data leakage or identifying the source of leakage), and Internet of Things (IoT) services. Various other services may also be implemented.
[0041] In some embodiments, the system 100 (e.g., malicious sample detector 170, security platform 140, etc.) trains detection models to detect vulnerability exploits (e.g., malicious samples), malicious traffic, application identities, or detect certain types of information (e.g., predefined categories of information such as financial information, GDPR data, PII, etc.). The security platform 140 can store blacklists, whitelists, etc. for data (e.g., mappings of signatures to malicious files, etc.). In response to processing traffic data, the security platform 140 can send updates to inline security entities (such as data device 102). For example, the security platform 140 provides updates to mappings of signatures to malicious files, updates to mappings of signatures to benign files, etc.
[0042] According to various embodiments, a machine learning process is used to obtain (one or more) models trained by the system 100 (e.g., the security platform 140). Examples of machine learning processes that can be implemented in conjunction with training (one or more) models include random forests, linear regression, support vector machines, naive Bayes, logistic regression, K-nearest neighbors, decision trees, gradient boosted decision trees, K-means clustering, hierarchical clustering, density-based spatial clustering of applications with noise (DBSCAN), principal component analysis, etc. In some embodiments, the system trains an XGBoost machine learning classifier model. As an example, the input to the classifier (e.g., the XGBoost machine learning classifier model) is a combined feature vector or a set of feature vectors, and based on the combined feature vector or the set of feature vectors, the classifier model determines whether the corresponding traffic (e.g., the input string) is malicious, or the likelihood that the traffic is malicious (e.g., whether the traffic is vulnerability exploit traffic).
[0043] According to various embodiments, the security platform 140 includes a DNS tunnel detector 138 and / or a malicious sample detector 170. Use of the malicious sample detector 170 is related to determining whether a sample (e.g., traffic data) is malicious. In response to receiving a sample (e.g., an input string (such as an input string whose input is related to a login attempt), a file, a traffic pattern), the malicious sample detector 170 analyzes the sample (e.g., the input string, etc.) and determines whether the sample is malicious. For example, the malicious sample detector 170 determines one or more feature vectors (e.g., a combined feature vector) of the sample and uses a model to determine (e.g., predict) whether the sample is malicious. The malicious sample detector 170 determines whether the sample is malicious based at least in part on one or more attributes of the sample. In some embodiments, the malicious sample detector 170 receives the sample, performs feature extraction (e.g., feature extraction on one or more attributes of the input string), and determines (e.g., predicts) whether the sample (e.g., a SQL or command injection string) is malicious based at least in part on the feature extraction results. For example, the malicious sample detector 170 uses a classifier (e.g., a detection model) to determine (e.g., predict) whether a sample is malicious based at least in part on the feature extraction results. In some embodiments, the classifier corresponds to a model (e.g., a detection model) for determining whether a sample is malicious, and the model is trained using a machine learning process.
[0044] In some embodiments, malicious sample detector 170 includes one or more of a traffic parser 172 , a prediction engine 174 , an ML model 176 , and / or a cache 178 .
[0045] The use of the traffic parser 172 is related to the following: determining (e.g., isolating) one or more attributes associated with the sample being analyzed. As an example, in the case of a file, the traffic parser 172 can parse / extract information from the file (such as from the file header). The information obtained from the file may include: the libraries, functions, or files called ("invoke" / "call") by the analyzed file, the order of the calls, etc. As another example, in the case of an input string, the traffic parser 172 determines the set of alphanumeric characters or values associated with the input string. In some embodiments, the traffic parser 172 obtains one or more attributes associated with (e.g., from) the sample. For example, the traffic parser 172 obtains one or more patterns (e.g., patterns of alphanumeric characters), one or more sets of alphanumeric characters, one or more commands, one or more pointers or links, one or more IP addresses, regular expression statements, etc. from the sample.
[0046] In some embodiments, one or more feature vectors corresponding to the sample are determined by the malicious sample detector 170 (e.g., the traffic parser 172 or the prediction engine 174). For example, one or more feature vectors are determined (e.g., populated) based at least in part on one or more characteristics or attributes associated with the sample (e.g., in the case where the sample is an input string, one or more attributes or a set of alphanumeric characters or values associated with the input string). As an example, the traffic parser 172 uses one or more attributes associated with the sample in connection with: determining one or more feature vectors. In some embodiments, the traffic parser 172 determines a combined feature vector based at least in part on one or more feature vectors corresponding to the sample. As an example, a set of one or more feature vectors is determined (e.g., set or defined) based at least in part on a model for detecting vulnerability exploits. The malicious sample detector 170 may use the set of one or more feature vectors to determine one or more attributes of the pattern (e.g., attributes of fields to be populated in the feature vector, etc.), the use of which is related to training or implementing the model. The model can be trained using a feature set that is obtained based at least in part on sample malicious traffic, such as a feature set corresponding to a predefined regular expression statement and / or a feature vector set determined based on (algorithm-based) feature extraction. For example, the model can be determined based at least in part on performing malicious feature extraction in connection with generating (e.g., training) a model to detect vulnerability exploits. The malicious feature extraction can include one or more of: (i) obtaining specific features from a file or SQL and command injection string using a predefined regular expression statement, and (ii) using an algorithm-based feature extraction to filter out the described features from the original input data set.
[0047] In response to receiving a sample for which the malicious sample detector 170 will determine whether the sample is malicious (or the likelihood that the sample is malicious), the malicious sample detector 170 determines one or more feature vectors (e.g., feature vectors corresponding to a set of predefined regular expression statements, feature vectors corresponding to attributes or patterns obtained using algorithm-based analysis of vulnerability exploits, and / or a combination of both, etc.). As an example, in response to determining (e.g., obtaining) the one or more feature vectors, the malicious sample detector 170 (e.g., the traffic parser 172) provides (or makes accessible) the one or more feature vectors to the prediction engine 174 (e.g., in connection with obtaining a prediction of whether the sample is malicious). As another example, the malicious sample detector 170 (e.g., the traffic parser 172) stores the one or more feature vectors in a cache 178 or database 160, for example.
[0048] In some embodiments, prediction engine 174 determines whether a sample is malicious based at least in part on one or more of: (i) a mapping of samples to indications of whether the corresponding samples are malicious, (ii) a mapping of identifiers of the samples (e.g., a hash or other signature associated with the samples) to indications of whether the corresponding samples are malicious, and / or (iii) a classifier (e.g., a model trained using a machine learning process). In some embodiments, determining whether a sample is malicious (e.g., based on a mapping of identifiers to indications that the sample is malicious) can be performed at data device 102, and for samples whose associated identifiers are not stored in the mapping(s), data device 102 offloads processing of the sample to security platform 140.
[0049] The prediction engine 174 is used to predict whether a sample is malicious. In some embodiments, the prediction engine 174 determines (e.g., predicts) whether a received sample is malicious. The prediction engine 174 determines whether a newly received sample is malicious based at least in part on features / attributes related to the sample (e.g., regular expression statements, information obtained from file headers, calls to libraries, APIs, etc.). For example, the prediction engine 174 applies a machine learning model to determine whether a newly received sample is malicious. Applying a machine learning model to determine whether a sample is malicious may include: the prediction engine 174 queries the machine learning model 176 (e.g., using information related to the sample, one or more feature vectors, etc.). In some embodiments, the machine learning model 176 is pre-trained, and the prediction engine 174 does not need to provide the training data set (e.g., sample malicious traffic and / or sample benign traffic) to the machine learning model 176 simultaneously with the query (indication information / determination of whether a particular sample is malicious). In some embodiments, the prediction engine 174 receives information associated with whether the sample is malicious (e.g., indication information that the sample is malicious). For example, prediction engine 174 receives the results of a determination or analysis by machine learning model 176. In some embodiments, prediction engine 174 receives an indication of the likelihood that a sample is malicious from machine learning model 176. In response to receiving the indication of the likelihood that a sample is malicious, prediction engine 174 determines (e.g., predicts) whether the sample is malicious based at least in part on the likelihood that the sample is malicious. For example, prediction engine 174 compares the likelihood that the sample is malicious to a likelihood threshold (e.g., a predetermined maliciousness threshold). In response to determining that the likelihood that the sample is malicious is greater than the likelihood threshold, prediction engine 174 may deem (e.g., determine) the sample to be malicious. Conversely, in response to determining that the likelihood that the sample is malicious is greater than the likelihood threshold, prediction engine 174 may deem (e.g., determine) the sample to be benign (e.g., non-malicious).
[0050] According to various embodiments, in response to the prediction engine 174 determining that a received sample is malicious, the security platform 140 sends an indication that the sample is malicious to a security entity (e.g., data device 102). For example, the malicious sample detector 170 can send an indication that the sample is malicious to an inline security entity (e.g., a firewall) or a network node (e.g., a client). The indication that the sample is malicious can correspond to an update to a sample blacklist (e.g., corresponding to malicious samples), such as when the received sample is deemed malicious, or to an update to a sample whitelist (e.g., corresponding to non-malicious samples), such as when the received sample is deemed benign. In some embodiments, the malicious sample detector 170 sends a hash or signature corresponding to the sample in conjunction with the indication that the sample is malicious or benign. The security entity or endpoint can calculate the hash or signature of the sample and perform a lookup (e.g., querying a whitelist and / or blacklist) for a mapping of the hash / signature to the indication of whether the sample is malicious / benign. In some embodiments, the hash or signature uniquely identifies the sample.
[0051] Prediction engine 174 is used in connection with determining whether a sample (e.g., an input string) is malicious (e.g., determining a likelihood or prediction of whether a sample is malicious). Prediction engine 174 uses information related to a sample (e.g., one or more attributes, patterns, etc.) in connection with determining whether the corresponding sample is malicious.
[0052] In response to receiving a sample to be analyzed, the malicious sample detector 170 may determine whether the sample corresponds to a previously analyzed sample (e.g., whether the sample matches a sample associated with historical information for which a maliciousness determination has been previously calculated). As an example, the malicious sample detector 170 determines whether an identifier or representative information corresponding to the sample is included in historical information (e.g., a blacklist, a whitelist, etc.). In some embodiments, the representative information corresponding to the sample is a hash or a signature of the sample. In some embodiments, the malicious sample detector 170 (e.g., the prediction engine 174) determines whether information related to a particular sample is included in a dataset of historical input strings and historical information associated with the historical dataset (e.g., a dataset such as VirusTotal). TMIn response to determining that information related to a particular sample is not included in or is not available in the dataset of historical input strings and the historical information, the malicious sample detector 170 may deem the sample as not yet analyzed, and the malicious sample detector 170 may invoke analysis (e.g., dynamic analysis) of the sample related to determining (e.g., predicting) whether the sample is malicious (e.g., the malicious sample detector 170 may query a classifier based on the sample, querying the classifier based on the sample related to determining whether the sample is malicious). An example of historical information associated with a historical sample indicating whether a particular sample is malicious corresponds to (VT) score. If the VT score of a particular sample is greater than 0, the third-party service considers the particular sample to be malicious. In some embodiments, the historical information associated with the historical sample that indicates whether the particular sample is malicious corresponds to a social score, such as a community-based score or rating (e.g., a reputation score) indicating that the sample is malicious or is likely to be malicious. Historical information (e.g., information from a third-party service, a community-based score, etc.) indicates whether other vendors or cybersecurity organizations consider the particular sample to be malicious.
[0053] In some embodiments, the malicious sample detector 170 (e.g., the prediction engine 174) determines that the received sample is newly analyzed (e.g., the sample is not within historical information / datasets, is not on a whitelist or blacklist, etc.). The malicious sample detector 170 (e.g., the traffic parser 172) can detect that the sample is newly analyzed in response to the security platform 140 receiving the sample from a security entity (e.g., a firewall) or an endpoint within the network. For example, the malicious sample detector 170 determines that the sample is newly analyzed concurrently with the security platform 140 or the malicious sample detector 170 receiving the sample. As another example, the malicious sample detector 170 (e.g., the prediction engine 174) determines that the sample is newly analyzed based on a predefined schedule (e.g., daily, weekly, monthly, etc.), such as in connection with a batch process. In response to determining that the received sample has not been analyzed for whether such sample is malicious (e.g., the system does not include historical information for such input string), the malicious sample detector 170 determines whether to use analysis of the sample (e.g., dynamic analysis) (e.g., querying a classifier to analyze the sample or one or more feature vectors associated with the sample, etc.), the analysis of the sample is relevant to determining whether the sample is malicious, and the malicious sample detector 170 uses a classifier for a set of feature vectors or combined feature vectors associated with relationships or characteristics of attributes or properties in the sample.
[0054] The machine learning model 176 predicts whether a sample (e.g., a newly received sample) is malicious based at least in part on the model. As an example, the model is pre-stored and / or pre-trained. Various machine learning processes can be used to train the model. According to various embodiments, the machine learning model 176 uses relationships and / or patterns, characteristics, or relationships between attributes of the sample and / or training set to estimate whether the sample is malicious, such as to predict the likelihood that the sample is malicious. For example, the machine learning model 176 uses a machine learning process to analyze a set of relationships between information indicating whether a sample is malicious (or benign) and one or more attributes (related to the sample), and uses the set of relationships to generate a predictive model for predicting whether a particular sample is malicious. In some embodiments, in response to predicting that a particular sample is malicious, an association between the sample and the information indicating that the sample is malicious is stored, such as at the malicious sample detector 170 (e.g., cache 178). In some embodiments, in response to predicting the likelihood that a particular sample is malicious, an association between the sample and the likelihood that the sample is malicious is stored, such as at the malicious sample detector 170 (e.g., cache 178). Machine learning model 176 can provide an indication of whether the sample is malicious or an indication of the likelihood that the sample is malicious to prediction engine 174. In some embodiments, machine learning model 176 provides the following indication to prediction engine 174: the analysis of machine learning model 176 is complete and the corresponding result (e.g., the prediction result) is stored in cache 178.
[0055] Cache 178 stores information related to samples (e.g., input strings). In some embodiments, cache 178 stores a mapping of information indicating whether an input string is malicious (or potentially malicious) to a specific input string, or a mapping of information indicating whether a sample is malicious (or potentially malicious) to a hash or signature corresponding to the sample. Cache 178 can store additional information related to a sample set, such as attributes of the sample, hashes or signatures corresponding to samples in the sample set, other unique identifiers corresponding to samples in the sample set, and the like. In some embodiments, an inline security entity (such as data device 102) stores a cache corresponding to or similar to cache 178. For example, the inline security entity can use a local cache to perform inline processing of traffic data, such as low-latency processing.
[0056] Return to Figure 1, assume that a malicious individual (using client device 120) has created malware or malicious input string 130. The malicious individual hopes that a client device (such as client device 104) will execute a copy of the malware or other exploit (e.g., malware or malicious input string) 130, thereby compromising the client device and causing the client device to become a zombie in a botnet. The compromised client device may then be instructed to perform tasks (e.g., cryptocurrency mining, or engaging in a denial of service attack) and / or report information to an external entity such as a command and control (C&C) server 150 (e.g., associated with such tasks, leaking sensitive corporate data, etc.), and receive instructions from the C&C server 150, if applicable.
[0057] Figure 1 The illustrated environment includes three Domain Name System (DNS) servers (122-126). As shown, DNS server 122 is under the control of ACME (for use by computing assets located within enterprise network 110), while DNS server 124 is publicly accessible (and can also be used by computing assets located within network 110 as well as other devices such as those located within other networks (e.g., networks 114 and 116). DNS server 126 is publicly accessible but under the control of a malicious operator of C&C server 150. Enterprise DNS server 122 is configured to resolve enterprise domain names into IP addresses and is further configured to communicate with one or more external DNS servers (e.g., DNS servers 124 and 126) to resolve domain names, where applicable.
[0058] In order to connect to a legitimate domain (e.g., www.example.com, depicted as website 128), a client device (such as client device 104) will need to resolve the domain to a corresponding Internet Protocol (IP) address. One way this resolution can occur is for client device 104 to forward a request to DNS servers 122 and / or 124 to resolve the domain. In response to receiving a valid IP address for the requested domain name, client device 104 can use the IP address to connect to website 128. Similarly, in order to connect to malicious C&C server 150, client device 104 will need to resolve the domain "kj32hkjqfeuo32ylhkjshdflu23.badsite.com" to a corresponding Internet Protocol (IP) address. In this example, malicious DNS server 126 is authoritative for *.badsite.com and the request from client device 104 will be forwarded (e.g.) to DNS server 126 for resolution, thereby ultimately allowing C&C server 150 to receive data from client device 104.
[0059] Data device 102 is configured to enforce policies regarding communications between client devices (such as client devices 104 and 106) and nodes external to enterprise network 110 (e.g., reachable via external network 118). Examples of such policies include policies governing traffic shaping, quality of service, and traffic routing. Other examples of policies include security policies, such as those requiring scanning for threats in incoming (and / or outgoing) email attachments, website content, information entered into a web interface such as a login screen, files exchanged via instant messaging programs, and / or other file transfers, and / or quarantining or deleting files or other exploits identified as malicious (or potentially malicious). In some embodiments, data device 102 is further configured to enforce policies with respect to traffic remaining within enterprise network 110. In some embodiments, security policies include instructions that network traffic (e.g., all network traffic, specific types of network traffic, etc.) be classified / scanned by a classifier stored in a local cache, or that certain detected network traffic be further analyzed (e.g., using a more refined detection model), such as by offloading processing to security platform 140.
[0060] In various embodiments, data device 102 includes a DNS module 134 configured to facilitate determining whether a client device (e.g., client devices 104-108) is attempting to participate in a malicious DNS tunnel, and / or to prevent connections (e.g., made by client devices 104-108) to malicious DNS servers. DNS module 134 may be integrated into data device 102 (e.g., Figure 1 and can also operate as a standalone device in various embodiments. Figure 1 Like the other components shown in , DNS module 134 can be provided by the same entity that provides data device 102 (or security platform 140), and can also be provided by a third party (e.g., a third party different from the provider of data device 102 or security platform 140). Additionally, in addition to preventing connections to malicious DNS servers, DNS module 134 can also take other actions, such as personalized logging of tunneling attempts made by clients (which indicates that a given client is compromised and should be quarantined or otherwise investigated by an administrator).
[0061] In various embodiments, when a client device (e.g., client device 104) attempts to resolve a domain, DNS module 134 uses the domain as a query to security platform 140. This query can be performed concurrently with the resolution of the domain (e.g., with requests being sent to DNS servers 122, 124, and / or 126 and security platform 140). As one example, DNS module 134 can send a query (e.g., in JSON format) to front end 142 of security platform 140 via a REST API. Using a process described in more detail below, security platform 140 will determine (e.g., using DNS tunnel detector 138, such as decision engine 152 of DNS tunnel detector 138) whether the queried domain indicates a malicious DNS tunnel attempt and provide a result back to DNS module 134 (e.g., "Malicious DNS Tunnel" or "Not Tunnel").
[0062] In various embodiments, when a client device (e.g., client device 104) attempts to parse an SQL statement, SQL command, or other command injection string, data appliance 102 uses the corresponding sample (e.g., input string) as a query to a local cache and / or security platform 140. This query can be executed concurrently with the parsing of the SQL statement, SQL command, or other command injection string. As one example, data appliance 102 sends the query (e.g., in JSON format) to front-end 142 of security platform 140 via a REST API. As another example, data appliance 102 sends the query directly from the data plane of data appliance 102 to security platform 140 (e.g., front-end 142 of security platform 140). For example, a process running on data appliance 102 (e.g., a daemon running on the data plane to facilitate data processing offload, such as WIFClient) transmits the query (e.g., a request message) to security platform 140 without first transmitting the query to the message plane of data appliance 102, which would then transmit the query to security platform 140. For example, data device 102 is configured to use a process running on the data plane to query security platform 140 without the mediation of the management plane of data device 102. Using a process described in more detail below, security platform 140 will determine (e.g., using malicious sample detector 170) whether the queried SQL statement, SQL command, or other command injection string indicates an exploit attempt, and provide a result back to data device 102 (e.g., "malicious exploit" or "benign traffic").
[0063] In various embodiments, when a client device (e.g., client device 104) attempts to open a file or input string (such as received via an email attachment, instant message, or otherwise exchanged over a network), or when the client device receives such a file or input string, DNS module 134 uses the file or input string (or a calculated hash or signature, or other unique identifier, etc.) as a query to security platform 140. This query can be performed simultaneously with the receipt of the file or input string, or in response to a request from a user to scan a file. As an example, data appliance 102 can send a query (e.g., in JSON format) to front end 142 of security platform 140 via a REST API. The query can be transmitted to the security platform via a process / connector implemented on the data plane of data appliance 102. Using the processing described in more detail below, security platform 140 will determine (e.g., using a malicious file detector that may be similar to malicious sample detector 170, such as by using a machine learning model to detect / predict whether a file is malicious) whether the queried file is a malicious file (or is likely to be a malicious file) and provide a result back to data device 102 (e.g., "malicious file" or "benign file").
[0064] In various embodiments, the DNS tunnel detector 138 (whether implemented on the security platform 140, on the data device 102, or on some other appropriate location / combination of locations) uses a two-pronged approach to identifying malicious DNS tunnels. The first approach uses an anomaly detector 146 (e.g., implemented using python) to build a real-time profile set of DNS traffic for the root domain (156). The second approach uses signature generation and matching (also referred to herein as similarity detection and implemented, for example, using Go). The two approaches are complementary. The anomaly detector serves as a general detector that can identify previously unknown tunnel traffic. However, the anomaly detector may need to observe multiple DNS queries before it can make a detection. To block the first DNS tunnel packet, the similarity detector 144 complements the anomaly detector 146 and extracts a signature from the detected tunnel traffic that can be used to identify a situation where an attacker registers a new malicious tunnel root domain, but uses tools / malware similar to those corresponding to the detected root domain.
[0065] When data device 102 receives DNS queries (e.g., from DNS module 134), data device 102 provides them to security platform 140, which performs both anomaly detection and similarity detection, respectively. In various embodiments, if either detector flags a domain, then the domain (e.g., as provided in the query received by security platform 140) is classified as a malicious DNS tunnel root domain.
[0066] The DNS tunnel detector 138 maintains a set of fully qualified domain names (FQDNs) for each device (from which it receives data) based on their root domain (in Figure 1 The queries are grouped by domain profile 156 (collectively illustrated in the figure as domain profile 156). Although grouping by root domain is generally described in this specification, it should be understood that the techniques described herein can be extended to any level of domain. In various embodiments, information about received queries for a given domain is retained in the profile for a fixed amount of time (e.g., a sliding time window of ten minutes).
[0067] As an example, DNS query information received from data device 102 for various foo.com sites is grouped (into a domain profile for the root domain foo.com) as follows: G(foo.com) = [mail.foo.com, coolstuff.foo.com, domain1234.foo.com]. A second root domain would have a second profile with similar applicable information (e.g., G(baddomain.com) = [lskjdf23r.baddomain.com, kj235hdssd233.baddomain.com]. Each root domain (e.g., foo.com or baddomain.com) is modeled using a set of characteristics unique to malicious DNS tunnels, such that even if benign DNS patterns are diverse (e.g., k2jh3i8y35.legitimatesite.com, xxx888222000444.otherlegitimatesite.com), such DNS patterns are highly unlikely to be misclassified as malicious tunnels. The following are example characteristics that may be extracted as features (eg, into feature vectors) for a given group of domains (ie, a shared root domain).
[0068] In some embodiments, the malicious sample detector 170 provides an indication of whether the sample is malicious to a security entity (such as the data device 102). For example, in response to determining that the sample is malicious, the malicious sample detector 170 sends an indication that the sample is malicious to the data device 102, and the data device may then enforce one or more security policies based at least in part on the indication that the sample is malicious. The one or more security policies may include isolating / quarantining input strings or files, deleting the sample, ensuring that the sample is not executed or parsed, warning or prompting the user about the maliciousness of the sample before the user opens / executes the sample, and the like. As another example, in response to determining that the sample is malicious, the malicious sample detector 170 provides the security entity with: an update to a mapping of the sample (or a hash, signature, or other unique identifier corresponding to the sample) to the corresponding indication of whether the sample is malicious, or an update to a blacklist of malicious samples (e.g., identifying samples) or a whitelist of benign samples (e.g., identifying samples that are not considered malicious).
[0069] In some embodiments, one or more feature vectors corresponding to a sample (such as a file, an input string, etc.) are determined by the system 100 (e.g., the security platform 140, the malicious sample detector 170, the pre-filter 135, etc.). For example, one or more feature vectors are determined (e.g., populated) based at least in part on one or more characteristics or attributes associated with the sample (e.g., in the case where the sample is an input string, one or more attributes or a set of alphanumeric characters or values associated with the input string). As an example, the system 100 uses features associated with a classifier of the malicious sample detector 170 (e.g., a machine learning model 176 such as a detection model), one or more attributes associated with the sample, which are related to determining one or more feature vectors. In some embodiments, the pre-filter 135 determines a combined feature vector based at least in part on one or more feature vectors corresponding to the sample. As an example, a set of one or more feature vectors is determined (e.g., set or defined) based at least in part on a pre-filter model (e.g., based on pre-filter features). The system 100 (e.g., the pre-filter 135) can use the set of one or more feature vectors to determine one or more attributes of a pattern (e.g., attributes of fields in the feature vectors to be populated, etc.), which can be used in connection with training or implementing a model. The pre-filter model can be trained using a feature set that is at least partially based on a feature set obtained in connection with obtaining the detection model.
[0070] According to various embodiments, an edge device (e.g., an inline security entity such as data device 102) receives traffic data such as a file and locally classifies the traffic data. The edge device may use a local classifier (e.g., a machine learning model) stored in a cache, etc. For example, the edge device performs feature extraction locally on a file or a subset of the file and classifies the file based on the feature extraction using a local classifier. In some embodiments, the edge device receives a data stream (e.g., a streaming file) and locally classifies the data stream based on an analysis of at least a subset of the data stream (e.g., one or more blocks of the streaming file). As an example, the edge device iteratively obtains blocks of the streaming data and predicts whether the streaming data (e.g., the streaming file) is malicious based at least in part on the blocks. The edge device may perform feature extraction on the blocks, query a local classifier based on the results from the feature extraction (e.g., using one or more feature vectors obtained from the feature extraction), and obtain a prediction of the classification of the streaming data from the local classifier. As an example, in the case of performing security analysis, the prediction may correspond to the likelihood that the streaming data is malicious. The edge device can compare the prediction of the classification of the streaming data with a corresponding likelihood threshold for such classification (e.g., a predetermined maliciousness threshold in the case of evaluating whether the streaming data is malicious). In response to comparing the prediction of the classification with the corresponding likelihood threshold (e.g., a threshold for GDPR classification, a threshold for PII classification, a threshold for financial information classification, etc.), if the prediction of the classification exceeds the likelihood threshold, the edge device can consider the streaming data to correspond to the classification, or conversely, if the prediction of the classification is less than (or equal to) the likelihood threshold, the streaming data can be considered not to correspond to the classification. The edge device can then process the streaming data according to the classification or other traffic, as applicable.
[0071] As an illustrative example, in the case of security analysis performed on streaming data, if the prediction of whether the streaming data is malicious exceeds a probability threshold, the edge device may deem the streaming data as malicious. Conversely, if the prediction of whether the streaming data is malicious is less than (or equal to) the probability threshold, the edge device may deem the streaming data as non-malicious (e.g., benign). In response to determining (e.g., predicting) the classification of the streaming data, the edge device may implement / enforce applicable policies.
[0072] According to various embodiments, the edge device performs classification of the streaming data (eg, the streaming file) before the entire set of streaming data (eg, the entire streaming file) has been processed / received at the edge device. Figure 1Data device 102 receives traffic data, such as streaming data. In response to receiving the traffic data, data device 102 may perform classification of the traffic data. For example, the data device may perform classification of the streaming file based at least in part on one or more blocks of the streaming file. As an example, a block may be a predefined number of bytes of data (e.g., 1500 bytes of data). Data device 102 may sequentially analyze (e.g., concurrently with the receipt of the blocks) a set (or subset) of blocks of the streaming file and perform a prediction of the classification of the streaming file before the entire streaming file has been received / processed by data device 102. In response to classifying the streaming file, the data device may enforce (e.g., on a block-by-block basis, as each block is sequentially analyzed for classification) a policy for handling the streaming data. Examples of enforcing a policy include performing proactive measures on the streaming file in response to a particular classification of a block of the streaming file (e.g., if a prediction that the file corresponds to a particular classification exceeds a predefined classification threshold, such as in response to a prediction indicating that the streaming data is malicious).
[0073] The data device 102 stores a classifier (e.g., a machine learning model) for locally classifying traffic data (such as streaming files). The classifier can be trained by a remote server (such as the security platform 140) and provided to the data device 102 for inline classification. In addition, the classifier can be updated or retrained by the remote server, and the updated classifier provided to the data device 102.
[0074] In some embodiments, the profile of traffic data / files received at an edge device (such as data device 102) is non-linear. Streaming files received at an edge device typically include header information in at least the first data block or data packet. The header information biases other data included in the streaming file. For example, the header information corresponds to an offset by which substantive information in the streaming file is shifted. In order to ensure that classification of streaming files using blocks (e.g., classifying streaming files on a block-by-block basis) is deterministic, various embodiments implement alignment of information included in the blocks. For example, the data device 102 aligns the information included in the blocks to ensure that the data device 102 (e.g., by using a classifier) analyzes the same type of bytes (e.g., substantive information, or bytes including non-header information, rather than header information). In connection with performing alignment of information included in a block, data device 102 retrieves a second data set from a first block (e.g., the last X bytes of the first block) and a first data set from a second block (e.g., the first Y bytes of the second block), treats the second data set from the first block and the first data set from the second block as data of a single specific block, and classifies the streaming file using the second data set from the first block and the first data set from the second block (e.g., using the single specific block). X and Y may be predefined positive integers. In some embodiments, data device 102 stores a mapping of file types to X and Y values, and in response to beginning to receive a streaming file, data device 102 determines the file type of the streaming file and retrieves the corresponding X and Y values. Data device 102 then uses the X and Y values to align the block data in the set of blocks of the streaming file (e.g., to account for header information).
[0075] In connection with performing classification of a streaming file based on blocks (e.g., block data), data device 102 obtains a prediction of a classification for the streaming file. The prediction may correspond to a likelihood that the streaming file corresponds to a particular classification (e.g., a likelihood that the streaming file is malicious). In various embodiments, data device 102 compares the prediction of the classification to a predefined classification threshold and, based on the result of the comparison, determines whether to deem the streaming file to meet the predicted classification. For example, if the likelihood that the streaming file is malicious is greater than a predefined maliciousness threshold, data device 102 deems the streaming file to be malicious and handles the traffic data (e.g., the streaming file) accordingly. According to various embodiments, when data device 102 obtains a prediction for a streaming file based on applicable block data, data device 102 handles the streaming data according to the predicted classification. If the prediction of the block classification indicates (e.g., is deemed to indicate based on the prediction meeting a predefined threshold) that the streaming file corresponds to a particular classification (e.g., a predefined classification such as malicious, PII, financial data, GDPR data, etc.), data device 102 handles the remaining portion of the streaming file according to that classification. For example, in the case of security processing at data device 102, when the first block has a predicted classification that meets a predefined maliciousness threshold, data device 102 treats the streaming file as malicious (e.g., treats the current block and all future blocks of the streaming file as malicious).
[0076] In some embodiments, in connection with classifying traffic data using blocks, the data device 102 uses dynamic classification thresholds. For example, a first classification threshold may be used to classify blocks received earlier (e.g., at the beginning of the streaming data), and a second classification threshold may be used to classify blocks received later (e.g., at the end of the streaming file). The first classification threshold and the second classification threshold are different. In some embodiments, the first classification threshold is higher than the second classification threshold. For example, as more and more data of the streaming file is processed, the data device 102 lowers the classification threshold (e.g., the maliciousness threshold). Various classification thresholds or changes to the dynamic classification threshold can be set based on empirical testing of the classification of the streaming file. The data device 102 can store a mapping of the sequence number of a particular block in the block set (or a percentile relative to the total number of blocks) to an applicable classification threshold to be used when classifying the particular block. As an example, the predefined maliciousness threshold for the first block is lower than the predefined maliciousness threshold for the jth block, and j is a positive integer greater than 1.
[0077] Figure 2 is a block diagram of a system for classifying files according to various embodiments. In some embodiments, system 200 is at least partially composed of Figure 1System 100 is implemented. System 200 can be implemented by an inline security entity. In various embodiments, system 200 is implemented in combination with the following: Figure 4 System 400, Figure 5 System 500, Figure 13 System 1300 and / or Figure 14 In various embodiments, the system 200 is implemented in conjunction with: Figure 3 The process of 300 Figure 7 The process of 700 Figure 8 The process of 800 Figure 9 The process of 900 Figure 10 Course 1000 and / or Figure 11 Process 1100, Figure 16 The process of 1600 Figure 17 Process 17 Figure 18 Process 1800 and / or Figure 19 The process 1900 of the system 200 can be implemented in one or more servers, security entities such as firewalls, and / or endpoints.
[0078] System 200 can be implemented by one or more devices such as servers. System 200 can be implemented at various locations on a network. In some embodiments, system 200 implements Figure 1 The data device 102 of the system 100 is configured as an edge device, such as a firewall or an inline security entity, that performs inline security processing on traffic data (e.g., the system 200 determines whether a file is malicious and processes the traffic data based on this maliciousness classification as a service). File classification can be implemented in conjunction with a locally stored classifier, such as a machine learning model. The system 200 can receive the classifier from the server and store it locally in its cache for inline file classification / processing.
[0079] According to various embodiments, in response to receiving traffic data to be analyzed (e.g., classified, such as determining whether a file is malicious), the system 200 performs feature extraction on a streaming file (e.g., performs feature extraction on a specific data block in the streaming file), and uses a classifier to classify the streaming file based on the results (e.g., feature vectors) obtained from the feature extraction. The system 200 processes the traffic data (e.g., the streaming file) according to the classification of the streaming file. The system 200 iteratively performs feature extraction on blocks in a sequence of received / processed blocks from a set of blocks of the streaming file, and classifies the streaming file using corresponding results from the feature extraction. For example, the system 200 sequentially receives blocks of the streaming file, and the system 200 sequentially performs classification of the streaming file using the specific data blocks of the streaming data (e.g., processes the blocks in the order in which they are received and uses the blocks to perform classification).
[0080] In the example shown, the system 200 implements one or more modules related to classifying (e.g., predicting a classification) a file, such as a streamed file, as malicious, determining a likelihood that a file corresponds to a particular classification, and / or providing a notification or indication of whether a file is malicious or performing proactive measures in response to determining that the classification of the file matches a predefined classification (e.g., in response to determining that the maliciousness prediction exceeds a predefined maliciousness threshold). The system 200 includes a communication interface 205, one or more processors 210, storage 215, and / or memory 220. The one or more processors 210 include one or more of: a communication module 225, a block acquisition module 227, a block alignment module 229, a feature extraction module 231, a model training module 233, a prediction module 235, a notification module 237, and a security enforcement module 239.
[0081] In some embodiments, the system 200 includes a communication module 225. The system 200 uses the communication module 225 to communicate with various nodes or endpoints (e.g., client terminals, firewalls, DNS resolvers, data devices, other security entities, etc.) or user systems such as administrator systems. For example, the communication module 225 provides information to be transmitted to the communication interface 205. As another example, the communication interface 205 provides information received by the system 200 to the communication module 225. The communication module 225 is configured to receive files to be analyzed, such as from network endpoints or nodes such as security entities (e.g., firewalls). The communication module 225 is configured to query (one or more) third-party services for information related to the file (e.g., services that disclose information about the file, such as: third-party ratings or assessments of the maliciousness of the file, community-based ratings, assessments, or reputations related to the file, blacklists of files and / or whitelists of files, etc.). The communication module 225 is configured to receive one or more settings or configurations from the administrator. Examples of one or more settings or configurations include: configuration of a process for determining whether a file is malicious, configuration related to a classifier or machine learning model used to classify files, (one or more) predefined classification thresholds (e.g., a predefined maliciousness threshold, a predefined financial data threshold, etc.), settings related to header information of files or file types (e.g., header information / characteristics of various types of streaming files), a format or process based on which a combined feature vector is to be determined, a set of feature vectors to be provided to a classifier used to classify files (e.g., determine whether a file is malicious), information related to a whitelist of files (e.g., files that are not considered suspicious and whose traffic or attachments are allowed), information related to a blacklist of files (e.g., files that are considered suspicious and whose traffic or attachments are to be restricted).
[0082] In some embodiments, system 200 includes a block acquisition module 227. System 200 uses block acquisition module 227 to receive traffic data, such as streaming files. Block acquisition module 227 determines and / or acquires blocks received by system 200 (e.g., in connection with monitoring network traffic). As an example, a block can be a predefined number of bytes of data (e.g., 1500 bytes of data). Block acquisition module 227 can use a specific block definition (e.g., a predefined number of bytes) for all file types, or block acquisition module 227 can use different block definitions based on the file type. For example, system 200 stores a mapping of file types to block definitions (e.g., a number of bytes considered a block), and block acquisition module 227 queries the mapping of file types to block definitions to determine the block definition for a particular file type. Block acquisition module 227 acquires a block set corresponding to a streaming file, and the blocks in the block set can be received sequentially (e.g., in the order in which the blocks are arranged in the streaming file) or can be appropriately sorted by block acquisition module 227 before processing. In response to receiving / processing the chunks, the chunk acquisition module 227 provides the chunks to the feature extraction module 231 for analysis and classification. In some embodiments, in environments where the profile of the streaming data is non-linear and the streaming file received by the system 200 includes header information in the first chunk or otherwise includes unaligned / non-aligned chunks, the chunk acquisition module 227 provides the chunk data (e.g., one or more chunks of the streaming file) to the chunk alignment module 229 for chunk alignment prior to classification.
[0083] In some embodiments, the system 200 includes a block alignment module 229. The system 200 uses the block alignment module 229 to perform alignment of block data in a set of blocks of a streaming file. The block data is aligned to ensure that the classification of the various blocks is deterministic. The block alignment module 229 aligns the block data to account for header information included in the first block or other misalignment of the block data relative to the block. As an example, a block can be a predefined number of bytes of data (e.g., 1500 bytes of data). The predetermined number of bytes can be configurable, such as by an administrator (e.g., the predetermined number of bytes can be preset in a policy such as a block policy or a file classification policy).
[0084] In some embodiments, the block alignment module 229 aligns the block data of the multiple blocks to ensure that the same type of data (e.g., non-header information) is analyzed during the classification of the streaming file. The block alignment module 229 aligns the block data of the multiple blocks by treating a subset of the data of each of two consecutive blocks as the block data of a single block for which classification will be performed. The block alignment module 229 can align the block data of the multiple blocks by obtaining a first predetermined number of bytes from a first block and a second predetermined number of bytes from a second block (e.g., the block immediately following the first block in the streaming file), and treating the combined data of the first predetermined number of bytes and the second predetermined number of bytes as a single block to be used in the classification of the streaming file. For example, if the block size is set to 1500 bytes and the first block of the streaming file includes 500 bytes of header information, the block alignment module 229 uses the last 1000 bytes of the first block and the first 500 bytes of the second block and treats this data as corresponding to a single data block. The first predetermined number of bytes and the second predetermined number of bytes can be configurable. In some embodiments, the first predetermined number of bytes is determined based on the number of bytes included in the header information of the streaming file. Aligning the chunks enables the system to analyze (e.g., run a machine learning model against) the same number of bytes, which is relevant for classifying streaming files based on chunks.
[0085] In some embodiments, the system 200 includes a feature extraction module 231. The system 200 uses the feature extraction module 231 to perform feature extraction on specific blocks. For example, the feature extraction module 231 performs feature extraction on blocks acquired by the block acquisition module 227. As another example, such as in the case where the streaming file has predetermined header information, the feature extraction module 231 performs feature extraction on aligned blocks (e.g., block data that is considered a block by the block alignment module 229 (e.g., based on information acquired from a subsequent block)).
[0086] In some embodiments, the system 200 uses a feature extraction module 231 to determine a set of feature vectors or a combination of feature vectors to use in connection with classifying a sample, such as determining whether a sample (e.g., a streamed file) is malicious (e.g., using a detection model). In some embodiments, the set of one or more feature vectors is determined based at least in part on information related to the sample. For example, the feature extraction module 231 determines feature vectors for (e.g., characterizing) one or more of: (i) a set of regular expression statements (e.g., predefined regular expression statements), and / or (ii) one or more characteristics or relationships determined based on (algorithm-based) feature extraction.
[0087] In some embodiments, the system 200 (e.g., the prediction module 235) uses a combined feature vector in connection with determining whether a sample is malicious or suspicious, or in connection with otherwise filtering (e.g., removing) benign traffic. In some embodiments, the system 200 (e.g., the prediction module 235) uses a combined feature vector in connection with classifying a streamed file (e.g., determining whether the file is malicious, or another classification such as GDPR data, financial data, export-controlled data, etc.). The feature extraction module 231 can determine (one or more) such combined feature vectors. The combined feature vector is determined at least in part based on: a set of one or more feature vectors (e.g., a model-based feature set, such as a set of detection features in a model used to determine whether a sample is malicious). For example, the combined feature vector is determined at least in part based on: a set of feature vectors for a set of predefined regular expression statements, and a set of feature vectors for a characteristic or relationship determined based on (algorithm-based) feature extraction. Feature extraction module 231 determines a combined feature vector by concatenating a set of feature vectors for a set of predefined regular expression statements and / or a set of feature vectors for characteristics or relationships determined based on (algorithm-based) feature extraction. Feature extraction module 231 concatenates the feature vector sets according to a predefined process (e.g., a predefined order, etc.).
[0088] In some embodiments, the system 200 includes a model training module 233. The system 200 uses the model training module 233 to determine a model (e.g., a classifier) for classifying a streamed file. The model training module 233 can determine multiple models for classifying files along different vectors, such as for classifying: (i) whether the file is malicious, (ii) whether the file includes financial information, (iii) whether the file includes PII information, (iv) whether the file includes GDPR data, (v) whether the file includes export controlled data, (vi) whether the file includes another type of characteristic based on which a disposal policy can be applied, etc. The model training module 233 can determine a relationship (e.g., a signature) between a characteristic of a file (e.g., a streamed file) and a particular classification, such as a relationship between a characteristic of a file and the maliciousness of the file (or the likelihood that the file is malicious). Examples of machine learning processes that can be implemented in connection with training models include random forests, linear regression, support vector machines, naive Bayes, logistic regression, K-nearest neighbors, decision trees, gradient boosted decision trees, K-means clustering, hierarchical clustering, density-based spatial clustering of applications with noise (DBSCAN), principal component analysis, etc. In some embodiments, the model training module 233 trains an XGBoost machine learning classifier model. The input to the classifier (e.g., the XGBoost machine learning classifier model) is a combined feature vector or a set of feature vectors, and based on the combined feature vector or the set of feature vectors, the classifier model determines whether the corresponding .NET file is malicious, or the likelihood that the .NET file is malicious.
[0089] In some embodiments, the model(s) implemented by the system 200 for classifying streaming files are trained by a server. The system 200 can use a model training module 231 to obtain a model from the server. For example, the model training module 233 can communicate with the server to determine whether a particular model is available or whether a particular model has been updated. The model training module 233 can query the server according to a preset frequency or otherwise according to a model training / update policy (which can be configured, for example, by an administrator).
[0090] In some embodiments, the system 200 includes a prediction module 235. The system 200 uses the prediction module 235 to predict a classification of a file. As an example, the prediction module 235 predicts whether the streamed file corresponds to a particular classification (e.g., malicious, PII, financial data, export controlled data, GDPR data, etc.). Predicting a particular classification of a file may include predicting a likelihood that the file corresponds to the particular classification, comparing the predicted likelihood to a predefined classification threshold, and determining whether the file corresponds to the particular classification based on the comparison. For example, if the predicted likelihood exceeds the predefined classification threshold, the prediction module 235 deems the file to correspond to the particular classification. As another example, if the predicted likelihood does not exceed (e.g., is less than or equal to) the predefined classification threshold, the prediction module 235 deems the file to not correspond to the particular classification.
[0091] The prediction module 235 uses a model (such as a machine learning model) trained by the model training module 233 (or obtained from a server) to determine whether a file corresponds to a specific classification (e.g., predicting the classification of a file, such as whether the file is malicious). For example, the prediction module 235 uses an XGBoost machine learning classifier model to analyze a combined feature vector obtained based on feature extraction of block data for a specific block to determine the classification of the streaming file. As another example, the prediction module 235 uses a convolutional neural network model to analyze features / characteristics (e.g., feature vectors) (based on feature extraction of block data for the streaming file).
[0092] In some embodiments, the prediction module 235 iteratively classifies the streaming file based on the next block to be analyzed. For example, the prediction module 235 performs classification of the streaming file based on a single block (or a single aligned block) of the streaming file. The prediction module 235 can successively analyze blocks in the streaming file and determine the classification of the file at each block so analyzed. For example, the prediction module 235 determines the predicted classification of the file for each block in the streaming file, or until the prediction module 235 determines that the predicted classification of a particular block exceeds a predetermined classification threshold (e.g., in which case the prediction module 235 treats the streaming file as corresponding to the predicted classification). At each block of the streaming file, the system 200 can determine how to handle the streaming file (e.g., whether to allow transmission of the file, processing of the file (such as rendering of the file), etc.).
[0093] In some embodiments, the system 200 includes a notification module 237. The system 200 uses the notification module 237 to provide an indication of the classification of a file (e.g., to provide an indication that a sample streaming file is malicious). For example, the notification module 237 obtains the indication of the classification of the file (or the likelihood that the sample corresponds to a particular classification) from the prediction module 235 and provides the indication of the classification to one or more security entities and / or one or more endpoints.
[0094] In some embodiments, the system 200 includes a security enforcement module 239. The system 200 uses the security enforcement module 239 to enforce one or more security policies for information such as network traffic, streaming files, etc. The security enforcement module 239 enforces one or more security policies based on the classification of the files. As an example, the system 200 stores policies corresponding to different classifications, and these policies indicate how the files are to be handled. Examples of policies that the security enforcement module 239 can enforce include policies for handling malicious files, policies for handling files containing financial information, policies for handling files containing GDPR data, policies for handling files containing export-controlled information, policies for handling files containing PII, etc.
[0095] As an example, in the case where the system 200 is a security entity or firewall, the system 200 includes a security enforcement module 239. Firewalls typically deny or allow network transmissions based on a set of rules. These rule sets are often referred to as policies (e.g., network policies, network security policies, security policies, etc.). For example, a firewall can filter inbound traffic by applying a set of rules or policies to prevent unwanted external traffic from reaching the protected device. A firewall can also filter outbound traffic by applying a set of rules or policies (e.g., allowing, blocking, monitoring, notifying, or logging, and / or specifying other actions in the firewall rules or firewall policies, which can be triggered based on various criteria such as those described herein). A firewall can also filter local network (e.g., intranet) traffic by similarly applying a set of rules or policies. Other examples of policies include security policies, such as a security policy requiring scanning for threats in the following: incoming (and / or outgoing) email attachments, website content, files exchanged via instant messaging programs, information obtained via a web interface or other user interface (such as an interface to a database system (e.g., a SQL interface)), and / or other file transfers.
[0096] According to various embodiments, storage 215 includes one or more of file system data 260, model data 262, and / or prediction data 264. Storage 215 includes shared storage (eg, a network storage system) and / or database data and / or user activity data.
[0097] In some embodiments, file system data 260 includes a database, such as one or more data sets (e.g., one or more data sets for files and / or file attributes, a mapping of indicators of maliciousness or other classifications to files or hashes, multiple signatures or other unique identifiers for files, a mapping of indicators of benign files to files or hashes, a single signature or other unique identifier for files, etc.). File system data 260 includes data such as historical information about files (e.g., the maliciousness of files), a whitelist of files that are considered safe (e.g., not suspicious), a blacklist of files that are considered suspicious or malicious (e.g., files whose likelihood of maliciousness is considered to exceed a predetermined / preset likelihood threshold), information associated with suspicious or malicious files, etc. File system data 260 includes one or more policies, such as a security policy for handling malicious files, or other policies for handling other classifications.
[0098] Model data 262 includes information related to one or more models (e.g., classifiers) that are used to classify files or predict the likelihood that a file matches a particular classification (e.g., the likelihood that a sample is malicious or suspicious). As an example, model data 262 includes a convolutional neural network model configured to classify streaming files. As another example, model data 262 stores a classifier (e.g., (one or more) XGBoost machine learning classifier models, such as a detection model, a pre-filter model, or both) used in conjunction with a set of feature vectors or combined feature vectors. Model data 262 may include feature vectors that are generated for each of one or more of the following: (i) a set of regular expression statements, and / or (ii) algorithm-based features (e.g., features extracted using TF-IDF, such as features extracted for sample vulnerability exploit traffic, etc.). In some embodiments, model data 262 includes a combined feature vector that is generated based at least in part on one or more feature vectors corresponding to each of one or more of: (i) a set of regular expression statements, and / or (ii) algorithm-based features (e.g., features extracted using TF-IDF, such as features extracted for sample vulnerability exploit traffic, etc.).
[0099] Prediction data 264 includes information related to a determination of whether a sample analyzed by system 200 corresponds to a particular classification (e.g., a prediction of whether the sample is malicious). For example, prediction data 264 stores an indication that the sample is malicious, an indication that the sample is benign, etc. Information related to the determination can be obtained and provided by notification module 237 (e.g., transmitted to an applicable security entity, endpoint, or other system). In some embodiments, prediction data 264 includes a hash or signature of a sample (such as a sample analyzed by system 200 to determine whether such a sample is malicious), or a historical data set that has been previously assessed for maliciousness (such as by a third party). Prediction data 264 may include a mapping of hash values to indications of maliciousness (e.g., an indication that the corresponding sample is malicious or benign, etc.).
[0100] According to various embodiments, the memory 220 includes executing application data 270. The executing application data 270 includes data obtained or used in connection with executing an application, such as an application that performs a hash function, an application for extracting information from a file, or an application for analyzing the execution of files within a sandbox. In an embodiment, the application includes one or more applications that perform one or more of the following operations: receiving and / or executing queries or tasks, generating reports and / or configuring information responsive to the queries or tasks executed, and / or providing information responsive to the queries or tasks to a user. Other applications include any other appropriate applications (e.g., index maintenance applications, communication applications, machine learning model applications, applications for detecting suspicious files, documentation applications, report compilation applications, user interface applications, data analysis applications, anomaly detection applications, user authentication applications, security policy management / update applications, etc.).
[0101] Figure 3 is a block diagram of a method for classifying a model. In some embodiments, process 300 is performed at least in part by Figure 1 system 100 and / or Figure 2 The process 300 may be implemented by an inline security entity.
[0102] At 310, the sample is transmitted. The sample may be transmitted across a network, or otherwise transmitted from one endpoint to another, or the like.
[0103] At 320, a sample is obtained by a security entity such as a firewall. A firewall is configured to monitor traffic across a network or between two endpoints. In some embodiments, the firewall can be an application running on a client system and monitoring traffic to / from the client system.
[0104] At 330, the sample is analyzed using a machine learning model (such as an XGBoost model). Prior art systems traditionally analyze samples after the sample has been fully received / processed. Prior art systems obtain the entire sample, perform feature extraction on the sample (or a portion thereof, such as header information), and analyze the sample using a machine learning model. In some embodiments, a system (e.g., a firewall) performs feature extraction on the file, generates (one or more) feature vectors, queries a model based at least in part on the feature vectors, and obtains a result from the model. As an example, the result can be an indication of whether the file is malicious or non-malicious. As another example, the result can be an indication of the likelihood that the file is malicious or non-malicious, and the system compares the predicted likelihood to a predefined maliciousness threshold to determine whether to consider the file malicious.
[0105] In response to determining at 330 that the file is not malicious (eg, is benign), process 300 proceeds to 340 where the sample is treated as non-malicious traffic. For example, a firewall allows the transfer or execution of the file.
[0106] In response to determining at 330 that the file is malicious, process 300 proceeds to 350 where the sample is treated as malicious traffic. For example, a firewall enforces one or more security policies against the sample. As another example, the firewall blocks the transmission or execution of the file.
[0107] Figure 4 A system for classifying a streaming file based on a subset of its blocks according to various embodiments is illustrated. In some embodiments, the system 400 is at least partially comprised of Figure 1 system 100 and / or Figure 2 The system 400 may be implemented by an inline security entity.
[0108] In some embodiments, system 400 is configured to provide a prediction of whether a file corresponds to a particular classification (e.g., whether the file is malicious) based on block data of a particular block of the file and before the entire file has been received / processed. As an example, system 400 is deployed in an environment in which streamed files are received. Predicting the classification of a streamed file enables system 200 to provide low-latency prediction / disposition decisions. For example, system 200 can decide to dispose of a file according to a particular policy for classification based on a determination that the prediction (obtained by using the blocks to predict the classification of the file) meets one / more classification criteria (e.g., the prediction (such as the predicted likelihood) exceeds a predefined classification threshold).
[0109] In some embodiments, the system 400 is configured to provide a prediction of whether a file corresponds to a particular classification locally at an edge device (such as a firewall, router, or other security entity). Due to the memory and computational constraints of the edge device, the system 400 is configured to use a relatively small model and a relatively small amount of data (e.g., very little data reserved for classification) to generate the prediction.
[0110] refer to Figure 4 At 410, acquisition / processing of the sample, or acquisition / processing of at least a portion of the sample, is initiated. For example, the system 400 begins receiving a streaming file corresponding to the sample. Streaming files are typically relatively large, and therefore, data of the streaming file is streamed over a relatively long period of time.
[0111] At 420, the system obtains one or more blocks of the streaming file. In some embodiments, the system 400 successively receives / processes blocks of the streaming file, such as block 421, block 422, block 423, block 424, block 425, and so on. The system 400 can use each specific block to determine a predicted classification for the streaming file, and on a block-by-block basis, can determine how to handle the streaming file (e.g., determine a policy to be enforced for handling files of a specific classification). The system obtains block data from the block to be analyzed (e.g., each block of the streaming file, or each block until a specific classification is made). The system 400 then provides the block data for the specific block to a convolutional neural network or other classifier to perform feature extraction and classification of the streaming file based on the block data.
[0112] At 430, the block data is input to the convolution layer. In the example shown, the block data of block 421 is input to convolution layer 431, block 422 is input to convolution layer 432, block 423 is input to convolution layer 433, block 424 is input to convolution layer 434, block 425 is input to convolution layer 435, and block 426 is input to convolution layer 436. The size of the block data can be adjusted to an optimal size and input to the corresponding convolution layer.
[0113] A convolutional layer includes filters or kernels that filter block data. For example, a kernel is placed over a block of data. The extent to which a kernel is placed over a block of data, or the extent to which a kernel is used to process a block of data, is based on the kernel size (e.g., the kernel's dimensions). In some embodiments, the performance (e.g., accuracy, speed, etc.) of the classification of a streaming file can be adjusted based on the kernel size used by the convolutional layer. A relatively large kernel size generally provides greater accuracy, but requires a relatively large machine learning model and typically takes longer to generate inferences / predictions. Therefore, it may be preferable to select a kernel size that corresponds to a model size suitable for a particular edge device and is fast enough to generate low-latency inferences, such as to comply with quality of service policies or other configurations. In some embodiments, the kernel size is between 8 and 12. In some embodiments, the kernel size is 8. Compared to a larger kernel (such as a kernel size of 12), if the kernel size is 8, the inferences generated by the model are relatively fast, and the performance improvement between embodiments when the kernel size is 12 and when the kernel size is 8 is relatively insignificant. Therefore, when balancing the trade-off between inference speed and accuracy, selecting a kernel size of 8 may be preferable / optimal.
[0114] In some embodiments, the kernel size used in the convolutional layer affects the number of characters used for lookback (e.g., the number of characters buffered). The convolutional layer compares the block data piece by piece. For example, the convolutional layer uses different features for different parts of the block data. The convolutional layer uses filters that are related to computing the match between the block data and the features (e.g., the features corresponding to the class that the classifier uses to generate a prediction for that class).
[0115] Various embodiments implement a pooling mechanism, which is relevant to analyzing block data (e.g., generating predictions for specific classifications). Pooling is a mechanism that takes a large amount of information and shrinks it, while typically retaining important information in the output. In the example shown, the outputs from convolutional layers 431-436 are input to corresponding ones of max pooling modules 441-446, respectively. For example, at 440, the output from convolutional layer 431 is input to max pooling module 441. The max pooling module is configured to perform a max pooling operation on the information output by the applicable convolutional layer. The max pooling operation is a pooling operation that calculates the maximum value (or maximum value) of each feature map. As an example, the max pooling operation is downsampling, which produces a downsampled feature map that highlights the most significant features in the data.
[0116] At 450 , the output from the max pooling operation (eg, the output from max pooling modules 441 - 446 ) is used to generate a max layer.
[0117] At 460, the max layer is processed by a dense layer. The dense layer may include applying an activation function (such as a softmax operation) to the input data. The dense layer may convert the feature map output from the max layer into a probability distribution.
[0118] At 470, the output from the dense layer is used to generate a prediction (e.g., inference). For example, the probability distribution output by the softmax operation is used to determine a prediction of the likelihood that the streaming file corresponds to a particular classification. The system can compare the prediction of the likelihood to a predefined classification threshold and determine the classification based on the result of the comparison. For example, if the prediction exceeds the corresponding predefined classification threshold, the system considers the streaming file to correspond to a particular classification.
[0119] Figure 5 A system for classifying a streaming file based on a subset of its blocks according to various embodiments is illustrated. In some embodiments, the system 500 is at least partially comprised of Figure 1 system 100 and / or Figure 2 The system 500 may be implemented by an inline security entity.
[0120] At 510 , acquisition / processing of a sample, or acquisition / processing of at least a portion of a sample, is initiated. For example, the system 500 begins receiving a streaming file corresponding to the sample.
[0121] At 520, the system obtains one or more blocks of the streaming file. In some embodiments, the system 500 successively receives / processes blocks of the streaming file, such as block 521, block 522, block 523, block 524, block 525, and so on. The system 500 can use each specific block to determine a predicted classification for the streaming file, and on a block-by-block basis, can determine how to handle the streaming file (e.g., determine the policy to be enforced for handling files of a specific classification). The system obtains block data from the block to be analyzed (e.g., each block of the streaming file, or each block until a specific classification is made). The system 500 then provides the block data for the specific block to a convolutional neural network or other classifier to perform feature extraction and classification of the streaming file based on the block data.
[0122] At 530, the chunk data of the received chunks of the streaming file is provided to the classifier for feature extraction and inference / prediction of a classification of the streaming file. The inference / prediction of the classification of the streaming file may include or correspond to a prediction of whether the streaming file corresponds to a particular classification. In some embodiments, the system 500 sequentially provides multiple chunks of the streaming file to the classifier for sequential feature extraction and inference.
[0123] Performing feature extraction and generating a classification prediction includes processing the corresponding block data using a convolutional layer (e.g., a convolutional neural network) at 531. The system obtains a feature map for the block based on the processing of the block data (using the convolutional layer). In some embodiments, the convolutional layer uses a kernel with a kernel size between 8 and 12. In some embodiments, the kernel size is 8.
[0124] In response to using the convolutional layer to process the block data, the system 500 provides the output from the convolutional layer to the pooling mechanism. At 532, the output from the convolutional layer (e.g., for the particular block processed by the convolutional layer) is input to the pooling mechanism, and the pooling mechanism performs a maximum pooling operation.
[0125] At 533, the system 500 provides the output from the pooling mechanism to generate a maximum layer for the features associated with the block. The maximum layer is cached in a cached maximum feature at 534, which can be used at 535 (this is related to the comparison of the features of the next block).
[0126] At 536, the max layer is processed by a dense layer. The dense layer may include applying an activation function (such as a softmax operation) to the input data. The dense layer may convert the feature map output from the max layer into a probability distribution.
[0127] At 540, the output from the dense layer is used to generate a prediction (e.g., inference). For example, the probability distribution output by the softmax operation is used to determine a prediction of the likelihood that the streaming file corresponds to a particular classification. The system 500 can compare the prediction of the likelihood with a predefined classification threshold and determine the classification based on the result of the comparison. For example, if the prediction exceeds the corresponding predefined classification threshold, the system considers the streaming file to correspond to a particular classification.
[0128] In the example shown, the system 500 successively generates predictions for successive blocks at 541, 542, 543, and 544. Predictions 541-544 can correspond to the likelihood that the streaming file corresponds to a particular classification. For example, where the model is used to predict whether a streaming file is malicious, the prediction output at 540 can correspond to the likelihood that the streaming file is malicious (e.g., based on an analysis of the corresponding block). In response to obtaining the prediction, the system 500 can compare the prediction with a predetermined classification threshold. If the predetermined classification threshold is 0.95 (or 95%), the system 500 determines that the streaming file does not correspond to a particular classification (e.g., not malicious) because predictions 541-543 are less than the predetermined classification threshold. Conversely, the system 500 determines that prediction 544 indicates that the streaming file corresponds to a particular classification because prediction 544 is greater than the predetermined classification threshold.
[0129] Figure 6A diagram illustrating the performance of file classification using a subset of blocks of a streaming file according to various embodiments. Figure 6 In the example shown, result 600 illustrates that malicious files are mostly identified based on classification using earlier blocks. Implementing various embodiments to classify streaming files based on classification of specific blocks enables the system to classify streaming files relatively early in processing the streaming files. Thus, a system according to various embodiments is able to identify that a streaming file corresponds to an applicable classification (e.g., malicious) early in processing the streaming file, thereby enabling the system to determine how to handle the streaming file early (e.g., before the entire streaming file has been received / processed), such as by enforcing applicable policies for files (deemed to correspond to the applicable classification). For example, 50% of malicious files are identified by the 36th block of the streaming file, and 75% of malicious files are identified by the 50th block of the streaming file.
[0130] Figure 7 is a flow chart of a method for classifying a streaming file before processing the entire contents of the streaming file according to various embodiments. In some embodiments, process 700 is at least partially performed by Figure 1 system 100 and / or Figure 2 The process 700 may be implemented by an inline security entity.
[0131] At 702, a sequence of integers corresponding to bytes of the block being analyzed is obtained. At 704, a gather operation is performed to provide a lookup for unsigned integers between 0 and 256, and its corresponding value with a dimension of 16 is extracted. At 706, an expand operation is performed to expand the dimension by 1. At 708, the system transposes the dimensions of the information to adapt the information to an applicable CNN format (e.g., a 1D CNN format). For example, the transpose operation may include determining the product of batch, channel, and length. At 710, the system performs a convolution operation. For example, the system performs a 1D convolution along the length of the bytes with a predefined kernel size. In some embodiments, the kernel size is between 8 and 12. In some embodiments, the kernel size is 8. At 712, a compression operation is performed to remove dimensions with single values. At 714, an addition operation is performed to add a bias to the output of the convolution. At 716, a rectified linear activation function (ReLU) is applied to perform a nonlinear activation operation. At 718, a global max pooling operation is performed. For example, the system obtains a single maximum value across all bytes (e.g., the length of the sequence). At 720, the system performs a compression operation to reduce the dimensionality of the maximum activation value. At 722, an external input tracks the maximum activation value. For example, the system caches the maximum activation value. At 724, the system performs a max pooling operation to obtain the maximum output of the global max pooling output and the maximum activation value. At 726, the system stores the output from the max pooling operation in the cache. At 728, the system performs a matrix multiplication operation using the weights of the linear layer. At 730, the system adds a bias to the output of the matrix multiplication operation. At 732, the system performs a nonlinear activation operation (e.g., applying a nonlinear activation function). At 734, the system performs a matrix multiplication using the weights of the linear layer. At 736, the system adds a bias to the output of the matrix multiplication operation. At 738, the system performs a softmax operation to obtain class probabilities (e.g., to obtain a probability distribution). At 740, the class probabilities are passed as output.
[0132] Figure 8 is a flow chart of a method for classifying a streaming file before processing the entire contents of the streaming file according to various embodiments. In some embodiments, process 800 is at least partially performed by Figure 1 system 100 and / or Figure 2 The process 800 may be implemented by an inline security entity.
[0133] At 805, streaming data of a file is obtained. Obtaining streaming data of a file (eg, a streaming file) includes successively receiving blocks of the streaming file. For example, the streaming file is obtained by an edge device.
[0134] At 810, a set of chunks associated with the streaming data of a file is processed using a machine learning model. The system performs feature extraction on the chunks and queries the model using the features (e.g., feature vectors / feature maps). The system obtains a prediction in response to querying the model. The prediction can correspond to a probability / likelihood that the streaming file corresponds to a particular classification.
[0135] At 815, the file is classified. In some embodiments, the system classifies the file based on the predictions obtained from the model. For example, the system compares the predictions generated based on the analysis block to a predefined classification threshold. If the prediction exceeds the predefined classification threshold, the system deems the streamed file to correspond to a particular classification. For example, if the model is used to detect malicious files, the system deems the streamed file to be malicious if the predicted likelihood of the streamed file exceeds a predefined maliciousness threshold.
[0136] At 820, a determination is made as to whether process 800 is complete. In some embodiments, process 800 is determined to be complete in response to a determination that no additional samples are to be analyzed (e.g., no additional predictions are needed for the samples), no additional traffic is to be analyzed, an administrator indicates that process 800 is to be paused or stopped, etc. In response to determining that process 800 is complete, process 800 ends. In response to determining that process 800 is not complete, process 800 returns to 805.
[0137] In some embodiments, the system can implement process 800 for each successive chunk of the streaming file until all chunks have been processed or the system deems the streaming file to correspond to a particular classification (eg, a classification for its deployment model).
[0138] Figure 9 is a flow chart of a method for classifying a streaming file before processing the entire contents of the streaming file according to various embodiments. In some embodiments, process 900 is at least partially performed by Figure 1 system 100 and / or Figure 2 The process 900 may be implemented by an edge device such as an inline security entity.
[0139] Process 900 is implemented to determine whether a streaming file is malicious. For example, process 900 analyzes each successive block (at least until the file is deemed malicious) and classifies the file based on the analysis of the block(s).
[0140] At 905 , stream data of the file is obtained. In some embodiments, 905 corresponds to or is similar to 805 of process 800 .
[0141] At 910 , a set of chunks associated with the streaming data of the file is processed using a machine learning model. In some embodiments, 910 corresponds to or is similar to 810 of process 800 .
[0142] At 915 , the file is classified. In some embodiments, 915 corresponds to or is similar to 815 of process 800 .
[0143] At 920, the system determines whether the file is malicious. Based on a comparison of whether the prediction obtained from the model exceeds a predefined classification threshold, the system determines whether the output of the comparison indicates that the file is malicious.
[0144] In response to determining at 920 that the file is malicious, process 900 proceeds to 925 where one or more security policies are applied to the file. Conversely, in response to determining at 920 that the file is not malicious, process 900 proceeds to 930 where the file is treated as non-malicious traffic.
[0145] At 935, a determination is made as to whether process 900 is complete. In some embodiments, process 900 is determined to be complete in response to a determination that no additional samples are to be analyzed (e.g., no additional predictions are needed for the samples), no additional traffic is to be analyzed, an administrator indicates that process 900 is to be paused or stopped, etc. In response to determining that process 900 is complete, process 900 ends. In response to determining that process 900 is not complete, process 900 returns to 905.
[0146] Figure 10 is a flow chart of a method for training a classification model according to various embodiments. In some embodiments, process 1000 is at least partially performed by Figure 1 system 100 and / or Figure 2 The process 1000 may be implemented by an inline security entity.
[0147] Process 1000 is implemented to train a model to detect malicious files. In some embodiments, process 1000 is implemented by a server that provides the trained model to an edge device for inline sample classification.
[0148] At 1005, information about a historical malicious sample collection is obtained. As an example, the system obtains information from a third-party service (e.g., VirusTotal TM ) obtain information related to a historical malicious sample set. As another example, the system obtains information related to a historical malicious sample set based on manual labeling by a human operator.
[0149] At 1010, information about a historical benign sample set is obtained. As an example, the system obtains information from a third-party service (e.g., VirusTotalTM ) obtain information related to the historical benign sample set. As another example, the system obtains information related to the historical benign sample set based on manual labeling by a human operator.
[0150] At 1015, one or more relationships between the characteristic(s) of the sample and the maliciousness of the sample are determined. In some embodiments, the system determines characteristics related to whether the streamed file is malicious or the likelihood that the streamed file is malicious. The characteristics can be determined based on a malicious feature extraction process performed on the sample.
[0151] In some embodiments, features may be determined for a set of regular expression statements (eg, predefined regular expression statements) and / or for use of algorithm-based feature extraction (eg, TF-IDF, etc.).
[0152] In some embodiments, the system divides the corresponding sample into blocks. As an example, a block corresponds to a predefined number of bytes. In response to obtaining a block from the historical sample, the system performs feature extraction on the information included in the block of the sample.
[0153] At 1020, a model is trained to determine whether a file is malicious. The model is a machine learning model trained using a machine learning process. In some embodiments, a convolutional neural network is used to train the model. Various other machine learning processes may be implemented. Examples of machine learning processes that may be implemented in connection with training the model include random forests, linear regression, support vector machines, naive Bayes, logistic regression, K-nearest neighbors, decision trees, gradient boosted decision trees, K-means clustering, hierarchical clustering, density-based spatial clustering of applications with noise (DBSCAN), principal component analysis, and the like.
[0154] At 1025, the model is deployed. In some embodiments, deploying the model includes storing the model in a model dataset for use in conjunction with analyzing samples to classify the samples (e.g., determining whether the samples are malicious if the model is a detection model that detects malicious samples). In some embodiments, deploying the model includes storing the model in a model dataset for use in conjunction with analyzing samples to determine whether the samples are malicious or suspicious (e.g., if the model is a pre-filtering model that pre-filters network traffic based on detection of malicious or suspicious samples). Deploying the model may include providing the model (or a location where the model can be called) to an edge device, such as a security entity.
[0155] At 1030, a determination is made as to whether process 1000 is complete. In some embodiments, process 1000 is determined to be complete in response to a determination that no additional samples are to be analyzed (e.g., no additional predictions are needed for the samples), no additional traffic is to be analyzed, an administrator indicates that process 1000 is to be paused or stopped, etc. In response to determining that process 1000 is complete, process 1000 ends. In response to determining that process 1000 is not complete, process 1000 returns to 1005.
[0156] Figure 11 is a diagram of a set of chunks associated with a streaming file according to various embodiments. In some embodiments, the alignment and / or classification of the chunk sets is determined at least in part by Figure 1 system 100 and / or Figure 2 The classification of blocks can be implemented by an inline security entity.
[0157] In the example shown, the streaming file includes a set of blocks, such as block 1105, block 1110, block 1115, and block 1120. Blocks 1105-1120 do not necessarily need to include the same type of information. For example, the profile of the received file may be nonlinear. The file may include header information, and the header information may be included in the first block. However, in order for the model to deterministically classify the file, the blocks analyzed include the same type of information. Various embodiments address issues caused by nonlinear file profiles by aligning the block data within the blocks.
[0158] In some embodiments, aligning the block data includes treating a first subset of the block data of a particular block and a second subset of the block data of a different block as block data of a single block. For example, the system retrieves a first predefined number of bytes from the particular block and a second predefined number of bytes from a subsequent block, and collectively treats this information as block data of a single block. In response to aligning the block data, the system classifies the streaming file based on analyzing various blocks in the streaming file (or until the system determines that the streaming file corresponds to a particular classification).
[0159] like Figure 11As shown, block 1105 includes header information 1130 and payload data 1140-1. Header information 1130 may be a predefined number of bytes based on the file type. To align the block data (e.g., to consider blocks that only include payload data for analysis to generate a classification prediction), the system retrieves payload data 1140-1 from the first block and payload data 1140-2 from the second block. The system treats the block data of block 1105 as including payload data 1140-1 and payload data 1140-2. Retrieving payload data 1140-1 from the first block may include retrieving a predefined number of bytes from the back of the block. As an example, the system may retrieve the last X bytes from a particular block and the first Y bytes from a subsequent block. In response to retrieving the block data for a particular block (e.g., the block data for block 1105), the system queries the model based on the block data and obtains a prediction of whether the streaming file corresponds to a particular classification (e.g., whether the streaming file is malicious, etc.).
[0160] Similarly, the system considers (i) payload 1150-1 of the second block and payload 1150-2 of the third block as block data for block 1110 (e.g., for processing to generate a predicted single block), and considers (ii) payload 1160-1 of the third block and payload 1160-2 included in the fourth block as block data for block 1115. Payload 1170-1 can be combined with a subset of data included in a subsequent block.
[0161] In response to the aligned chunks, the system uses the corresponding chunk data to make a prediction as to whether the streaming file corresponds to a particular classification.
[0162] Figure 12 is a flow chart of a method for detecting malicious files according to various embodiments. In some embodiments, the alignment and / or classification of the block sets is at least partially determined by Figure 1 system 100 and / or Figure 2 The classification of blocks can be implemented by an inline security entity.
[0163] At 1205, a byte is input to the block alignment mechanism. In some embodiments, the input byte is a byte converted to an unsigned integer value between 0 and 255.
[0164] At 1210, the system (e.g., a block alignment mechanism) obtains bytes from the block data of the streaming file and processes the block set to align the block data. For example, the system assumes a variable length block size and splits the block data based on an offset from the first byte.
[0165] Referring to 1230, block alignment includes obtaining a subset of bytes from a first block and a subset of bytes from a second block. The byte subsets obtained from the first block and / or the second block may be predefined, such as in a block alignment policy or based on the type of file being processed. In the example shown, a block comprises 1500 bytes. In the example shown, 1000 bytes of the first block may correspond to header information, which may be discarded (or ignored for purposes of generating a prediction of file classification). For example, byte subset 1232 of the first block corresponds to payload data. To ensure that the same amount of payload data is analyzed to generate a prediction, the system obtains byte subsets from subsequent blocks. Therefore, the system treats byte subset 1232 of the first block and byte subset 1236 of the second block 1234 as block data for a single block to be analyzed to generate a prediction. As illustrated, because the byte subset obtained from the first block is 500 bytes and because the predefined block size is 1500 bytes, byte subset 1236 comprises 1000 bytes. The second subset of bytes 1238 of the second block 1234 is then used in conjunction with a subset of bytes from a subsequent block (eg, the third block).
[0166] If the expected block size is 1500 bytes and the first block is also 1500 bytes (payload data), then the offset is 0 and all packets will be aligned, and no alignment of the block data is required. Conversely, if the expected block size is 1500 bytes and the first block includes an offset (e.g., header information that results in the offset), then that offset is used to segment subsequent blocks to obtain information from the alignment block. The offset can be predefined based on the file type, or can be determined by the system when the streaming file is received. In the example shown, the first block is 500 bytes (payload data), and therefore the offset is 1000 bytes, so the second block 1234 is segmented into a first byte subset 1236 and a second byte subset 1238. The first byte subset 1236 and the second byte subset 1238 can be used in different blocks for block alignment and prediction generation.
[0167] Returning to 1210, the system (e.g., a block alignment mechanism, such as the block alignment module 229 of the system 200) obtains the extraction results of a set of bytes of size k-1 1215, where k is the maximum kernel size used in the model (e.g., a convolutional neural network). The extraction results of the set of k-1 bytes can be, for example, the previous k-1 bytes from the previous block. The system (e.g., the block alignment mechanism) obtains the maximum activation value of each block 1220. For example, the system keeps track of the maximum activation of the aligned blocks (e.g., every 1500 bytes in the example above).
[0168] According to various embodiments, a system (e.g., a block alignment mechanism) implements a predefined algorithm to align block data. The algorithm includes:
[0169] Initialize alignment block size = m;
[0170] Calculate block size = n;
[0171] Get offset k=mn;
[0172] If m=0,
[0173] Pass the block to the model and obtain the maximum activation and prediction for document classification; and
[0174] otherwise,
[0175] Split the block into blocks of size k and size nk;
[0176] Pass the previous maximum activation along with k bytes to obtain the maximum activation and prediction for document classification; and
[0177] Pass the remaining nk bytes to update the maximum activation.
[0178] Figure 13 is a block diagram of a system for classifying streaming files based on block data according to various embodiments. In some embodiments, the system 1300 may be composed at least in part of Figure 1 system 100 and / or Figure 2 The classification of blocks can be implemented by an inline security entity.
[0179] System 1300 classifies files based on performing feature extraction on the block data and classification. Feature extraction includes processing the block data using a convolutional neural network 1305 and processing the output from the convolutional neural network 1305 using a global max layer 1310. In the example shown, the convolutional neural network 1305 is a one-dimensional convolutional neural network with a kernel size of 12. However, various other kernel sizes can be implemented. In some embodiments, the kernel size is 8. In response to passing the block data through the convolutional neural network 1305, the system provides the output from the convolutional neural network 1305 to the global max layer 1310 to obtain the maximum activation (e.g., (one or more) maximum activation values). For example, the global max layer 1310 performs a max pooling operation on the output from the convolutional neural network 1305.
[0180] In response to performing feature extraction, the system performs classification on the feature vector / feature map to obtain a prediction of whether the streaming file corresponds to a specific classification. Classifying the streaming file based on a specific data block of the streaming file includes: using a dense layer and a softmax module 1315 to process the output from the global maximum layer 1310.
[0181] Figure 14 is a block diagram of a classification of a set of blocks obtained from streaming data of a file according to various embodiments. In some embodiments, the classification of the set of blocks is at least partially determined by Figure 1 system 100 and / or Figure 2 The classification of blocks can be implemented by an inline security entity.
[0182] System 1400 performs classification of a streaming file based on block data of a particular block. System 1400 obtains a streaming file based on receiving consecutive blocks. A first block includes header information 1405 and payload information 1410-1; a second block includes payload information 1410-2 (e.g., a first byte subset of the second block) and payload information 1415-1 (e.g., a second byte subset of the second block); a third block includes payload information 1415-2 (e.g., a first byte subset of the third block) and payload information 1420-1 (e.g., a second byte subset of the third block); and a fourth block includes payload information 1420-2 (e.g., a first byte subset of the fourth block) and payload information 1425-1 (e.g., a second byte subset of the fourth block). Because the first block includes header information 1405, the payload data of the blocks of the streaming file have an offset equal to the number of bytes of header information 1405.
[0183] The system 1400 applies the convolutional neural network model to successive blocks in succession. For example, at 1440, the system 1400 passes the payload information 1410-1 of the first block through the convolutional neural network; at 1442, the system 1400 passes the payload information 14010-2 and 1415-1 through the convolutional neural network; at 1444, the system 1400 passes the payload information 1415-2 and 1420-1 through the convolutional neural network; and at 1446, the system 1400 passes the payload information 1420-2 and 1425-1 through the convolutional neural network.
[0184] In response to passing the block (e.g., payload information) through the convolutional neural network, system 1400 passes the output from the convolutional neural network through a pooling layer. For example, at 1450, system 1400 passes the output from 1440 through a global max layer; at 1452, system 1400 passes the output from 1442 through a global max layer; at 1454, system 1400 passes the output from 1444 through a global max layer; and at 1456, system 1400 passes the output from 1446 through a global max layer. As an example, the global max layer performs a max pooling operation on the output from the convolutional layer.
[0185] At 1460, for the first virtual block (e.g., a block aligned with the payload information from the first received block and the second received block), the system 1400 passes the output from the global maximum layer (e.g., the maximum activation value) for payload information 1410-1 and payload information 1410-2 through a dense layer and a softmax operation. Similarly, at 1465, for the second virtual block, the system 1400 passes the output from the global maximum layer (e.g., the maximum activation value) for payload 1415-1 and payload 1452-2 through a dense layer and a softmax operation.
[0186] At 1470, the system 1400 uses the output from the dense layer and the softmax operation for the first virtual block to determine a classification of the streaming file based on the first virtual block. For example, the system 1400 generates a prediction of the likelihood that the streaming file corresponds to a particular classification. Similarly, at 1475, the system 1400 uses the output from the dense layer and the softmax operation for the second virtual block to determine a classification of the streaming file based on the second virtual block.
[0187] The system 1400 uses successive block data (e.g., successive aligned blocks or virtual blocks) to successively perform classification and generate predictions of whether the streaming file corresponds to a particular classification. The system 1400 can perform successive classifications and prediction generation until the earlier of: (i) all blocks in the streaming file have been processed for classification, and (ii) the system 1400 determines that the streaming file corresponds to a particular classification based on the prediction (e.g., the system 1400 determines that the prediction exceeds a predefined classification threshold).
[0188] Figure 15 is a block diagram of a classification of a set of blocks obtained from streaming data of a file according to various embodiments. In some embodiments, the classification of the set of blocks is at least partially determined by Figure 1 system 100 and / or Figure 2 The classification of blocks can be implemented by an inline security entity.
[0189] The system 1500 classifies streaming files based on analysis of successive blocks. For example, when the system 1500 receives a block of a streaming file, the system 1500 processes the corresponding block data and generates a prediction of whether the streaming file is a malicious succession.
[0190] In the example shown, the first block includes header information 1505 (or other offset) and payload information 1510-1; the second block includes payload information 1510-2 and 1515-1; the third block includes payload information 1515-2 and 1520-1; and the fourth block includes payload information 1520-2 and 1525-1.
[0191] System 1500 performs feature extraction 1530-1 on payload information 1510-1 obtained from the first block, and performs feature extraction 1530-2 on payload information 1510-2 obtained from the second block. In response to performing feature extraction 1530-1 and 1530-2, system 1500 performs classification 1535 of the streaming file (e.g., a prediction based on analysis of the first virtual block or payload information 1510-1 and 1510-2). For example, system 1500 generates a prediction of whether the streaming file corresponds to a particular classification.
[0192] System 1500 performs feature extraction 1540-1 on payload information 1515-1 obtained from the first block, and performs feature extraction 1540-2 on payload information 1515-2 obtained from the second block. In response to performing feature extraction 1540-1 and 1540-2, system 1500 performs classification 1545 of the streaming file (e.g., a prediction based on analysis of the first virtual block or payload information 1515-1 and 1515-2). For example, system 1500 generates a prediction of whether the streaming file corresponds to a particular classification.
[0193] System 1500 performs feature extraction 1550-1 on payload information 1520-1 obtained from the first block, and performs feature extraction 1550-2 on payload information 1520-2 obtained from the second block. In response to performing feature extraction 1550-1 and 1550-2, system 1500 performs classification 1555 of the streaming file (e.g., a prediction based on analysis of the first virtual block or payload information 1520-1 and 1520-2). For example, system 1500 generates a prediction of whether the streaming file corresponds to a particular classification.
[0194] In some embodiments, if system 1500 does not consider the streaming file to correspond to a particular classification based on classification 1535, system 1500 only performs classification 1545. Similarly, if system 1500 does not consider the streaming file to correspond to a particular classification based on classification 1545, system 1500 may only perform classification 1555.
[0195] Figure 16 is a flow chart of a method for classifying stream data of a file according to various embodiments. In some embodiments, process 1600 is at least partially performed by Figure 1 system 100 and / or Figure 2 The process 1600 may be implemented by an inline security entity.
[0196] At 1605 , streaming data of a file is obtained. Obtaining streaming data of a file (eg, a streaming file) includes successively receiving blocks of the streaming file. For example, the streaming file is obtained by an edge device.
[0197] In some embodiments, in response to receiving a first data chunk of a streaming file, the system determines whether chunk alignment is to be performed. For example, the system determines whether the first chunk includes an offset (e.g., header information). The system can determine whether the first chunk includes an offset based on analysis of the chunk, based on the file type, etc.
[0198] At 1610, the system aligns a predetermined amount of data in blocks associated with streaming data of the file. In some embodiments, in response to determining an offset associated with the streaming file (e.g., the extent to which payload information is offset in the blocks), the system performs block alignment to account for the offset. For example, the system determines virtual blocks (also referred to herein as alignment blocks) and sequentially performs sorting of the streaming file based on successive virtual blocks.
[0199] In some embodiments, a virtual block includes a subset of bytes of a particular block and a subset of bytes of another block (such as a subsequent block). The subset of bytes of the particular block may be a predefined number of bytes at the end of the first block, and the subset of bytes in the subsequent block may be a predefined number of bytes at the beginning of the subsequent block.
[0200] At 1615, the plurality of aligned blocks are processed using the machine learning model. In some embodiments, the system performs feature extraction on the aligned blocks, and in response to performing the feature extraction, performs classification for the aligned blocks.
[0201] At 1620, the file is classified. In some embodiments, the file is classified on a block-by-block basis. For example, the system classifies the file based on the processing of a particular block by a machine learning model. The system can sequentially classify the file as the blocks are processed. In some embodiments, the system stops processing blocks of the streaming file when the system determines that the predicted classification (based on the particular block) exceeds a predefined classification threshold.
[0202] At 1625, a determination is made as to whether process 1600 is complete. In some embodiments, process 1600 is determined to be complete in response to a determination that no additional samples are to be analyzed (e.g., no additional predictions are needed for the samples), no additional traffic is to be analyzed, an administrator indicates that process 1600 is to be paused or stopped, etc. In response to determining that process 1600 is complete, process 1600 ends. In response to determining that process 1600 is not complete, process 1600 returns to 1605.
[0203] In response to classifying the streaming file with each successive block, the system can determine the manner in which the streaming file (eg, the particular block(s)) is to be handled, such as whether a policy is to be enforced for the streaming file.
[0204] Figure 17 is a flow chart of a method for detecting malicious files according to various embodiments. In some embodiments, process 1700 is performed at least in part by Figure 1 system 100 and / or Figure 2 The process 1700 may be implemented by an inline security entity.
[0205] Process 1700 is implemented to determine whether a streaming file is malicious. For example, process 1700 analyzes each successive block (at least until the file is deemed malicious) and classifies the file based on the analysis of the block(s).
[0206] At 1705 , stream data of the file is obtained. In some embodiments, 1705 corresponds to or is similar to 1605 of process 1600 .
[0207] At 1710 , the system aligns a predetermined amount of data in a block associated with the streaming data of the file. In some embodiments, 1710 corresponds to or is similar to 1610 of process 1600 .
[0208] At 1715 , the plurality of aligned blocks are processed using a machine learning model. In some embodiments, 1715 corresponds to or is similar to 1615 of process 1600 .
[0209] At 1720, the files are classified. In some embodiments, the files are classified on a block-by-block basis. In some embodiments, 1720 corresponds to or is similar to 1620 of process 1600.
[0210] At 1725, the system determines whether the file is malicious. Based on a comparison of whether the prediction obtained from the model exceeds a predefined classification threshold, the system determines whether the output of the comparison indicates that the file is malicious.
[0211] In response to determining that the file is malicious at 1725, process 1700 proceeds to 1730 where one or more security policies are applied to the file. The system can handle malicious traffic / information based at least in part on one or more policies, such as one or more security policies.
[0212] According to various embodiments, the handling of malicious sample traffic / information may include performing proactive measures. The proactive measures may be performed in accordance with (e.g., based at least in part on) one or more security policies. As an example, one or more security policies may be preset by a network administrator, a client (e.g., an organization / company) to a service (which provides detection of malicious input strings or files, etc.). Examples of proactive measures that may be performed include: isolating a sample (e.g., isolating the sample), deleting the sample (e.g., deleting one or more blocks of block data), alerting a user that a malicious sample has been detected, providing a prompt to a user when a device attempts to open or execute the sample, blocking the transmission of the sample, updating a blacklist of malicious input strings (e.g., a mapping of a hash of the sample to information indicating that the sample is malicious), etc.
[0213] In response to determining at 1725 that the traffic does not include a malicious sample, process 1700 proceeds to 1735 where the sample (e.g., streamed file) is handled as a non-malicious sample (e.g., non-malicious traffic / information) at 1735. For example, the system can handle the non-malicious sample according to normal operation (e.g., allowed transfer / communication of files, etc.).
[0214] At 1740, a determination is made as to whether process 1700 is complete. In some embodiments, process 1700 is determined to be complete in response to a determination that no additional samples are to be analyzed (e.g., no additional predictions are needed for the samples), no additional traffic is to be analyzed, an administrator indicates that process 1700 is to be paused or stopped, etc. In response to determining that process 1700 is complete, process 1700 ends. In response to determining that process 1700 is not complete, process 1700 returns to 1705.
[0215] Figure 18 is a flow chart of a method for detecting malicious files according to various embodiments. In some embodiments, process 1800 is performed at least in part by Figure 1 system 100 and / or Figure 2 The process 1800 may be implemented by an inline security entity.
[0216] Process 1800 is implemented to determine whether a streaming file is malicious. For example, process 1800 analyzes each successive block (at least until the file is deemed malicious) and classifies the file based on the analysis of the block(s).
[0217] Process 1800 illustrates an example in which the system determines how to handle a streaming file with each subsequent analysis / classification of chunks in the streaming file.
[0218] At 1805 , stream data of the file is obtained. In some embodiments, 1805 corresponds to or is similar to 1605 of process 1600 .
[0219] At 1810, n is set equal to 1. n is a positive integer used as a counter during processing of a block of streaming data for a file.
[0220] At 1815 , a predefined subset of the nth block is obtained.
[0221] At 1820 , a predefined subset of the (n+1)th block is obtained.
[0222] At 1825 , feature extraction is performed on a predefined subset of the nth block and a predefined subset of the (n+1)th block.
[0223] At 1830 , the model is queried based on the feature extraction.
[0224] At 1835, a prediction is obtained from the model. The model may provide a predicted classification, or the likelihood that the file corresponds to a particular classification.
[0225] At 1840 , the system determines whether the prediction is greater than a maliciousness threshold.
[0226] In response to determining at 1840 that the prediction is greater than the maliciousness threshold, process 1800 proceeds to 1855 where one or more security policies are applied to the file (or at least any future received / processed blocks of the file) at 1855. As an example, if the prediction is greater than the maliciousness threshold, the system deems the file malicious.
[0227] In contrast, in response to determining at 1840 that the prediction is not greater than the maliciousness threshold, process 1800 proceeds to 1845 where the system determines whether the file is complete. The system can determine whether the file is complete based on determining whether the most recently processed block (e.g., the (n+1)th block) is the last block of the streaming file.
[0228] In response to determining at 1845 that the file is intact, process 1800 proceeds to 1860. Conversely, in response to determining at 1845 that the file is not intact, process 1800 proceeds to 1850, where n is incremented (e.g., n=n+1). Process 1800 then returns to 1815, and process 1800 iterates through 1815-1840 until file processing is complete (e.g., a classification of the file is predicted for each block), or the system classifies the file as malicious based on analyzing the block data for a particular block.
[0229] At 1860, a determination is made as to whether process 1800 is complete. In some embodiments, process 1800 is determined to be complete in response to a determination that no additional samples are to be analyzed (e.g., no additional predictions are needed for the samples), no additional traffic is to be analyzed, an administrator indicates that process 1800 is to be paused or stopped, etc. In response to determining that process 1800 is complete, process 1800 ends. In response to determining that process 1800 is not complete, process 1800 returns to 1805.
[0230] Figure 19 is a flow chart of a method for detecting malicious files according to various embodiments. In some embodiments, process 1900 is at least partially performed by Figure 1 system 100 and / or Figure 2 The process 1900 may be implemented by an inline security entity.
[0231] Process 1900 is implemented to determine whether a streamed file is malicious. For example, process 1900 analyzes each successive block (at least until the file is deemed malicious) and classifies the file based on the analysis of the block(s). Although process 1900 illustrates an example for classifying whether a file is malicious, various other embodiments may be implemented to determine other classifications for streamed files.
[0232] Process 1900 illustrates an example in which the system determines how to handle a streaming file with each subsequent analysis / classification of chunks in the streaming file.
[0233] At 1905 , stream data of the file is obtained. In some embodiments, 1905 corresponds to or is similar to 1805 of process 1800 .
[0234] At 1910 , n is set equal to 1. In some embodiments, 1910 corresponds to or is similar to 1805 of process 1800 .
[0235] At 1915, the last X bytes of the nth block are obtained. X is a positive integer. In some embodiments, X is predefined. In some embodiments, X is determined based on the number of bytes included in the header information of the streaming file.
[0236] At 1920, the first Y bytes of the nth block are obtained. Y is a positive integer. In some embodiments, Y is predefined. In some embodiments, Y is determined based on the number of bytes included in the header information of the streaming file.
[0237] At 1925 , feature extraction is performed on the last X bytes of the nth block and the first Y bytes of the (n+1)th block.
[0238] At 1930 , the model is queried based on the feature extraction. In some embodiments, 1930 corresponds to or is similar to 1830 of process 1800 .
[0239] At 1935 , a prediction is obtained from the model. In some embodiments, 1935 corresponds to or is similar to 1835 of process 1800 .
[0240] At 1940 , the system determines whether the prediction is greater than a maliciousness threshold. In some embodiments, 1940 corresponds to or is similar to 1840 of process 1800 .
[0241] In response to determining at 1940 that the prediction is greater than the maliciousness threshold, process 1900 proceeds to 1955 where one or more security policies are applied to the file (or at least any future received / processed blocks of the file) at 1955. As an example, if the prediction is greater than the maliciousness threshold, the system deems the file malicious.
[0242] In contrast, in response to determining at 1940 that the prediction is not greater than the maliciousness threshold, process 1900 proceeds to 1945 where the system determines whether the file is complete. The system can determine whether the file is complete based on determining whether the most recently processed block (e.g., the (n+1)th block) is the last block of the streaming file.
[0243] In response to determining at 1945 that the file is intact, process 1900 proceeds to 1960. Conversely, in response to determining at 1945 that the file is not intact, process 1900 proceeds to 1950, where n is incremented (e.g., n=n+1). Process 1900 then returns to 1915, and process 1900 iterates through 1915-1940 until file processing is complete (e.g., a classification of the file is predicted for each block), or the system classifies the file as malicious based on analyzing the block data for a particular block.
[0244] At 1960, a determination is made as to whether process 1900 is complete. In some embodiments, process 1900 is determined to be complete in response to a determination that no additional samples are to be analyzed (e.g., no additional predictions are needed for the samples), no additional traffic is to be analyzed, an administrator indicates that process 1900 is to be paused or stopped, etc. In response to determining that process 1900 is complete, process 1900 ends. In response to determining that process 1900 is not complete, process 1900 returns to 1905.
[0245] Various examples of the embodiments described herein are described in conjunction with flow charts. Although the examples may include certain steps performed in a specific order, according to various embodiments, the various steps may be performed in various orders and / or the various steps may be combined into a single step or performed in parallel.
[0246] Although the above embodiments have been described in detail for the purpose of clarity of understanding, the present invention is not limited to the details provided. There are many alternative ways to implement the present invention. The disclosed embodiments are illustrative rather than restrictive.
Claims
1. A system for performing classification at an edge device, comprising: One or more processors configured to: Obtaining streaming data of a file at the edge device; processing a set of chunks associated with the stream data of the file using a machine learning model; as well as classifying the file at the edge device before processing the entire contents of the file; as well as A memory is coupled to the one or more processors and configured to provide instructions to the one or more processors. The system of claim 1 , wherein the edge device is a network device. The system of claim 1 , wherein the edge device is an inline security entity.
4. The system of claim 1 , wherein the machine learning model is configured to classify whether the file is malicious.
5. The system of claim 1 , wherein the machine learning model is configured to classify whether the file is copyrighted material.
6. The system of claim 1, wherein the machine learning model is configured to classify whether the file is health data or financial data.
7. The system of claim 1 , wherein the file is determined to be malicious if the prediction obtained from the machine learning model exceeds a predefined maliciousness threshold.
8. The system of claim 7, wherein the file is determined to be malicious after processing an nth block using the machine learning model, n corresponding to a positive integer less than a total number of blocks in the file.
9. The system of claim 7, wherein the predefined maliciousness threshold is constant for each block in the file.
10. The system of claim 7, wherein the predefined maliciousness threshold is dynamic across classifications of blocks in the file. 11 . The system of claim 10 , wherein the predefined malicious threshold of the first block is lower than the predefined malicious threshold of the j-th block, and j is a positive integer greater than 1.
12. The system of claim 7, wherein in response to determining that the file is malicious, proactive measures are implemented against the malicious file.
13. The system of claim 12, wherein the proactive measures include discarding or blocking remaining blocks associated with the file.
14. The system of claim 1, wherein each block corresponds to m bytes, and m is a positive integer.
15. The system of claim 1, wherein the machine learning model is trained using a deep learning process.
16. The system of claim 15, wherein the deep learning process comprises a convolutional neural network.
17. The system of claim 15, wherein the machine learning model is trained at least in part based on a recurrent neural network, and a max pooling operation is performed to maintain state information across at least a subset of the blocks associated with the file.
18. The system of claim 1, wherein the model is trained on the entire document.
19. A method for performing classification at an edge device, comprising: Obtaining, by one or more processors, streaming data of a file at the edge device; processing a set of chunks associated with the stream data of the file using a machine learning model; as well as The file is classified at the edge device before processing the entire contents of the file.
20. A computer program product embodied in a non-transitory computer-readable medium, the computer program product for performing classification at an edge device, the computer program product comprising computer instructions for: Obtaining, by one or more processors, streaming data of a file at the edge device; processing a set of chunks associated with the stream data of the file using a machine learning model; as well as The file is classified at the edge device before processing the entire contents of the file.
21. A system for performing classification at an edge device, comprising: One or more processors configured to: Obtaining streaming data of a file at the edge device; aligning a predetermined amount of data in a block associated with the stream data of the file; processing a plurality of aligned chunks associated with the stream data of the file using a machine learning model; as well as classifying the file at the edge device based at least in part on the classification of the plurality of aligned blocks; as well as A memory is coupled to the one or more processors and configured to provide instructions to the one or more processors.
22. The system of claim 21, wherein the edge device is a network device.
23. The system of claim 21, wherein the edge device is an inline security entity.
24. The system of claim 21, wherein the machine learning model is configured to classify whether the file is malicious.
25. The system of claim 21, wherein the file is classified using the machine learning model before processing the entire contents of the file.
26. The system of claim 21, wherein the first block comprises overhead associated with the file.
27. The system of claim 21, wherein: Aligning the predetermined amount of data in blocks associated with the streaming data of the file comprises: determining a first file segment based at least in part on associating a predetermined amount of the first blocks with a predetermined amount of the second blocks; Processing the plurality of aligned blocks includes: querying the machine learning model based on the first file segment; and The document is classified based at least in part on the classification of the first document segment.
28. The system of claim 21, wherein: Aligning the predetermined amount of data in blocks associated with the streaming data of the file comprises: determining an nth file segment based at least in part on associating a predetermined amount of an i-th block with a predetermined amount of a j-th block; i and j are positive integers, and j is greater than i; Processing the plurality of aligned blocks includes: querying the machine learning model based on the nth file segment; and The file is classified based at least in part on the classification of the nth file segment.
29. The system of claim 28, wherein the nth file segment comprises a predetermined number of bytes.
30. The system of claim 28, wherein the nth file segment comprises 1500 bytes.
31. The system of claim 28, wherein the one or more processors are configured to select the predetermined number of bytes from a set of preset numbers of bytes, the predetermined number of bytes being selected based on a packet size of the file.
32. The system of claim 21, wherein the predetermined amount of data in a block is aligned to file overhead in a first block, and ensuring sorting includes processing the same number of bytes in the file for each aligned block.
33. The system of claim 21 , wherein classifying the file using the machine learning model based on a predefined number of bytes is deterministic.
34. The system of claim 21, wherein the file is determined to be malicious if the prediction obtained from the machine learning model exceeds a predefined maliciousness threshold.
35. The system of claim 34, wherein the file is determined to be malicious after processing an nth file segment using the machine learning model, n corresponding to a positive integer less than a total number of blocks in the file.
36. The system of claim 34, wherein the predefined maliciousness threshold is constant for each file segment in the file.
37. The system of claim 34, wherein the predefined maliciousness threshold is dynamic across classifications of file segments in the file.
38. The system of claim 34, wherein in response to determining that the file is malicious, proactive measures are implemented against the malicious file.
39. The system of claim 21, wherein the machine learning model is trained using a deep learning process.
40. The system of claim 39, wherein the deep learning process comprises a convolutional neural network.
41. The system of claim 40, wherein the convolutional neural network uses a kernel size less than 12.
42. The system of claim 40, wherein the convolutional neural network uses a kernel of size 8.
43. The system of claim 21 , wherein the number of characters for which state information is stored for evaluation of aligned blocks is based at least in part on a kernel size of a convolutional neural network.
44. The system of claim 21, wherein the machine learning model is an XGBoost model.
45. A method for performing classification at an edge device, comprising: Obtaining, by one or more processors, streaming data of a file at the edge device; aligning a predetermined amount of data in a block associated with the stream data of the file; processing a plurality of aligned chunks associated with the stream data of the file using a machine learning model; as well as The file is classified at the edge device based at least in part on the classification of the plurality of aligned blocks.
46. A computer program product embodied in a non-transitory computer-readable medium, the computer program product for performing classification at an edge device, the computer program product comprising computer instructions for: Obtaining, by one or more processors, streaming data of a file at the edge device; aligning a predetermined amount of data in a block associated with the stream data of the file; processing a plurality of aligned chunks associated with the stream data of the file using a machine learning model; as well as The file is classified at the edge device based at least in part on the classification of the plurality of aligned blocks.