Detect unknown malicious content in a computer system
By applying a feature-weighted deep learning model in malicious content detection, the problems of low prediction accuracy and high error rate in the prior art are solved, and more accurate and efficient malicious content detection is achieved.
Patent Information
- Application Number
- CN202080078166.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-17
- Filing Date
- 2020-10-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-10-21
AI Technical Summary
The prior art has low prediction accuracy and high error rate when detecting malicious content. Especially when facing unknown or new variants, it is difficult to effectively identify and distinguish malicious content.
Deep learning models such as conjoined neural networks (SNNs) or deep structured semantic models (DSSMs) are used to detect malicious content through feature-weighting methods. This method learns the importance of different features and gives them different weights to improve the accuracy of similarity scores.
Improve the prediction accuracy of malicious content detection, reduce the error rate, and effectively identify and distinguish malicious content, including unknown or new variants.
Smart Images

Figure CN114730339B_ABST
Abstract
Description
Background Art
[0001] Computer systems can be infected with malicious content (e.g., malware), which can cause damage or allow cyber attackers to access these computer systems without authorization. There are various known families and sub - families of malicious content, such as viruses, Trojans, worms, ransomware, etc. Detecting malicious content remains a significant challenge for the prior art, especially when there are unknown or new variants. Cyber attackers continuously change and evolve malicious content over time to evade detection. The amount of this change varies by family, making it difficult to detect the presence of malicious behavior. Summary of the Invention
[0002] This Summary of the Invention is provided to introduce some concepts in a simplified form that will be further described in the detailed description below. This Summary of the Invention is not intended to identify the main features or essential features of the claimed subject matter, nor is it intended to be used in isolation to assist in determining the scope of the claimed subject matter.
[0003] Various embodiments discussed herein enable the detection of malicious content. Some embodiments do this by determining a similarity score between known malicious content or an indication representing malicious content (e.g., a vector, file hash, file signature, code, etc.) and other content (e.g., an unknown file) or an indication representing other content, based on feature weighting. During various training phases, certain feature characteristics can be learned for each labeled content or indication. For example, for a first malware family, the most prominent feature may be a specific URL, while for different iterations of the first family, other features change significantly (e.g., due to modifications by cyber attackers). Thus, the specific URL can be weighted to determine a specific output classification. In this way, the embodiments learn the weights corresponding to different features such that important features found in similar malicious content and from the same family contribute positively to the similarity score, while features that distinguish malicious content from benign (non - malicious) content contribute negatively to the similarity score. Thus, even if a cyber attacker introduces unknown or new variants of malicious content, the malicious content can be detected. Additionally, this allows some embodiments to determine which family the malicious content belongs to.
[0004] In some embodiments, a unique deep - learning model, such as a variant of a Siamese Neural Network (SNN) or a variant of a Deep Structured Semantic Model (DSSM), can be used to detect unknown malicious content. Certain embodiments train a model that learns to give different weights to features based on their importance. In this way, deep - learning model embodiments are useful for taking unknown content or indications and mapping them into a feature space to determine the distance or similarity to known malicious files or indications based on the specific features of the unknown file and the training weights associated with the features of known files or indications.
[0005] The prior art has various drawbacks, resulting in low prediction accuracy, high error rates, etc. For example, existing tools use the Jaccard Index to implement similarity scores between files. However, the Jaccard Index and other techniques require that all features in a file have the same weight. Various embodiments of the present disclosure improve these prior arts by increasing prediction accuracy and reducing error rates, as described herein, for example, with respect to experimental results. The embodiments also improve these techniques because they learn certain key features that are most important for detecting whether the content is malicious or belongs to a specific malicious code or file family and weight them accordingly. Some embodiments also improve the functionality of the computer itself by reducing the consumption of computing resources such as memory, CPU, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The present invention will be described in detail below with reference to the accompanying drawings, in which:
[0007] Figure 1 is a block diagram of an example system according to some embodiments;
[0008] Figure 2 is a block diagram of an example computing system architecture according to some embodiments;
[0009] Figure 3 is a block diagram of a system for training a machine learning model on various malware content and predicting whether one or more specific unknown content sets contain malware according to some embodiments;
[0010] Figure 4 is a block diagram of an example system for using a trained model to determine whether new content is malicious according to some embodiments;
[0011] Figure 5 is a schematic diagram of an example deep learning neural network (DNN) used in a specific embodiment;
[0012] Figure 6 is a schematic diagram of an example deep learning neural network (DNN) used in a specific embodiment;
[0013] Figure 7A is a flowchart of an example process for training a machine learning model according to some embodiments;
[0014] Figure 7B is a flowchart of an example process for evaluating new or unknown content according to some embodiments;
[0015] Figure 8 is a block diagram of a computing device according to some embodiments;
[0016] Figure 9 An example table illustrating the paired count breakdown for each family according to some embodiments;
[0017] Figure 10 A chart showing the Jaccard index similarity score distribution for similar and dissimilar files according to some embodiments;
[0018] Figure 11 A chart showing the SNN similarity score distribution for similar and dissimilar files according to some embodiments;
[0019] Figure 12 An example table illustrating the performance measurements of KNN for different highly prevalent malware families; and
[0020] Figure 13 An example visualization chart illustrating the separability of the latent vectors of malware classes using the t-sne method according to some embodiments. Detailed Description
[0021] The subject matter of aspects of the present disclosure is described herein in detail to meet statutory requirements. However, the description itself is not intended to limit the scope of this patent. Instead, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, to include different steps or combinations of steps similar to those described herein, and in conjunction with other existing or future technologies. Additionally, although the terms "step" and / or "block" may be used herein to denote different elements of the methods employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless and except when the order of individual steps is explicitly described. Each method described herein may include a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in a memory. The method may also be embodied as computer-usable instructions stored on a computer storage medium. These methods may be provided by a stand-alone application, service, or hosted service (stand-alone or in combination with another hosted service) or a plug-in of another product, to name a few.
[0022] As used herein, the term "set" may be used to refer to an ordered (i.e., sequential) or unordered (i.e., non-sequential) collection of objects (or elements), such as, but not limited to, data elements (e.g., events, event clusters, etc.). A set may include N elements, where N is any non-negative integer that is 1 or greater. That is, a set may include 1, 2, 3, … N objects and / or elements, where N is a positive integer with no upper limit. A set may contain only one element. In other embodiments, a set may include a plurality of elements significantly greater than one, two, or three elements. As used herein, the term "subset" is a set that is included in another set. A subset may or may not be a proper or strict subset of the other set that contains the subset. That is, if set B is a subset of set A, then in some embodiments, set B is a proper or strict subset of set A. In other embodiments, set B is a subset of set A but not a proper or strict subset of set A.
[0023] The various embodiments described herein are capable of detecting malicious content or malicious computer objects. As described herein, "content" or "computer object" is any suitable unit of information, such as a file, a code / instruction set, one or more messages, one or more database records, and / or one or more data structures, or certain behaviors or functions that the content or computer object performs or is associated with. "Malicious" content or a malicious computer object may refer to malicious code / instructions, malicious files, malicious behaviors (e.g., malicious code known to inject code at a specific timestamp or known to be inactive prior to its launch activity), messages, database records, data structures, and / or any other suitable function that causes damage, harm, has an adverse effect, and / or results in unauthorized access to a computing system. Although various examples are described herein in terms of files, it should be understood that this is merely representative, and any computer object or content may be used in place of a file. It should be understood that the terms "content" and "computer object" may be used interchangeably when described herein. Some embodiments perform detection by determining a similarity score (i.e., similarity metric) between known malicious content (or malicious indicators) and unknown content (or unknown indicators) based on feature weighting. As used herein, an "indicator" refers to any identifier or data set that represents content. For example, an indicator may be a vector, file hash, file signature, or code that represents malicious content.
[0024] As used herein, "features" represent specific attributes or attribute values of content. For example, a first feature can be the length and format of a file, a second feature can be a specific URL of the file, a fourth feature can be an operational characteristic, such as being written in short blocks, and a fifth feature can be a registry key pattern. In various cases, "weights" represent the importance or significance of a feature or feature value for classification or prediction. For example, each feature can be associated with an integer or other real number, where the higher the real number, the more important the feature is for prediction or classification. In some embodiments, the weights in a neural network or other machine learning application can represent the connection strength between nodes or neurons from one layer (input) to the next layer (output). A weight of 0 can mean that the input does not change the output, while a weight above 0 changes the output. The higher the value of the input or the closer the value is to 1, the greater the change or increase in the output. Similarly, there can be negative weights. Negative weights proportionally decrease the output value. For example, the more the input value increases, the more the output value decreases. Negative weights can result in negative scores, which will be described in more detail below. In many cases, only a selected set of features is primarily responsible for determining whether the content belongs to a particular malicious family and is thus malicious.
[0025] Various embodiments learn the key features of content and weight them accordingly during training. For example, some embodiments learn embedding vectors based on deep learning to detect similar computer objects or indicators in a feature space using a distance metric such as cosine distance. In these embodiments, each computer object is converted from a string or other form into a vector (e.g., a set of real numbers), where each value or set of values represents a separate feature of the computer object or an indicator in the feature space. The feature space (or vector space) is a collection of vectors (e.g., each representing a malicious or benign file), and each vector is oriented or embedded in the space based on the similarity of vector features. At different training stages, certain feature characteristics of each labeled computer object or indicator can be learned. For example, for a first malware family (e.g., a certain class or type of malware), the most prominent feature can be a specific URL, while other features vary widely for different iterations of the first malware family. Thus, the specific URL can be weighted to determine a specific output classification. In this way, embodiments learn the weights corresponding to different features such that important features found in similar malicious content and from the same family contribute positively to the similarity score, while features that distinguish malicious content from benign content (non-malicious) contribute negatively to the similarity score.
[0026] In some embodiments, a unique deep learning model, such as a variant of a Siamese neural network (SNN) or a variant of a deep structured semantic model (DSSM), can be used to detect unknown malicious content. Embodiments train a model that learns to assign different weights to features based on the importance of the features. Some deep learning model embodiments include two or more identical sub-networks or branches, meaning that the sub-networks have the same configuration, which has the same or bound parameters and weights. Each sub-network receives distinct inputs (e.g., two different files), but is connected by an energy function at the top of the two identical sub-networks, which determines the degree of similarity between the two inputs. Weight binding ensures or increases the likelihood that two extremely similar sets of malicious content or indicators cannot be mapped by their respective identical networks to very dissimilar locations in the feature space, because each network computes the same function. In this way, deep learning model embodiments are useful for obtaining unknown content or indicators and mapping them into the feature space to determine the distance or similarity to a set of known malicious content or indicators based on the specific features of the unknown file and the training weights associated with the features of known files or indicators.
[0027] The prior art has various functional deficiencies, resulting in reduced prediction accuracy and increased error rates, etc. A key component of some malware detection techniques is to determine similar content in a high-dimensional input space. For example, instance-based malware classifiers (such as the K-Nearest Neighbor (KNN) classifier) rely on similarity scores or distances between two files. In some cases with an infinite amount of training data, the K-Nearest Neighbor classifier can be the best classifier. Malware clustering that identifies groups of malicious content may also rely on calculating similarity scores between sets of content. Most of the prior art uses the Jaccard index as the similarity score. For example, some techniques create behavior profiles based on execution traces. Locality-Sensitive Hashing schemes are used to reduce the number of pairs considered. Malware files are grouped by hierarchical clustering, where the Jaccard index is used as the similarity metric. Compared with the previously proposed single-CPU systems, other techniques (such as BITSHRED) adopt feature hashing, bit vectors, and MapReduce implementation schemes on Hadoop to speed up calculations and reduce memory consumption. BITSHRED also uses the Jaccard index as the similarity metric for its co-clustering algorithm. Other techniques compare the Jaccard indices between behavior files generated from multiple analysis instances of a single file in multiple Anubis sandboxes to detect similar malware. However, one problem with these techniques (among others) is that the Jaccard index requires all features in a file to have the same weight. As mentioned above, certain key features (e.g., specific URLs or registry entries) or patterns may be the most important for detecting whether a single piece of content is associated with malware or may belong to a highly prevalent family or other malicious content families. Therefore, the prior art cannot dynamically learn key features to assign them higher prediction weights. As a result, the prediction accuracy is low, and the error rate is relatively high.
[0028] Various embodiments of the present disclosure improve upon these prior arts via new functions not currently employed by these prior arts or computer security systems. As described herein with respect to the experimental results, these new functions improve prediction accuracy and reduce error rates. This improvement occurs because some embodiments do things that computer security systems have not done before, such as learning certain key features that are most important for detecting whether content contains malware or belongs to a specific malware family and weighting them accordingly. Thus, embodiments perform the new function of learning the weights corresponding to different features such that important features found in similar content from the same family have high similarity scores, while other features that distinguish malware content from benign content have lower similarity scores. The new functions include learning the feature space embeddings of specific known malware families (and / or benign files) via a deep learning system such that any new or unknown content indication can be mapped into the same feature space embedding to determine a specific distance or similarity between the new or unknown indication and a specific known malware family indication, whereby malicious content can be detected and / or new or unknown content can be grouped or mapped to a specific malware family of malicious content.
[0029] Prior arts also consume unnecessary computing resources such as memory and CPU. For example, prior arts may require training on millions of files to detect malicious content to obtain acceptable prediction results. Storing millions of such files not only consumes a large amount of memory, but also has a high CPU utilization because prediction data points may be compared with each trained file. This may result in bottlenecks in fetching, decoding, or executing operations, or otherwise affect throughput or network latency, etc. Some techniques also only store malware signatures or other strings representing malicious content, which may consume unnecessary memory, especially when storing thousands or millions of files.
[0030] Specific embodiments improve the functions of the computer itself and improve other technologies because they do not consume unnecessary computing resources. For example, some embodiments use deep learning models with shared or tied weights or other parameters for two or more inputs. This means fewer parameters to train, which means less data is needed and there is also less tendency to overfit. Thus, less memory is consumed and the CPU utilization is also lower because there is less data to compare prediction data points with. Therefore, embodiments can improve metrics such as throughput and network latency. Additionally, some embodiments perform a kind of compression function on data by converting strings and other content into vectors and performing calculations on the vectors in memory (e.g., similarity scores based on cosine distance), rather than performing calculations on strings or malware signatures that consume a relatively large amount of memory compared to vectors. Thus, embodiments save computing resource utilization such as CPU and memory.
[0031] Turning now to Figure 1 FIG. 1 provides a block diagram showing an example operating environment 100 in which some embodiments of the present disclosure may be employed. It should be understood that such and other arrangements described herein are presented only by way of example. Other arrangements and elements (e.g., machines, interfaces, functions, orders, and groupings of functions) may be used to supplement or replace those shown, and some elements may be omitted altogether for clarity. In addition, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components and implemented in any suitable combination and location. The various functions described herein as being performed by entities may be performed by hardware, firmware, and / or software. For example, some functions may be performed by a processor executing instructions stored in a memory.
[0032] Among other components not shown, the example operating environment 100 includes a plurality of user devices, such as user devices 102a and 102b through 102n; a plurality of data sources (e.g., databases or other data stores), such as data sources 104a and 104b through 104n; a server 106; sensors 103a and 107; and a network 110. It should be understood that Figure 1 the environment 100 shown in FIG. 1 is an example of a suitable operating environment. Figure 1 Each of the components shown in FIG. 1 may be implemented via any type of computing device, such as, for example, the computing device 800 described in conjunction with Figure 8 FIG. 8. These components may communicate with each other via a network 110, which may include, but is not limited to, a local area network (LAN) and / or a wide area network (WAN). In an exemplary implementation, the network 110 includes the Internet and / or a cellular network, a network in any of a variety of possible public and / or private networks.
[0033] It should be understood that within the scope of the present disclosure, any number of user devices, servers, and data sources may be employed within the operating environment 100. Each may include a single device or multiple devices that cooperate in a distributed environment. For example, the server 106 may be provided via a plurality of devices arranged in a distributed environment that together provide the functions described herein. Additionally, other components not shown may also be included in the distributed environment.
[0034] User devices 102a and 102b through 102n can be client devices on the client side of the operating environment 100, while the server 106 can be a server on the server side of the operating environment 100. The server 106 can include server-side software designed to work with client software on user devices 102a and 102b through 102n to implement any combination of the features and functions discussed in this disclosure. This division of the operating environment 100 is provided to illustrate an example of a suitable environment, and for each implementation, it is not required that any combination of the server 106 and user devices 102a and 102b through 102n remain separate entities. In some embodiments, one or more servers 106 represent one or more nodes in a cloud computing environment. Consistent with various embodiments, a cloud computing environment includes a network-based distributed data processing system that provides one or more cloud computing services. Additionally, a cloud computing environment can include many computers, hundreds or thousands or more, arranged within one or more data centers and configured to share resources over the network 110.
[0035] In some embodiments, the user device 102a or the server 106 can alternatively or additionally include one or more web servers and / or application servers to facilitate the delivery of web pages or online content to a browser installed on the user device 102b. Generally, the content may include static content and dynamic content. When a client application (such as a web browser) requests a website or web application via a URL or search term, the browser typically contacts a web server to request static content or the basic components of the website or web application (e.g., HTML pages, image files, video files, etc.). The application server typically delivers any dynamic part of the web application or the business logic part of the web application. Business logic can be described as the function that manages the communication between the user device and a data store (such as a database). Such functions can include business rules or workflows (e.g., code indicating conditional if / then statements, while statements, etc. to represent the sequence of a process).
[0036] User devices 102a and 102b through 102n can include any type of computing device that can be used by a user. For example, in one embodiment, user devices 102a through 102n can be as described herein with respect to Figure 8Types of computing devices described. By way of example and not limitation, a user device may be embodied as a personal computer (PC), laptop computer, mobile or portable device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), music player or MP3 player, global positioning system (GPS) or device, video player, handheld communication device, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, camera, remote control, barcode scanner, computerized measurement device, household appliance, consumer electronic device, workstation, or any combination of these delineated devices, or any other suitable computing device.
[0037] In some embodiments, user device 102a and / or server 106 may include any of the components described herein or any other functionality (e.g., as described with respect to Figure 2 , 3 or 4). For example, user device 102 may detect malicious content, as described in Figure 2 . In some embodiments, server 106 may assist in detecting malicious behavior such that user device 102a and server 106 are used in combination to detect malicious code. For example, a web application may be opened on user device 102a. As a background task, or based on an explicit request from user device 102a, user device 102a may participate in a communication session or otherwise contact server 106, at which time server 106 uses one or more models to detect malicious behavior, etc., such as described with respect to Figure 2 .
[0038] Data sources 104a and 104b through 104n may include data sources and / or data systems that are configured to make data available to any of the various components of the operating environment 100 or system 200 described in connection with Figure 2 . Examples of one or more of data sources 104a through 104n may be one or more of a database, file, data structure, or other data storage. Data sources 104a and 104b through 104n may be discrete from user devices 102a and 102b through 102n and server 106, or may be combined and / or integrated into at least one of these components. In one embodiment, data sources 104a through 104n include sensors (such as sensors 103a and 107) that may be integrated into or associated with one or more of user devices 102a, 102b, or 102n or server 106.
[0039] The operating environment 100 may be used to implement one or more components of system 200, as described in Figure 2 . The operating environment 100 may also be used to implement in connection withFigure 7A and 7B aspects of process flows 700 and 730 described, and any other functionality as described in Figure 2-13 the following.
[0040] Now referring to Figure 2 and in conjunction with Figure 1 , a block diagram is provided that illustrates aspects of an example computing system architecture that is suitable for implementing embodiments of the present disclosure and is generally designated as system 200. Generally, embodiments of system 200 enable or support the detection of malicious content (e.g., code, functionality, features, etc.) and / or mapping of malicious content to one or more families (e.g., types, categories, or titles). System 200 is not intended to be limiting and represents only one example of a suitable computing system architecture. Other arrangements and elements may be used in addition to or in place of those shown, and some elements may be omitted altogether for clarity. Further, like the operating environment 100 of Figure 1 , many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components and implemented in any suitable combination and location. For example, the functionality of system 200 may be provided via a software-as-a-service (SAAS) model, e.g., cloud- and / or web-based services. In other embodiments, the functionality of system 200 may be implemented via a client / server architecture.
[0041] The emulator 203 is generally responsible for running or simulating the content (e.g., applications, code, files, or other objects) in the labeled data 213 and / or the unknown data 215 and extracting raw information from the labeled data 213 and / or the unknown data 215. The labeled data 213 includes samples of files or other objects that are marked or indicated with labels or classifications for use in training a machine learning system. For example, the labeled data 213 may include multiple files where the files have been labeled as benign (or family / sub-family) based on a label for a particular malicious code or file family (and / or sub-family) and / or a benign file. For example, the labeled data 213 may include several iterations or sub-families (sub-families are labels) of files infected with Rootkit malware, as well as other malicious code families. In this way, a machine learning model may be trained to identify patterns or associations indicated in the labeled data 213 for prediction purposes, as described in more detail herein. The unknown data 215 includes files or other content that do not have a predetermined label or classification. For example, the unknown data 215 may be any incoming file (e.g., a test file) that is analyzed after a machine learning model has been deployed, or used to train or test the labeled data 213.
[0042] The labeled data 213 and the unknown data 215 can generally be represented as storage. Storage generally stores information, including data used in embodiments of the techniques described herein, computer instructions (e.g., software program instructions, routines, or services), content, data structures, training data, and / or models (e.g., machine learning models). By way of example and not limitation, the data included in the labeled data 213 and the unknown data 215 can generally be referred to as data throughout. Some embodiments store computer logic (not shown) including rules, conditions, associations, classification models, and other criteria to perform the functions of any of the components, modules, analyzers, generators, and / or engines of the system 200.
[0043] In some embodiments, the emulator 203 (or any component described herein) runs in a virtualized (e.g., virtual machine or container) or sandboxed environment. In this way, any running malicious content does not infect the host or other applications. In some embodiments, specific raw information is extracted from the labeled data 213 and / or the unknown data 215. For example, the raw information can be an unpacked file string (or a string that makes a function call to decompress the file string) and API calls and their associated parameters. This is because malicious content is often packed, compressed, or encrypted, and thus may require calls to decrypt or otherwise unpack the data. Regarding API calls, certain malicious content may have a specific API call pattern, and thus this information may also need to be extracted.
[0044] The feature selector 205 is generally responsible for selecting specific features of the labeled data 213 and / or the unknown data 215 (e.g., selected features of the information extracted by the emulator 203) for training, testing, and / or making predictions. In various cases, there may be hundreds or thousands of features in the raw data generated by the emulator 203. Training a model using all of these features may require a large amount of computing resources, and thus a selected set of features can be used for training, testing, or prediction. Features can be selected according to any suitable technique. For example, features can be selected based on features that produce the most discriminative features using the mutual information criterion. That is, A(t,c) is calculated as the expected “mutual information (MI) of term t and class c. MI measures how much information the presence / absence of the term contributes to making a correct classification decision on c. Formally:
[0045] where U is a random variable that takes the values e t = 1 (the document contains the term t) and e t = 0 (the document does not contain t), and C is a random variable that takes the values e c = 1 (the document is of class c) and e c= 1 (the document is not of category c). If it is not clear from the context which term t and category c refer to, write U t and U c For the MLE of probability, Equation 1 is equivalent to:
[0046]
[0047] where N is an e with two subscripts t The count of documents with the value. For example, N 10 is the inclusion of t(e t =1) and not in c(e c =0). N1 = N 10 +N 11 is the inclusion of t(e t = 1) and the number of documents independent of class membership is counted as (e c ∈{0,1}). N=N 00 +N 01 +N 10 +N 11 is the total number of documents.
[0048] The training and / or test set construction component 207 is generally responsible for selecting malicious content that is known to be similar (e.g., content from the same family / sub-family) and / or selecting benign content that may not be similar in preparation for training and / or testing. In some embodiments, the training and / or test set construction component 207 receives user instructions to make these selections. In some embodiments, the training and / or test set construction component 207 generates a unique identifier (e.g., a signature ID) for each group of content in the tag data 213, and the embodiments can then responsively group or select similar malicious content pairs that belong to the same malicious family. For example, from a virus family, two sub-families of boot sector viruses can be paired together in preparation for training.
[0049] The model training component 209 is generally responsible for training the machine learning model by using the data selected via the training set construction component 207 and / or other components. In some embodiments, the model training component 209 additionally converts the selected features generated by the feature selector 205 into vectors in preparation for training and orienting them in feature space. In some embodiments, the weights are adjusted during training based on the cosine or other distance between the embedding of known malicious content and other known malicious content and / or benign content. For example, after several training stages, it can be determined that the set of vectors representing malicious files are within a threshold distance from each other. Back propagation and other techniques can be used to calculate the gradient of the loss function of the weights to fine-tune the weights based on the error rate of the previous training stage. In this way, specific features that are considered more important for specific malicious content can be learned.
[0050] The unknown construction component 220 is generally responsible for selecting known similar malicious content (e.g., content from the same family / sub-family) and / or selecting benign content and pairing them together or pairing them with known malicious content for testing or prediction after model deployment. For example, from a virus family, a sub-family of boot sector viruses can be paired with newly incoming files for prediction.
[0051] The unknown evaluator 211 is generally responsible for determining which sets of malicious content (within the labeled data 213) are similar to the set of unknown content (i.e., whether the unknown set contains malicious content) (within the unknown data 215) and scoring the similarity accordingly. For example, after model deployment, features can be extracted from incoming malware files or code. Then the malware file / code or features can be converted into vectors and oriented in a feature space. Then the distance between each vector representing the labeled data 213 and the incoming malware file or code can be calculated, and this distance represents the similarity score. If the similarity score exceeds a specific threshold (the vectors are within the threshold distance), the incoming file or code can be automatically mapped to a specific family or sub-family of malicious content. This is because the model training component 209 has presumably trained and / or tested the model to reflect the importance of certain features for a given malware file or object, meaning the trained vectors are oriented in the vector space based on the correct weights. Thus, if the vectors representing the incoming file or code are close within a certain distance (e.g., cosine distance), this may mean that the incoming file or code with unknown malicious content has the same or similar features compared to the labeled malicious files or codes or their family members.
[0052] The presentation component 217 is generally responsible for presenting an indication of whether malicious content has been detected based on the similarity score determined by the unknown evaluator 211. In some embodiments, the presentation component 217 generates a user interface or report that indicates whether a particular content set (e.g., the unknown data 215) is likely to be malicious. In some embodiments, the user interface or report additionally indicates each family (as assigned in the labeled data 213) (or certain families within the distance threshold) and the confidence score or likelihood that the particular content set belongs to that family. For example, the user interface can return a sorted list of items (e.g., indications of files in a telemetry report) that have a similarity score higher than the threshold compared to a particular file (e.g., the unknown data 215). In some embodiments, the presentation component 217 generates a new file or otherwise generates structured data (e.g., via tagging, entering in columns or rows or fields) (e.g., storable) to indicate whether a particular content snippet is likely to be malicious and / or belongs to a particular family of malicious content. For example, such a new file can be an attachment report of the investigation results.
[0053] Figure 3It is a block diagram of a system 300 for training a machine learning model on various malware content and predicting whether one or more specific unknown files (files whose malicious nature is unknown) contain malware, particularly with respect to highly prevalent malware families. The system 300 includes a training component 340 and an evaluation component 350. The training component 340 illustrates the functionality for training the underlying model, while the evaluation component 350 illustrates the steps followed when evaluating a set of unknown files to automatically predict whether they belong to one of the highly prevalent malware families.
[0054] In some embodiments, file emulation 303 represents the functionality performed by Figure 2 emulator 203. In some embodiments, the first component for model training and known file evaluation is lightweight file emulation 303. In some embodiments, a modified version of a production anti-malware engine is used to perform file emulation 303. This anti-malware engine extracts raw high-level data and generates logs during the emulation process. These logs can be consumed by downstream processing (e.g., other components within the training component 340). Embodiments utilize dynamic analysis to extract raw data from labeled files 313 to train the model and from unknown files 315 to predict the family (or to predict which family / sub-family a specific file belongs to based on its vector representation).
[0055] Some embodiments of file emulation 303 employ two types of data extracted from files, including unpacked file strings and application programming interface (API) calls and their associated parameters. Malware and other malicious code / files are often packed or encrypted. When the anti-malware engine emulates the file emulation 303 functionality, the malware unpacks or decrypts itself and often writes null-terminated objects into the emulator's memory. Typically, these null-terminated objects are strings that have been recovered during unpacking or decryption, and these strings can be good indicators of whether the file is malicious.
[0056] In some embodiments, in addition to unpacking file strings or alternative unpacking file strings, a second set of high-level features is constructed from an API call sequence and its parameter values via file emulation 303. The API stream consists of function calls from different sources, including user-mode operating systems and kernel-mode operating systems. For example, there are a specific number of WINDOWS APIs available for reading registry key values, including the user-mode functions RegQueryValue() and RegQueryValueEx(), and the kernel-mode function RtlQueryRegistryValues(). In some embodiments, functions that perform the same logical operation are mapped to a single API event. In these embodiments, calls to RegQueryValue(), RegQueryValueEx(), and RtlQueryRegistryValues() are all mapped to the same API event ID (EventID). Alternatively or additionally, important API parameter values, such as key names or key values, are also captured by file emulation 303. In some embodiments, using this data, a second feature set is constructed from a series of API call events and their parameter values. In some embodiments, to handle the case where multiple API calls are mapped to the same event but have different parameters, only the two (or other threshold number) most important parameters shared by different API calls are considered.
[0057] In some embodiments, file emulation 303 includes low-level feature encoding. Each malware sample may generate thousands of raw unpacked file strings or API call events and their parameters. Since in addition to non-polymorphic malware, certain embodiments also detect polymorphic malware (malware that constantly changes its identifiable features to evade detection), some embodiments do not directly encode potential features as sparse binary features (usually binary zero values). For example, if each variant in a family discards or removes a second temporary file with a partially random name or contacts a command and control (C&C) server using a partially random URL, then in some embodiments, the file name or URL is not explicitly represented. Instead, some embodiments encode the raw unpacked file strings and API calls and their parameters as a set of N-Grams characters. In some embodiments, triples of characters of all values are used (i.e., N-Grams with N = 3).
[0058] One limitation of the similarity system based on the Jaccard index (as described above) is that it cannot distinguish or determine the importance among multiple types of features (e.g., EventID, parameter value 1, parameter value 2) within the same set. Additionally, short values such as EventID (e.g., 98) have less impact on the Jaccard index than longer features including parameter values (e.g., registry key names). To improve the performance of the Jaccard index baseline system, some embodiments overcome these limitations by expanding the EventID to the full API and encoding the entire API name as a string using character-level triples (or other N-Gram configurations). Thus, representing the API name using their triples (or other N-Gram configurations) allows the API name to make a greater contribution to the Jaccard index of the file pair. In some embodiments, the triple representation of the API name is used for all models to fairly compare the results of the SNN model with those of the Jaccard index-based model, the results of which will be described in more detail below.
[0059] Certain embodiments described herein are not affected by these limitations suffered by the Jaccard index-based models. Some embodiments of the file emulation 303 encode the event ID or API name as a single categorical feature because certain deep learning networks, such as two-layer deep neural networks, can learn to assign greater weights to the most important API calls for similar file pairs. Thus, the entire call can be encoded individually as (EventID, parameter 1 N-Grams, parameter 2 N-Grams). This improves the performance of the learning models (including the SNN model or DSSM) because it learns a specific representation for each combination of the EventID and the N-Grams of the parameter values.
[0060] In some embodiments, the feature selection 305 includes functions performed by the feature selector 205. In various cases, there may be hundreds of thousands of potential N-Gram features in the raw data generated during the dynamic analysis process, and using all of these features to train a model may be computationally prohibitive. Thus, feature selection can be performed by class, which can produce the most discriminative features using, for example, the mutual information criterion, as described above. To handle production-level input data streams, embodiments implement the functions that may be required for preprocessing big data, such as in MICROSOFT's COSMOS MapReduce system.
[0061] In some embodiments, the training set construction 307 is performed by Figure 2The functions performed by the training set construction component 207. In some embodiments, prior to training, embodiments of the training set construction 307 first construct a training set that includes N-Gram features selected from known similar malware file pairs and dissimilar benign files (for labeling purposes, even if they may actually be similar). Embodiments determine pairs of similar malware files or other content for the training set based on a number of criteria. For example, to properly train the model, pairs of similar files need to be carefully selected first. Randomly selecting two files that match in family may not work well in practice. The problem is that some of these families may have many different variants. To solve this problem, the detection signatures of malware files can be utilized. An anti-malware engine can use a specific signature to determine whether an unknown file or content is malicious or benign. Each signature is typically very specific and has a unique identifier (signature ID). Thus, in some embodiments, the first step in determining similar pairs for training is to group malware files or content pairs that have equivalent signature IDs detected. While most malicious files or other content with the same signature ID belong to the same malware family, this is not always the case, which is why further analysis may be needed to map or label the unique identifier to the correct family. This can be a problem with existing technologies that detect whether a file is malicious mainly or only based on the signature ID of the malware. Thus, some embodiments improve these technologies by labeling candidate file or other content pairs as belonging to the same family.
[0062] Various embodiments of the training set construction 307 construct malware file pairs based on the signature ID and / or malware family. In certain embodiments, the benign files all belong to one class or label and do not have a signature ID assigned because they are not malicious. As a result, similar pairs of benign files are not constructed in a particular embodiment. Thus, to overcome this, some embodiments also construct "dissimilar" pairs that are constructed by randomly selecting a unique malware file and a benign file to form a pair. According to some embodiments, the format of the training set is as shown in Table 1 below:
[0063]
[0064] Table 1: Training and test set instance format.
[0065] The training set ID is composed of the concatenation of the SHA1 file hashes of malware 1 (M1) and malware 2 (M2) or a benign file (B) (i.e., SHA1 M1 -SHA1 M2,B)。This ID allows embodiments to identify which files were used to construct the training instances. The next field in the training set provides a label, where 1 indicates that two files are similar (M1, M2), and -1 indicates that they are dissimilar (M1, B). The third field provides N-Gram features selected from the primary malware file M1. N-Grams from the matching malware file M2 or a randomly selected benign file (B) are provided in the last field.
[0066] For the held-out test set used to evaluate all models, as described in more detail below, embodiments of the training set construction 307 ensure that the file pairs in the training and test sets are unique or distinct. To this end, a first set of malware files can be randomly selected for the training and test sets, followed by a pair of files including a malware file and a benign file (e.g., the first set of files). In response, a second pair of similar malware files is selected. If any of the files in the second pair match a file in the first set, the second pair of malware is added to the training set. If it is not in the training set, the embodiment adds it to the test set. Similarly, in some embodiments, the same process can be performed on the second pair of dissimilar malware and benign pairs such that the second pair of dissimilar malware and benign pairs is compared to the first set. Certain embodiments continue the process of randomly selecting malware pairs and adding them to the training or test set until each is complete.
[0067] In some embodiments, model training 309 includes functions performed by Figure 2 the model training component 209. In some embodiments, after the training set has been constructed via the training set construction 307, the model (e.g., SNN) is trained, e.g., with respect to Figure 5 or Figure 6 as depicted. In some embodiments, the weights are adjusted during training based on the cosine distance between the vector embedding of the known malware file M1 on the left (analyzed by the training component 340) and the vector embedding of the malware or benign file on the right (analyzed by the evaluation component 350). The combined set of similar and dissimilar files can be represented as F ∈ {M2, B}. Some embodiments use backpropagation with stochastic gradient descent (SGD and Adam optimizer) to train the model parameters.
[0068] When testing a test file or other content, some embodiments input known malware files or content into the left side (training component 340), and then use the right side (evaluation component 350) to evaluate new or unknown files or content (e.g., the test file). The test file or content can represent new or different files or content that the model has not been trained on. Thus, according to some embodiments, during testing, the output of the model can represent the cosine distance between the embedding or vector of the known malware content on the left side and the embedding or vector of the malicious or benign file on the right side. For example, a first set of files in the unknown file 315 can first undergo file emulation 333 (which can be the same function performed by the file emulation 303 and executed by the emulator 203), such that the first set of files is emulated and specific information, such as API calls and packed strings, is extracted and then unpacked. Then features can be selected from the first set of files for each 335 (e.g., via the same or similar functionality as described with respect to feature selection 305 or via the feature selector 205). In response, an unknown pair construction 320 can be performed (e.g., via the same or similar functionality as the training set construction 307 or via the unknown construction component 220), such that similar malicious test files are grouped together and any other benign test files are grouped together. This functionality can be the same as the training set construction 307, except that the files are not training data, but rather test data used to test the accuracy or performance of the model. In response, in some embodiments, an unknown pair evaluation 311 is performed on each file in the first set of files. This evaluation can be done using the model training 309 of the training component 340. For example, a first test file in the first set of files can be converted to a first vector and mapped into the feature space, and the similarity score between the first vector and one or more other vectors represented in the labeled file 313 (i.e., other malware files represented as vectors) can be determined by determining the distance between the vectors in the feature space. In various embodiments, an indication of the similarity score result is then output via the unknown file prediction 317. For example, any suitable structured format, user interface, file, etc. can be generated, as described with respect to the presentation component 217.
[0069] In some embodiments, the evaluation component 350 illustrates how to evaluate or predict a file after the data has been trained and tested, or after the model has been otherwise deployed in a particular application. For example, after the model has been trained and tested, it can be deployed in a web application or other application. Thus, for example, a user can upload a particular file (e.g., the unknown file 315) to a particular web application during a session to request a prediction result regarding whether the particular file is likely to be associated with malware or belong to a particular malware family. Thus, all the processes described regarding the evaluation component 350 can be performed at runtime in response to a request, such that the unknown file prediction 317 can indicate whether the file uploaded by the user is associated with malware. As shown in system 300, such a prediction can be based on the model training 309 (and / or testing) of other files.
[0070] The format of the evaluation set can be provided as shown in Table II below:
[0071]
[0072]
[0073] Table II: Evaluation Set Instance Format.
[0074] Similar to the training set ID, the evaluation set ID includes the SHA1 file hashes of known malware files and unknown files (i.e., AHA1 M1 _SHAI U ), so that it can be determined which malware file is similar to the unknown file. The other two fields include the N-Grams from the known malware files and unknown files. In some embodiments, to evaluate an unknown file, the selected N-Gram features are first included from all known variants of highly prevalent families in the training set (e.g., the labeled files 313). In some embodiments, these features correspond to Figure 4 the left side of the deep learning model. Then the selected N-Gram features can be included from all unknown files that arrive for processing within a particular time period (e.g., a particular day, week, month, etc.).
[0075] Depending on the number of known variants of the prevalent families and the incoming rate of unknown files, it may be necessary to further pre-filter the number of file pairs to be considered. In some embodiments, this pre-filtering includes using the MinHash algorithm to reduce the number of file pairs used during training or the number of file pairs included during evaluation. Alternatively or additionally, a locality-sensitive hashing algorithm can be used. The MinHash algorithm is approximately O(n), and only identifies a small number of samples that need to be compared with each unknown file being evaluated.
[0076] In an embodiment, after constructing known file pairs, they can be evaluated via unknown pair evaluation 311. That is, the known file pairs can be evaluated using model training 309 and compared with the training pairs. If the similarity score exceeds a specified threshold, the embodiment automatically determines that the file belongs to the same family (or sub-family) as the known malware file in the evaluation pair.
[0077] Some embodiments alternatively determine the similarity score or otherwise detect whether an unknown file is malicious by replacing the cosine() distance with an optional K-Nearest Neighbor (KNN) classifier and assign the unknown file to the majority malware family of the votes or the K benign class known files with one or more of the highest similarity scores. Assigning the label of the nearest single file (K = 1) may perform well. Thus, some embodiments only need to find the single file (e.g., stored into the labeled file 313) that is most similar to the unknown file (in the unknown file 315).
[0078] In some embodiments, to process production-level input data streams, other functional blocks for preprocessing data for training and testing are required (e.g., in MICROSOFT's COSMOS MapReduce system). These functional blocks can include feature selection for training and training set construction, as well as an evaluation function for selecting features to create an unknown pair data set. In some embodiments, once the data set is constructed, the model can be trained and the results of the evaluation or test set can be evaluated on a single computer. In practice, the prediction scores of the unknown file set and the KNN classifier can also be evaluated from the trained model in a platform such as a MapReduce platform.
[0079] Figure 4 is a block diagram of an example system 400 for using a trained model to determine whether a new file is malicious according to some embodiments. System 400 includes a file repository 402, a detonation and extraction module 404, a label database 406, a combination and construction module 408, a train similarity module 410, index data 412, a KNN index construction module 414, a similarity model 416, a new unlabeled file 418, a detonation and extraction module 420, a new file classification module 422, and a similar file with a KNN classification output 424. It should be understood that any component of system 400 can be replaced Figure 2 and / or Figure 3 with any component described in or combined with the system of
[0080] In some embodiments, the detonation and extraction module 404 first detonates and extracts string and behavior features from the file repository 402. In some embodiments, the detonation and extraction module 404 includes information about Figure 2emulator 203 and / or Figure 3 the functionality described by file emulation 303. In an illustrative example, the detonation and extraction module can extract the packed file strings (and then unpack them) and API calls and their associated parameters from the file repository 402. In various embodiments, the file repository represents a data store of files that have not been labeled such that it is not known whether the files are associated without malicious content.
[0081] In some embodiments, in response to the detonation and extraction module 404 performing its function, the combination and construction module 408 combines the features with the labels from the label database 406 and combines them into a similarity training data set, where similar files are paired (with a label called "similar") and dissimilar files are paired (with a label called "dissimilar"). In some embodiments, the combination and construction module 408 includes a feature selector 205 and / or a training set construction component 207 and / or Figure 2 the functionality described by feature selection 305 and / or training set construction 307. In an illustrative example, the computing device can receive a user selection of different members of the same family, which are paired together for training and labeled as "similar", and other members are combined with benign files or members of different families and labeled as "dissimilar". Figure 3 In some embodiments, in response to the function performed by the combination and construction module 408, the training similarity module 410 uses a machine learning architecture such as an SNN or DSSM architecture to train a similarity metric to produce a similarity model. In some embodiments, the training similarity module 410 includes a model training component 209 and / or
[0082] the functionality described by model training 309. In an illustrative example, the training similarity module 410 can obtain the pairs generated by the combination and construction module 408, convert the pairs to vectors, and embed each pair in a feature space, and at different training stages, the weights of specific important features can be adjusted as described herein such that the final training output is each file represented as a vector in the feature space, which is embedded with the minimum possible loss based on changing the weights during the training iterations. Figure 2 a model training component 209 and / or Figure 3 In some embodiments, in response to the function performed by the combination and construction module 408, the training similarity module 410 uses a machine learning architecture such as an SNN or DSSM architecture to train a similarity metric to produce a similarity model. In some embodiments, the training similarity module 410 includes a model training component 209 and / or
[0083] In some embodiments, in response to the trained similarity model, the KNN (K-Nearest Neighbor) index construction module 414 receives the extraction strings and behavioral features generated by the detonation and extraction module 404, and further receives labels from the label database 406, and further receives a portion of the similarity model 416 to construct the KNN index 412. In various embodiments, the KNN index 412 is used to map or quickly index incoming or new files (e.g., after the model is deployed) to the training data so that appropriate classification can be performed.
[0084] In some embodiments, after the KNN index 412 is constructed, new unlabeled files 418 (which are not part of the label database 406) are tested and / or otherwise used for making predictions, such as after the model is deployed. As Figure 4 shown, the detonation and extraction module 420 processes the new unlabeled files 418 by detonating and extracting strings and behavioral features. In some embodiments, the new unlabeled files 418 are new files that were previously in the file repository 402 and have not been analyzed for maliciousness. Alternatively, in some embodiments, the new unlabeled files 418 are brand new files that are not in the file repository 402. In some embodiments, the detonation and extraction module 420 represents the same module as the detonation and extraction module 404. Alternatively, these can be separate modules. In some embodiments, the detonation and extraction module 420 includes functions described as Figure 3 the emulator 203 and / or file emulation 333 with respect to
[0085] In some embodiments, in response to the detonation and extraction module 420 performing its function, the new file classification module 422 classifies the new files in the new unlabeled files 418. In some embodiments, this can occur by using the similarity model 416 and the KNN index 412 to find similar labeled files to produce a set of similar files with a KNN classification 424. In some embodiments, the label (or the classification performed by the new file classification module 422) determines the label or classification of the new unlabeled files 418 by majority voting.
[0086] Figure 5 is a schematic diagram of an example deep learning neural network (DNN) 500 used in a particular embodiment of the present disclosure. In some embodiments, the deep learning neural network 500 represents Figure 4 the similarity model 416 of Figure 3 or the model training 309 of Figure 5As shown, each branch includes an input layer, two hidden layers, and an output layer. Branches 501 and 503 are connected at the top by function 503 to determine the similarity between two inputs (e.g., two contents, such as two files). It should be understood that although there are only two branches 501 and 503 and a specific DNN configuration, there may be an appropriate number of branches or configurations. Compared with other layers, each layer can perform a linear transformation and / or a squashing non-linear function. The DNN can effectively have an input layer that assigns weighted inputs to the first hidden layer, which transforms its inputs and sends them to the second hidden layer. The second hidden layer transforms the output received from the first hidden layer and passes it to the output layer, which performs a further transformation and produces an output classification or similarity score.
[0087] In a particular embodiment, DNN 500 represents a trained double deep neural network, where cosine similarity scores are used to learn the parameters of branches 501 and 503. In some embodiments, during training, the left branch 501 and the right branch 503 are trained with known similar and dissimilar content pairs (e.g., representing training components 340). For example, in the first training phase, the first file is converted to a first vector and input at the input layer of the first branch 501, and the second similar file is converted to a second vector and input at the input layer of the second branch 503 and fed through the layers (hidden layer 1, hidden layer 2, output layer) such that the output layer learns the distance function between the two vectors. Then, via function 505, such as an energy function, the cosine similarity between the two similarity content sets is calculated. In some embodiments, the similarity score emerges as a result of combining two vectors representing the first content and the second content into a single vector, which is achieved by taking the element-wise absolute difference between the two vectors (|h(X1) - h(X2)|). In a particular embodiment, the single vector is then passed through a sigmoid function to output a similarity score between 0 and 1. This process can be repeated in the first training phase (or other training phases) for dissimilar content pairs (e.g., malicious files and benign files) such that the content sets are converted to corresponding vectors and combined into a single vector for which the similarity score is calculated. In this way, the DNN model 500 is configured to receive 2 inputs or input pairs and 1 output of the similarity score of the inputs, and can adjust the weights over time.
[0088] In some embodiments, during evaluation (e.g., by Figure 3The function performed by the evaluation component 350), the set of unknown content is input to the right branch 503 and compared with the known malicious content in the left branch 501 (e.g., as shown in the label database 406 or the tagged data 213). In some embodiments, the output is the cosine similarity score between the set of unknown content and the known malicious content. Since the content analysis is performed pairwise, the output process can be repeated for different malicious contents until the score is within the distance threshold between the known malicious content and the new content (e.g., the file is close enough to ensure that the unknown file to be classified is malicious and / or belongs to a specific family of malicious content). For example, in the first iteration, the first known malware file of the first family is input at the first branch 501 and the first unknown file is input at the second branch 503. After determining that the distance between the first known malware file and the first unknown file is outside the threshold (they are not similar) via the function 505, the second known malware file of the second family can be input at the first branch 501 and the first unknown file can be input at the second branch to calculate the similarity score again. This process can be repeated until the similarity score between the file pairs is within the threshold. However, in some embodiments, multiple iterations are not required. Instead, each trained set of malicious content can be represented in the feature space in a single view or analyzed at a single time, such that the embodiments can determine which vector representing the malicious content is closest or has the highest similarity score compared to the new unknown content.
[0089] In an example illustration of how the evaluation works, a new unknown file (not knowing whether it contains malicious content) is converted into a first vector and input at the input layer of the first branch 501, and a second known malicious file is converted into a second vector and input at the input layer of the first branch 501 and fed through the layers (hidden layer 1, hidden layer 2, output layer) such that the output layer learns the distance function between the two vectors. Then the cosine similarity between the two similarity files is calculated via the function 505. In some embodiments, the similarity score results from combining the two vectors representing the first file and the second file into a single vector, which is achieved by taking the element-wise absolute difference between the two vectors. In a particular embodiment, the single vector is then passed through a sigmoid function to output a similarity score between 0 and 1.
[0090] Figure 6 is a schematic diagram of an example deep learning neural network (DNN) 600 used in a particular embodiment of the present disclosure. In some embodiments, the deep learning neural network 600 represents Figure 4 the similarity model 416, or Figure 3Model training 309. The DNN 600 includes branches 603 (M1), 605 (M2), and 607 (B). Branch 603 indicates processing of first malware content. Branch 605 indicates processing of second malware content. And branch 607 indicates processing of benign content (or new / unknown content). As Figure 6 shown, each branch includes an input layer, three hidden layers, and an output layer. Branches 603, 605, and 607 are connected at the top by function 609 and function 611 to determine the similarity between two input contents. It will be appreciated that although there are only three branches 603, 605, and 607 and a specific DNN configuration, any suitable number of branches or configurations may exist. Each layer can perform a linear transformation and / or a squashing non-linear function compared to other layers.
[0091] In some embodiments, the DNN 600 represents a variant of the Deep Structured Semantic Model (DSSM), although the DNN 600 is a new model that improves upon existing models. The DSSM can address the high-level goal of training a model that learns to assign different weights to different features based on importance. However, when adopting a particular embodiment, a typical DSSM may not work. The input of a typical DSSM system consists of a very large collection of query-document pairs, which are known to result in high click-through rates. Radically different query-document pairs are usually unrelated, especially when the dataset is large and randomly shuffled beforehand. A typical DSSM generates a set of non-matching documents for each query by randomly selecting documents from other instances in the original training set. This idea is known as negative sampling. While negative sampling typically works in the context of web search, it is not applicable to the task of identifying sets of similar malware content. One problem is due to the fact that many pairs in the dataset belong to the same malware family. Additionally, since benign content is typically not polymorphic, there may not be matches for pairs of benign content. Thus, if features from two sets of malware content (known to be similar) are used, negative sampling will typically end up generating content that should be negative (i.e., non-matching) but instead matches. Therefore, the algorithm does not learn to encourage a large distance between malware content and benign content. Thus, some embodiments introduce a variant that explicitly accepts as input a pair of matching malware files (M1 and M2) and a non-matching benign file (B) for training.
[0092] Some embodiments modify DSSM training to require an additional item that is known to be dissimilar from the match in the pair. In this context, the third item is a set of N-Gram features selected from randomly chosen benign content. Performing feature selection on the dataset can result in corresponding to Figure 6The sparse binary features of the input layer size 14067 shown in []. In an embodiment, the first hidden layer is h1 = f(W1x + b1), where x is the input vector, W1 and b1 are the learning parameters of the first hidden layer, and f() is the activation function of the hidden layer. In some embodiments, the outputs of the remaining N - 1 hidden layers are:
[0093] h i = f(W i h i-1 + b i ), i = 2,..., N Equation 3
[0094] In some embodiments, the output layer for each DNN (or branches 603, 605, and 607) used to implement the model is y = f(W N h N-1 + b N ). In some embodiments, the tanh() function is used as the activation function for all hidden layers and the output layer of each individual DNN (or branch), where the tanh() function is defined in some embodiments as
[0095] Although Figure 6 the conceptual model shown in [] depicts three deep neural networks or branches (603, 607, and 607), in some embodiments, the DSSM is actually implemented as two branches, using one DNN for the first malware content and using another DNN for the second malware content and benign content. A DNN with any architecture can be supported by one constraint - the two branches or neural networks have the same number of output nodes.
[0096] In some embodiments, it is assumed that the correlation score between the first malware content and the second malware content or benign content (represented by feature vectors F1 and F2 respectively) is proportional to the cosine similarity score of their corresponding semantic concept vectors Y F1 and Y F2 :
[0097]
[0098] where Y F1 is the output of the first DNN (or the first branch), and Y F2 is the output of the second DNN (or the second branch).
[0099] In a particular embodiment, DSSM training seeks to maximize the conditional likelihood of the second malware content given the first malware content, while minimizing the conditional likelihood of the benign content given the first malware content. To do this for training instances, the posterior probability can first be calculated as:
[0100]
[0101] where γ is the smoothing parameter of the softmax function. The loss function is equivalently minimized during training as:
[0102]
[0103] where Λ are the model parameters W i and b i . In some embodiments, backpropagation is used together with stochastic gradient descent (SGD) to train the model parameters.
[0104] In some embodiments, the unknown content is input into the DNN 600, and the similarity score is determined in a manner similar to that during training. For example, the format can be implemented as shown in Table II. Similar to training, the evaluation set ID includes the SHA1 file hashes of known malware content, which allows determining which malware content is similar to the unknown content. The other two fields include the N-Grams from the known malware files and the unknown files.
[0105] To evaluate the unknown content, the selected N-Gram features can be included from all known family variants. Then the selected N-Gram features can be included from all unknown content arriving for processing within a specific time period (e.g., one day). Depending on the number of known family variants and the incoming rate of unknown files, it may be useful to further pre-filter the number of file pairs to consider. This can be done using the MinHash algorithm to reduce the number of content pairs used during training or the number of file pairs included during evaluation. Once the set of unknown content pairs is constructed, it can be evaluated using the trained DNN model 600. If the similarity score exceeds a specified threshold, it can be automatically determined that the content belongs to the same family as the known malware content in the evaluation pair.
[0106] Figure 7Ais a flowchart of an example process 700 for training a machine learning model according to some embodiments. Process 700 (and / or any functionality described herein, such as process 730) may be performed by processing logic that includes hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions that run on a processor to perform hardware emulation), firmware, or a combination thereof. Although the specific blocks described in this disclosure are referenced in a particular order and in a particular number, it should be understood that any block may occur substantially in parallel with any other block or before or after any other block. Additionally, there may be more (or fewer) blocks than illustrated. Such additional blocks may include blocks that embody any functionality described herein. Computer-implemented methods, systems (including at least one computing device having at least one processor and at least one computer-readable storage medium), and / or computer storage media as described herein may perform or be caused to perform these processes 700, 730, and / or any other functionality described herein. In some embodiments, process 700 represents the functionality described with respect to Figure 3 the training component 340.
[0107] According to block 702, a collection of computer objects (e.g., files or other content) is emulated (e.g., via file emulation 303). In some embodiments, the emulation includes extracting information from the collection of computer objects. For example, an unpacking or decrypting of a file string may be invoked to obtain a particular value, and an API may be invoked to obtain a particular value. In some embodiments, block 702 represents or includes functionality performed by Figure 2 the emulator 203 of Figure 4 the detonation and extraction module 404 of Figure 3 and / or the file emulation 303 of
[0108] In some embodiments, the collection of computer objects is labeled or pre-classified prior to analyzing features. For example, multiple known malicious files may be labeled as similar or dissimilar prior to training, and files labeled as benign files may be used to train a deep learning model, such as described with respect to the label database 406 or the labeled files 313 or the training component 340. Figure 2 the feature selector 205 of Figure 4 the combination and construction module 408 ofFigure 3 The functions described by feature selection 305.
[0109] According to block 706, a training set construction pair of computer objects is identified (e.g., by training set construction component 207). For example, similar file pairs can be paired (e.g., malware files marked as belonging to the same family), and dissimilar files can be paired (e.g., a benign file can be paired with any malware file or two malware files that are members of different malware families can be paired). In some embodiments, block 706 represents or includes the functions described with respect to Figure 3 training set construction component 207, combination and construction module 408, and / or training set construction 307.
[0110] According to block 708, a machine learning model (e.g., a deep learning model) is trained at least in part based on learning weights associated with significant feature values of a feature set. For example, using the figure above, a particular malware file can be associated with a particular URL or URL value and a particular registry key value (e.g., a container object with a particular bit value, such as a LOCAL_MACHINE or CURRENT_CONFIG entry value). These weights can be learned for each labeled malware file of a particular family member such that features can be learned that are most important for files classified as malware or within a particular family member.
[0111] In some embodiments, pairs of similar and dissimilar computer objects (or computer objects described with respect to block 706) of a set of computer objects are processed or run through a deep learning model by comparing the set of computer objects to a feature space and mapping them into the feature space. And at least in part based on this processing, the weights associated with the deep learning model can be adjusted to indicate the importance of certain features of the set of computer objects for prediction or classification. In some embodiments, the adjustment includes changing the embedding of a first computer object of similar computer objects in the feature space. For example, after the first round or multiple rounds of training of a set, it may not be known which features of the set of computer objects are important for making a particular classification or prediction. Thus, each feature can have an equal weight (or a weight that is nearly equal within a threshold, such as a 2% change in weight) such that all indications of the set of computer objects are substantially close to or within a distance threshold in the feature space. However, after several rounds of training or any threshold amount of training, the indications can be adjusted or changed based on feature similarity to be closer or farther apart from each other. The more features two computer objects match or are within a threshold, the closer the two computer objects are to each other, and when the features do not match or are not within a threshold, the farther apart the two computer objects are from each other.
[0112] In various embodiments, a deep learning model is trained at least in part based on identifying labels of pairs of computer objects as similar or dissimilar during preparation for training. Training can include adjusting weights associated with the deep learning model to indicate the importance of certain features of the set of computer objects for prediction or classification. In some embodiments, training includes learning an embedding in a feature space of a first computer object (or set of computer objects) of similar computer objects. Learning the embedding can include learning a distance between two or more indicators (e.g., files) representing two or more computer objects based on feature similarity between values of two or more indicators and adjusting the weights of the deep learning model. For example, as described above, the more features that match between two files or are within a threshold feature vector value, the closer the two files are to each other in the feature space, while when features do not match or are not within a feature vector value threshold, the farther apart the two files are from each other in the feature space. Thus, in response to different training phases, the connection strength between nodes or neurons in different layers can be weighted or strengthened based on the corresponding learned feature values that are most prominent or important for a particular family of malicious content. In this way, for example, the entire feature space can include embeddings of vectors or other indicators that are all learned or embedded in the feature space based on learned weights corresponding to different features, such that indicators of computer objects having important features found in similar computers are within a threshold distance of each other in the feature space, while indicators corresponding to dissimilar computer objects or computer objects having unimportant features are not within a threshold distance of each other in the same feature space.
[0113] In some embodiments, block 708 represents or includes functionality as described with respect to Figure 2 model training component 209 of Figure 4 training similarity module 410 of Figure 3 model training 309 of Figure 5 and / or DNN 500 of
[0114] Figure 7B is a flow diagram of an example process 730 for evaluating a new or unknown file according to some embodiments. In some embodiments, process 730 occurs after Figure 7A process 700 of Figure 7A such that the similarity score or prediction made at block 709 is based on using the learned model described with respect to Figure 3The functionality described in the evaluation component 350 of the process 730. In some embodiments, the "computer object" used in process 730 is a new or unknown computer object (e.g., a file) representing a test computer object t, such that process 730 represents testing a machine learning model. Alternatively, in some embodiments, the computer object is analyzed at runtime based on the user's interaction with the web application or other application after the machine learning model has been trained, tested, and deployed. In some embodiments, the computer object used in process 730 may alternatively be any content or computer object, such as a code sequence, function, data structure, and / or behavior. Thus, for example, any time the word "computer object" is used in process 730, it may be replaced by the word "content" or "computer object." Similarly, any time the term "file" is used herein, the term "file" may be replaced by the term "computer object."
[0115] According to box 703, a request is received to determine whether the computer object contains malicious content (e.g., via unknown construction component 220). In some embodiments, the malicious content includes known malware signatures or other indications of malware, such as rootkits, Trojan horses, etc. In some embodiments, the malicious content includes known functionality that known malware is known to exhibit, such as waiting for a specific amount of time before attacking, or injecting specific code sequences at different times, or other suitable behavior. In some embodiments, the request is received at box 703 based on a user uploading a file or indication to a web application or app at runtime (e.g., after training and model deployment) to determine whether the computer object contains any malicious content. In other embodiments, the request is received at box 703 based on a user uploading a new or unknown computer object that the machine learning model has not yet been trained on to test the machine learning model (e.g., before the learning model is deployed).
[0116] According to block 705, one or more features of the computer object are extracted (e.g., by emulator 203). In some embodiments, the extraction includes extracting unpacked file strings and extracting API calls, such as, for example, Figure 3described by file emulations 303 or 333. In some embodiments, block 705 includes encoding the unpacked file string and API calls and associated parameters into a set of N-Gram characters, as described, for example, with respect to file emulations 303 and / or 333. As described herein, one limitation of the similarity system based on the Jaccard index is that it cannot distinguish between multiple types of features of the same set or computer object (e.g., event ID, parameter value 1, parameter value 2, etc.). For example, these limitations can be overcome by expanding the EventID to the full API name and encoding the entire API name as a string using character-level triples. In some embodiments, block 705 represents or includes functionality described with respect to file emulations 303, 333, Figure 4 detonation and extraction module 404 and / or Figure 2 emulator 203 described.
[0117] According to block 709, based on the respective features (extracted at block 705), a similarity score is generated between a computer object and each computer object of a plurality of computer objects known to contain malicious content via a deep learning model (e.g., by unknown evaluator 211). In some embodiments, the deep learning model is associated with a plurality of indicators representing known malicious computer objects. The plurality of indicators can be compared with an indicator representing the computer object. In some embodiments, the similarity score is generated between the computer object and each known malicious computer object of the plurality of known malicious computer objects, at least in part based on processing or running the indicator of the computer object through the deep learning model. In some embodiments, the similarity score represents or indicates a distance metric (e.g., cosine distance) between the computer object and each known malicious computer object of the plurality of known malicious computer objects. In this way, for example, it can be determined whether the indicator of the computer object is within a threshold distance of the set of known malicious computer objects in the feature space based on a plurality of features. The embedding or orientation of the set of known malicious computer objects in the feature space can be learned via training based on learned weights for different features of the known malicious computer objects (e.g., as described with respect to Figure 7A block 708). Thus, in some embodiments, the similarity score can represent the specific distance of the computer object from other known malicious computer objects (which may each belong to distinct families). This distance can be specifically based on the exact feature values by which the computer object is compared to the known malicious computer objects. For example, if the computer object has the exact feature values that have been weighted to prominence or importance during training as some known malware computer objects (e.g., as described with respect to Figure 7AIf so (as described by block 708 of the box 708), the distance between these two computer objects will be close within the threshold of the feature space, resulting in a high similarity score. In fact, in the training of known malware computer objects, the more eigenvalue features of a computer object that have been weighted for importance, the closer the computer object is to the known malware computer object in the feature space. Vice versa. In the training of known malware computer objects, the more eigenvalue features of a computer object that have not been weighted for importance, the farther the computer object is from the known computer object in the feature space.
[0118] In some embodiments, the learning model used at block 709 is a deep learning model that includes two identical sub-networks with shared weights and connected by a distance learning function. For example, in some embodiments, the deep learning model represents or includes Figure 5 DNN 500 and the related functions described herein. In some embodiments, the deep learning model includes two identical sub-networks that share weights during training and process a pair of similar known malware computer objects and a pair of dissimilar computer objects (e.g., benign files) (e.g., as described with respect to block 708 of FIG. 7). For example, these deep learning model embodiments may include Figure 5 DNN 500. In some embodiments, the deep learning model explicitly accepts a pair of matching malware computer objects and a non-matching benign computer object as inputs for training, such as described with respect to Figure 6 DNN 600. In various embodiments, the deep learning model is associated with multiple known malicious computer objects embedded in the feature space (e.g., with respect to Figure 7A the learning model described by block 708). In some embodiments, block 709 represents or includes the functions described with respect to Figure 3 unknown evaluator 211, new file classification module 422, and / or unknown pair evaluation 311.
[0119] In some embodiments, the similarity score generated at block 709 additionally or alternatively includes or represents a classification or prediction score (e.g., and an associated confidence), which indicates the likelihood that a classified or predicted computer object belongs to each of a plurality of known malicious computer objects. For example, each known malicious computer object may include three distinct malware families, such as a first Rootkit type, a second Rootkit type, and a third Rootkit type. The similarity score may indicate a 0.96 confidence or likelihood that the computer object belongs to the first Rootkit type, a 0.32 confidence or likelihood that the file belongs to the second Rootkit type, and a 0.12 confidence or likelihood that the file belongs to the third Rootkit type. In some embodiments, the classification is based on the distance in the feature space, as described above with respect to block 709.
[0120] In some embodiments, the similarity score of a first computer object is set at least in part based on one or more of a plurality of features matching or being close (within a threshold distance) to a modified embedding of the first computer object. In some embodiments, a "modified" embedding represents a learned or trained embedding with appropriate weights based on feature importance, e.g., as described with respect to Figure 7A the final trained embedding or model in block 708.
[0121] In some embodiments, the feature space includes an indication of benign computer objects such that the determination of whether an indication of a computer object is within a threshold distance is at least in part based on an analysis of the indication of benign computer objects. For example, the computer object received at block 703 may actually be a benign file that does not contain malware. Thus, when the file is run through the learning model, the file may be closer in the feature space to other benign files based on eigenvalue matching, or closer to other features of benign files analyzed during training. Thus, it can be determined that the computer object is outside the threshold distance of known malware computer objects (and within the threshold distance of benign computer objects). In various embodiments, the indication includes a vector in the feature space of embeddings of two branches of a deep learning model based on a distance function between two inputs (computer objects), where the first input is the vector and the second input is another vector representing a first malware computer object of a set of known malware computer objects. For example, this is described with respect to the Figure 5 functionality and branches of the deep learning model 500. In some embodiments, block 709 represents or includes functionality described with respect to Figure 2 the unknown evaluator 211, the unknown pair evaluation 311, and / or Figure 4 the new file classification module 422.
[0122] According to block 713, one or more identifiers (e.g., names of specific malware families or files) representing at least one of a plurality of known malicious computer objects are provided (e.g., by the rendering component 217) or generated on the computing device. The one or more identifiers can indicate that the computer object may be malicious and / or that the computer object may belong to a specific malicious family. In some embodiments, in at least partial response to a similarity score being higher than a threshold for a set of a plurality of known malicious computer objects and the computer object, a set of identifiers is provided to the computing device, the set of identifiers representing the set of a plurality of known malicious computer objects in a ranked order. The highest ranked identifier indicates that the computer object may contain malicious content and / or that the computer object may belong to a specific family or type of malicious content relative to other families associated with lower ranked identifiers. For example, the server 106 may provide a user interface or other display to the user device 102a that displays the identifiers for each rank (e.g., different malware family types such as virus XYZ, virus TYL, and virus BBX), as well as the probability or likelihood (e.g., confidence) that the computer object belongs to or is classified as one of the different malware family types (e.g., virus XYZ-.96, virus TYL-.82, and virus BBX-.46, which indicates that the computer object may contain the virus XYZ variety). Thus, this describes a set of indicators for the set of known malicious computer objects, which are scored and provided to the user interface such that the provision indicates that the computer object belongs to a specific malware family.
[0123] In some embodiments, in at least partial response to a similarity score being higher than a threshold for at least one of a plurality of known malicious computer objects, the computing device is provided with at least one identifier that represents at least one known malicious computer object of the plurality of known malicious computer objects. In some embodiments, the at least one identifier indicates that the computer object may be malicious and that the content may belong to the same family as at least one known malicious computer object of the plurality of known malicious computer objects, as described above with respect to virus XYZ-.96, virus TYL-.82, and virus BBX-.46, which indicates that the computer object may contain the virus XYZ variety. In some embodiments, the similarity score higher than the threshold is at least partially based on a learned embedding of the first computer object (e.g., as described with respect to Figure 7A(as described by block 708). For example, an indication of a computer object can be mapped into a feature space with trained embeddings such that each known malware computer object indication within a threshold distance (e.g., cosine distance) from the computer object in the feature space can be provided to the computing device. Being within the threshold distance can indicate that computer objects are similar based on computer objects that include eigenvalue(s) weighted by the importance of the classification in the known computer object as well. Thus, one or more identifiers representing the known computer objects can be provided to the computing device because they indicate similarity to the computer object based on the learned embeddings.
[0124] In some embodiments, in at least partial response to determining whether an indication of a computer object is within a threshold distance (e.g., as described by the similarity score regarding block 709), an indicator indicating whether the computer object contains malicious content is provided to a user interface of the computing device. For example, the embodiment of block 713 can alternatively or additionally include a function that states the probability or likelihood that the computer object contains malware (e.g., without considering which malware family the file belongs to or without considering the candidate malware family or the ranked list of computer objects associated with the computer object). In some embodiments, block 713 represents or includes a function as described regarding Figure 2 the rendering component 217 of Figure 3 the unknown file prediction 317 and / or the similar files with KNN classification data 424.
[0125] Embodiments of the present disclosure may be described in the general context of computer code or machine - available instructions, including computer - available or computer - executable instructions executed by a computer or other machine (such as a smart phone, a tablet computer, or other mobile devices, a server, or a client device), such as program modules. Generally, program modules including routines, programs, objects, components, data structures, etc. refer to code that performs particular tasks or implements particular abstract data types. Embodiments of the present disclosure may be practiced in various system configurations, including mobile devices, consumer electronics, general - purpose computers, more specialized computing devices, etc. Embodiments of the present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including memory storage devices.
[0126] Some embodiments may include an end-to-end software-based system that may operate within the system components described herein to operate computer hardware to provide system functionality. At a low level, a hardware processor may execute instructions selected from a machine language (also referred to as machine code or native) instruction set for a given processor. The processor identifies native instructions and performs corresponding low-level functions related to, for example, logic, control, and memory operations. Low-level software written in machine code may provide more complex functionality for high-level software. Thus, in some embodiments, computer-executable instructions may include any software, including low-level software written in machine code, high-level software such as application software, and any combination thereof. In this regard, the system components may manage resources and provide services for system functionality. Embodiments of the present disclosure cover any other variations and combinations thereof.
[0127] Reference Figure 8 , computing device 800 includes a bus 10 that directly or indirectly couples the following devices: a memory 12, one or more processors 14, one or more presentation components 16, one or more input / output (I / O) ports 18, one or more I / O components 20, and an exemplary power supply 22. Bus 10 represents which may be one or more buses (such as an address bus, a data bus, or a combination thereof). Although, for clarity, Figure 8 the individual boxes are shown connected by lines, in reality, these boxes represent logical components and not necessarily actual components. For example, a presentation component such as a display device may be considered an I / O component. Additionally, a processor has memory. The inventors recognize this as being in the nature of the art and reiterate Figure 8 the schematic diagram only illustrates an exemplary computing device that may be used in conjunction with one or more embodiments of the present disclosure. There is no distinction between categories such as "workstation", "server", "laptop computer", "handheld device", or other computing devices because all of these categories are within Figure 8 the scope of and are considered with reference to "computing device".
[0128] Computing device 800 generally includes a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 800 and includes volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by computing device 800. Computer storage media does not itself include signals. Communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" refers to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.
[0129] Memory 12 includes computer storage media in the form of volatile and / or non-volatile memory. The memory can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid state memory, hard disk drives, optical disk drives, or other hardware. Computing device 800 includes one or more processors 14 that read data from various entities such as memory 12 or I / O components 20. In some embodiments, processor 14 executes instructions in memory to perform any of the operations or functions described herein, such as those described with respect to Figure 7A and 7B processes 700 and / or 730. One or more presentation components 16 present data indications to a user or other device. Exemplary presentation components include display devices, speakers, printing components, vibration components, etc.
[0130] The I / O port 18 allows the computing device 800 to be logically coupled to other devices, including the I / O component 20, some of which may be built-in. Illustrative components include microphones, joysticks, gamepads, satellite dishes, scanners, printers, wireless devices, etc. The I / O component 20 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some cases, the input can be transmitted to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, touch and stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display on the computing device 800. The computing device 800 can be equipped with a depth camera, such as a stereo camera system, an infrared camera system, an RGB camera system, and combinations thereof, for gesture detection and recognition. In addition, the computing device 800 can be equipped with an accelerometer or gyroscope capable of detecting motion. The output of the accelerometer or gyroscope can be provided to the display of the computing device 800 to present immersive augmented reality or virtual reality.
[0131] Some embodiments of the computing device 800 may include one or more radios 24 (or similar wireless communication components). The radio 24 transmits and receives radio or wireless communications. The computing device 800 can be a wireless terminal adapted to receive communications and media via various wireless networks. The computing device 800 can communicate via wireless protocols such as Code Division Multiple Access (“CDMA”), Global System for Mobile Communications (“GSM”), or Time Division Multiple Access (“TDMA”) to communicate with other devices. The radio communication can be a short-range connection, a long-range connection, or a combination of short-range and long-range radio telecommunications connections. When we refer to “short” and “long” types of connections, we do not mean the spatial relationship between two devices. Instead, we generally refer to short distance and long distance as different categories or types of connections (i.e., primary connections and secondary connections). By way of example and not limitation, a short-range connection can include a connection to a device that provides access to a wireless communication network (e.g., a mobile hotspot), such as a WLAN connection using the 802.11 protocol; a Bluetooth connection to another computing device is a second example of a short-range connection or a near-field communication connection. By way of example and not limitation, a long-range connection can include a connection using one or more of the CDMA, GPRS, GSM, TDMA, and 802.16 protocols.
[0132] The various components used herein have been identified. It should be understood that any number of components and arrangements can be employed to achieve the desired functions within the scope of the present disclosure. For example, for clarity of concept, the components in the embodiments depicted in the figures are shown with lines. Other arrangements of these and other components can also be implemented. For example, although some components are depicted as single components, many of the elements described herein can be implemented as discrete or distributed components or in combination with other components and in any suitable combination and location. Some elements can be omitted entirely. Additionally, the various functions described herein as being performed by one or more entities can be performed by hardware, firmware, and / or software, as described below. For example, various functions can be performed by a processor executing instructions stored in a memory. Thus, other arrangements and elements (e.g., machines, interfaces, functions, order, and groupings of functions, etc.) can be used in addition to or in place of those shown.
[0133] The embodiments of the present disclosure have been described with an illustrative rather than a restrictive intent. The embodiments described in the above paragraphs can be combined with one or more of the specifically described alternatives. In particular, the claimed embodiments can incorporate references to more than one other embodiment as alternatives. The claimed embodiments can specify further limitations of the claimed subject matter. Alternative embodiments will become apparent to the reader of the present disclosure after and because of reading the present disclosure. The foregoing alternatives can be implemented without departing from the scope of the appended claims. Certain features and subcombinations are useful and can be employed without reference to other features and subcombinations and are contemplated to be within the scope of the claims.
[0134] Example experimental result embodiments.
[0135] Several example experiments are described below to evaluate the performance (e.g., accuracy and error rate) of the embodiments of the present disclosure (e.g., which use DNN 500 and / or 600) compared to the prior art. First, the settings and hyperparameters used in the experiments are described. Next, the performance of the embodiments is compared to the performance of a baseline system based on the Jaccard index for file similarity tasks - the Jaccard index has been used in several previously proposed malware detection systems. Finally, the performance of the k-nearest neighbor classifier is studied based on features from the SNN and the Jaccard index calculated from a highly popular malware file set.
[0136] Dataset. Analysts at MICROSOFT Corporation provided researchers with the raw data logs extracted during the dynamic analysis of 190,546 Windows Portable Executable (PE) files from eight highly prevalent malware families and 208,061 logs from benign files. These files were scanned by the company's production file scanning infrastructure using its production anti-malware engine over several weeks in April 2019.
[0137] Following the above process (e.g., as Figure 3 shown), a training set of 596,478 pairs was created from two similar malware files (296,871 pairs) or malware files and benign files (299,607 pairs). Each row in the training set is unique. Similar malware pairs were chosen to be distinct, while dissimilar pairs were formed by a combination of randomly selected malware files and benign files. For testing, a separate held-out dataset was constructed, which consisted of 311,787 distinct pairs, including 119,562 unique similar pairs and 192,225 unique dissimilar pairs. Figure 9 The detailed classification for each family is shown in Table 900 of
[0138] Experimental setup. The SNN was implemented and trained using the PyTorch deep learning framework. The deep learning results were computed on an NVidia P100 GPU (Graphics Processing Unit). The Jaccard index calculation was also implemented in Python. For learning, the mini-batch size was set to 256 and the learning rate was set to 0.01. In some embodiments, the network architecture parameters are as Figure 5 shown.
[0139] File similarity. Some embodiments of the present disclosure that employ an SNN variant to learn the files in the embedded feature space are compared with a baseline technical system based on the Jaccard index. The Jaccard index (JI) for sets A and B is: JI(A,B) = A ∩ B) / (A ∪ B). The elements of the sets correspond to the above N-Gram encodings (e.g., regarding file emulation 303) of the corresponding feature types (e.g., strings, API events, and their parameters).
[0140] First, the Jaccard indices of similar and dissimilar files can be compared as a baseline, as Figure 10 shown. Figure 10 represents the Jaccard index similarity score distribution for similar and dissimilar files. Generally, the Jaccard index is typically higher for similar files and lower for dissimilar files, as expected. However, Figure 10The Jaccard index indicating a reasonable large number of similar documents is less than 0.9. There is also a small peak in the Jaccard index of dissimilar documents around the value of 0.65.
[0141] Figure 11 Illustrates the SNN similarity score distribution of similar and dissimilar documents. For Figure 11 , the SNN similarity scores of similar and dissimilar documents are compared. Since the range of SNN scores is [-1, 1], while the Jaccard index varies from [0, 1], the SNN scores are transformed to new values so that all curves can be fairly compared. Compared with the Jaccard index similarity scores, the SNN similarity performance is as expected, where the model learns to emphasize the weights from similar documents to produce cosine similarity scores very close to 1.0, and emphasizes the weights from dissimilar documents to produce cosine similarity values, which are well separated from similar documents (e.g., closer to 0). This behavior makes it easier to set the value of the SNN threshold (e.g., similarity score threshold) to automatically predict whether two documents are truly similar.
[0142] In these tests, it was found that calculating the Jaccard index is very slow. To address this constraint and to complete the Jaccard index experiment within 2 - 3 days, the MinHash algorithm described in this paper was implemented, where the MinHash filtering threshold is variable. Three different tests were conducted, including: (1) training set pairs = 50,000 and test set pairs = 10,000, with a threshold of 0.5; (2) training set pairs = 10,000 and test set pairs = 10,000, with a threshold of 0.8; (3) training set pairs = 100,000 and test set pairs = 10,000, with a threshold of 0.8. Figure 10 and Figure 12 The results reported in Table 1200 of
[0143] Family classification. Although Figure 10 and Figure 11 show that the embodiments of the present disclosure produce significantly improved similarity scores compared to the Jaccard index technique, it is useful to understand whether this leads to an improved detection rate. To conduct the research, the K - nearest neighbor classification results (K = 1) of the two similarity models (Jaccard index and SNN) were compared. The FACEBOOK AI similarity search library was used.
[0144] In some cases, evaluating the KNN classifier requires revisiting the test set. As described herein, pairs of similar malware files with matching feature IDs and families can be formed, as well as dissimilar pairs where the second file is known to be benign. The unknown file is then compared to a set of known malware files (e.g., Figure 5 the left branch 501) using the example model to determine if it is similar to any previously detected malware files. However, to evaluate the output of the example model of the KNN classifier, the unknown file is compared to a set of known malware and known benign files. For the test data set, this can be the set of files represented as F and processed through the right branch 503. Substantially, the set of known test files is swapped from Figure 5 the left branch 501 to Figure 5 the right branch 503. When evaluating an unknown file using the KNN classifier, the distances from the unknown file to the set of known malware and benign files can be computed, and the label can be determined based on the majority vote of the K closest files. It should be noted that in the following tests, the test files are all malicious and an attempt is made to determine the family of the malicious files.
[0145] Since the data set does not allow for the formation of similar benign file pairs, the false positive rate and true negative rate are not measured. If a file is benign, the score for a similar benign file from the data may not be computed. For example, Chrome.exe is not similar to AcroRd32.exe, even though both files are benign. If a file is a non-matching file, some embodiments do not compute this because the KNN results are already generated from a combination of matching and non-matching pairs.
[0146] Figure 12 Several performance metrics for malware family classification are summarized. Table 1200 shows that the SNN or the specific embodiments described herein are better than or equivalent to the JI for most families. Overall, the error rate of the SNN is 0.11%, while the error rate of the JI is 0.420%, which is a significant difference. The results indicate that the individual malware families are well separated in the latent vector space, where the latent vector is the final embedding, which is output, for example, by Figure 5 each branch of the model in. The experiment also used a visualization chart (as shown in Figure 13 ) that illustrates the visualization of the separability of the latent vectors of malware classes using the t-sne method. The t-sne method is used to project the vectors into a two-dimensional space, which confirms that the classes are indeed well separated.
[0147] To avoid detection, malware sometimes uses obfuscation, which occurs when it detects that it might be executing in an emulation or virtualization environment commonly used by dynamic analysis systems and does not perform any malicious operations. To overcome these attacks, systems that execute unknown files in multiple analysis systems and search for differences between the outputs can be used, which may indicate that malware is using obfuscation.
[0148] Malware may also attempt to delay the execution of its malicious code, hoping that the emulation engine will eventually stop after some time. To prevent this evasion technique, some systems may change the system time. This, in turn, causes the malware to check the time from the Internet. If the malware cannot connect to the Internet to check the time, it may choose to stop all malicious activities.
[0149] Some adversarial learning systems create examples that are misclassified by deep neural networks. For malware analysis, white-box attacks may require access to the model parameters of the DNN, which means that if an unknown file is analyzed in the DNN analysis backend service, the corporate network of an anti-malware company can be successfully invaded. Therefore, successfully using such an attack on the backend service can be very challenging. If the DNN is to be run locally on a client computer, the attacker may be able to reverse-engineer the model and successfully run a variant of the attack on the model embodiments described herein (e.g., Figure 5 and Figure 6 ). To counter this attack, embodiments run the DNN or other models in a Software Guard Extensions (SGX) enclave. This is a secure instruction code set built into the CPU. The enclave is only dynamically decrypted within the CPU itself.
[0150] When not successfully subverted by emulation and virtualization detection techniques, the embodiments of the dynamic analysis described herein have shown excellent results in detecting unknown malicious content. As described herein, a dynamic analysis system for learning malware file similarity based on deep learning has been proposed. The performance of the embodiments on highly prevalent families (which may require a large amount of support from analysts and automated analysis) has been described. It has been shown that, compared to the Jaccard-index-based similarity systems previously proposed in several other technologies, the embodiments described herein provide a significant improvement in terms of the error rate. The results show that the embodiments reduce the KNN classification error rate of these highly prevalent families from 0.420% of the Jaccard index to 0.011% of the SNN with two hidden layers. Therefore, it is believed that the embodiments described herein can be an effective technical tool for reducing the amount of automation costs and the time spent by analysts in combating these malware families.
[0151] The following embodiments represent exemplary aspects of the concepts envisioned herein. Any of the following embodiments may be combined in a multiple - dependent manner to depend on one or more other clauses. Additionally, any combination of dependent embodiments (e.g., clauses that explicitly depend on a previous clause) may be combined while remaining within the scope of the aspects envisioned herein. The following clauses are exemplary in nature and not restrictive:
[0152] Clause 1. A computer - implemented method, comprising: receiving a request to determine whether a computer object contains malicious content; extracting a plurality of features from the computer object; generating, at least in part based on the plurality of features, via a deep - learning model, a similarity score between the computer object and each of a plurality of computer objects known to contain malicious content, the deep - learning model being associated with a plurality of indicators representing the plurality of computer objects, the plurality of indicators being embedded in a feature space; and providing, at least in part in response to the similarity score being higher than a threshold for the set of the plurality of computer objects and the computer object, a set of identifiers to a computing device, the set of identifiers representing the set of the plurality of computer objects in a ranked order, wherein the identifier with the highest rank indicates that the computer object may belong to a particular family of malicious content.
[0153] Clause 2. The method according to clause 1, wherein the extraction of the plurality of features includes extracting unpacked file strings and extracting API calls.
[0154] Clause 3. The method according to clause 2, further comprising encoding the unpacked file strings and the API calls and associated parameters as a set of N - Gram characters.
[0155] Clause 4. The method according to clause 1, wherein the deep - learning model includes two identical sub - networks that share weights and are connected by a distance - learning function.
[0156] Clause 5. The method according to clause 1, wherein the deep - learning model is explicitly trained with a pair of matching malware files and a non - matching benign file as inputs.
[0157] Clause 6. The method according to clause 1, further comprising training the deep - learning model before receiving the request, the training of the deep - learning model including: simulating a set of files, the simulation including extracting information from the set of files; processing pairs of similar and dissimilar files of the set of files through the deep - learning model; adjusting, at least in part based on the processing, the weights associated with the deep - learning model to indicate the importance of certain features of the set of files for prediction or classification, wherein the adjustment includes changing the embedding of a first file of the similar files in the feature space.
[0158] Clause 7. The method according to Clause 6, wherein the similarity score of the computer object is set based at least in part on one or more of the plurality of features that match or are close to the altered embedding of the first document.
[0159] Clause 8. One or more computer storage media having computer-executable instructions thereon that, when executed by one or more processors, cause the one or more processors to perform a method comprising: receiving a request to determine whether content is malicious; extracting a plurality of features from the content; generating, based on the plurality of features, a similarity score between the content and each of a plurality of known malicious contents via a deep learning model, each of the plurality of known malicious contents belonging to a distinct malicious family; and providing, at least in part in response to the similarity score being higher than a threshold of known malicious contents of the content and the plurality of known malicious contents, an identifier representing the known malicious content to a computing device, wherein the identifier indicates that the content may be malicious or that the content may belong to the same malicious family as the known malicious content.
[0160] Clause 9. The computer storage media according to Clause 8, wherein the plurality of known malicious contents are labeled as similar or dissimilar during training, and wherein the deep learning model is further trained with content labeled as benign.
[0161] Clause 10. The computer storage media according to Clause 8, further comprising expanding the EventID feature of the content to a full API name and encoding the full API name as a string using character-level triples.
[0162] Clause 11. The computer storage media according to Clause 8, wherein the deep learning model comprises two identical sub-networks that share weights and process a pair of similar known malware files and a pair of dissimilar files during training, wherein the pair of dissimilar files includes a benign file.
[0163] Clause 12. The computer storage media according to Clause 8, wherein the deep learning model is explicitly trained with a pair of matching malware files and a non-matching benign file as input.
[0164] Clause 13. The computer storage medium according to Clause 8, wherein the method further comprises: training the deep learning model before receiving the request, the training of the deep learning model comprising: receiving a set of files marked as belonging to distinct malware families or benign files; emulating the set of files, the emulation comprising extracting information from the set of files; identifying the labels of pairs of the set of files as similar or dissimilar for training; and training the deep learning model at least in part based on the identification, the training comprising adjusting weights associated with the deep learning model to indicate the importance of certain features of the set of files for prediction or classification, wherein the training comprises learning an embedding of a first file of the similar files in a feature space.
[0165] Clause 14. The computer storage medium according to Clause 13, wherein the similarity score above the threshold is at least in part based on the learned embedding of the first file.
[0166] Clause 15. A system comprising: one or more processors; and one or more computer storage media storing computer-usable instructions that, when used by the one or more processors, cause the one or more processors to perform a method comprising: receiving a request to determine whether a computer object contains malicious content; extracting a plurality of features from the computer object; determining whether an indication of the computer object is within a threshold distance in a feature space to a set of known malicious computer objects, wherein an embedding of the set of known malicious computer objects in the feature space is learned via training based on learned weights of different features of the set of known malicious computer objects; and providing, at least in part in response to determining whether the indication of the computer object is within the threshold distance, an identifier indicating whether the computer object contains malicious content to a user interface of a computing device.
[0167] Clause 16. The system according to Clause 15, wherein a set of indicators of the set of known malicious computer objects is scored and provided to the user interface, the providing indicating the likelihood that the computer object belongs to a particular malware family.
[0168] Clause 17. The system according to Clause 15, wherein the feature space further comprises an indication of benign files, and wherein the determination of whether the indication of the computer object is within the threshold distance is at least in part based on analyzing the indication of the benign files.
[0169] Clause 18. The system according to Clause 15, wherein the indication includes two branch embeddings of a deep learning model based on learning a distance function between two inputs of vectors in the feature space, wherein the first input is the vector and the second input is another vector representing a first malware file of the set of known malicious computer objects.
[0170] Clause 19. The system according to Clause 15, wherein the feature space is associated with a deep learning model, the deep learning model including two identical sub-networks sharing weights and connected by a distance learning function.
[0171] Clause 20. The system according to Clause 15, wherein the feature space is associated with a deep learning model, the deep learning model being explicitly trained with a pair of matching malware files and a non-matching benign file as inputs.
Claims
1. A computer-implemented method, comprising: Receiving a request to determine whether a computer object contains malicious content; Extracting a plurality of features from the computer object; Generating, at least in part based on the plurality of features, via a deep learning model, a similarity score between the computer object and each of a plurality of computer objects known to contain malicious content, the deep learning model being associated with a plurality of indicators representing the plurality of computer objects, the plurality of indicators being embedded in a feature space, wherein the deep learning model is explicitly trained with a pair of matching malware files and a non-matching benign file as input; and Providing, at least in part in response to the similarity score being higher than a threshold for the set of the plurality of computer objects and the computer object, a set of identifiers representing the set of the plurality of computer objects to a computing device in a ranked order, wherein the identifier with the highest rank indicates that the computer object may belong to a specific malicious content family.
2. The method according to claim 1, wherein the extraction of the plurality of features includes extracting unpacked file strings and extracting API calls.
3. The method according to claim 2, further comprising: Encoding the unpacked file string, the API calls, and related parameters into a combination of N-Gram characters.
4. The method according to claim 1, wherein the deep learning model includes two identical sub-networks that share weights and are connected by a distance learning function.
5. The method according to claim 1, wherein the deep learning model includes a Siamese neural network (SNN).
6. The method according to claim 1, further comprising: Training the deep learning model before the receiving of the request, the training of the deep learning model including: Simulating a set of files, the simulation including extracting information from the set of files; Processing, by the deep learning model, pairs of similar and dissimilar files of the set of files; Adjusting, at least in part based on the processing, weights associated with the deep learning model to indicate the importance of certain features of the set of files for prediction or classification, wherein the adjustment includes: changing the embedding of a first file of the similar files in the feature space.
7. The method according to claim 6, wherein the similarity score for the computer object is set based on one or more of the plurality of features that match or are close to the changed embedding of the first file.
8. One or more computer storage media having computer-executable instructions thereon that, when executed by one or more processors, cause the one or more processors to execute a method, the method including: Receiving a request to determine whether content is malicious; Extracting a plurality of features from the content; Expanding the EventID feature of the content into a full API name and encoding the full API name as a string using character-level triples Generating a similarity score between the content and each of a plurality of known malicious contents based on the plurality of features and the extension, each of the plurality of known malicious contents belonging to a distinct malicious family, wherein the deep learning model is explicitly trained with a pair of matching malware files and a non-matching benign file as inputs; and At least in part in response to the similarity score being higher than a threshold of the known malicious content of the content and the plurality of known malicious contents, causing an identifier representing the known malicious content to be provided to a computing device, wherein the identifier indicates that the content may be malicious or that the content may belong to the same malicious family as the known malicious content.
9. The computer storage medium according to claim 8, wherein, The plurality of known malicious contents are labeled as similar or dissimilar during training, and wherein the deep learning model is further trained with content labeled as benign.
10. The computer storage medium according to claim 8, wherein the deep learning model comprises a Siamese Neural Network (SNN).
11. The computer storage medium according to claim 8, wherein the deep learning model comprises two identical sub-networks that share weights and during training process a pair of similar known malware files and a pair of dissimilar files, wherein, The pair of dissimilar files includes a benign file.
12. The computer storage medium according to claim 8, wherein the deep learning model is explicitly trained with a pair of matching malware files and a non-matching benign file as inputs.
13. The computer storage medium according to claim 8, wherein the method further comprises: Training the deep learning model before receiving the request, the training of the deep learning model comprising: receiving a set of files labeled as belonging to distinct malware families or a set of benign files; simulating the set of files, the simulation including extracting information from the set of files; labeling the set of file pairs as similar or dissimilar in preparation for training; and at least partially based on the labeling, training the deep learning model, the training including adjusting weights associated with the deep learning model to indicate the importance of certain features of the set of files for prediction or classification, wherein the training includes learning an embedding of a first file of the similar files in a feature space.
14. The computer storage medium according to claim 13, wherein, The similarity score higher than the threshold is at least in part based on the learned embedding of the first file.
15. A system, comprising: One or more processors; And One or more computer storage media storing computer-usable instructions that, when used by the one or more processors, cause the one or more processors to perform a method, the method comprising: Receiving a request to determine whether a computer object contains malicious content; Extracting a plurality of features from the computer object; Based on the plurality of features, determine whether an indication of the computer object is within a threshold distance from a set of known malicious computer objects in a feature space, wherein an embedding of the set of known malicious computer objects in the feature space is learned via training based on learned weights for different features of the set of known malicious computer objects, wherein the indication includes a vector that is embedded in the feature space based on two branches of a deep learning model that learns a distance function between two inputs, and wherein the deep learning model is explicitly trained with a pair of matching malware files and a non-matching benign file as inputs; and At least partially in response to determining that the indication of the computer object is within the threshold distance, provide an identifier indicating whether the computer object contains malicious content to a user interface of a computing device.
16. The system of claim 15, wherein a set of indicators of the set of known malicious computer objects is scored and provided to the user interface, the providing indicating a likelihood that the computer object is of a particular malware family.
17. The system of claim 15, wherein the feature space further includes an indication of a benign file, and wherein the determination of whether the indication of the computer object is within the threshold distance is at least partially based on analyzing the indication of the benign file.
18. The system of claim 15, wherein the first input is the vector and the second input is another vector representing a first malware file of the set of known malicious computer objects.
19. The system of claim 15, wherein the feature space is associated with a deep learning model that includes two identical sub-networks that share weights and are connected by a distance learning function.
20. The system of claim 15, wherein the feature space is associated with a deep learning model that is explicitly trained with a pair of matching malware files and a non-matching benign file as inputs.
Citation Information
Patent Citations
Dalvik instruction abstraction-based Android malicious code detection method
CN106096405A
Population-based connectivity architecture for spiking neural networks
US20180174033A1
Malware detection using local computational models
US20190026466A1