Systems and methods for machine learning-based malware detection
A machine learning-based system analyzes network data using communication metrics to differentiate between benign and malware traffic, effectively identifying and blocking malicious communication channels.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- BLACKBERRY LTD
- Filing Date
- 2023-04-25
- Publication Date
- 2026-05-26
Smart Images

Figure 0007865919000001 
Figure 0007865919000002 
Figure 0007865919000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to machine learning, and more particularly to systems and methods for machine learning-based malware detection.
Background Art
[0002] Command and control is a post-exploitation tactic that enables an attacker to maintain persistence, communicate with infected hosts, steal data, and issue commands. Once a host is infected, the malware establishes a command and control channel to the attacker. To avoid detection, agents often remain dormant for long periods of time and communicate with the server periodically for further instructions. These intermittent communications are referred to as malware beacons.
[0003] Detecting the presence of malware beacons in network data is difficult for several reasons. For example, the check-in intervals for embedded agents vary, and most command and control systems have built-in techniques for avoiding detection, such as adding random jitter to the callback times. As another example, malware beacons often masquerade as normal communications, such as DNS or HTTP requests, to blend in with network data.
Summary of the Invention
Means for Solving the Problems
[0004] The present invention provides, for example, the following items. (Item 1) Obtaining a training set of network data including benign network data and malware network data; Causing a feature extraction engine to generate a pair of two sets for each source-destination pair in the training set of the network data; Using the above pair of sets, train a machine learning engine to distinguish between the above harmless network data and the above malware network data. Methods that include... (Item 2) Obtain network data that identifies at least one new event relating to at least one source-destination pair, Engaging the feature extraction engine to generate a pair of at least one source-destination pairs associated with at least one new event, Send the above set of two, relating to the above at least one source-destination pair associated with the above at least one new event, to the above machine learning engine for classification. The method described in the above items, further including the method described in the above items. (Item 3) The machine learning engine further includes receiving data classifying the above at least one source-destination pair as either harmless or malware. The method described in any one of the above items. (Item 4) The machine learning engine receives data classifying at least one of the above source-destination pairs as malware, Adding at least one of the above source or destination Internet Protocol addresses to a blacklist The method described in any one of the above items, further including: (Item 5) The process further includes using a training malware server to generate the malware network data such that the malware network data includes an Internet Protocol address known to be associated with the training malware server, The method described in any one of the above items. (Item 6) The above malware network data is generated in a manner that mimics a malware beacon by varying at least one of the following: communication interval, jitter rate, or data channel, as described in any one of the above items. (Item 7) The above set of two for each source-destination pair includes the communication interval skewness and the communication interval kurtosis, as described in any one of the above items. (Item 8) The above malware network data includes a more consistent communication interval skew than the above harmless network data, as described in any one of the above items. (Item 9) The above malware network data includes communication interval kurtosis that is more uniformly clustered than the above harmless network data, as described in any one of the above items. (Item 10) The above set of two items is the method described in any one of the above items, including the number of flow events, the byte-down mean, the byte-down standard deviation, the byte-up mean, the byte-up standard deviation, the communication interval mean, the communication interval standard deviation, the communication interval skewness, the communication interval kurtosis, the number of local endpoints connected to the destination, and the number of remote endpoints to which the local endpoints are connected. (Item 11) It is a system, At least one processor, A memory connected to the at least one processor and storing instructions, wherein the instructions are executed by the at least one processor. Obtain a training set of network data that includes both harmless network data and malware network data, The feature extraction engine is engaged to generate a set of two for each source-destination pair in the training set of the above network data. Using the above pair of sets, train a machine learning engine to distinguish between the above harmless network data and the above malware network data. To perform the above, configure at least one processor, memory and A system that includes these features. (Item 12) The above instruction is executed by at least one of the above processors, Obtain network data that identifies at least one new event relating to at least one source-destination pair, Engaging the feature extraction engine to generate a pair of at least one source-destination pairs associated with at least one new event, Send the above machine learning engine for classification a set of two, each containing at least one source-destination pair associated with the above at least one new event. The system described in any one of the above items, further comprising the configuration of at least one of the above processors to perform the following: (Item 13) The above instruction is executed by at least one of the above processors, The system according to any one of the above items, further configuring the at least one processor to receive data from the machine learning engine that classifies the at least one source-destination pair as either harmless or malware. (Item 14) The above instruction is executed by at least one of the above processors, The machine learning engine receives data classifying at least one of the above source-destination pairs as malware, Adding at least one of the above source or destination Internet Protocol addresses to a blacklist The system described in any one of the above items, further comprising the configuration of at least one of the above processors to perform the following: (Item 15) The above instruction is executed by at least one of the above processors, The system according to any one of the above items, further configuring the at least one processor to generate the malware network data using a training malware server such that the malware network data includes an Internet protocol address known to be associated with the training malware server. (Item 16) The above malware network data is generated to mimic a malware beacon by varying at least one of the following: communication interval, jitter rate, or data channel, as described in any one of the above items. (Item 17) The above set of two for each source-destination pair is a system as described in any one of the above items, including communication interval skewness and communication interval kurtosis. (Item 18) The malware network data includes at least one of the following: a more consistent communication interval skewness than the harmless network data, or a more uniformly clustered communication interval kurtosis than the harmless network data, as described in any one of the above items. (Item 19) The above pair of sets is a system as described in any one of the above items, including the number of flow events, the byte-down mean, the byte-down standard deviation, the byte-up mean, the byte-up standard deviation, the communication interval mean, the communication interval standard deviation, the communication interval skewness, the communication interval kurtosis, the number of local endpoints connected to the above destination, and the number of remote endpoints to which the local endpoints are connected. (Item 20) A non-transient computer-readable medium having processor-executable instructions stored thereon, wherein when the processor-executable instructions are executed by the processor, the processor provides the processor with Obtain a training set of network data that includes both harmless network data and malware network data, Cause the feature extraction engine to generate a set of two for each source-destination pair in the training set of the network data, and Use the set of two to train a machine learning engine to distinguish the harmless network data from the malware network data, and A non-transitory computer-readable medium that causes the above to be performed. (Abstract) The method includes obtaining a training set of network data including harmless network data and malware network data, causing a feature extraction engine to generate a set of two for each source-destination pair in the training set of the network data, and using the set of two to train a machine learning engine to distinguish the harmless network data from the malware network data.
Brief Description of the Drawings
[0005] Here, as an example, reference will be made to the accompanying drawings showing exemplary embodiments of the present application.
[0006] [Figure 1] FIG. 1 shows a high-level block diagram of a system for machine learning-based malware detection according to an embodiment.
[0007] [Figure 2] FIG. 2 provides a flowchart illustrating a method for training a machine learning engine for malware detection according to an embodiment.
[0008] [Figure 3A] FIG. 3A is a graph showing data transmission in malware network data used to train a machine learning engine according to the method of FIG. 2.
[0009] [Figure 3B]Figure 3B is a graph showing data transmission within harmless network data used to train a machine learning engine using the method in Figure 2.
[0010] [Figure 4A] Figure 4A is a graph showing the communication intervals within malware network data used to train a machine learning engine using the method in Figure 2.
[0011] [Figure 4B] Figure 4B is a graph showing the communication intervals in harmless network data used to train a machine learning engine using the method in Figure 2.
[0012] [Figure 5] Figure 5 provides a flowchart illustrating a method for machine learning-based malware detection.
[0013] [Figure 6] Figure 6 provides a flowchart illustrating a method for adding an Internet Protocol address to a blacklist.
[0014] [Figure 7] Figure 7 shows a high-level block diagram of an exemplary computing device according to one embodiment.
[0015] Similar reference figures are used within the drawing to refer to similar elements and features. [Modes for carrying out the invention]
[0016] Detailed description of exemplary embodiments Therefore, in one respect, a method is provided which includes obtaining a training set of network data containing harmless network data and malware network data; engaging a feature extraction engine to generate pairs of sets for each source-destination pair in the training set of network data; and training a machine learning engine using the pairs of sets to distinguish between harmless network data and malware network data.
[0017] In one or more embodiments, the method further includes obtaining network data that identifies at least one new event relating to at least one source-destination pair; engaging a feature extraction engine to generate pairs of at least one source-destination pair associated with at least one new event; and sending the pairs of at least one source-destination pair associated with at least one new event to a machine learning engine for classification.
[0018] In one or more embodiments, the method further includes receiving data from a machine learning engine that classifies at least one source-destination pair as either harmless or malware.
[0019] In one or more embodiments, the method further includes receiving data from a machine learning engine that classifies at least one source-destination pair as malware, and adding at least one Internet Protocol address of either the source or destination to a blacklist.
[0020] In one or more embodiments, the method further includes using a training malware server to generate malware network data such that the malware network data includes an Internet Protocol address known to be associated with the training malware server.
[0021] In one or more embodiments, malware network data is generated to mimic malware beacons by varying at least one of the following: communication interval, jitter rate, or data channel.
[0022] In one or more embodiments, each set of two for each source-destination pair includes a communication interval skewness and a communication interval kurtosis.
[0023] In one or more embodiments, malware network data contains more consistent communication interval skew than harmless network data.
[0024] In one or more embodiments, the malware network data includes more uniformly clustered communication interval kurtosis than the harmless network data.
[0025] In one or more embodiments, a pair of sets includes the number of flow events, the byte-down mean, the byte-down standard deviation, the byte-up mean, the byte-up standard deviation, the communication interval mean, the communication interval standard deviation, the communication interval skewness, the communication interval kurtosis, the number of local endpoints connected to the destination, and the number of remote endpoints to which the local endpoints are connected.
[0026] In another aspect, a system is provided, comprising at least one processor and memory coupled to at least one processor and storing instructions, the instructions comprising at least one processor, which, when executed by at least one processor, obtain a training set of network data including harmless network data and malware network data; engage a feature extraction engine to generate pairs of sets for each source-destination pair in the training set of network data; and train a machine learning engine using the pairs of sets to distinguish between harmless network data and malware network data.
[0027] In one or more embodiments, the instruction further configures at least one processor to perform the following actions when executed by at least one processor: to obtain network data identifying at least one new event relating to at least one source-destination pair; to engage a feature extraction engine to generate a pair of sets relating to at least one source-destination pair associated with at least one new event; and to send the pair of sets relating to at least one source-destination pair associated with at least one new event to a machine learning engine for classification.
[0028] In one or more embodiments, the instruction further configures at least one processor to receive data from a machine learning engine classifying at least one source-destination pair as either harmless or malware, when executed by at least one processor.
[0029] In one or more embodiments, the instruction further configures at least one processor to, when executed by at least one processor, receive data from a machine learning engine classifying at least one source-destination pair as malware, and blacklist at least one Internet Protocol address of either the source or the destination.
[0030] In one or more embodiments, the instruction further configures at least one processor to generate malware network data using a training malware server, such that the malware network data includes an Internet Protocol address known to be associated with the training malware server, when the instruction is executed by at least one processor.
[0031] In one or more embodiments, malware network data is generated to mimic malware beacons by varying at least one of the following: communication interval, jitter rate, or data channel.
[0032] In one or more embodiments, each set of two for each source-destination pair includes a communication interval skewness and a communication interval kurtosis.
[0033] In one or more embodiments, the malware network data includes at least one of the following: a more consistent interval skewness than that of the harmless network data, or a more uniformly clustered interval kurtosis than that of the harmless network data.
[0034] In one or more embodiments, a pair of sets includes the number of flow events, the byte-down mean, the byte-down standard deviation, the byte-up mean, the byte-up standard deviation, the communication interval mean, the communication interval standard deviation, the communication interval skewness, the communication interval kurtosis, the number of local endpoints connected to the destination, and the number of remote endpoints to which the local endpoints are connected.
[0035] In another aspect, a non-transient computer-readable medium is provided, the non-transient computer-readable medium having processor-executable instructions stored thereon, which, when executed by the processor, cause the processor to: obtain a training set of network data including harmless network data and malware network data; engage a feature extraction engine to generate pairs of sets for each source-destination pair in the training set of network data; and use the pairs of sets to train a machine learning engine to distinguish between harmless network data and malware network data.
[0036] Other exemplary embodiments of this disclosure will be apparent to those skilled in the art from a close examination of the following detailed description in conjunction with the drawings.
[0037] In this application, the term "and / or" is intended to encompass all possible combinations and secondary combinations of the enumerated elements, including any one of the enumerated elements, any secondary combination, or all of the elements, and not necessarily excluding additional elements.
[0038] In this application, the phrase "at least one of ~ or ~" is intended to encompass any one or more of the listed elements, and includes only one of the listed elements, any secondary combination, or all of the elements, without necessarily excluding any additional elements, and without necessarily requiring all of the elements.
[0039] Figure 1 is a high-level block diagram of a system 100 for machine learning-based malware detection according to one embodiment. System 100 includes a server computer system 110 and a data store 120.
[0040] The data store 120 may contain various data records. At least some of the data records may contain network data. The network data may include a training set of network data, which includes harmless network data and malware network data. As will be described, harmless network data may include network data that is known to be harmless, i.e., known not to contain malware. Malware network data may include network data that is known to be malware. Each network log may be referred to as a flow event.
[0041] In one or more embodiments, the network data stored in the data store 120 may include network logs, where each network log is a flow event between a specific destination and a specific remote endpoint. Each network log may include a timestamp, source Internet Protocol (IP) address, destination IP address, source port, destination port, etc. In addition, the network log may include data transmission information such as packet size and byte up / down.
[0042] Datastore 120 may also maintain one or more whitelists, including IP addresses known to be trusted, and one or more blacklists, including IP addresses known to be malware. As will be described, one or more whitelists may be viewed, and therefore any flow events associated with whitelisted IP addresses may not be classified. Similarly, one or more blacklists may also be viewed, and therefore any flow events associated with blacklisted IP addresses may be blocked or alerted.
[0043] The data store 120 may store only network data that occurred within a threshold period. For example, the data store 120 may store network data only for the past 7 days, and therefore may discard, erase, or otherwise delete network data older than 7 days.
[0044] In one or more embodiments, system 100 includes a malware engine 130, a feature extraction engine 140, and a machine learning engine 150. The malware engine 130 and the feature extraction engine 140 communicate with a server computer system 110. The malware engine 130 may log network data locally and export the logged network data to a data store 120 via the server computer system 110. The machine learning engine 150 communicates with the feature extraction engine 140 and the server computer system 110. The malware engine 130, the feature extraction engine 140, and the machine learning engine 150 may be separate computing devices in different environments.
[0045] The malware engine 130 may be configured to generate malware network data that can be used to train the machine learning engine 150. The malware network data may include malware beacons that are communicated between source-destination pairs. In one or more embodiments, the malware engine may include a virtual server, such as a training malware server, and one or more virtual computing devices, and communication between the virtual server and one or more virtual computing devices may be logged as malware network data.
[0046] The malware engine 130 may be configured to generate malware network data by, for example, varying the beacon interval, jitter amount, data channel, etc., in which extensive beacon obfuscation is obtained. The malware network data may include the IP address of a training malware server, which may be used to train the machine learning engine 150. For example, any network log containing the IP address of a training malware server may be identified as malware network data.
[0047] As will be explained, malware network data generated by the malware engine 130 may be stored in the data store 120 and used to train the machine learning engine 150.
[0048] The feature extraction engine 140 is configured to analyze network data received from the data store 120 and generate pairs of data for each source-destination pair in the network data. To generate pairs of data for each source-destination pair in the network data, the feature extraction engine 140 may analyze the network data and classify the network data by source-destination pair. For each source-destination pair, the pairs of data may include the number of flow events, the byte-down mean, the byte-down standard deviation, the byte-up mean, the byte-up standard deviation, the communication interval mean, the communication interval standard deviation, the communication interval skewness, the communication interval kurtosis, the number of local endpoints connected to the destination, and the number of remote endpoints to which the local endpoints are connected.
[0049] The number of flow events may include the count of flow events occurring within the network data relating to source-destination pairs. Since each network log is a flow event, the feature extraction engine 140 may also count the number of network logs relating to source-destination pairs in the network data, which can determine the number of flow events per source-destination pair.
[0050] The average byte-down for each source-destination pair may be generated by calculating the average size of the byte-down for each source-destination pair for all flow events in the network data relating to that source-destination pair.
[0051] The byte-down standard deviation for each source-destination pair may be generated by calculating the byte-down standard deviation for each source-destination pair for all flow events in the network data relating to that source-destination pair.
[0052] The average byte-up for each source-destination pair may be generated by calculating the average size of the byte-up for each source-destination pair for all flow events in the network data relating to that source-destination pair.
[0053] The byte-up standard deviation for each source-destination pair may be generated by calculating the byte-up standard deviation for each source-destination pair for all flow events in the network data relating to that source-destination pair.
[0054] The average communication interval may include the average number of seconds between flow events, or it may be generated by calculating the average number of seconds between flow events. It will be understood that the number of seconds between flow events can be the amount of time between adjacent flow events.
[0055] The standard deviation of the communication interval may include the standard deviation of the number of seconds between flow events, or it may be generated by calculating the standard deviation of the number of seconds between flow events.
[0056] The communication interval skewness may include a metric that indicates the degree to which the distribution is distorted toward one end.
[0057] The communication interval kurtosis may include a metric that indicates the tapering nature of the probability distribution. The communication interval kurtosis may be generated by determining a measure of the combined weight of the tails of the distribution relative to the center of the distribution.
[0058] The number of local endpoints connected to the destination may include the count of local endpoints that have flow events associated with the destination within the network data.
[0059] The number of remote endpoints to which a local endpoint is connected may include the count of remote endpoints that have one or more flow events with a source in the network data.
[0060] The machine learning engine 150 may include or utilize one or more machine learning models. For example, the machine learning engine 150 may be a classifier, such as a random forest classifier, which may be trained to classify network data as either malware network data or harmless network data. Other machine learning methods that may be used include support vector machines and decision tree-based boosting methods, such as AdaBoost(TM) and XGBoost(TM).
[0061] In one or more embodiments, a pair of sets of data generated by a feature extraction engine using training network data may be used to train a machine learning engine 150 for malware detection.
[0062] Figure 2 is a flowchart illustrating an operation performed by the server computer system 110 to train a machine learning engine for malware detection, according to one embodiment. This operation may be included in Method 200, which can be performed by the server computer system 110. For example, computer executable instructions stored in the memory of the server computer system 110 may be configured to perform Method 200 or a part thereof when executed by the processor of the server computer system. It will be understood that the server computer system 110 may offload at least some of the operations to the malware engine 130, the feature extraction engine 140, and / or the machine learning engine 150.
[0063] Method 200 includes obtaining a training set of network data, which includes harmless network data and malware network data (step 210).
[0064] In one or more embodiments, the server computer system 110 may obtain a training set of network data from the data store 120. As stated, malware network data may be generated by the malware engine 130. Harmless network data includes network data known to be harmless, and malware network data includes network data known to be malware.
[0065] Figure 3A is a graph showing data transmission within malware network data used to train a machine learning engine.
[0066] Figure 3B is a graph showing data transmission within harmless network data used to train a machine learning engine.
[0067] Comparing Figure 3A and Figure 3B, it can be seen that malware network data contains very consistent packet sizes, while harmless network data has inconsistent data patterns and packet sizes.
[0068] Figure 4A is a graph showing the communication intervals within malware network data used to train a machine learning engine.
[0069] Figure 4B is a graph showing the communication intervals within harmless network data used to train a machine learning engine.
[0070] Comparing Figure 4A and Figure 4B, it can be seen that malware network data includes consistent and regular communication intervals, while harmless network data includes communication intervals with long periods of inactivity and has a large number of communication intervals near zero.
[0071] In one or more embodiments, the malware network data may include more consistent interval skewness than the harmless network data, and / or more uniformly clustered interval kurtosis than the harmless network data.
[0072] Method 200 includes engaging a feature extraction engine to generate pairs of sets for each source-destination pair in the training set of network data (step 220).
[0073] As described, in order to generate pairs of data for each source-destination pair in the network data, the feature extraction engine 140 can analyze the network data and classify the network data by source-destination pair. For each source-destination pair, the pairs of data may include the number of flow events, the byte-down mean, the byte-down standard deviation, the byte-up mean, the byte-up standard deviation, the communication interval mean, the communication interval standard deviation, the communication interval skewness, the communication interval kurtosis, the number of local endpoints connected to the destination, and the number of remote endpoints to which the local endpoints are connected.
[0074] Method 200 includes training a machine learning engine using pairs of data to distinguish between harmless network data and malware network data (step 230).
[0075] The pair of data points is fed into a machine learning engine, which is used to train the engine and classify network data as either harmless network data or malware network data. Once trained, the machine learning engine can classify network data as either harmless network data or malware network data.
[0076] In embodiments where the machine learning engine 150 includes a random forest classifier, pairs of data can be labeled with zero (0) to indicate harmless network data, or with one to indicate malware network data. Furthermore, package functions may be used to fit the model to the data. For example, a fitting method may be used for each decision tree associated with the random forest classifier, and this fitting method may include selecting features and numerical values, as when the network data is split based on features. In this form, the purity of each split data chunk is maximized. This may be repeated multiple times for each decision tree, so that the classifier is trained for the prediction task.
[0077] Figure 5 is a flowchart illustrating an operation performed for machine learning-based malware detection according to one embodiment. This operation may be included in Method 500, which can be performed by the server computer system 110. For example, computer executable instructions stored in the memory of the server computer system 110 may be configured to perform Method 500 or a part thereof when executed by the processor of the server computer system.
[0078] Method 500 includes obtaining network data (step 510) that identifies at least one new event relating to at least one source-destination pair.
[0079] In one or more embodiments, the data store 120 may receive new flow events in the form of network data, which may occur periodically, for example, every minute, every 5 minutes, every 30 minutes, every hour, every 24 hours, etc. Specifically, the server computer system 110 may send a request for a new flow event to one or more source or destination computer systems connected to it, and the new flow event may be received in the form of network data. The server computer system 110 may send the received network data to the data store 120 for storage.
[0080] The server computer system 110 may analyze at least one new event and determine whether at least one new event is associated with a source-destination pair that is known to be trusted. For example, the server computer system 110 may look up a whitelist stored in the data store 120, which includes a list of IP addresses that are known to be trusted, and determine that at least one new event is associated with a source-destination pair that is known to be trusted. In response to determining that at least one new event is associated with a source-destination pair that is known to be trusted, the server computer system 110 may abandon the at least one new event and take no further action.
[0081] In response to determining that at least one new event is not associated with a known trusted source-destination pair, the server computer system 110 may send a request to the data store 120 for all available network data relating to at least one source-destination pair. In other words, the server computer system 110 does not merely request network data associated with at least one new event, but rather requests all available network data relating to at least one source-destination pair associated with at least one new event.
[0082] Method 500 includes engaging a feature extraction engine to generate a pair of sets of at least one source-destination pairs associated with at least one new event (step 520).
[0083] Network data acquired by the server computer system 110 is sent to the feature extraction engine to generate a pair of data sets relating to at least one source-destination pair associated with at least one new event. As described, the pair of data sets may include the number of flow events, the byte-down mean, the byte-down standard deviation, the byte-up mean, the byte-up standard deviation, the communication interval mean, the communication interval standard deviation, the communication interval skewness, the communication interval kurtosis, the number of local endpoints connected to the destination, and the number of remote endpoints to which the local endpoints are connected.
[0084] Method 500 includes sending pairs of data to a machine learning engine for classification (step 530).
[0085] The pairs are sent to a machine learning engine for classification.
[0086] Method 500 includes receiving data from a machine learning engine that classifies source-destination pairs as either harmless or malware (step 540).
[0087] As stated, the machine learning engine is trained to classify network data as either harmless network data or malware network data. Specifically, the machine learning engine analyzes pairs of network data to classify them as either harmless network data or malware network data.
[0088] The machine learning engine may classify source-destination pairs as either harmless or malware, which may be based on classifying network data as either harmless network data or malware network data. For example, in embodiments where network data is classified as malware network data, at least one source-destination pair may be classified as malware.
[0089] In embodiments where at least one source-destination pair is classified as harmless, the server computer system 110 may determine that no further action is required.
[0090] In embodiments where at least one source-destination pair is classified as malware, the server computer system 110 may perform one or more corrective actions. For example, the server computer system 110 may issue a flag or alarm indicating that the source-destination pair is malware.
[0091] In another embodiment, the server computer system 110 may blacklist at least one of the sources or destinations in a source-destination pair. Figure 6 is a flowchart illustrating the actions taken to blacklist an Internet Protocol address according to one embodiment. This action may be included in Method 600, which can be performed by the server computer system 110. For example, computer executable instructions stored in the memory of the server computer system 110 may be configured to perform Method 600 or a part thereof when executed by the processor of the server computer system.
[0092] Method 600 includes receiving data from a machine learning engine that classifies source-destination pairs as malware (step 610).
[0093] The machine learning engine may perform operations similar to those described herein with reference to Method 500, and may classify source-destination pairs as malware. In response, the server computer system 110 may receive data from the machine learning engine classifying source-destination pairs as malware.
[0094] Method 600 includes adding at least one Internet Protocol address, either a source or a destination, to a blacklist (step 620).
[0095] The server computer system 110 can determine at least one IP address, either a source or a destination, by analyzing the network data associated with it. The server computer system 110 can send a signal to the data store 120 and add the IP address to the blacklist maintained by it.
[0096] In addition to identifying the destination IP address as malware, or as an alternative thereto, in one or more embodiments, it will be understood that a fully qualified domain name (FQDM) may be identified as malware.
[0097] As stated, the server computer system 110 is a computing device. Figure 7 shows a high-level block diagram of an exemplary computing device 700. As shown, the exemplary computing device 700 includes a processor 710, memory 720, and an I / O interface 730. The aforementioned modules of the exemplary computing device 700 communicate with each other via a bus 740 and are coupled in a communicative manner.
[0098] Processor 710 includes hardware processors, and may include one or more processors, for example, using ARM, x86, MIPS, or PowerPC™ instruction sets. For example, Processor 710 may include Intel™ Core™ processors, Qualcomm™ Snapdragon™ processors, or equivalents.
[0099] Memory 720 comprises physical memory. Memory 720 may include random access memory, read-only memory, persistent storage such as flash memory, a solid-state drive, or equivalent. Read-only memory and persistent storage are computer-readable media, and more specifically, can be considered non-transient computer-readable storage media, respectively. The computer-readable media may be organized using a file system that can be managed by software governing the overall operation of the exemplary computing device 700.
[0100] The I / O interface 730 is an input / output interface. The I / O interface 730 enables the exemplary computing device 700 to receive inputs and provide outputs. For example, the I / O interface 730 may enable the exemplary computing device 700 to receive inputs from a user or provide outputs to a user. In another embodiment, the I / O interface 730 may enable the exemplary computing device 700 to communicate with a computer network. The I / O interface 730 may serve to interconnect the exemplary computing device 700 with one or more I / O devices, such as a keyboard, display screen, pointing device like a mouse or trackball, fingerprint reader, communication module, hardware security module (HSM) (e.g., Trusted Platform Module (TPM)), or equivalent. Virtual counterparts of the I / O interface 730 and / or devices accessed via the I / O interface 730 may be provided, for example, by a host operating system.
[0101] The software containing the instructions is executed by the processor 710 from a computer-readable medium. For example, software corresponding to a host operating system may be loaded into random-access memory from the persistent storage or flash memory of memory 720. In addition, or alternatively, the software may be executed by the processor 710 directly from the read-only memory of memory 720. In another embodiment, the software may be accessed via the I / O interface 730.
[0102] It will be understood that the malware engine 130, the feature extraction engine 140, and the machine learning engine 150 may also be computing devices similar to those described herein.
[0103] It will be understood that some or all of the actions of the various exemplary methods described above may be carried out in an order other than that shown, and / or in parallel without altering the overall operation of those methods.
[0104] The various embodiments presented above are merely examples and are not intended to limit the scope of this application. Modifications of the innovations described herein will be obvious to those skilled in the art, and such modifications will fall within the intended scope of this application. In particular, features from one or more of the exemplary embodiments described above may be selected to create alternative exemplary embodiments, including secondary combinations of features not expressly described above. In addition, features from one or more of the exemplary embodiments described above may be selected and combined to create alternative exemplary embodiments, including combinations of features not expressly described above. Features suitable for such combinations and secondary combinations will be readily apparent to those skilled in the art upon closer examination of this application. The subject matter described herein and in the enumerated claims is intended to encompass and include all suitable modifications in the art.
Claims
1. A method, The method involves obtaining a training set of network data, which includes harmless network data and malware network data, wherein the malware network data is generated to mimic malware beacons by varying the amount of jitter. The feature extraction engine is engaged to generate a set of two for each source-destination pair in the training set of the aforementioned network data. Using the aforementioned set of two, a machine learning engine is trained to distinguish between the harmless network data and the malware network data. Obtain network data that identifies at least one new event relating to at least one source-destination pair, Obtaining all available network data relating to the aforementioned at least one source-destination pair, Using the acquired network data, the feature extraction engine is engaged to generate a pair of at least one source-destination pairs associated with the at least one new event. Send the set of two, relating to the at least one source-destination pair associated with the at least one new event, to the machine learning engine for classification. Methods that include...
2. The machine learning engine further includes receiving data classifying the at least one source-destination pair as either harmless or malware. The method according to claim 1.
3. The machine learning engine receives data classifying the at least one source-destination pair as malware, Adding at least one Internet Protocol address of the source or destination to a blacklist The method according to claim 1, further comprising:
4. The process further includes using a training malware server to generate the malware network data such that the malware network data includes an Internet Protocol address known to be associated with the training malware server. The method according to claim 1.
5. The method according to claim 1, wherein the malware network data is further generated to mimic the malware beacon by varying the communication interval or data channel.
6. The method according to claim 1, wherein the set of two for each source-destination pair includes communication interval distortion and communication interval kurtosis.
7. The method according to claim 1, wherein the malware network data includes a more consistent communication interval skew than the harmless network data.
8. The method according to claim 1, wherein the malware network data includes communication interval kurtosis that is more uniformly clustered than the harmless network data.
9. The method according to claim 1, wherein the set of two includes the number of flow events, the byte-down average, the byte-down standard deviation, the byte-up average, the byte-up standard deviation, the communication interval average, the communication interval standard deviation, the communication interval skewness, the communication interval kurtosis, the number of local endpoints connected to the destination, and the number of remote endpoints to which the local endpoints are connected.
10. It is a system, At least one processor, A memory coupled to the at least one processor and storing instructions, wherein the instructions are executed by the at least one processor. The method involves obtaining a training set of network data, which includes harmless network data and malware network data, wherein the malware network data is generated to mimic malware beacons by varying the amount of jitter. The feature extraction engine is engaged to generate a set of two for each source-destination pair in the training set of the aforementioned network data. Using the aforementioned set of two, a machine learning engine is trained to distinguish between the harmless network data and the malware network data. Obtain network data that identifies at least one new event relating to at least one source-destination pair, Obtaining all available network data relating to the aforementioned at least one source-destination pair, Using the acquired network data, the feature extraction engine is engaged to generate a pair of at least one source-destination pairs associated with the at least one new event. Send the set of two, relating to the at least one source-destination pair associated with the at least one new event, to the machine learning engine for classification. To perform the above, the at least one processor is configured with memory and A system equipped with these features.
11. When the instruction is executed by the at least one processor, The system according to claim 10, further comprising the configuration of the at least one processor to receive data from the machine learning engine classifying the at least one source-destination pair as either harmless or malware.
12. When the instruction is executed by the at least one processor, The machine learning engine receives data classifying the at least one source-destination pair as malware, Adding at least one Internet Protocol address of the source or destination to a blacklist The system according to claim 11, further comprising configuring the at least one processor to perform the following:
13. When the instruction is executed by the at least one processor, The system according to claim 10, further comprising configuring the at least one processor to generate the malware network data using a training malware server such that the malware network data includes an Internet protocol address known to be associated with the training malware server.
14. The system according to claim 10, wherein the malware network data is further generated to mimic the malware beacon by varying the communication interval or data channel.
15. The system according to claim 10, wherein each set of two for each source-destination pair includes communication interval distortion and communication interval kurtosis.
16. The system according to claim 10, wherein the malware network data includes at least one of a more consistent communication interval skewness than the harmless network data, or a more uniformly clustered communication interval kurtosis than the harmless network data.
17. The system according to claim 10, wherein the set of two includes the number of flow events, the byte-down average, the byte-down standard deviation, the byte-up average, the byte-up standard deviation, the communication interval average, the communication interval standard deviation, the communication interval skewness, the communication interval kurtosis, the number of local endpoints connected to the destination, and the number of remote endpoints to which the local endpoints are connected.
18. A non-transient computer-readable medium having processor-executable instructions stored thereon, wherein when the processor executes the processor, the processor executes the instructions. The method involves obtaining a training set of network data, which includes harmless network data and malware network data, wherein the malware network data is generated to mimic malware beacons by varying the amount of jitter. The feature extraction engine is engaged to generate a set of two for each source-destination pair in the training set of the aforementioned network data. Using the aforementioned set of two, a machine learning engine is trained to distinguish between the harmless network data and the malware network data. Obtain network data that identifies at least one new event relating to at least one source-destination pair, Obtaining all available network data relating to the aforementioned at least one source-destination pair, Using the acquired network data, the feature extraction engine is engaged to generate a pair of at least one source-destination pairs associated with the at least one new event. Send the set of two, relating to the at least one source-destination pair associated with the at least one new event, to the machine learning engine for classification. A non-transient, computer-readable medium that enables the following action.