Systems and methods for automated threat detection

A dynamically trained security threat detection system using supervised machine learning models learns from analyst workflows to efficiently identify and prioritize potential breaches and suspicious activities, improving the efficiency of threat detection.

JP7717173B2Active Publication Date: 2025-08-01SECUREWORKS CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023549576
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-18
Filing Date
2021-12-17
Publication Date
2025-08-01
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

The overwhelming amount of security logs and the increasing sophistication of malicious actors make it difficult for security analysts to identify attack patterns and suspicious activities efficiently.

Method used

A dynamically trained security threat detection system that learns from the workflow of human analysts, using supervised machine learning models to automatically tag, prioritize, and organize scan results, generating an automated threat hunting playbook for efficient threat detection.

Benefits of technology

The system enhances the efficiency of threat detection by automatically identifying potential breaches and prioritizing suspicious activities, allowing analysts to focus on the most relevant data, thus streamlining the threat detection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717173000001
    Figure 0007717173000001
  • Figure 0007717173000002
    Figure 0007717173000002
  • Figure 0007717173000003
    Figure 0007717173000003
Patent Text Reader

Abstract

A system and method for dynamically training a threat detection system includes monitoring security analyst workflow data from security analysts analyzing scans of security logs. The workflow data includes rules applied to security log scan results, rule results selected for further analysis, tags applied to rule results, filters applied to rule results, rankings applied to rule results, or actions associated with pivots by the security analyst. A tagging classifier is then trained based on tags assigned to the scan results. A review classifier is trained based on scan results previously reviewed by the security analyst.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Identifying attack patterns and suspicious activities from malicious actors is an important aspect of computer network security. Highly trained security analysts spend a lot of time scrutinizing scans of security logs to identify and investigate potential breach events. The vast amount of security logs that can potentially be scrutinized can overwhelm the resources of security analysts. As malicious actors become more sophisticated and adjust their attack strategies, it becomes increasingly difficult to identify attack patterns and suspicious activities, and the limited resources of trained analysts are increasingly spread thinly.

[0002] Therefore, it can be seen that there is a need for systems and methods that can not only automatically detect potential breach events and suspicious activities, but also organize and prioritize scan results for more efficient scrutiny by security analysts.

[0003] This disclosure is directed to the foregoing and other related or unrelated problems or issues in the relevant art.

Summary of the Invention

Means for Solving the Problems

[0004] Briefly described, according to various aspects, the present disclosure includes systems and methods for dynamically training a security threat detection system. According to one aspect, a method for dynamically training a security threat detection system is disclosed. The method includes monitoring security analyst workflow data from one or more security analysts analyzing a scan of security logs. The workflow data includes one or more rules applied to the security log scan results, rule results selected for further analysis, tags applied to the rule results, filters applied to the rule results, rankings applied to the rule results, or one or more actions associated with a pivot by one or more security analysts, and / or combinations thereof. The method also includes training a tagging classifier based on tags assigned to rule results from the workflow data, training a scrutiny classifier based on rule results selected for further analysis, training a filter and ranking method based on filters and rankings applied to rule results from one or more security analysts, generating an automated threat hunting playbook including the tagging classifier, the scrutiny classifier, and the filter and ranking method, and generating one or more scripts for automatically analyzing incoming security data using the automated threat hunting playbook. In one embodiment, the method also includes training a pivot sequence model based on actions performed by one or more security analysts, and the automated threat hunting playbook also includes the pivot sequence model. In one embodiment, the tagging classifier, the scrutiny classifier, the filter and ranking method, and the pivot sequence are each supervised machine learning models trained based on the workflow data of one or more security analysts. In one embodiment, the script for automatically analyzing incoming security data generates several tags, each tag being an indicator of an intrusion within a computer network.In one embodiment, the method also includes receiving a tag update from one or more security analysts and dynamically updating the tagging classifier based on the tag update. In one embodiment, a script for automatically analyzing incoming security data generates a selected item of results regarding an audit. In one embodiment, the method also includes receiving analyst feedback regarding the selected item of results regarding the audit and dynamically updating the audit classifier based on the analyst feedback regarding the selected item of results regarding the audit. In one embodiment, a script for automatically analyzing incoming security data generates a selected item of prioritized results. In one embodiment, the method also includes receiving analyst feedback regarding the selected item of prioritized results and dynamically updating the filtering and ranking method based on the analyst feedback. In one embodiment, a script for automatically analyzing incoming security data generates one or more pivot chains, where a pivot chain is a series of rule results that track a potential attack. In one embodiment, the method also includes receiving pivot chain feedback from one or more security analysts and dynamically updating the pivot sequence model based on the pivot chain feedback.

[0005] According to another aspect, a dynamically trained threat detection system includes a computing system for monitoring and storing security analyst workflow data from one or more security analysts analyzing security log scans. The workflow data includes rules applied to security log scan results, rule results selected for further analysis, tags applied to rule results, filters applied to rule results, rankings applied to rule results, or one or more actions associated with a pivot by one or more security analysts, and / or combinations thereof. The system also includes a tagging classifier trained based on tags assigned to rule results from workflow data, a scrutiny classifier trained based on rule results selected for further analysis, a filter and ranking method trained based on filters and rankings applied to rule results from one or more security analysts, and an automated threat hunting playbook including the tagging classifier, the scrutiny classifier, and the filter and ranking method. The automated threat hunting playbook is configured to generate one or more scripts for automatically analyzing incoming security data. In one embodiment, the system also includes a pivot sequence model based on actions performed by one or more security analysts, where the pivot chain is a series of rule results tracking a potential attack, and the automated threat hunting playbook also includes the pivot sequence model. In one embodiment, the tagging classifier, the scrutiny classifier, the filter and ranking method, and the pivot sequence are each supervised machine learning models trained based on workflow data of one or more security analysts. In one embodiment, the pivot sequence model generates a pivot chain when applied to raw scan data from security logs.In one embodiment, when the tagging classifier is applied to raw scan data from a security log, it generates tags, and each tag is an indicator of an intrusion within a computer network. In one embodiment, it generates a selected item of the result regarding the scrutiny when applied to the raw scan data from the security log. In one embodiment, when the filter and ranking method are applied to the raw scan data from the security log, they generate a selected item of the prioritized result.

[0006] According to another aspect, a system for dynamically training a security threat detection system includes one or more processors and at least one memory storing instructions. When executed, the instructions cause the system to monitor and record workflow data from one or more security analysts for analyzing security logs within a computer network. The workflow data includes rules applied to security log scan results, rule results selected for further analysis, tags applied to rule results, filters applied to rule results, rankings applied to rule results, or one or more actions associated with pivots by one or more security analysts, and / or combinations thereof. The instructions also cause the system to train a tagging classifier based on tags applied to rule results from the workflow data, train a scrutiny classifier based on rule results selected for further analysis by one or more security analysts, train a filter and ranking method based on filters and rankings applied to rule results from security analysts, and train a pivot sequence model based on actions performed by one or more security analysts. The tagging classifier, scrutiny classifier, filter and ranking method, and pivot sequence are each supervised machine learning models trained based on the workflow data of one or more security analysts. The instructions also cause the system to generate an automated threat hunting playbook that includes the tagging classifier, scrutiny classifier, and filter and ranking method. In one embodiment, the instructions also cause the system to analyze incoming security data using one or more scripts to generate several tags, a curated set of results for scrutiny, a curated set of ranked results, and one or more pivot chains. Each tag is an indicator of an intrusion within the computer network, and a pivot chain is a series of rule results that track a potential attack.The command also causes the system to receive tags, selected items of inspection results, selected items of prioritized results, and analyst feedback regarding one or more pivot chains, and based on the analyst feedback, dynamically update the tagging classifier, inspection classifier, filters and ranking methods, and pivot sequence model.

[0007] For example, in some aspects, the process for threat detection and training of an automated threat detection system can include a first step where a scan is initiated. For example, for an organization with 1,000 computers, the detection rules will be launched on all 1,000 computers and will collect evidence regarding selected events, actions, indicators, etc. These logs can then be loaded for analysts to review. In some embodiments, the logs can include raw logs, information aggregated from multiple logs, information about files, or any other information collected from hosts, systems, or services. However, assuming a substantial amount, potentially millions of logs can be made available for analysts to review. Analysts then need to determine which logs they will actually review since they cannot review all of the incoming logs. The system can start the scan while applying a selected set of rules on one or more computing devices and collect the relevant security logs. Such detection rules can start the processing of security logs by a host computer or server and then resend the results. For example, some of the rules can be configured to scan for known types of malware, and other detection rules can simply collect all of the logon events regarding a certain time or location, etc. Analysts can create search queries or start scan searches, evaluate whether the logs can be triaged and / or sorted, narrow down the logs to a certain hit or specific scan that should be more closely evaluated, and confirm whether an attack exists.

[0008] In one embodiment, the analyst begins to initiate or generate a scan by selecting a set of rules for investigation startup or an organization being investigated for infringement. The scan can be a set such as 10 rules, 20 rules, 100 rules, 1,000 rules, etc. The scan then starts on those computers of the selected set or network-connected group and can return the results to a central location for the analyst to review. Thus, the analyst has a set of rule results. In some embodiments, the results can be organized by the rule from which they originated. For example, the analyst can be presented with 20 rule results and, once approved, one or more specific rule numbers, such as rules for collecting, inspecting, and reviewing all the various services installed on a host, can be selected. The analyst can filter the results in order to be presented with more results. Thus, they can apply a filter, for example, they can search for a specific host, a specific username, internal, external IP addresses, or they can order it based on criteria such as ordering the results by rarity or whether any annotation exists. Some of the rules can return all the files available on a host, and these files can be scanned by a set of virus scanners and can result in annotations such as whether any known malware exists. Alternatively, the analyst can focus more specifically on one or more of the rule results.

[0009] These rule results can be organized to enable an analyst to click on one of the rule results for further details and make a determination as to whether there is evidence of infringement and whether to tag the result. If they find that it is evidence of infringement, they will assign a malware tag to it. If they find that it is just commercial software, they can apply an appropriate tag to it. Additionally, if the analyst finds evidence of infringement, they may perform a pivot. For example, if the rule result is related to a particular computer, they may obtain further information about that computer. In other embodiments, the analyst may pivot by time, by username, towards different hosts that were or are connected to the relevant computer, or towards other relevant results.

[0010] The system can record the actions of the analyst, which can also be described as workflow data, and these actions can be organized into a playbook or a set of analyzed actions. A playbook can be constructed from multiple rules, including any rules selected for a scan and one or more pivot sequences. A playbook can include a tagging classifier that is trained based on the actions observed when the analyst tags. In one embodiment, the deployed tagging classifier can be trained using supervised learning techniques with a group of labeled or tagged results. Starting from a binary classification that distinguishes malicious from non-malicious, the system can focus on all rule results tagged as malicious and all rule results tagged as non-infected, and then provide the tagged rule results as supervised learning input. If an unknown rule result is given or if a rule result is not tagged, the classifier can automatically tag it.

[0011] In some embodiments, the playbook can filter and reduce the number of results that an analyst focuses on. This filtering can be ordered by what is most likely to be clicked on by the analyst. For example, the system can have a list of all rule results that the analyst clicked on, and all those that they did not actually use, and a record of the filters they used. For example, if an analyst received one million results, the system can monitor which filters were applied, whether any regular expressions of any type were used during filtering, whether any ranking was applied, which results the analyst pivoted on, whether the pivot was by host name or user name, and which pivot sequences were followed.

[0012] To execute the playbook, a scan can be launched to obtain a set of rule results. All rule results are logs derived from the scan. Those rule results are fed through tagged classifiers that exist within the playbook and can provide a set of tags. The resulting tags provide a list of malicious results and a list of evidence of infringement. Another output of the playbook can be a scrutiny classifier that can predict which rule results are likely to be focused on by the analyst. For example, from one million scan results, the scrutiny classifier can provide 1,000 results that are likely to be focused on. Yet another output of the playbook can be a ranking of criteria that an analyst may wish to scrutinize the results based on, for example, which filters and ranking methods were applied. For example, one million results can be ranked in order of predicted importance to the analyst. Yet another output of the playbook can be a pivot sequence, which can provide a list of pivot chains when launched on the results. Thus, from the playbook, part of the output or scan is fully automated, such as processing the rule results through a selected tag classifier for a set of evidence of infringement.

[0013] In some embodiments, each of the operations described herein can be configured to be launched in parallel to generate a series of results that can be used to create scripts based on the learned behavior / actions from the observed analysts, and these scripts can be used as part of a threat hunting playbook or a set of rules for detecting security threats at an early stage and applied to future incoming security information / data / logs.

[0014] The various objectives, features, and advantages of the present disclosure will become apparent to those skilled in the art upon a review of the following detailed description when taken in conjunction with the accompanying drawings. This specification also provides, for example, the following items. (Item 1) A method for dynamically training a security threat detection system, monitoring security analyst workflow data from one or more security analysts analyzing a scan of security logs, the workflow data including one or more rules applied to security log scan results, rule results selected for further analysis, tags applied to the rule results, filters applied to the rule results, rankings applied to the rule results, or one or more actions associated with a pivot by the one or more security analysts, and / or combinations thereof, training a tagging classifier based on the tags assigned to the rule results from the workflow data, training an auditing classifier based on the rule results selected for the further analysis, training a filter and ranking method based on filters and rankings applied to the rule results from one or more security analysts, generating an automated threat hunting playbook including the tagging classifier, the auditing classifier, and the filter and ranking method, generating one or more scripts for automatically analyzing incoming security data using the automated threat hunting playbook comprising a method. (Item 2) training a pivot sequence model based on actions performed by the one or more security analysts, the automated threat hunting playbook also including the pivot sequence model, further comprising the method according to item 1. (Item 3) The method according to item 2, wherein the tagging classifier, the auditing classifier, the filter and ranking method, and the pivot sequence are each supervised machine learning models trained based on the workflow data of one or more security analysts. (Item 4) The method according to item 3, wherein the one or more scripts for automatically analyzing incoming security data generate a plurality of tags, and each tag is an indicator of an intrusion within a computer network. (Item 5) Receiving tag updates from one or more security analysts, dynamically updating the tagging classifier based on the tag updates The method according to item 4, further comprising. (Item 6) The method according to item 3, wherein the one or more scripts for automatically analyzing incoming security data generate a selected item of the result regarding the audit. (Item 7) Receiving analyst feedback regarding the selected item of the result regarding the audit, dynamically updating the audit classifier based on the analyst feedback regarding the selected item of the result regarding the audit The method according to item 6, further comprising. (Item 8) The method according to item 3, wherein the one or more scripts for automatically analyzing incoming security data generate a selected item of the prioritized result. (Item 9) Receiving analyst feedback regarding the selected item of the prioritized result, dynamically updating the filtering and ranking method based on the analyst feedback regarding the selected item of the prioritized result The method according to item 8, further comprising. (Item 10) The method according to item 3, wherein the one or more scripts for automatically analyzing incoming security data generate one or more pivot chains, and the pivot chain is a series of rule results for tracking potential attacks. (Item 11) Receiving pivot chain feedback from one or more security analysts, dynamically updating the pivot sequence model based on the pivot chain feedback The method according to item 10, further comprising. (Item 12) A dynamically trained threat detection system, One or more computing systems configured to monitor and store security analyst workflow data from one or more security analysts analyzing security log scans, wherein the workflow data includes rules applied to security log scan results, rule results selected for further analysis, tags applied to rule results, filters applied to rule results, rankings applied to rule results, or one or more actions associated with a pivot by the one or more security analysts, and / or combinations thereof, one or more computing systems, and A tagging classifier trained based on the tags assigned to the rule results from the workflow data, and An inspection classifier trained based on the rule results selected for the further analysis, and A filter and ranking method trained based on the filters and rankings applied to the rule results from the one or more security analysts, and An automated threat hunting playbook including the tagging classifier, the inspection classifier, and the filter and ranking method, wherein the automated threat hunting playbook is configured to generate one or more scripts for automatically analyzing incoming security data, an automated threat hunting playbook A system comprising. (Item 13) A pivot sequence model based on actions performed by one or more security analysts, wherein the pivot chain is a series of rule results tracking potential attacks, and the automated threat hunting playbook also includes the pivot sequence model, a pivot sequence model The system according to item 12, further comprising. (Item 14) The system according to item 13, wherein the tagging classifier, the inspection classifier, the filter and ranking method, and the pivot sequence are each a supervised machine learning model trained based on the workflow data of one or more security analysts. (Item 15) The system according to item 13, wherein the pivot sequence model is configured to generate one or more pivot chains when applied to raw scan data from a security log. (Item 16) The system according to item 12, wherein the tagging classifier is configured to generate a plurality of tags when applied to raw scan data from a security log, and each tag is an indicator of an intrusion within a computer network. (Item 17) The system according to item 12, wherein the auditing classifier is configured to generate a selected item of results regarding auditing when applied to raw scan data from a security log. (Item 18) The system according to item 12, wherein the filtering and ranking method is configured to generate a selected item of prioritized results when applied to raw scan data from a security log. (Item 19) A system for dynamically training a security threat detection system, comprising one or more processors and at least one memory and, wherein the at least one memory stores instructions therein, and when the instructions are executed by the one or more processors, the system is caused to monitor and record workflow data from one or more security analysts analyzing security logs in a computer network, the workflow data including rules applied to security log scan results, rule results selected for further analysis, tags applied to rule results, filters applied to rule results, rankings applied to rule results, or one or more actions associated with pivots by the one or more security analysts, and / or combinations thereof, train a tagging classifier based on the tags applied to the rule results from the workflow data, train an auditing classifier based on the rule results selected for further analysis by the one or more security analysts, train a filtering and ranking method based on the filters and rankings applied to the rule results from the one or more security analysts. Training a pivot sequence model based on actions performed by one or more security analysts, wherein the tagging classifier, the scrutiny classifier, the filter and ranking method, and the pivot sequence are each a supervised machine learning model trained based on the workflow data of one or more security analysts, and generating an automated threat hunting playbook including the tagging classifier, the scrutiny classifier, and the filter and ranking method A system that causes it to perform. (Item 20) The instructions further cause the system to Use the one or more scripts to analyze incoming security data and generate a plurality of tags, a selected item of results regarding scrutiny, a selected item of prioritized results, and one or more pivot chains, wherein each tag is an indicator of an intrusion within a computer network, and a pivot chain is a series of rule results for tracking potential attacks, and Receiving analyst feedback regarding the plurality of tags, the selected items of results regarding scrutiny, the selected items of prioritized results, and the one or more pivot chains; Dynamically updating the tagging classifier, the scrutiny classifier, the filter and ranking method, and the pivot sequence model based on the analyst feedback The system according to item 19, which causes it to perform.

Brief Description of the Drawings

[0015] For simplicity and clarity of illustration, it will be understood that the elements illustrated in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements. Embodiments incorporating the teachings of the present disclosure are shown and described herein with respect to the drawings.

[0016]

Figure 1

[0017]

Figure 2

[0018]

Figure 3

[0019]

Figure 4

[0020]

Figure 5

[0021] Detailed Description The following description, in combination with the figures, is provided to assist in understanding the teachings disclosed herein. This description focuses on specific implementations and embodiments of the present teachings and is provided to assist in explaining the present teachings. This focus should not be construed as a limitation on the scope or availability of the present teachings.

[0022] In one embodiment, the present disclosure relates to a system for automated threat detection that learns a threat hunting playbook or threat detection sequence based on analyzing the behavior of human security experts during threat hunting. In some embodiments, the system can include a supervised machine learning (ML) algorithm trained on the behavior and workflow of trained security analysts, which can automatically detect malicious attacks. Such a threat detection system can also increase the efficiency of analysts during threat hunting by allowing analysts to focus their time on the information most likely to be associated with malicious activities or by discovering new attack techniques or new suspicious behavior.

[0023] As used herein, a host describes one or more than one computer within an organization or group that is scanned as part of threat hunting.

[0024] As used herein, threat hunting describes a process for forensic analysis of criminalistics information to search for evidence of malicious attacks.

[0025] As used herein, a rule describes a detection rule that includes logic or programming that is stored on and / or executable on a host machine, and that is configured to seek or identify patterns associated with potentially malicious behavior.

[0026] As used herein, a rule result describes information collected from a host machine when a detection rule finds at least one indicator of potentially malicious activity.

[0027] As used herein, a scan describes the execution of a series of detection rules on a host to collect information about potentially malicious activity as part of threat hunting.

[0028] As used herein, a false positive describes a rule result from a detection rule that is determined not to be associated with malicious activity.

[0029] As used herein, a true positive describes a rule result from a detection rule that is proven to be associated with an actual instance of malicious activity.

[0030] Threat hunting is a process by which one or more security analysts examine available evidence such as security logs and the output of threat detection rules to determine whether a security breach has occurred, where it occurred, and how it occurred. Threat hunting often requires hours of attention from highly skilled security analysts. According to embodiments of the present disclosure, an automated threat detection system can learn from the threat hunting behavior of security experts to automatically detect potential intrusion events and also to generate a prioritized list of potential suspicious activities in order to streamline the threat detection process.

[0031] FIG. 1 shows a schematic diagram of a workflow 100 for identifying attack patterns or suspicious activities according to one aspect of the present disclosure. According to this embodiment, threat hunting begins by selecting rules for scan 101 and initiating scan 103. The detection rules can be selected to be launched on a host computing system and can include, for example, tactical rules designed to detect specific types of attacks and strategic rules designed to pick up generalized patterns. The scan results can, in some embodiments, be uploaded to a database and then can be evaluated by an analyst performing threat hunting.

[0032] Not all rule results are positive detections for malicious attacks. Strategic rules can be designed to pick up broad events or file artifacts and find traces left behind by malicious attacks, but can also include a large number of results related to legitimate use. For large organizations, a scan can return tens of millions of results. Searching through these results for attacks generally requires both time and skill.

[0033] Workflow 100 can continue in operation 105, which involves selecting rule results. For each rule result, in some embodiments, there is a vector of columns. The columns vary for different types of rules. For example, a rule that returns installed services may include columns such as "Service Name", "Service Type", "Service Start Type", "Service Path", "Service Account", "Host Name", "IP Address", "Detection Timestamp", "Installation Timestamp", "Log", "Record Number", "Category", and "Event Description". The rule results can be annotated with additional information such as how many times similar results have been previously confirmed (on how many hosts and within how many organizations). If the result is from a file, it can be annotated with virus scan information. These annotations are added as additional columns appended to the rule results.

[0034] In one embodiment, threat hunting can be performed on a platform designed to search, view, and tag rule results. Once a rule result is selected, several different actions can be selected in operation 107. A non-exclusive list of types of actions that an analyst can perform on a threat hunting platform includes, for example, changing the sorting criteria (109), applying a filter (111), and viewing the results (113). Selecting a rule can include selecting a rule for browsing the returned scan results. Applying a filter can include applying a regular expression to one or more columns. Changing the sorting criteria can include sorting and / or ordering the results based on a particular column. Viewing the rule results can include viewing the results in more detail and examining the original log information returned by the detection rule.

[0035] In some embodiments, in response to reviewing the results at operation 113, the analyst may decide to apply a tag to the results at operation 115 and record whether the results indicate a truly malicious attack or a false positive. If a tag is to be applied, workflow 100 may include tagging the results (117). In some embodiments, the tagging scheme can be binary (e.g., "malicious" or "not infected"), or a list of categories (e.g., "commercial", "potentially undesirable", "adware", "potentially malicious", "commodity malware", "APT malware"). If a tag is not to be applied, or after the tag has been successfully applied, workflow 100 may continue to pivot to other results (119). For example, threat hunting may pivot at operation 119, continue to select pivot rule results (120), and then, at 113, review the rule results again. In one embodiment, threat hunting may pivot to results from other rules having the same or similar values for a particular attribute (such as a host, user, or source IP address).

[0036] Workflow 100 may continue to determine whether to continue browsing the current rule results at operation 121. If "yes", workflow 100 may return to selecting a different action at 107. If "no", workflow 100 may continue to determine whether to continue the analysis at operation 123. If "yes", workflow 100 may return to selecting different or new rule results at operation 105. If "no", workflow 100 ends.

[0037] In one embodiment, the threat hunting workflow 100 can be performed by one or more trained security analysts and can generate workflow data that can be used to train an automated threat detection system. The workflow data can include, for example, a list of security log scan results selected for further analysis. The workflow data can also include types of filters, rankings, sorting criteria, or tags applied to different types of results. The workflow data can also include one or more security log items or actions associated with a pivot by one or more security analysts, and which pivot sequences were executed. For example, if a rule result related to an intrusion event is associated with a particular computer, the pivot may involve obtaining additional information about that computer. The pivot may also include collecting additional information about the time when the suspicious activity occurred, or the username or different hosts connected to a particular computer.

[0038] FIG. 2 is a block diagram of a method for training a threat detection system according to one aspect of the present disclosure. In one embodiment, the threat detection system can execute an automated threat hunting playbook 201. The threat hunting playbook 201 can include, in some embodiments, several trained elements including a pivot sequence model 205 and rules 203, and a tagging classifier 207, an auditing classifier 209, a filter 211, and a ranking method 213. The playbook 201 can be constructed or taught based on workflow data collected by monitoring the actions of trained security analysts, such as the actions described above with reference to FIG. 1.

[0039] According to the embodiment shown in FIG. 2, the automated threat hunting playbook 201 includes a rule set 203, a tagging classifier 207, an auditing classifier 209, a filter 211, a ranking method 213, and a pivot sequence model 205.

[0040] The rule set 203 can include a set of detection rules that are to be triggered in a scan. In one embodiment, a clustering algorithm may be applied to obtain a rule set selected by a security analyst during threat hunting. In this way, the set of rules can be automatically maintained based on the workflow data discussed above.

[0041] The tagging classifier 207 can include a classifier that analyzes rule results and automatically tags them regardless of whether the results are malicious (positive detections) or false detections. Each rule can have one associated tagging classifier. In one embodiment, for each rule, a set of tagging results can be used to train the automated tagging classifier 207. This can be achieved, for example, using a supervised ML algorithm. What adds an annotation column to the values of the rule column can be used as an input feature vector for such an algorithm. Tags can be used as classification labels. Exemplary classification algorithms that can be used include, but are not limited to, instance-based classifiers, decision trees, decision forests, support vector machines, or neural networks. To ensure that the classifiers are not overfitted, in some embodiments, they may be trained on data obtained from multiple organizations. In some embodiments, the tags can be indicators of intrusions within a computer network.

[0042] The review classifier 207 can include a classifier that filters rule results to recommend a subset of the results to be manually reviewed by a security analyst. Each rule can have one associated review classifier. In some embodiments, for each rule, the review classifier 209 can be trained to automatically identify results that are of high interest to the analyst and should be manually reviewed. Since the analyst has time to view only a small subset of the full set of rule results, a log recording which results were viewed by the analyst can be used as a training set for the review classifier 209. In one embodiment, the feature vector for the training set can be rule column data and annotations. The classification label can be binary and can indicate whether the result was viewed by the analyst. Any classification-based supervised learning algorithm can be used to train the review classifier 209 such as those algorithms enumerated above with reference to training the tagging classifier 207.

[0043] The filter 211 can include a filter that is automatically applied to the rule results to reduce the set of results to be reviewed. The rules can have zero or multiple filters in various embodiments. Multiple filters can be applied in an AND combination (only results that satisfy all filters are retained) or an OR combination (results that satisfy any filter are retained). The ranking method 213 can include the order in which the rule results should be viewed, in order from the highest priority result to the lowest priority one. Each rule can have one associated ranking method. In one embodiment, the filter and the ranking method can be associated with the rule using, as a training set, the recorded filter and sort methods and columns from the security analyst's workflow data.

[0044] The pivot sequence model 205 can include an automated sequence of results to be viewed that tracks possible attacks. In one embodiment, the pivot sequence model can be constructed and trained using the actions of an analyst as a training set during threat hunting. The pivot sequence can include a series of actions taken when investigating potential breach events such as those recorded within the workflow data discussed above.

[0045] FIG. 3 is a block diagram of a method of applying a threat detection system to raw scan results according to one aspect of the present disclosure. In this embodiment, the raw scan results 301 can be applied to a trained tagging classifier 303, a scrutiny classifier 305, a filter 307, and a ranking method 309, and a pivot sequence model 311. The tagging classifier 3, scrutiny classifier 305, filter 307, and ranking method 309, and pivot sequence model 311 can be trained using one or more ML algorithms and based on the recorded actions and workflow data of trained security analysts, as discussed above with reference to FIG. 2. The output of the threat detection system can include, for example, automated tags 313, results 315 regarding manual scrutiny, prioritized results 317, and a pivot chain 319.

[0046] In one embodiment, running the automated threat detection system can include inputting the raw results 301 (i.e., the results from scan 103 described with reference to FIG. 1) into the trained tagging classifier 303 and generating automated tags 313. These automated tags 313 can be automatically generated according to the assigned maliciousness level using the tagging classifier 303. In one embodiment, the tagging classifier can scrutinize 10 million results and automatically generate 20 malicious tags. In such an example, this provides a limited number of potential breach events from the 10 million results scrutinized.

[0047] In one embodiment, executing an automated threat detection system may also include inputting the raw results 301 into the scrutiny classifier 305 to generate a selection of results regarding the scrutiny 315. These results can be a subset of the complete scan results and can include a set of results that should be manually scrutinized by a security analyst because they are likely to be related to malicious attacks. In one embodiment, the scrutiny classifier may automatically analyze 10 million results and provide a list of 10,000 results that are selected for further manual scrutiny. In such an example, these 10,000 results are what the scrutiny classifier 305 predicts and are the most important for the analyst to scrutinize.

[0048] In one embodiment, executing an automated threat detection system may also include inputting the raw results 301 into the filter 307 and the ranking method 309 to generate a list of prioritized results 317. In some embodiments, this master list of prioritized results 317 can include the complete set of scanned results that are filtered and ranked if a security analyst desires to manually inspect the results.

[0049] In one embodiment, executing an automated threat detection system may also include inputting the raw results 301 into the pivot sequence model 311 to generate a pivot chain. The pivot chain 319 can include a set of rule results that track potential attacks. In some embodiments, these results may be used for an automated solution or may be manually scrutinized by a security analyst.

[0050] As discussed herein, the actions recorded by an analyst can be used as a basis for creating automated threat hunting scripts and training various models and classifiers. In some embodiments, these processes, as illustrated in FIG. 3, each start in parallel and can generate tag 313, results 315 regarding manual review, prioritized results 317, and pivot chain 319.

[0051] FIG. 4 is a schematic diagram of a workflow 400 for dynamically training a threat detection system according to an aspect of the present disclosure. In some embodiments, the accuracy of an automated threat detection system can be continuously improved by retraining the model according to how an analyst reviews and corrects the results.

[0052] In one embodiment, workflow 400 begins with monitoring workflow data 401. As discussed above, the workflow data can include, for example, a list of security log scan results selected for further analysis. The workflow data can also include types of filters, rankings, sorting criteria, or tags applied to different types of results. The workflow data can also include one or more security log items or pivot chains associated with a pivot by a security analyst, and actions performed by the security analyst. In some embodiments, the pivot chain includes a series of rule results that track potential attacks.

[0053] This workflow data can then be used to train the ML models 403 discussed above, including tagging classifiers, review classifiers, filtering and ranking methods, and pivot sequence models. The training of these ML models is discussed in more detail with reference to FIG. 2.

[0054] Once the model is trained at 403, an automated threat hunting playbook, including a tagging classifier, a scrutiny classifier, filters, and a ranking method, can be generated at 405. The playbook can also include a pivot sequence model, as discussed above.

[0055] Workflow 400 continues at 407 to generate one or more scripts for automatically analyzing incoming security data using the automated threat hunting playbook. Once the playbook is generated, the trained models within the playbook can be applied at 409 to raw scan data to generate tags, selected items of results regarding scrutiny, ranked results, and / or a pivot chain. These operations, and the generation of the output of the ML model, are discussed in detail with reference to FIG. 3.

[0056] Workflow 400 continues at 411 to receive analyst feedback on the output of the ML model generated at 409. In some embodiments, the analyst feedback can include edits or changes to the automated tags generated by the tagging classifier. For example, when an analyst corrects a tag, or finds other results and assigns tags, this information can be used to further train the tagging classifier.

[0057] The analyst feedback can also include a list of the actual results scrutinized by the analyst from the results regarding scrutiny generated by the scrutiny classifier. For example, if the scrutiny classifier generates a focused or curated list of 2,000 results regarding scrutiny and the analyst only scrutinizes a subset of 800 results, this information can be stored for further training of the scrutiny classifier.

[0058] In some embodiments, the analyst feedback can include a list of results reviewed by the analyst from the ranked results generated by the filter and ranking method. For example, if the filter and ranking method sorts the results in a particular ranking or applies a particular filter and the analyst reviews the results in an order different from the automatically generated ranking, this information can be stored for further training of the filter and ranking method.

[0059] In some embodiments, the analyst feedback can include an alternative pivot sequence executed by the analyst that is different from the pivot chain generated by the pivot sequence model. If the analyst executes a pivot different from what is recommended by the pivot sequence model, this can be used as additional input for further training of the pivot sequence model.

[0060] The workflow can continue at 403 by applying the analyst feedback to the training of the ML method to dynamically update the model and increase the accuracy of the automated threat detection system.

[0061] For purposes of the present disclosure, an information handling system 80 (FIG. 5) may include any means or aggregation of means operable to calculate, compute, determine, classify, process, transmit, receive, read, devise, switch, store, display, communicate, manifest, detect, record, copy, handle, or utilize any form of information, intelligence, or data for commercial transactions, science, control, or other purposes. For example, the information handling system may be a personal computer (e.g., desktop or laptop), tablet computer, mobile device (e.g., personal digital assistant (PDA) or smartphone), server (e.g., blade server or rack server), network storage device, or any other suitable device, and may vary in size, shape, performance, functionality, and price. The information handling system may include random access memory (RAM), one or more processing resources such as a central processing unit (CPU) or hardware or software control logic, read only memory (ROM), and / or other types of non-volatile memory. Additional components of the information handling system may include one or more disk drives, one or more network ports for communicating with external devices, and various input and output (I / O) devices such as a keyboard, mouse, touch screen, and / or video display. The information handling system may also include one or more buses operable to transmit communications between the various hardware components.

[0062] As shown in FIG. 5, in some embodiments, client 12 can manage or otherwise include one or more network-connected systems 82 of information handling system / device 80 or other communicable systems / devices. Network 84 can provide data communication between information handling systems / devices 80, which can include workstations, personal computers, smart mobile phones, personal digital assistants, laptop computers, servers, and other suitable devices. Network 84 can include a private or public network such as a local area network, or the Internet or another wide area network, virtual personal network, peer-to-peer file sharing system, or other suitable network, and / or other suitable communication lines, or combinations thereof. FIG. 5 also shows that one or more monitoring devices 86 to which linked or network-connected information handling systems 80 are communicably coupled to network 84 can be included. Monitoring device 86 can be managed by a managed security service provider (MSSP).

[0063] In one embodiment, the monitoring device 86 may include a server or a sequence analyzer, or a computing device suitable for other clients having a processor and a memory or other suitable storage device. The memory can include random access memory (RAM), read only memory (ROM), and / or other non-transitory computer-readable media. The monitoring device 86 is further typically operable to store and execute computer-readable instructions and will continuously monitor activities, e.g., the activities of the information handling system 80 connected to the network 84, in each network-connected system in real time. The monitoring device 86 can capture or aggregate information or data logs related to the activities of the information handling system 80 by the automated threat detection system described herein and can provide these captured / aggregated data logs or information, or information related thereto. Additionally, or alternatively, the automated threat detection system described herein can include, for example, one or more servers 90 accompanied by at least one memory 92 and one or more processors 94 for receiving information or data logs related to the activities of the information handling system 80 of the system 82, including a data center 88 such as a data center 88 managed by an MSSP accompanied by a plurality of network-connected information handling systems 80. These information / data logs can be part of the raw logs 14 provided to the automated threat detection system described herein.

[0064] One or more components of the systems described herein may reside on or be accessible by a device 80, a server 90, or other devices or information handling systems in communication therewith. One or more processors of device 80, of one or more processors, may process or execute instructions, workflows, etc. stored in at least one memory (e.g., the memory of device 90 or memory 92) to facilitate performance of the various processes, functions, etc. of the automated threat detection system described herein.

[0065] The foregoing description generally illustrates and describes various embodiments of the present disclosure. However, it will be understood by those skilled in the art that various changes and modifications can be made to the above-discussed structures of the present disclosure without departing from the spirit and scope of the present disclosure as disclosed herein, and that all matters contained in the foregoing description or shown in the accompanying drawings are to be construed as illustrative and not in a limiting sense. Further, the scope of the present disclosure is to be construed to cover various modifications, combinations, additions, alterations, etc., and such should be considered to be within the scope of the present disclosure with respect to the embodiments described above. Thus, the various features and characteristics of the present disclosure as discussed herein may be selectively exchanged and applied to other illustrated and unillustrated embodiments of the present disclosure, and numerous variations, modifications, and additions can further be made thereto without departing from the spirit and scope of the invention as described in the appended claims.

Claims

1. A method for dynamically training a security threat detection system, comprising: monitoring security analyst workflow data from one or more security analysts analyzing a scan of security logs, said workflow data including one or more rules applied to security log scan results, rule results selected for further analysis, tags applied to rule results, filters applied to rule results, rankings applied to rule results, or one or more actions associated with a pivot by said one or more security analysts, and / or combinations thereof; training a tagging classifier based on the tags assigned to rule results from said workflow data; training a scrutiny classifier based on rule results selected for said further analysis; training a filter and ranking method based on filters and rankings applied to rule results from one or more security analysts; training a pivot sequence model based on actions performed by said one or more security analysts; generating an automated threat hunting playbook including said tagging classifier, said scrutiny classifier, said pivot sequence model, and said filter and ranking method; generating one or more scripts for automatically analyzing incoming security data using said automated threat hunting playbook; A method comprising the above.

2. The method according to claim 1, wherein the tagging classifier, the scrutiny classifier, the filter and ranking method, and the pivot sequence model are each supervised machine learning models trained based on the workflow data of one or more security analysts.

3. The method according to claim 2, wherein the one or more scripts for automatically analyzing incoming security data generate a plurality of tags, each tag being an indicator of an intrusion within a computer network.

4. Receiving tag updates from one or more security analysts; dynamically updating the tagging classifier based on the tag update The method according to claim 3, further comprising this.

5. The method according to claim 2, wherein the one or more scripts for automatically analyzing incoming security data generate selected items of results regarding the scrutiny.

6. receiving analyst feedback regarding the selected items of results regarding the scrutiny dynamically updating the scrutiny classifier based on the analyst feedback regarding the selected items of results regarding the scrutiny The method according to claim 5, further comprising this.

7. The method according to claim 2, wherein the one or more scripts for automatically analyzing incoming security data generate selected items of prioritized results.

8. receiving analyst feedback regarding the selected items of the prioritized results dynamically updating the filtering and ranking method based on the analyst feedback regarding the selected items of the prioritized results The method according to claim 7, further comprising this.

9. The method according to claim 2, wherein the one or more scripts for automatically analyzing incoming security data generate one or more pivot chains, and the pivot chain is a series of rule results for tracking potential attacks.

10. receiving pivot chain feedback from one or more security analysts dynamically updating the pivot sequence model based on the pivot chain feedback The method according to claim 1, further comprising this.

11. A dynamically trained threat detection system One or more computing systems configured to monitor and store security analyst workflow data from one or more security analysts analyzing security log scans, the workflow data including rules applied to security log scan results, rule results selected for further analysis, tags applied to rule results, filters applied to rule results, rankings applied to rule results, or one or more actions associated with a pivot by the one or more security analysts, and / or combinations thereof, one or more computing systems, and A tagging classifier trained based on the tags assigned to rule results from the workflow data, An auditing classifier trained based on the rule results selected for the further analysis, A pivot sequence model trained based on actions performed by one or more security analysts, A filter and ranking method trained based on the filters and rankings applied to rule results from one or more security analysts, An automated threat hunting playbook including the tagging classifier, the auditing classifier, the pivot sequence model, and the filter and ranking method, the automated threat hunting playbook configured to generate one or more scripts for automatically analyzing incoming security data, an automated threat hunting playbook and Comprising a system.

12. The system according to claim 11, wherein the pivot sequence model generates a pivot chain, and the pivot chain is a series of rule results that track potential attacks.

13. The system according to claim 11, wherein the tagging classifier, the auditing classifier, the filter and ranking method, and the pivot sequence model are each a supervised machine learning model trained based on the workflow data of one or more security analysts.

14. The system according to claim 11, wherein when the pivot sequence model is applied to raw scan data from a security log, it is configured to generate one or more pivot chains.

15. The system according to claim 11, wherein when the tagging classifier is applied to raw scan data from a security log, it is configured to generate a plurality of tags, and each tag is an indicator of an intrusion within a computer network.

16. The system according to claim 11, wherein when the scrutiny classifier is applied to raw scan data from a security log, it is configured to generate a selected item of the results regarding the scrutiny.

17. The system according to claim 11, wherein when the filter and ranking method is applied to raw scan data from a security log, it is configured to generate a selected item of the prioritized results.

18. A system for dynamically training a security threat detection system, comprising one or more processors and at least one memory and, the at least one memory stores instructions therein, and when the instructions are executed by the one or more processors, the system is caused to monitor and record workflow data from one or more security analysts analyzing security logs within a computer network, the workflow data including rules applied to security log scan results, rule results selected for further analysis, tags applied to the rule results, filters applied to the rule results, rankings applied to the rule results, or one or more actions associated with pivots by the one or more security analysts, and / or combinations thereof, train a tagging classifier based on the tags applied to the rule results from the workflow data, train a scrutiny classifier by one or more security analysts based on the rule results selected for the further analysis, train a filter and ranking method based on the filters and rankings applied to the rule results from one or more security analysts. Training a pivot sequence model based on actions performed by one or more security analysts, wherein the tagging classifier, the scrutiny classifier, the filter and ranking method, and the pivot sequence model are each a supervised machine learning model trained based on the workflow data of one or more security analysts, and, Generating an automated threat hunting playbook including the tagging classifier, the scrutiny classifier, the pivot sequence model, and the filter and ranking method A system for causing to perform.

19. The instructions further cause the system to Use the threat hunting playbook to analyze incoming security data and generate a plurality of tags, a selection of results related to scrutiny, a selection of prioritized results, and one or more pivot chains, wherein each tag is an indicator of an intrusion within a computer network, and a pivot chain is a series of rule results for tracking potential attacks, and, Receive analyst feedback regarding the plurality of tags, the selection of results related to scrutiny, the selection of prioritized results, and the one or more pivot chains Dynamically update the tagging classifier, the scrutiny classifier, the filter and ranking method, and the pivot sequence model based on the analyst feedback The system according to claim 18, for causing to perform.

20. The system according to claim 19, wherein the analyst feedback includes an alternative pivot sequence performed by the analyst that is different from the pivot chain generated by the pivot sequence model.

21. The pivot by the one or more security analysts is one or more of the pivots based on a time-based, username-based, followed pivot sequence to different hosts connected to a computer at a specific time, or towards relevant results, according to the system of claim 18.

22. The system according to claim 21, wherein the pivot sequence includes a series of actions taken when investigating potential intrusion events.

23. The system according to claim 18, wherein the workflow data is monitored and recorded during threat hunting, and the threat hunting includes a process for forensic analysis of criminal science information for searching for evidence of malicious attacks.

Citation Information

Patent Citations

  • Malware detection systems and methods

    JP2012501504A

  • Data processing apparatus, data processing method, and program

    JP2016024506A

  • Cyber terrorism security simulator of nuclear power plant

    JP2017198836A

  • Analyzer and method for analysis

    JP2020113216A

  • System and method for the detection of malware

    US20100058474A1