Dynamic control of cyber-threat hunting pace
A dynamic cyber-threat hunting system using heuristic and reinforcement learning optimizes hunting pace based on real-time environmental factors, addressing resource constraints and evolving threats to enhance threat identification efficiency and resource utilization.
Patent Information
- Application Number
- PCT/IB2024/057377
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-05
AI Technical Summary
Cyber-threat hunting processes in Security Operations Centers (SOCs) face challenges due to limited resources, fluctuating constraints, and the continuous evolution of the threat landscape, leading to inefficiencies in resource allocation and timely threat identification, with existing automation methods failing to account for dynamic changes in threat landscapes and resource availability.
A system that dynamically controls the hunting pace of cyber-threat hunting processes using a combination of heuristic and reinforcement learning to optimize quality and minimize costs, adjusting the pace based on real-time environmental factors such as security event importance, volume, and resource availability.
The system ensures efficient use of resources, maintains a proactive security posture, and adapts to evolving threats by dynamically adjusting the hunting pace, enhancing the SOC's ability to identify threats promptly and accurately.
Smart Images

Figure IB2024057377_05022026_PF_FP_ABST
Abstract
Description
DYNAMIC CONTROL OF CYBER-THREAT HUNTING PACE Technical Field
[0001] The present disclosure relates to cyber-threat hunting. Background
[0002] In the dynamic and ever-evolving landscape of cybersecurity, threat hunting has emerged as a pivotal proactive defense line practiced in identifying sophisticated attacks that stealthily bypass traditional detection mechanisms. According to a recent interview with threat hunters at the U.S. Department of Homeland Security (see W. P. M. Iii and J. C. Davis, “An Interview Study on Third-Party Cyber Threat Hunting Processes in the U.S. Department of Homeland Security,” in USENIX Security Symposium, 2024, which is hereinafter referred to as the “Interview Study on Third- party Cyber Threat Hunting Processes”), threat hunting is recognized as a crucial capability within cybersecurity for both security practitioners and academia. In fact, threat hunting is mandated for Federal Civilian Executive Branch Agencies, such as the National Aeronautics and Space Administration (NASA) and the Department of the Treasury.
[0003] Figure 1 illustrates the process of threat hunting based on the recent Interview Study of Third-Party Cyber Threat Hunting Processes. As illustrated in Figure 1, the threat hunting team needs to plan the mission and prepare all hunting activities by using internal and external threat intel (e.g., collect threat intel, trends, Indicator of Compromise (IoC), Indicator of Attack (IoA), Tactics, Techniques, and Procedures (TTPs)). The hunting process is a cumbersome task where security experts have to manually scrutinize a tedious number of events and logs to generate threat hypotheses. Automated alert analysis tools might also be used to assist the threat hunter. Those hypotheses are then investigated to validate the attack's existence before it is fully materialized (see F. K. Kaiser et al., “Attack Hypotheses Generation Based on Threat Intelligence Knowledge Graph,” IEEE Trans. Dependable Secur. Comput., pp. 1–17, 2023, doi: 10.1109 / TDSC.2022.3233703, which is hereinafter referred to as the “Kaiser Paper”). The task of threat hunting is intricate; it demands extensive knowledge and substantial time and effort making it a challenging endeavor for the Security Operations Center (SOC).Summary
[0004] Systems and methods are disclosed for controlling a hunting pace of a cyber- threat hunting process. In one embodiment, a computer-implemented method comprises obtaining information indicative of a current state of a cyber-threat hunting process, selecting a hunting pace for the cyber-threat hunting process based on the current state of the cyber-threat hunting process and a machine learning scheme for hunting pace selection that maximizes a quality of the cyber-threat hunting process while minimizing a cost of operation and processing time for the cyber-threat hunting process, and configuring the cyber-threat hunting process with the selected hunting pace. The method further comprises repeating the steps of obtaining and selecting, responsive to a change in the current state of the cyber-threat hunting process, thereby dynamically controlling the hunting pace for the cyber-threat hunting process. In this manner, the hunting pace of the cyber-threat hunting process is dynamically controlled in a manner that optimizes the quality of the cyber-threat hunting process while minimizing the cost of operation and processing time.
[0005] In one embodiment, the machine learning scheme is a reinforcement learning scheme that uses states, actions, and a reward function. The states represent different states of the cyber-threat hunting process and include one or more metrics indicative of or otherwise related to the quality of the cyber-threat hunting process, the cost of operation for the cyber-threat hunting process, and the processing time for the cyber- threat hunting process. The actions represent selection of different hunting paces, and the reward is a function that quantifies a degree to which use of a particular hunting pace maximizes the quality of the cyber-threat hunting process while minimizing a cost of operation and processing time for the cyber-threat hunting process. In one embodiment, selecting the hunting pace for the cyber-threat hunting process based on the current state of the cyber-threat hunting process and a machine learning scheme comprises, for each hunting pace from among at least a subset of a set of possible hunting paces, computing an action-value that represents an expected cumulative reward for using that hunting pace given the current state of the cyber-threat hunting process. Selecting the hunting pace for the cyber-threat hunting process based on the current state of the cyber-threat hunting process and a machine learning scheme further comprises selecting one of the set of possible hunting paces, based on the computed action-values, as a candidate hunting pace to be used for the cyber-threathunting process and determining an expected maximum cumulative reward for selecting the candidate hunting pace as a new hunting pace for the cyber-threat hunting process when in the current state. In one embodiment, selecting the hunting pace for the cyber-threat hunting process further comprises determining that the expected maximum cumulative reward for selecting for selecting the candidate hunting pace as a new hunting pace for the cyber-threat hunting process when in the current state is greater than or equal to a predefined or configured reward threshold and, responsive to the expected maximum cumulative reward being greater than or equal to the predefined or configured threshold, selecting the candidate hunting pace as the hunting pace for the cyber-threat hunting process.
[0006] In one embodiment, the current state of the cyber-threat hunting process comprises the current quality of the cyber-threat hunting process or one or more parameters from which the current quality of the cyber-threat hunting process can be derived, the cost of operation for the cyber-threat hunting process for a current hunting pace, average security event arrival time, and security event importance, or one or more parameters from which the current quality of the cyber-threat process can be derived, the processing time for the cyber-threat hunting process at the current hunting pace, an average security event arrival time, and the current hunting pace.
[0007] In one embodiment, the quality of the cyber-threat hunting process and the cost of operation for the cyber-threat hunting process are each a function of average security event arrival time for a set of security events, a security event importance computed for the set of security events, and a current hunting pace used for the cyber- threat hunting process. In one embodiment, the quality of the cyber-threat hunting process is defined as: ^^^^^, ^, ^^ℋ ^, ^, ^^ =^^^^, ^, ^^ + ^^^^, ^, ^^where ℋ^^, ^, ^^ is the quality of the cyber-threat hunting process for a given averagearrival rate ^, security event importance ^, and hunting pace ^, ^^ represents true positives for the hunting pace ^ for the average security event arrival rate ^ with the security event importance ^, and ^^ represents false positives for the hunting pace ^ for the average security event arrival rate ^ with the security event importance ^.
[0008] In one embodiment, the cost of operation for the cyber-threat hunting process is defined as: ^^^, ^, ^^ = ^^^^^^^, ^, ^^ + ^^^^^^, ^, ^^where ^^^, ^, ^^ is the cost of operation for the cyber-threat hunting process for a givenaverage security event arrival rate ^, security event importance ^, and hunting pace ^,^^^^^^^, ^, ^^ is a computational cost needed to perform cyber-threat hunting at thehunting pace of ^ for security events received with the average event arrival rate ^ andsecurity event importance ^, and ^^^^^^, ^, ^^ is an overall manpower cost needed toperform cyber-threat hunting at the hunting pace of ^ for security events received with the average event arrival rate ^ and the security event importance ^.
[0009] In one embodiment, the method further comprises, prior to the aforementioned steps, selecting an initial hunting pace for the cyber-threat hunting process via a heuristic process. In one embodiment, selecting the initial hunting pace for the cyber-threat hunting process comprises, for each candidate hunting pace from a set of candidate hunting paces, determining a quality of the cyber-threat hunting process for the candidate hunting pace for an average arrival rate of a set of received security events and an importance of the set of received security events, a cost of operation of the cyber-threat hunting process for the candidate hunting pace for the average arrival rate of the set of received security events and the importance of the set of received security events, and a processing time needed for the cyber-threat hunting process for the candidate hunting pace for the average arrival rate of the set of received security events and the importance of the set of received security events. Selecting the initial hunting pace for the cyber-threat hunting process further comprises selecting the candidate hunting pace that maximizes the quality of the cyber-threat hunting process while minimizing the cost of operation and processing time, as the initial hunting pace.
[0010] Corresponding embodiments of a computing system are also disclosed. In one embodiment, a computing system comprises processing circuitry and memory storing software instructions executable by the processing circuitry whereby the computing system is configured to obtain information indicative of a current state of a cyber-threat hunting process, select a hunting pace for the cyber-threat hunting process based on the current state of the cyber-threat hunting process and a machine learning scheme for hunting pace selection that maximizes a quality of the cyber-threat hunting process while minimizing a cost of operation and processing time for the cyber-threathunting process, configure the cyber-threat hunting process with the selected hunting pace, and repeat the steps of obtaining and selecting, responsive to a change in the current state of the cyber-threat hunting process, thereby dynamically controlling the hunting pace for the cyber-threat hunting process.
[0011] In one embodiment, a non-transitory computer-readable medium is provided that stores software instructions executable by processing circuitry of a computing system, whereby the computing system is caused to obtain information indicative of a current state of a cyber-threat hunting process, select a hunting pace for the cyber- threat hunting process based on the current state of the cyber-threat hunting process and a machine learning scheme for hunting pace selection that maximizes a quality of the cyber-threat hunting process while minimizing a cost of operation and processing time for the cyber-threat hunting process, configure the cyber-threat hunting process with the selected hunting pace, and repeat the steps of obtaining and selecting, responsive to a change in the current state of the cyber-threat hunting process, thereby dynamically controlling the hunting pace for the cyber-threat hunting process. Brief Description of the Drawings
[0012] The accompanying drawing figures incorporated in and forming a part of this specification illustrate several aspects of the disclosure, and together with the description serve to explain the principles of the disclosure.
[0013] Figure 1 illustrates the process of threat hunting based on the recent Interview Study of Third-Party Cyber Threat Hunting Processes;
[0014] Figures 2A, 2B, and 2C illustrate a motivating example;
[0015] Figure 3 illustrates a system that implements an embodiment of the proposed solution to adaptively determine an optimal hunting pace for effective threat hunting in a dynamic environment;
[0016] Figure 4 is workflow diagram that illustrates the operation of the Initial Optimization Agent and the Dynamic Optimization Agent of Figure 3, in accordance with an embodiment of the present disclosure;
[0017] Figure 5 illustrates one example of the grid search of steps 402, 404, and 406 of Figure 4;
[0018] Figure 6 illustrates another embodiment of the operation of the Dynamic Optimization Agent of Figure 3;
[0019] Figure 7 shows the results of the adaptive change of hunting pace (^) according to the change of event arrival rate and security event importance;
[0020] Figure 8 shows the results of the effect of hunting pace (^) on the quality of threat hunting, cost of operation, and processing time;
[0021] Figure 9 shows the results of the effect of events arrival rate (^) on the quality of hunting, cost of operation, and processing time; and
[0022] Figure 10 is a schematic block diagram of a computing system in which the Initial Optimization Agent and / or the Dynamic Optimization Agent of Figure 3 may be implemented, according to an example embodiment of the present disclosure. Detailed Description
[0023] The embodiments set forth below represent information to enable those skilled in the art to practice the embodiments and illustrate the best mode of practicing the embodiments. Upon reading the following description in light of the accompanying drawing figures, those skilled in the art will understand the concepts of the disclosure and will recognize applications of these concepts not particularly addressed herein. It should be understood that these concepts and applications fall within the scope of the disclosure.
[0024] The existing cyber-threat hunting solutions have a number of challenges. The Security Operations Center (SOC) often grapples with the daunting challenge of: (i) limited resources / manpower; (ii) constraints that fluctuate unpredictably, such as geopolitical intentions, and (iii) continuous evolution of threat landscape leading to new attack techniques (see M. Vielberth, F. Böhm, I. Fichtinger, and G. Pernul, “Security operations center: A systematic study and open challenges,” IEEE Access, vol. 8, pp. 227756–227779, 2020, which is hereinafter referred to as the “Vielberth Paper”). This variability is compounded by the sheer volume and diversity of events that SOC must monitor and analyze. On any given day, the number of events streaming into a SOC can range from a few thousand to several million. Micro Focus reported that SOC receives millions of alerts (20,000 events per second) that need to be deeply investigated by security experts (see “Technical Requirements for the ArcSight Platform,” 2021, micro Focus ArcSight Platform). Each event potentially leads to critical security implications and high-priority incidents (see “APT41: A Dual Espionage and Cyber Crime Operation,” Mandiant, Tech. Rep., 2022 and G. C. A. Team, “Threat Horizons - April 2023 ThreatHorizons Report,” Google Cloud, Tech. Rep., 2023, online: https: / / services.google.com / fh / files / blogs / gcat_threathorizons_full_apr2023.pdf). Such a vast and variable influx of data and scarcity of resources necessitates a strategic approach to timing the hunting. This approach would navigate the delicate balance between the thoroughness of the hunting activity, the efficient allocation of scarce resources, and the urgency of timely attack identification.
[0025] The Interview Study of Third-Party Cyber Threat Hunting Processes in the U.S. Department of Homeland Security revealed that one of the main obstacles to threat hunting efficiency is the lack of automation. Toward this, various efforts have been presented to automate the process of threat hunting, encompassing a wide range of approaches from hypothesis generation (see e.g., B. Nour, M. Pourzandi, R. K. Qureshi, and M. Debbabi, “AUTOMA: Automated Generation of Attack Hypotheses and Their Variants for Threat Hunting Using Knowledge Discovery,” IEEE Trans. Netw. Serv. Manag., 2024, doi: 10.1109 / TNSM.2024.3378972, which is hereinafter referred to as the “first Nour Paper”, and the Kaiser Paper) and the creation of new testflows (e.g., B. Nour, M. Pourzandi, R. K. Qureshi, and M. Debbabi, “Accurify: Automated New Testflows Generation for Attack Variants in Threat Hunting,” in Foundations and Practice of Security (FPS), 2023, which is hereinafter referred to as the “second Nour Paper”, and “The Threat Hunter’s Hypothesis - A case for structured threat hunting and how to make it work in the real world.” 2023. [Online]. Available: https: / / www.cyborgsecurity.com / library / guides / the-threat-hunters-hypothesis-2 / , Access March 2024, which is hereinafter referred to as the ”Threat Hunter’s Hypothesis”), to the extraction of attack story (e.g., A. Alsaheel et al., “ATLAS: A sequence-based learning approach for attack investigation,” in USENIX Security Symposium, 2021, pp. 3005–3022, which is hereinafter referred to as the “Alsaheel Paper”). However, these methods are typically guided by the judgment of security experts, neglecting the critical consideration of available resources and the ever- changing threat landscape. This oversight can result in overlooking alerts or missing critical events.
[0026] In a similar vein, numerous industrial security tools, including Splunk's PEAK, Sqrrl, TaHiTTI, and TheHive Project, strive to enhance threat hunting methodology by incorporating automation into certain aspects of the process. Despite these advancements, the effective execution of threat hunting still demands considerableexpertise and effort from security analysts, without adequately accounting for the strategic timing of hunting activities in alignment with hunting performance, dynamic change of threats, security event performance, event volumes, and the available resources.
[0027] Motivating Example: Cyber defense can operate with a dual-layered defense approach: police patrols and police controls. Police controls are tasked with monitoring and controlling access serving as the first line of defense. On the other hand, police patrols deploy patrols within the camp, whose role is to actively search for and neutralize any adversaries that might have penetrated the initial defenses (see the Interview Study on Third-Party Cyber Threat Hunting). The police controls act as SOC's Threat Detection and Monitoring and the police patrol function is comparable to the process of threat hunting. The scheduling of patrol rotations directly influences the effectiveness of identifying adversaries. Therefore, the strategic timing of threat hunting activities (i.e., hunting pace) within the SOC necessitates a nuanced understanding of the interplay between the change in threat landscape, importance of security events, volume of incoming security events, available hunting resources, and the density of security analyst coverage. Figures 2A, 2B, and 2C illustrates a motivating example for such a problem. Indeed, real-time hunting (Figure 2A) – where the hunting is conducted for each event, is characterized by frequent and rapid investigations. This approach proves advantageous when event volumes are low and there is an abundance of both security analysts and computational resources which in return facilitates meticulous event scrutiny without the risk of overload. However, this approach can falter under high event volumes and resource constraints, leading to compromised hunting quality, alert fatigue, resource exhaustion, and an uptick in false positives as the security team faces pressure to manage a higher volume of alerts. Batch hunting (Figure 2B) – where the hunting is conducted at a longer pace, involves accumulating events over time for collective analysis. This approach can be efficient when event volumes are low and there is an abundance of both security analysts and computational resources but it risks inefficient resource use, delayed attack identification / detection, and loss of real-time context in more demanding environments.
[0028] To tackle these challenges, systems and methods are disclosed herein that leverage a heuristic algorithm and reinforcement learning (see, e.g., Figure 2C) to decide the threat hunting pace (i.e., initial optimization) and then dynamically adjust thehunting pace in response to the evolving cybersecurity landscape and environmental conditions (i.e., dynamic optimization). This method allows for an initial pace setting based on current conditions, followed by real-time adjustments informed by ongoing environmental feedback (security event importance, received event volumes, resource availability) which will ensure an optimal balance between thorough analysis and resource management, thereby enhancing the SOC's ability to promptly and accurately hunt threats.
[0029] Embodiments of the systems and methods disclosed herein provide a novel solution that automatically and adeptly adjusts the threat hunting pace in alignment with one or more parameters including any one or more of the following: importance of security events, volume of security events, evolving threat landscape, and available resources. Embodiments of the present disclosure integrate two optimizations: (i) an initial heuristic optimization using, e.g., grid search to establish a baseline hunting pace, and (ii) a dynamic optimization using reinforcement learning to fine-tune the hunting pace, e.g., in response to changes in the environment. The changes in the environment are captured based on changes in one or more parameters such as, e.g., volume of events, importance of received events (e.g., based on cybersecurity knowledge base), available hunting resources (computational and / or manpower resources), and the evolving threat landscape.
[0030] Embodiments of the present disclosure may include any one or more of the following aspects: • Dynamic Hunting Pace Adjustment: The threat hunting pace is dynamically adjusted based on various environmental factors (such as, e.g., importance of security events, the threat landscape, security event volume, resource availability, and manpower). o Note that traditional threat hunting approaches do not account for such dynamic adjustments, often relying on static schedules or manually adjusted plans. o Embodiments are also disclosed herein related to ways of computing event importance and threat hunting performance from a security point of view of threat hunting, and their usage in a reinforcement learning approach. • Adaptive Learning: Heuristic methods are combined with reinforcement learning to provide adaptive learning.o Heuristic methods rely on experience-based techniques that can quickly identify hunting pace, o Reinforcement learning can adapt and optimize the decision-making process over time based on feedback from the environment. • Context-Aware Threat Hunting: A wide range of environmental factors are considered to assist the threat hunter including importance of security events. o Adapting to the SOC's current state and external threat conditions makes the solution highly context-aware.
[0031] Embodiments of the solution described herein may provide a number of advantages over existing solutions. In particular, embodiments of the solution disclosed herein may provide any one or more of the following advantages: • From the functionality perspective: o Dynamic Adjustment: Dynamically adjusting the threat hunting pace based on environmental changes ensures that the system can adapt its operations in real time and optimize for current conditions. o Context Awareness: By considering various environmental factors (e.g., security event importance, security event volume, the threat landscape, resource, and manpower), the system demonstrates a high degree of context awareness and enables informed and relevant hunting activities. o Adaptability: The system's ability to adapt to the evolving threat landscape and internal resource changes ensures its long-term effectiveness and relevance, as it can evolve alongside the threats it aims to mitigate. • From the Security Operations (SecOps) perspective: o Efficiency and Resource Optimization: Dynamically adjusting the hunting pace can lead to more efficient use of resources and manpower and ensure that efforts are concentrated when and where they are most needed. o Proactive Security Posture: The solution supports a proactive security posture by continuously adjusting to new threats and changing conditions. o Improved Responsiveness: By being able to quickly respond to the changing threat landscape, the system enhances the SecOps team's ability to deal with emerging threats.o Sustainability: Taking into account the availability of resources and manpower in hunting activities ensures the sustainability of SecOps efforts, preventing burnout, and ensuring that security measures can be maintained over time.
[0032] In addition, threat hunting has been mandated by the United States (U.S.) and is now exercised in many federal agencies. Some companies hire practitioners to manually frame the hunting based on the organization's resources and available intelligence. Embodiments of the solution disclosed herein may help the SOC and threat hunting team to automatically frame the hunt rather than automate it. Currently, there is no tool that sets the hunting pace or defines how the hunt should be conducted. All processes are defined manually by security practitioners, driven by their knowledge and judgment. This challenge becomes even more pronounced in dynamic network environments. Embodiments of the solution disclosed herein are not merely an improvement of existing hunting activities; they provide automation of the entire threat hunting process. Embodiments of the solution disclosed herein may determine how the hunting should be conducted, when it should start, for how long it should continue, and when it should stop. This is extremely important, as security practitioners often become bored and demotivated when they do not find anything, leading to wasted time and resources. Embodiments of the solution disclosed herein may be used as a tool to plan hunting resources and efforts based on available manpower, computational resources, and the threat landscape. Embodiments of the solution disclosed herein are distinct from traditional detection methods; they fall under proactive threat hunting and aim to automate a process that is currently performed manually.
[0033] The threat hunting process can be interpreted into three main functions: • Quality of Threat Hunting (^): This function measures the quality of generated threat hypotheses during the hunting pace ^ for security events that have importance ^ (i.e., criticality / severity level) and an average event arrival rate ^. The quality of threat hunting might be sensitive to how frequently new events occur. The function encompasses previous performance and might include factors like accuracy, false positive rate, and the ability to generate relevant hypotheses for ongoing sophisticated threats. The Quality of Threat Hunting (ℋ) may be calculated as:ℋ^^, ^, ^^ = ^^^^, ^, ^^^^^^, ^, ^^ + ^^^^, ^, ^^where ^^ represents the true ^ for the received event at rate ^ with importance ^,and ^^ represents the false positives, i.e., non-threats incorrectly identified as threats. • Cost of Operation (^): This function represents the operational cost of running the threat hunting for a hunting pace of ^ for the received security events with average event arrival rate ^ and event importance ^. It includes computational costs and manpower. ^^^, ^, ^^ = ^^^^^^^, ^, ^^ + ^^^^^^, ^, ^^where ^^^^^^^, ^, ^^ is the computational cost needed to perform threat huntingat hunting pace of ^ for the received security events with average event arrival rate ^ and event importance ^, e.g., in terms of resources (hardware and software cost). Similarly, ^^^^^^, ^, ^^ is the overall manpower cost needed toperform threat hunting at hunting pace of ^ for the received security events with average event arrival rate ^ and event importance ^, e.g., in terms of wages (e.g., salaries, shift differentials, and overtime). •Processing Time (^): The Processing Time function ^^^, ^, ^^ measures thetime taken by the SOC team for threat hunting at hunting pace of ^ for the received security events with average event arrival rate ^ and event importance ^, including processing logs, analyzing events, and generating and texting threat hypotheses.
[0034] Embodiments of the systems and methods disclosed herein operate to solve the following problem. For a given set of security events received with an average arrival rate ^ and importance ^, and hunting pace ^, the objective function ℱ, expressed below, aims at maximizing the quality of threat hunting ℋ while minimizing resource usage ^ and processing time ^. The objective function ℱ, which is also referred to herein as “threat hunting performance”, can be formulated as: ℱ= ^^ℋ^^, ^, ^^ − ^^^^^, ^, ^^ − ^^^^^, ^, ^^subject to: ^^^, ^, ^^ ≤ ^^^ ,^^^, ^, ^^ ≤ ^^^ ,^^!^ ≤ ^ ≤ ^^^ ,∑! #! = 1.Note that ^^^is a predefined or configured maximum allowable resource usage, ^^^is a or configured maximum allowable processing time, and ^^!^and ^^^are predefined or configured minimum or maximum allowable hunting pace. The #!values are predefined or configured weighting factors.
[0035] Figure 3 illustrates a system 300 that implements an embodiment of the proposed solution to adaptively determine an optimal hunting pace for effective threat hunting in a dynamic environment. The system 300 implements a two-stage approach including a first stage (“Stage 1”) in which a heuristic procedure (e.g., grid search) is initially used to find an optimal hunting pace for the threat hunting (^∗) and then a second stage (“Stage 2”) in which reinforcement learning is employed to dynamically adjust and optimize this hunting pace over time based on real-world feedback.
[0036] In Stage 1, an Initial Optimization Agent 302 uses a heuristic procedure, which in the illustrated example is a Grid Search, to systematically explore a range of hunting paces according to their respective performances (Step 1). In particular, the Initial Optimization Agent 302 builds a grid search space containing a set of possible hunting paces ^ for given events with an arrival rate (^) and importance (^). The Initial Optimization Agent 302 performs a grid search in which the Initial Optimization Agent 302 evaluates the threat hunting's performance ^ (step 2) in order to maximize theobjective function ℱ (i.e., to maximize the quality of threat hunting ℋ^^, ^, ^^ whileminimizing cost of operation ^^^, ^, ^^ (e.g., resource usage) and processing time ^^^, ^, ^^) for the set of^ (additional details are provided below). Based onthese results, the Initial Optimization Agent 302 selects the hunting pace ^ from the set of hunting paces that provides the best threat hunting performance (i.e., maximum value of ℱ) (step 3) as an initial hunting pace (^∗) (step 4). In other words, the Initial Optimization Agent 302 performs the grid search by calculating the overall cost (e.g.,cost of operation ^^^, ^, ^^ and processing time ^^^, ^, ^^) for each hunting pace ^ inthe set of hunting paces, for the given events with the arrival rate (^) and importance (^). The Initial Optimization Agent 302 then finds, or selects, the hunting pace ^, fromamong the set, that maximizes the quality of threat hunting ℋ^^, ^, ^^ while minimizingcost of operation ^^^, ^, ^^ (e.g., resource usage) and processing time ^^^, ^, ^^ (i.e.,maximizes the threat hunting performance ℱ).
[0037] The initial hunting pace (^∗) is provided to the SOC (e.g., via a Monitoring Agent 306 in the illustrated example of Figure 3). At the SOC, the initial hunting pace (^∗) is utilized in a threat hunting process, which includes, in this example, creating hypotheses, testing hypotheses, validating hypotheses, and uncovering new patterns.
[0038] In Stage 2, a Dynamic Optimization Agent 304 employs reinforcement learning (RL) to adaptively adjust the hunting pace utilized in the hunting process according to real-world feedback coming from the SOC environment (i.e., feedback from the hunting process). The state of the environment includes recent performance metrics (e.g., quality of hunting, cost, time), historical data (e.g., event arrival rate and the importance of these events), and current hunting pace (step 5). Based on the observed status, the Dynamic Optimization Agent 304 quantifies the success of the chosen hunting pace ^ in the form of a reward (i.e., a hunting performance based reward) (step 6). The reward reflects the objective function ℱ, i.e., maximize the quality of threat hunting while minimizing the cost of operation (e.g., resource usage) and processing time (see below for more details) (step 7). The Dynamic Optimization Agent 304 starts with the initial hunting pace (^∗) selected by the Initial Optimization Agent 302 (e.g., via Grid Search) and explores different intervals by maximizing the cumulative rewards in response to changing conditions and performance feedback in order to take action in the form of adjusting the hunting pace (τ') (step 8), which is sent to the SOC (step 9).
[0039] It should be noted that while the Initial Optimization Agent 302 is described herein, it is optional. For example, in an alternative embodiment, the initial hunting pace initial hunting pace (^∗) may be provided by different means such as, for example, by a security expert associated with the SOC.
[0040] Designing a solution for dynamically determining threat hunting pace raises a set of unique challenges. The main challenges are: (1) Alert Fatigue: Increasing the hunting pace could lead to a higher volume of alerts, which can cause alert fatigue among security analysts. This can reduce the overall quality of threat response as critical alerts might be overlooked or not acted upon promptly. (2) Speed vs Accuracy: There is always a trade-off between the speed of threat hunting and the accuracy of the hunting results. Increasing the hunting pace could lead to higher false positive rates while slowing down could improve accuracy but potentially miss timely threats.(3) Customization and Context Awareness: Different environments and contexts may require different hunting paces. It is always challenging to adjust the hunting pace and ensure the threat hunting strategies are contextually aware without compromising the hunting quality. (4) Dynamic Threat Landscapes: The threat landscape is constantly evolving, with new vulnerabilities, attack vectors, and tactics emerging regularly. Ensuring high- quality threat hunting in such a dynamic environment requires the system to continuously learn and adapt, which can be difficult to manage at a varied pace.
[0041] Embodiments of the solution disclosed herein adeptly navigate the intricate balance between speed and accuracy in threat hunting by intelligently adjusting the hunting pace. Embodiments of the solution disclosed herein achieve this by leveraging real-time data on available resources, current threat landscape insights, and the ongoing effectiveness of the threat hunting activities. In response to continuous performance evaluations and environmental feedback, embodiments of the solution disclosed herein dynamically adjust the hunting speed, ensuring an adaptive approach that mitigates alert fatigue and aligns with the evolving context of the environment. This strategic adjustment fosters an optimized threat hunting process that is both swift and precise, tackling the nuanced demands of evolving cybersecurity landscapes.
[0042] Figure 4 is workflow diagram that illustrates the operation of the Initial Optimization Agent 302 and the Dynamic Optimization Agent 304 in accordance with an embodiment of the present disclosure. In Stage 1, the Initial Optimization Agent 302 may determine the initial hunting pace (^∗) using different techniques, such as: • Empirical Testing: In this case, the Initial Optimization Agent 302 may determine the initial hunting pace (^∗) by conducting threat hunting for different values of hunting pace ^ (e.g., 1 sec, 1 min, 5 mins) and measuring the associated costs and benefits. • Machine Learning strategies: In this case, the Initial Optimization Agent 302 may determine the initial hunting pace (^∗) by predicting the optimal hunting pace based on historical data. • Statistical Analysis: In this case, the Initial Optimization Agent 302 may determine the initial hunting pace (^∗) by quantifying the quality of hunting and estimating the optimal hunting pace using a statistical model.Given the scarcity of comprehensive datasets, which are critical for conducting empirical studies or for training sophisticated machine learning models, using a statistical analysis through carefully constructed statistical models presents a viable option. This approach is particularly relevant in scenarios where data is limited, either in quantity or quality. Statistical models, unlike many machine learning algorithms, do not inherently require large volumes of data to yield insightful and reliable results.
[0043] Thus, in the procedure of Figure 4, the Initial Optimization Agent 302 is responsible for finding a near-optimal initial hunting pace (^∗) based on predefined performance criteria. The Initial Optimization Agent 302 systematically explores a range of possible hunting paces and evaluates the performance of threat hunting at each hunting pace. The selection of the initial hunting pace (^∗) can be performed using different techniques such as Grid Search, Genetic Algorithm, and Bayesian Optimization. However, in the example of Figure 4, Grid Search is used due to the small parameter space that facilitates the evaluation of every combination in a predefined grid of parameters (i.e., security event importance, arrival rate, hunting pace).
[0044] As illustrated in Figure 4, a security event importance (^) is computed for a set of received security events (400). The event importance (^) may be computed by an entity separate from the Initial Optimization Agent 302 or alternatively computed by the Initial Optimization Agent 302. Given the large number of security events received daily, an effective event triage is critical for guiding threat hunters in identifying and prioritizing potential threats. In order to take into account the security importance of an event, the following new measures are introduced: (i) the criticality of an event in the possible attack kill chain, (ii) attack conductivity as the event degree to lead to other security events / attacks, (iii) time difference of time between consecutive security events, and (iv) entropy to quantify uncertainty or randomness in the distribution of security events. The event criticality and attack conductivity are calculated based on an existing security knowledge base (e.g., MITRE ATT&CK). Each of these metrics offers unique insights into the nature and potential impact of security events, making them indispensable in a comprehensive cybersecurity strategy.
[0045] Event Criticality ((): Event Criticality (C) measures the importance of a security event (i.e., technique) within an attack chain (i.e., path or sequence of security events involved in an attack on an organization’s environment or system). To reflect this, we retrieve all attack chains associated with the received event, calculate thecentrality of the event in question in each attack chain, and then compute the average or weighted average depending on the impact of the attack chain. High criticality indicates that an event is pivotal within the path of known attack patterns (e.g., how close we are to an end node / breach success). Such events often act as critical bottlenecks that, if successful, can facilitate the spread of an attack (e.g., fast attack advance). By assessing the criticality of security events, threat hunters can prioritize those that are likely to have a significant impact on the overall security posture. This focus ensures that resources are directed towards mitigating the most potentially damaging threats.
[0046] Attack Conductivity ()): Th attack conductivity refers to the number of connections (i.e., degree) an event has within the attack chain. Events with a high attack conductivity are highly connected and can influence or be influenced by many other security events. These highly interconnected events indicate a potential attack, leading to damage of multiple systems and services. Identifying events with a high attack conductivity allows threat hunters to recognize and contain potential spread vectors, thereby preventing widespread disruption.
[0047] Time Difference (*+): The time difference between the arrival of security events is crucial for understanding the temporal dynamics of a potential ongoing attack. A rapid succession of events can indicate an ongoing attack or a coordinated effort to breach systems. By monitoring the time intervals between events, threat hunters can identify patterns that may signify an escalation or the presence of a sophisticated attack strategy. This temporal analysis helps in quickly identifying and responding to critical incidents before they cause significant damage.
[0048] Entropy (,): The entropy quantifies the uncertainty or randomness in the distribution of security events. ^ -^.^ = − / 0^1!^2340^1!^where: (a) . is the metric1!are the possible values of ., and (c) 0^1!^ is the probability of 1!. High entropy indicates a high level of disorder and unpredictability, which can beof sophisticated attacks that employ diverse techniques to evade detection. By incorporating entropy into event analysis, threat hunters can identify suspicious activities and unusual patterns that might not beevident through traditional metrics alone. This measure of unpredictability is important for identifying new or evolving threats that do not conform to established patterns.
[0049] In one example embodiment, the event importance (^) is computed as a weighted function over those metrics: ^= ^^6 + ^^7 + ^^- + ^9∆^,where ∑! #! = 1 (or depends on the impact of the attack chain). Note that theweighting factors #!used here for computing the event importance (^) are not the same as the weighting factors used above in the equation for the objective function ^ even though the same notation (i.e., alpha) is used. A higher value of the event importance (^) indicates events are more significant, considering event criticality, attack conductivity, timing, and overall uncertainty (entropy). A lower score indicates that events that are less significant based on the same criteria.
[0050] The Initial Optimization Agent 302 builds a grid search space based on available resources (e.g., manpower and / or computing resources), event volume of the received events, and the computed event importance (^) (402). The grid search spaceincludes, in this example, a set of parameters <^, ^, ^> for each (candidate) huntingpace ^ in a set of (candidate) hunting paces to be searched (i.e., each hunting pace in the range of ^^!^to ^^^), where ^ is the event arrival rate for the received security events, ^ is the computed event importance, and ^^!^and ^^^may be predefined or configured (e.g., by a security expert).
[0051] Then, using the built grid search space, the Initial Optimization Agent 302 calculates, for each hunting pace ^, the threat hunting performance (ℱ) (e.g., using the equation above) for the computed event importance (^) and average event arrival rate (^) (404). The Initial Optimization Agent 302 then finds a (near) optimal hunting pace that maximizes the threat hunting performance (ℱ) (i.e., maximizes the threat hunting quality ; while minimizing the cost of operation < and processing time ^ (406). This optimal hunting pace is selected as the initial hunting pace (^∗).
[0052] Figure 5 illustrates one example of the grid search of steps 402, 404, and 406 as “Algorithm 1”. Here, this process is used to tune the hyperparameters of a given model by systematically working through multiple combinations of parameter values, measuring the performance based on a predefined metric, and then determining which combination gives the best performance. For a given range of the hunting paces(^^!^, ^^^ ), arrival rate (^), and events importance (^), the grid search space is builtand then the performance of threat hunting is evaluated. The objective function, i.e.,the threat hunting performance ℱ^^, ^, ^^, is computed for each ^ in the grid searchspace, and the hunting pace ^∗that yields the maximum threat hunting performanceℱ^^, ^, ^^ (i.e., maximizes threat hunting quality ℋ^^, ^, ^^ while minimizing costs^^^, ^, ^^ and ^^^, ^, ^^) is selected. The selected hunting pace ^∗ is then used as theinitial hunting pace for the threat hunting.
[0053] Returning to Figure 4, the initial hunting pace (^∗) is provided to the SOC, where the SOC utilizes the initial hunting pace (^∗) in the threat hunting process. Upon detecting a change in the environment, the SOC triggers the Dynamic Optimization Agent 304 (step 408).
[0054] The Dynamic Optimization Agent 304 dynamically adjusts the hunting pace in response to changing conditions of the arrival rate of security events, threat landscape, and ongoing performance feedback (e.g., in the form of a reward).
[0055] In the preferred embodiments described herein, the Dynamic Optimization Agent 304 is built using RL and consists of states, actions, and rewards: • State: A representation of the current status of the system, which includes performance metrics (quality of hunting ℋ, cost of operation ^, processing time ^), historical data (event arrival rate ^, event importance ^, and current hunting pace ^); • Action: The selection of a new hunting pace ^̃; • Reward: A quantifiable measure of the success of threat hunting during the used ^. The reward function shares the same criteria used in the initial optimization's objective function (maximizing the quality of hunting ℋ, while minimizing the cost of operation ^ and processing time ^).
[0056] The Dynamic Optimization Agent 304 observes the state of the system and receives feedback (reward) based on the threat hunting performance. The Dynamic Optimization Agent 304 uses this feedback to learn a policy for selecting the hunting pace.
[0057] The Dynamic Optimization Agent 304 applies a Q-learning (i.e., a model-free reinforcement learning) to learn the value of an action (i.e., new hunting pace ^̃) in aparticular state. An action-value function ;^>, #^ is used to represent the expectedcumulative reward for taking action # in a state >. The cumulative reward ?@, at time A, is expressed asF B@ = / CD ?@EDwhere ?@ED is the reward received H, C is a discount factor (between 0and 1), and ^ is the time horizon
[0058] The Q-function is then updated as follows: ;^IJ^>@, #@^ = ^1 − ^^ ;^>@, #@^ + ^ ^?@ + C m^ax ;^>@E^, #^^where >@is the current state, #@is the current action (i.e., current ^), ?@is the reward received after performing the action #@in state >@, >@E^is the next state, ^ is thelearning rate, C is the discount factor, and m^ax ;^>@E^, #^ is the maximum expectedfuture reward given the new state and all possible actions.
[0059] Over time, the agent adjusts its policy to maximize cumulative rewards (Algorithm 2, see Figure 6), effectively finding the best dynamic interval adjustment strategy.
[0060] Looking specifically at Figure 4, the Dynamic Optimization Agent 304 calculates a threat hunting performance based on a defined reward function and the current environment state >@(410). In other words, in one embodiment, the DynamicOptimization Agent 304 calculates ℱ^^, ^, ^^ for the current environment state >@,defined above. The Dynamic Optimization Agent 304 selects a new hunting pace ^̃, which is also referred to herein as an “execution interval” (412). Next, the Dynamic Optimization Agent 304 takes the selected hunting pace ^̃ and, using the Q-learning function, maximizes the expected cumulative reward (i.e., the expected cumulative threat detection performance (ℱ)) responsive to using the selected hunting pace ^̃ (step 414). The Dynamic Optimization Agent 304 determines whether the expected cumulative reward satisfies a threshold performance level, which may be predefined or configured (step 416). If not, the process returns to step 412 and is repeated. Once a new hunting pace ^̃ has been selected for which the expected cumulative reward is at or above the threshold, the Dynamic Optimization Agent 304 provides this new hunting pace ^̃ to the SOC environment (step 418). The Dynamic Optimization Agent 304 may be again triggered, e.g., as needed, in response to further change in the SOC environment.
[0061] Figure 6 illustrates another embodiment of the operation of the Dynamic Optimization Agent 304. This embodiment is similar to steps 410 to 418 of Figure 4. Asillustrated, in line 1 of the procedure of Figure 6, the values of ;^>, #^ for each possiblestate (i.e., each possible state of the SOC environment) and action (each hunting pace) are initialized to, in this example, arbitrary values. Then, for each episode of the Q- learning procedure, the Dynamic Optimization Agent 304 initializes the state > to the current state of the SOC environment (see lines 2-3). More specifically, at a currenttime A + 1, the current state > = >@E^ is the state of the SOC environment as a result ofthe selection of action (i.e., hunting pace) #@when the SOC environment was in state >@at time A.
[0062] Then, for each possible action # (i.e., for each possible hunting pace ^), theDynamic Optimization Agent 304 determines (e.g., from a Q-table storing the ;^>, #^values for all > and all #) the value of ;^>, #^, which represents the expected cumulativereward for taking that action # in the current state > (see lines 5-7). The Dynamic Optimization Agent 304 then selects one of the actions (i.e., candidate new hunting pace ^̃ ) in accordance with an epsilon-greedy selection policy (see line 8). The Dynamic Optimization Agent 304 then determines the expected maximum cumulative reward if the selected action is chosen as the new action (i.e., the new hunting pace for the SOC environment) (see lines 9-12 noting the repetition of lines 5-12 until reaching a terminating state). Once the expected maximum cumulative reward is determined, the Dynamic Optimization Agent 304 determines whether the expected maximum cumulative reward is greater than or equal to a predefined threshold reward (see line 14). If so, the selected action (i.e., the candidate new hunting pace ^̃ ) selected for the first iteration of lines 5-12 is returned (i.e., provided to the SOC environment) as the new action (i.e., the new hunting pace ^̃) (see line 15). The procedure of Figure 6 may be repeated, e.g., each time there is a change in the SOC environment or each time the dynamic optimization is otherwise triggered. Example Implementation
[0063] One example implementation of an embodiment of the present disclosure was done using Python 3 programming language. Experiments were performed on Intel Xeon E312xx, 6 vCPU @ 2.69Ghz, with 32GB RAM running Ubuntu 20.04. In the environment, the minimum number of events arrival rate was configured to 100000 events / sec and the maximum was configured to 400000 events / sec, mimicking a realistic SOC environment. The hunting pace ranges from real-time hunting (expressedin the form of 1 min) to an hour (60 min). For reinforcement learning implementation, and to balance between exploration and exploitation, the learning rate was fixed to 0.1 and the discount factor to 0.99.
[0064] Figure 7 shows the results of the adaptive change of hunting pace (^) according to the change of event arrival rate and security event importance. It can be seen that, when the arrival rate increases from 10K to 30k, the hypothesis generation needs to reduce the hunting pace to process this data more frequently. One can notice a drop of the ℱ function from 80% to 50%, and then an adaptive selection of ^ from 15 to 1. This selection is performed according to the environment state handled by stage 2 – dynamic optimization. Conversely, when the arrival rate decreases, indicating less incoming data, the hunting pace needs to be changed accordingly and not necessarily increased as the selection is based not only on the arrival rate but other metrics such as security event importance, quality of hunting, cost of operation, and processing time. In order to perform in depth processing of each data batch, potentially improving the quality of hunting, one can see an increase in interval rate from 1 to 10 min due to changes in the event arrival and importance.
[0065] If the performance is enhanced at shorter intervals, stage 2 would learn to favor these intervals in similar future scenarios, and vice versa, which is witnessed in the 4th change of arrival rate back to 30K leading to a hunting pace of 15 min. This ensures that the hunting process remain responsive to new changes in the system environment.
[0066] Figure 8 shows the results of the effect of hunting pace (^) on the quality of threat hunting, cost of operation, and processing time. We can notice that a longer hunting pace affords the SOC team more time to conduct a comprehensive analysis that potentially leads to more detailed and quality hunting. For hunting pace ^=60 min (^ of 30K events / sec), the quality of threat hunting is around 0.85. This increased duration allows for a deeper examination of data, and a comprehensive understanding of potential threats, which can enhance the accuracy possibly uncovering stealthy and more complex threats, thereby enhancing the overall quality and completeness of the hunting.
[0067] On the other side, a shorter hunting pace enables the SOC to respond swiftly to new and emerging data, which is crucial in environments where threats are dynamic and evolve rapidly, yet it may limit the depth of analysis in each hunting cycle andtherefore reach fewer quality of hunting. For a pace ^= 1 (^ of 30K events / sec), the quality of hunting was only 0.6 These results also show that as we accelerate the hunting pace, the operational costs decrease disproportionately to the quality of threat hunting. This is due to the increased demand for resources as the SOC has more time and resources for hunting.
[0068] While a longer hunting pace allows for more in-depth analysis, the SOC requires more time to perform quality hunting as they need to deeply analyze and scrutinize a large volume of data before generating relevant hypotheses, which might not be suitable for critical scenarios when the organization is under extensive cyber- attacks. For a pace ^ = 1 (^ of 40K events / sec), the processing time is 0.4 while for a pace ^ = 60 min, the processing time is half lower (i.e., 0.2) this is due to the fact that the SOC faces pressure to manage a higher volume of alerts, including both true and false positives. However, the cost of operation has been decreased as we increase the hunting pace. This is due to the fact that operation cost is more impacted by the number of events received and hunting pace (the SOC has more resources to investigate more events in longer intervals). For a pace ^ = 1 (^ of 40K events / sec), the cost of operation was 0.9, and when ^ = 60 min (^ of 40K events / sec), the cost of operation increased only to 2.3.
[0069] Table 1 shows the effect of events’ importance on the quality of hunting, cost of operation, and processing time for different events’ arrival rates. For the sake of simplicity, the results are presented as a ratio. For optimal threat hunting, balancing event arrival rates with resource allocation and prioritizing high-importance events ensures the best performance (in terms of quality of hunting, cost, and time) and in return better security outcomes. Specifically, for low-importance events, we can notice a significant degradation in hypothesis quality as the number of received events increases due to the high noise of events (most of them are not important), which significantly higher costs due to the need for more extensive resources and eventually long processing time to process the high event volume. The hypothesis quality is maintained for medium importance events regardless of the number of received events. The costs are higher, and the processing time is increased but still manageable. For high- importance events, we can see that despite high event rates, high-quality hypotheses are generated which ensures that critical events are well investigated. This will lead tothe highest cost due to intensive processing needs for critical events and a longer processing time due to the need to prioritize and analyze critical events thoroughly. Table 1. Effect of events importance (quality of hunting, cost of operation, and processing time presented as ratio).
[0070] Figure 9 shows the results of the effect of events arrival rate (^) on the quality of hunting, cost of operation, and processing time.
[0071] Higher event rates lead to an increased volume of data that needs to be analyzed and scrutinized by the hypothesis generation tool, yet it can potentially enhance the quality of hunting by providing more comprehensive and varied data about ongoing threats. With more comprehensive data, the hypothesis generation tool is better positioned to hunt intricate attacks and perform subtle correlations and therefore improve the accuracy and robustness of the threat hypotheses it generates. For instance, for a pace of ^ = 30 min, we can notice that the quality of the hypothesis increased from 0.5 for ^ of 10000 events / sec to 0.65 for ^ of 40000 events / sec. This shows that the more information we have, the better hypotheses we generate.
[0072] On the other side, as the volume of data to be processed increases, the threat hunting requires more time to analyze the data and produce quality hunting. This has been witnessed in the increase of processing time from 0.2 (for ^ of 10000 events / sec) to 0.4 (^ of 40000 events / sec), which is not recommended during sensitive periods when an organization is under cyber-attacks. We also notice that the operational cost is influenced by higher event rates. Increased data processing demands more computational resources, which can escalate the cost of operation. Discussion
[0073] Embodiments of the solution disclosed herein may use the grid search methodology as the initial optimization to establish an initial hunting pace baseline. The solution subsequently refines this pace in response to the adaptive recommendations from the dynamic optimization using RL, which is attuned to environmental changes. This iterative feedback loop, wherein our solution incessantly assimilates performance metrics and environmental changes into the RL module, facilitates the nuanced adjustment of the hunting strategy over time and according to the attack landscape and available resources. Nevertheless, substantial changes in the environmental dynamics (e.g., number of received events and their importance) or operational prerequisites (e.g., computation resources, manpower) necessitate a re-execution of the initial optimization process, involving a re-iteration of the Grid Search to ascertain a new baseline and then a re-training of the RL agent. Yet, identifying such a pivotal change might require the security expert’s judgment. Moreover, the effectiveness of grid search is limited by available computational resources and the size of the parameter space, which might be sensitive to the values of the parameters.
[0074] Figure 10 is a schematic block diagram of a computing system 1000 in which the Initial Optimization Agent 302 and / or the Dynamic Optimization Agent 304 may be implemented, according to some embodiments of the present disclosure. The computing system 1000 includes one or more processors 1004 (e.g., Central Processing Units (CPUs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), and / or the like), memory 1006, and a network interface 1008 (e.g., a wired or wireless local or wide area network interface). The one or more processors 1004 are also referred to herein as processing circuitry. The one or more processors 1004 operate (e.g., in accordance with software stored in memory 1006) to cause the computing system 1000 to perform one or more functions of the Initial Optimization Agent 302 and / or the Dynamic Optimization Agent 304, as described herein.
[0075] Any appropriate steps, methods, features, functions, or benefits disclosed herein may be performed through one or more functional units or modules of one or more virtual apparatuses. Each virtual apparatus may comprise a number of these functional units. These functional units may be implemented via processing circuitry, which may include one or more microprocessor or microcontrollers, as well as other digital hardware, which may include Digital Signal Processors (DSPs), special-purpose digital logic, and the like. The processing circuitry may be configured to executeprogram code stored in memory, which may include one or several types of memory such as Read Only Memory (ROM), Random Access Memory (RAM), cache memory, flash memory devices, optical storage devices, etc. Program code stored in memory includes program instructions for executing one or more telecommunications and / or data communications protocols as well as instructions for carrying out one or more of the techniques described herein. In some implementations, the processing circuitry may be used to cause the respective functional unit to perform corresponding functions according to one or more embodiments of the present disclosure.
[0076] While processes in the figures may show a particular order of operations performed by certain embodiments of the present disclosure, it should be understood that such order is exemplary (e.g., alternative embodiments may perform the operations in a different order, combine certain operations, overlap certain operations, etc.).
[0077] Those skilled in the art will recognize improvements and modifications to the embodiments of the present disclosure. All such improvements and modifications are considered within the scope of the concepts disclosed herein.
Claims
Claims 1. A computer-implemented method, comprising: obtaining (410) information indicative of a current state of a cyber-threat hunting process; selecting (412 to 416) a hunting pace for the cyber-threat hunting process based on the current state of the cyber-threat hunting process and a machine learning scheme for hunting pace selection that maximizes a quality of the cyber-threat hunting process while minimizing a cost of operation and processing time for the cyber-threat hunting process; configuring (418) the cyber-threat hunting process with the selected hunting pace; and repeating the steps of obtaining and selecting, responsive to a change in the current state of the cyber-threat hunting process, thereby dynamically controlling the hunting pace for the cyber-threat hunting process.
2. The method of claim 1, wherein the machine learning scheme is a reinforcement learning scheme that uses states, actions, and a reward function, wherein: the states represent different states of the cyber-threat hunting process and include one or more metrics indicative of or otherwise related to the quality of the cyber-threat hunting process, the cost of operation for the cyber-threat hunting process, and the processing time for the cyber-threat hunting process; the actions represent selection of different hunting paces; and the reward is a function that quantifies a degree to which use of a particular hunting pace maximizes the quality of the cyber-threat hunting process while minimizing a cost of operation and processing time for the cyber-threat hunting process.
3. The method of claim 2, wherein selecting (412 to 416) the hunting pace for the cyber-threat hunting process based on the current state of the cyber-threat hunting process and a machine learning scheme comprises: for each hunting pace from among at least a subset of a set of possible hunting paces, computing (412; Fig. 6, lines 5-7) an action-value that represents an expected cumulative reward for using that hunting pace given the current state of the cyber- threat hunting process;selecting (412; Fig. 6, line 8) one of the set of possible hunting paces, based on the computed action-values, as a candidate hunting pace to be used for the cyber- threat hunting process; determining (414; Fig. 6, lines 9-13) an expected maximum cumulative reward for selecting the candidate hunting pace as a new hunting pace for the cyber-threat hunting process when in the current state.
4. The method of claim 3, wherein selecting (412 to 416) the hunting pace for the cyber-threat hunting process further comprises: determining (416; Fig. 6, line 14) that the expected maximum cumulative reward for selecting for selecting the candidate hunting pace as a new hunting pace for the cyber-threat hunting process when in the current state is greater than or equal to a predefined or configured reward threshold; and responsive to the expected maximum cumulative reward being greater than or equal to the predefined or configured threshold, selecting the candidate hunting pace as the hunting pace for the cyber-threat hunting process.
5. The method of any of claims 1 to 4, wherein the current state of the cyber-threat hunting process comprises: the current quality of the cyber-threat hunting process or one or more parameters from which the current quality of the cyber-threat hunting process can be derived; the cost of operation for the cyber-threat hunting process for a current hunting pace, average security event arrival time, and security event importance, or one or more parameters from which the current quality of the cyber-threat process can be derived; the processing time for the cyber-threat hunting process at the current hunting pace; an average security event arrival time; the current hunting pace.
6. The method of any of claims 1 to 5, wherein the quality of the cyber-threat hunting process and the cost of operation for the cyber-threat hunting process are eacha function of average security event arrival time for a set of security events, a security event importance computed for the set of security events, and a current hunting pace used for the cyber-threat hunting process.
7. The method of claim 6, wherein the quality of the cyber-threat hunting process is defined as: ^^^^^, ^, ^^ℋ ^, ^, ^^ =^^^^, ^, ^^ + ^^^^, ^, ^^where ℋ^^, ^, ^^ is the quality of the cyber-threat hunting process for a given averagesecurity event arrival rate ^, security event importance ^, and hunting pace ^, ^^ represents true positives for the hunting pace ^ for the average security event arrival rate ^ with the security event importance ^, and ^^ represents false positives for the hunting pace ^ for the average security event arrival rate ^ with the security event importance ^.
8. The method of claim 6 or 7, wherein the cost of operation for the cyber-threat hunting process is defined as: ^^^, ^, ^^ = ^^^^^^^, ^, ^^ + ^^^^^^, ^, ^^where ^^^, ^, ^^ is the cost of operation for the cyber-threat hunting process for a givenaverage security event arrival rate ^, security event importance ^, and hunting pace ^,^^^^^^^, ^, ^^ is a computational cost needed to perform cyber-threat hunting at thehunting pace of ^ for security events received with the average event arrival rate ^ andsecurity event importance ^, and ^^^^^^, ^, ^^ is an overall manpower cost needed toperform cyber-threat hunting at the hunting pace of ^ for security events received with the average event arrival rate ^ and the security event importance ^.
9. The method of any of claims 1 to 8, further comprising, prior to the steps of claim 1, selecting (402 to 406) an initial hunting pace for the cyber-threat hunting process via a heuristic process.
10. The method of claim 9, wherein selecting (402 to 406) the initial hunting pace for the cyber-threat hunting process comprises:for each candidate hunting pace from a set of candidate hunting paces, determining (404) a quality of the cyber-threat hunting process for the candidate hunting pace for an average arrival rate of a set of received security events and an importance of the set of received security events, a cost of operation of the cyber-threat hunting process for the candidate hunting pace for the average arrival rate of the set of received security events and the importance of the set of received security events, and a processing time needed for the cyber-threat hunting process for the candidate hunting pace for the average arrival rate of the set of received security events and the importance of the set of received security events; and selecting (406) the candidate hunting pace that maximizes the quality of the cyber-threat hunting process while minimizing the cost of operation and processing time, as the initial hunting pace.
11. A computing system (1000), comprising: processing circuitry (1004); and memory (1006) storing software instructions executable by the processing circuitry (1004) whereby the computing system (1000) is configured to: obtain (410) information indicative of a current state of a cyber-threat hunting process; select (412 to 416) a hunting pace for the cyber-threat hunting process based on the current state of the cyber-threat hunting process and a machine learning scheme for hunting pace selection that maximizes a quality of the cyber- threat hunting process while minimizing a cost of operation and processing time for the cyber-threat hunting process; configure (418) the cyber-threat hunting process with the selected hunting pace; and repeat the steps of obtaining and selecting, responsive to a change in the current state of the cyber-threat hunting process, thereby dynamically controlling the hunting pace for the cyber-threat hunting process.
12. The computing system of claim 11, wherein the machine learning scheme is a reinforcement learning scheme that uses states, actions, and a reward function, wherein:the states represent different states of the cyber-threat hunting process and include one or more metrics indicative of or otherwise related to the quality of the cyber-threat hunting process, the cost of operation for the cyber-threat hunting process, and the processing time for the cyber-threat hunting process; the actions represent selection of different hunting paces; and the reward is a function that quantifies a degree to which use of a particular hunting pace maximizes the quality of the cyber-threat hunting process while minimizing a cost of operation and processing time for the cyber-threat hunting process.
13. The computing system of claim 12, wherein, in order to select (412 to 416) the hunting pace for the cyber-threat hunting process based on the current state of the cyber-threat hunting process and a machine learning scheme, the processing circuitry, via execution of the software instructions, is further configured to cause the computing system to: for each hunting pace from among at least a subset of a set of possible hunting paces, compute (412; Fig. 6, lines 5-7) an action-value that represents an expected cumulative reward for using that hunting pace given the current state of the cyber- threat hunting process; select (412; Fig. 6, line 8) one of the set of possible hunting paces, based on the computed action-values, as a candidate hunting pace to be used for the cyber-threat hunting process; determine (414; Fig. 6, lines 9-13) an expected maximum cumulative reward for selecting the candidate hunting pace as a new hunting pace for the cyber-threat hunting process when in the current state.
14. The computing system of claim 13, wherein, in order to select (412 to 416) the hunting pace for the cyber-threat hunting process based on the current state of the cyber-threat hunting process and a machine learning scheme, the processing circuitry, via execution of the software instructions, is further configured to cause the computing system to: determine (416; Fig. 6, line 14) that the expected maximum cumulative reward for selecting for selecting the candidate hunting pace as a new hunting pace for thecyber-threat hunting process when in the current state is greater than or equal to a predefined or configured reward threshold; and responsive to the expected maximum cumulative reward being greater than or equal to the predefined or configured threshold, select the candidate hunting pace as the hunting pace for the cyber-threat hunting process.
15. The computing system of any of claims 11 to 14, wherein the current state of the cyber-threat hunting process comprises: the current quality of the cyber-threat hunting process or one or more parameters from which the current quality of the cyber-threat hunting process can be derived; the cost of operation for the cyber-threat hunting process for a current hunting pace, average security event arrival time, and security event importance, or one or more parameters from which the current quality of the cyber-threat process can be derived; the processing time for the cyber-threat hunting process at the current hunting pace; an average security event arrival time; the current hunting pace.
16. The computing system of any of claims 11 to 15, wherein the quality of the cyber-threat hunting process and the cost of operation for the cyber-threat hunting process are each a function of average security event arrival time for a set of security events, a security event importance computed for the set of security events, and a current hunting pace used for the cyber-threat hunting process.
17. The computing system of claim 16, wherein the quality of the cyber-threat hunting process is defined as: ^^ ^^^^, ^, ^^ℋ ^, ^, ^ =^^^^, ^, ^^ + ^^^^, ^, ^^where ℋ^^, ^, ^^ is the quality of the cyber-threat hunting process for a given averagesecurity event arrival rate ^, security event importance ^, and hunting pace ^, ^^ represents true positives for the hunting pace ^ for the average security event arrivalrate ^ with the security event importance ^, and ^^ represents false positives for the hunting pace ^ for the average security event arrival rate ^ with the security event importance ^.
18. The computing system of claim 16 or 17, wherein the cost of operation for the cyber-threat hunting process is defined as: ^^^, ^, ^^ = ^^^^^^^, ^, ^^ + ^^^^^^, ^, ^^where ^^^, ^, ^^ is the cost of operation for the cyber-threat hunting process for a givenaverage security event arrival rate ^, security event importance ^, and hunting pace ^,^^^^^^^, ^, ^^ is a computational cost needed to perform cyber-threat hunting at thehunting pace of ^ for security events received with the average event arrival rate ^ andsecurity event importance ^, and ^^^^^^, ^, ^^ is an overall manpower cost needed toperform cyber-threat hunting at the hunting pace of ^ for security events received with the average event arrival rate ^ and the security event importance ^.
19. The computing system of any of claims 11 to 18, wherein the processing circuitry, via execution of the software instructions, is further configured to cause the computing system to, prior to the steps of claim 1, selecting (402 to 406) an initial hunting pace for the cyber-threat hunting process via a heuristic process.
20. The computing system of claim 19, wherein, in order to select (402 to 406) the initial hunting pace for the cyber-threat hunting process, the processing circuitry, via execution of the software instructions, is further configured to cause the computing system to: for each candidate hunting pace from a set of candidate hunting paces, determine (404) a quality of the cyber-threat hunting process for the candidate hunting pace for an average arrival rate of a set of received security events and an importance of the set of received security events, a cost of operation of the cyber-threat hunting process for the candidate hunting pace for the average arrival rate of the set of received security events and the importance of the set of received security events, and a processing time needed for the cyber-threat hunting process for the candidate hunting pace for the average arrival rate of the set of received security events and the importance of the set of received security events; andselect (406) the candidate hunting pace that maximizes the quality of the cyber- threat hunting process while minimizing the cost of operation and processing time, as the initial hunting pace.
21. A non-transitory computer-readable medium storing software instructions executable by processing circuitry of a computing system, whereby the computing system is caused to: obtain (410) information indicative of a current state of a cyber-threat hunting process; select (412 to 416) a hunting pace for the cyber-threat hunting process based on the current state of the cyber-threat hunting process and a machine learning scheme for hunting pace selection that maximizes a quality of the cyber-threat hunting process while minimizing a cost of operation and processing time for the cyber-threat hunting process; configure (418) the cyber-threat hunting process with the selected hunting pace; and repeat the steps of obtaining and selecting, responsive to a change in the current state of the cyber-threat hunting process, thereby dynamically controlling the hunting pace for the cyber-threat hunting process.
Citation Information
Patent Citations
System and Method for Cyber Security Threat Detection
US20180255077A1
Breach prediction via machine learning
US20230370495A1