A data-efficient method for threat detection in computer networks
The data-efficient threat detection method using generative models filters out normal events, addressing EDR's inefficiencies in data management to reduce costs and maintain effective threat detection.
Patent Information
- Application Number
- JP2020159741
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2020-09-24
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2040-09-24
AI Technical Summary
EDR systems face challenges with large data volumes leading to increased costs, resource consumption, and reduced service quality due to inefficient data management, which can compromise threat detection capabilities.
Implementing a data-efficient threat detection method using generative models to replicate normal behavior, allowing only anomalous events to be transmitted for further processing, thereby reducing data volume while maintaining effective threat detection.
This approach reduces data processing costs and resource consumption while ensuring high-quality threat detection by filtering out normal events, enabling efficient and adaptive anomaly detection.
Smart Images

Figure 0007738985000001 
Figure 0007738985000002
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and apparatus for highly data-efficient threat detection in computer networks. [Background technology]
[0002] Computer network security systems are widespread. Known examples of such security systems include endpoint detection and response (EDR) and managed detection and response (MDR) products and services. EDR focuses on detecting and monitoring breaches both during and after they occur, helping to determine the best response method. The growth of efficient and robust EDR solutions has been made possible in part by the rise of machine learning, big data, and cloud computing. MDR, on the other hand, is a managed cybersecurity service that provides threat detection, response, and remediation services.
[0003] EDR or other applicable systems deploy data collectors on selected network endpoints (which can be any element of an IT infrastructure). The data collectors monitor activity occurring on the endpoints and then send the collected data to a central backend system (the "EDR backend"), often located in the cloud. Once the EDR backend receives the data, it is processed (e.g., aggregated and enriched) and then analyzed and scanned by the EDR provider for signs of a security breach or anomaly. Summary of the Invention [Problem to be solved by the invention]
[0004] However, the challenge with EDR is that the volume of data generated by data collectors can be quite large. Data volume is typically proportional to the activity occurring at a particular EDR endpoint, so if there is a lot of activity at that EDR endpoint, there will also be a lot of data generated. The direct impact of this large volume of data generation includes reduced service quality, increased service costs, and increased resource consumption associated with managing large volumes of data. For example, if large volumes of data need to be processed and made available in a usable format, the associated resource overhead and financial costs can sometimes be prohibitive for EDR providers, thereby increasing the cost of providing EDR to customer organizations. For this reason, many organizations choose not to simply implement EDR and continue to rely solely on endpoint protection (EPP) solutions, which pose security risks because they cannot protect organizations from advanced fileless threats.
[0005] Some EDR systems propose reducing data overhead by selecting the data they collect (i.e., selective data collection restriction strategies). However, this solution presents challenges because efficient monitoring, detection, and forensic analysis often require as complete a data picture as possible. In many cases, it is not possible to know in advance what data may be needed to monitor or track malicious actors. The realization that important information is not being collected often halts any investigation, potentially rendering such EDR systems ineffective.
[0006] There is a need to reduce the costs associated with managing large amounts of data and improve how data is collected and processed behind EDR systems, while avoiding significant risks to threat detection capabilities and reducing resource consumption and scalability issues caused by the ever-increasing volume of data. [Means for solving the problem]
[0007] According to an aspect of the present invention, there is provided a method for data-efficient threat detection as specified in claims 1, 10 and 14.
[0008] According to another aspect of the present invention there is provided an apparatus in a computer network security system as specified in claim 18.
[0009] According to another aspect of the present invention there is provided a computer program product comprising a computer storage medium having stored thereon computer code which, when executed on a computer system, causes the system to operate as a server according to the above aspect of the present invention. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 illustrates a schematic diagram of a network architecture. [Figure 2] FIG. 1 is a flow diagram illustrating a method according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] FIG. 1 is a schematic diagram of a portion of a first computer network 1 on which a computer system, e.g., an EDR system, is installed. Any other computer system capable of implementing embodiments of the present invention may be used in place of or in addition to the EDR system used in this example. The first computer network is connected to a security service network, here a security backend / server 2, via a cloud 3. The backend / server 2 forms a node on this security service computer network for the first computer network. This security service computer network is managed by an EDR system provider and may be separated from the cloud 3 by a gateway or other interface (not shown) or other network element appropriate to the backend 2. The first computer network 1 may be further separated from the cloud 3 by a gateway 4 or other interface. Other network configurations are also contemplated.
[0012] The first computer network 1 is formed by a plurality of interconnected nodes 5a-5g, each representing a component of the computer network 1, such as a computer, smartphone, tablet computer, laptop computer, or other network-enabled hardware component. The illustrated network nodes 5a-5g each represent an EDR endpoint on which a data collector (or “sensor”) 6a-6h is installed. Data collectors may similarly be installed on any other component of the computer network, such as a gateway or other interface. The data collector 4a is installed on the gateway 4 in FIG. 1. The data collectors 6a-6h, 4a collect various types of data at the nodes 5a-5h or the gateway 4, including, for example, hash values of programs or files stored on the nodes 5a-5h, network traffic logs extracted from memory, process logs, binaries or files (e.g., DLLs, EXEs, or memory forensic artifacts), and / or logs (e.g., TCP dumps) obtained from monitoring actions performed by programs or scripts running on the nodes 5a-5h or the gateway 4.
[0013] Any kind of data useful for detecting and monitoring security threats, such as malware, security breaches or system intrusions, may be collected during the lifecycle by the data collectors 6a-6h, 4a, and the types of data that may be monitored and collected may be configured according to rules defined by the EDR system provider at the time of installation of the EDR system or in response to instructions from the EDR backend 2.
[0014] For example, data collectors 6a-6h, 4a may collect data about the behavior of programs running on the EDR endpoints and may monitor when new programs are started. If appropriate resources are available, the collected data may be stored permanently or temporarily by data collectors 6a-6h, 4a in suitable storage locations on their respective nodes or on first computer network 1 (not shown).
[0015] The data collectors 6a-6h, 4a may further perform pre-processing steps on the collected data that are limited by the computing and network resources available at each node 5a-5h or gateway 4.
[0016] The data collectors 6a-6h, 4a are configured to transmit information such as collected data and to send and receive commands to and from the EDR backend 2 via the cloud 3. This allows the EDR system provider to remotely manage the EDR system without having personnel on-site at the organization managing the first computer network 1.
[0017] In one embodiment, the data collectors 6a-6h may be further configured to implement an internal swarm intelligence network including data collector modules of multiple interconnected network nodes 5a-5h in the computer network 1. The modules 6a-6h may be further configured to collect data related to each of the network nodes 5a-5h and share information based on the collected data in the established internal swarm intelligence network. The swarm intelligence network may include multiple semi-independent security nodes (security agent modules) that can function equally well on their own. Thus, the number of instances in a swarm may vary considerably. Multiple connected swarms may also exist within a local computer network, working together.
[0018] The modules 6a-6h, 4a may be further configured to use the collected data and information received from the internal swarm intelligence network to generate and adapt models associated with the respective network nodes 5a-5h. For example, if a known security threat is detected, the modules 6a-6h, 4a may be configured to generate and transmit a security alert to the internal swarm intelligence network and to a local central node (not shown) in the local computer network to activate security countermeasures in response to the detected security threat. Furthermore, if an anomaly is identified that is highly likely to be a new threat, the modules 6a-6h, 4a may be configured to verify and contain the threat, generate a new threat model based on the collected data and received information, and share the new threat model with the internal swarm intelligence network and the local central node.
[0019] FIG. 2 is a flow diagram illustrating a method according to one embodiment.
[0020] At S201, raw data related to a network node is received. This raw data may be received / collected and collated from multiple network nodes (5a-5h), where different data types are collated as input events. The raw transmission processing component is responsible for the initial pre-processing of all data transmissions received from various types of endpoint sensors. Its purpose is to collate all the different data types so that the next level components in the data processing pipeline can interpret / process the data blocks (further referred to as events).
[0021] The raw data associated with each network node may be collected by a security server backend or from or by multiple network nodes of a computer network. The events monitored associated with a network node are something that can be effectively measured and are caused by a multitude of potential processes / actors. Such actors can be, for example, actual users or operating systems.
[0022] At S202, one or more local behavior models associated with the network nodes are generated based on the received input events. The local behavior models characterize normal behavior associated with each network node, and the local behavior models associated with the network nodes are generated locally by each network node. At S203, the generated one or more local behavior models associated with each network node may be shared with one or more other network nodes in the computer network / swarm intelligence network and / or a security server backend of the computer network.
[0023] Most potential processes / actors associated with monitored events have some normal behavior that can be modeled with a well-functioning model. In one embodiment, such behavior is at least partially shared across hosts and partially shared locally, with local behaviors sharing commonalities even if they are not completely identical. For example, all versions of the same operating system will exhibit similar background behavior, but all developers will have slightly different practices and still tend to use some similar tools and flows. That is, similarities between background behaviors can be detected between them, but the instances will be different.
[0024] In one embodiment, normal behavior modeling is designed via one or more generative models. One or more such models may be generated in association with each network node, depending on their complexity, and these models may take very different forms, for example, recurrent neural networks (RNNs) such as long short-term memories (LSTMs), although many other models are equally viable.
[0025] At S204, at least one common model of normal behavior is generated based on local behavior models associated with the plurality of network nodes, and the common model of normal behavior may be generated by a security server backend and / or by any network node of the computer network.
[0026] In one embodiment, local behavioral models associated with multiple network nodes are used to understand the behavior of individual network nodes, and information spanning multiple hosts is used to build one or more common models of normal behavior, and these common learnings are then redistributed to address, for example, operating system updates or latest versions of Chrome that are global but may fluctuate and otherwise cause problems for such models (which utilize distributed / federated learning techniques).
[0027] In one embodiment, if at least one common model of normal behavior is generated by the network nodes of the computer network, the process may include cooperation among at least some of these network nodes to learn common behaviors associated with those network nodes. This type of implementation may be feasible, for example, when the same user controls multiple different computers and / or within the same organization.
[0028] At S205, one or more of these input events are filtered using a measure that estimates the likelihood that the input event was generated by the generated common model of normal behavior and / or by the generated one or more local behavior models. Only input events that have a likelihood generated by one of the models (the common model of normal behavior and the one or more local behavior models) below a predefined or adaptable threshold are passed through the filtering. A suitable generative model may also take into account the volume of events and / or statistics may be collected to ensure that the model can be retrained.
[0029] Thus, after at least one common generative model (or set of models) of normal behavior is constructed, this model can be used to compare what is observed at network nodes / sensors with what is expected to be observed (i.e., what the model generates). To make such a comparison, a probability measure can be established that is used to estimate the likelihood that the event will be generated by the model. If this likelihood is very low, the event can be determined to be anomalous, and appropriate further action can be taken to protect the computer network. However, this is not considered here as an obvious use case for anomaly detection, where all such anomalies are expected to be meaningful, but rather as a form of highly efficient data reduction. By sharing a common generated model of normal behavior, all commonly occurring events can be represented in just one large "event" that contains model parameters describing normal behavior. This can also be done in a privacy-preserving manner, since the model does not contain any actual events.
[0030] In one embodiment, only events that differ from those that can be reproduced by the model(s) need to be shared for further processing. As a result, all events that were not expected to be generated by the model(s) are displayed (passed through filtering), resulting in thorough data reduction while still maintaining full anomaly / threat detection capabilities on the backend. This, in turn, allows for efficient detection of new types of attacks.
[0031] In one embodiment, the filtering of one or more of the input events may be further based on one or more of a self-learning rule set, a decision tree, a deep learning neural network, or other machine learning models, and may be performed by a backend of the security server and / or by a network node of the computer network.
[0032] At S206, the input events passed through filtering are processed to generate security-related decisions. This process may include an event enrichment process. Other processes, such as aggregation, may also be used in preparing the data for generating security-related decisions.
[0033] Processing input events passed through filtering to generate security-related decisions may include analyzing the facts using any rules, heuristics, machine learning models, fuzzy logic-based models, statistical inference-based models, etc., and coming up with appropriate decisions and recommendations (findings) that positively impact the state of the protected IT infrastructure in real time.
[0034] If a security threat is detected based on the results of the event analysis component, further actions may be taken, such as immediately taking action to change the configuration of one or more network nodes to ensure that the attacker is stopped and that any traces of the attacker's movements are not erased. Such configuration changes may include, for example, preventing one or more nodes (which may be computers or other devices) from being turned off in order to store information in RAM, turning on a firewall on one or more nodes to immediately shut off the attacker, throttling or blocking one or more network connections of the network nodes, deleting or quarantining suspicious files, collecting logs from the network nodes, executing a set of commands on the network nodes, alerting users of one or more nodes that a breach has been detected and that their workstations are under investigation, and / or sending a system update or software patch to the nodes from the EDR backend 2 in response to detecting the security threat. It is envisioned that one or more of these actions may be automatically initiated by the algorithm described above. For example, data may be collected and sent from nodes of computer network 1 to the EDR backend 2 using the method described above. The analysis algorithm determines that a security threat has been detected. Upon determining that a security threat has been detected, the algorithm can generate and issue commands to the relevant network nodes, without human intervention, to automatically initiate one or more of the above actions at the nodes. This can be done very quickly and automatically, without human intervention, to shut down the threat and / or minimize damage.
[0035] Overall, the proposed approach brings many improvements to the traditional EDR backend data processing pipeline scheme, including an improved filtering component that can, for example, filter out the most important parts of common events and / or clean up the unnecessary parts of events that do not need to be passed on to the next element in the pipeline.
[0036] Overall, the present invention aims to overcome one of the significant challenges described above: reducing the amount of data processed with minimal compromise to the accuracy of detecting known or unknown threats. Embodiments of the present invention provide a flexible and adaptive data selection approach driven entirely by an analysis engine that may utilize machine learning, statistics, heuristics, and any other decision mechanism. Embodiments of the present invention further enable flexible filtering of events along with the definition of associated filtering logic. Embodiments of the present invention provide an integrated data processing pipeline capable of achieving both efficient detection and data reduction.
[0037] An embodiment of the present invention allows for cost reduction by performing data processing without significant risk to detection capabilities, for example, in EDR systems. Building a sustainable security system without the risk of data collection requires balancing both cost and efficiency. One embodiment of the present invention is based on the recognition that generative models can be used to reduce data. This also requires trust that the generated behavioral models are indeed valid representations of normal behavior, and if these models are trained locally, they can be sent retroactively to a backend to indicate to the backend what is needed from normal behavior. For example, only events that are determined to be sufficiently "interesting" (based on the use of the model) will require further processing by the backend.
[0038] Prior art solutions have typically used almost exclusively back-end trained models for anomaly detection, i.e., alerting based on unexpected behavior. Additionally, standard approaches to data reduction tend to focus on either removing "known good" events or attempting to model detections to pass only events expected to generate detections (removing unwanted events or processing only events known to be needed). However, neither approach addresses the challenge when the need is unknown. For example, event data may contain information that could potentially help detect new attacks, and the standard approach is to sample the entire data stream.
[0039] The solution of the present invention is a novel method for transforming techniques typically used for anomaly detection into data reduction filters that only work when a common generative model is shared to replicate normal behavior. The present invention enables an improved, data-efficient method for threat detection by using generative models to model normal data and transmitting only events that cannot be predicted by these models because they are deemed unique enough to merit attention. Data efficiency in this context refers to the efficiency of one or more processes applied to threat detection that achieve good results while not requiring large amounts of data.
[0040] Using an approach according to one embodiment, instead of filtering out events known to be clean or attempting to send events known to be malicious, the process is modeled so that the underlying machine learning model learns what is normal. This model is shared, and data statistics may be replicated, for example, at the backend. Ultimately, only events that significantly deviate from what is expected (as modeled) are sent for processing. This solution actually results in greater data reduction than any other known solution, while still allowing the possibility of generating entirely new detections at the backend to remain feasible because unexpected events (and only unexpected events) are sent to the backend, while at the same time we are not relying on what is known to be malicious or generating detections. Anomaly-based approaches can also return everything that "might be interesting," with adjustable parameters for how much is returned, due to the inherent calculation of similarity or likelihood in the observed data from the modeled distribution.
[0041] An embodiment of the present invention provides a generative model to control the amount of data allowed to pass through, thereby keeping process costs and resource usage reasonable, and also optimizes the detection of the most relevant data to be processed automatically.
[0042] Machine learning may be used to estimate the normal operation of the system, including rules and other machine learning models. The nature of the models used in the EDR system may be or incorporate elements from one or more of the following: neural networks trained using a training dataset, rules of thumb or heuristic rules (e.g., hard-coded logic), fuzzy logic-based modeling, and statistical inference-based modeling. The models may be defined to take into account specific patterns, files, processes, connections, and dependencies between processes.
[0043] Although the present invention has been described with reference to the above preferred embodiments, it should be understood that these embodiments are merely examples and that the claims are not limited to these embodiments. Those skilled in the art will be able to make modifications and substitutions in light of this disclosure that are believed to fall within the scope of the appended claims. Each feature disclosed or illustrated herein can be incorporated into the present invention alone or in any suitable combination with other features disclosed or illustrated herein.
Claims
1. receiving raw data associated with a network node, wherein different data types are aligned as input events; generating one or more local behavior models associated with the network node characterizing normal behavior associated with the network node based on the received input events; generating at least one common model of normal behavior based on local behavior models associated with a plurality of said network nodes; filtering one or more of the input events by using a measure that estimates the likelihood that the input event will be generated by the generated common model of normal behavior and / or by the generated one or more local behavior models, wherein only input events that have a likelihood of being generated by the generated common model of normal behavior and / or by the generated one or more local behavior models that is below a predefined threshold are passed through the filtering; processing input events passed through said filtering to generate security-related decisions; sharing the generated one or more local behavior models associated with the network node with one or more other network nodes in a computer network and / or a security server backend of the computer network; A method for data-efficient threat detection in computer networks.
2. 2. The method of claim 1, wherein the one or more local behavioral models associated with the network node are generated by the network node, and the at least one common model of normal behavior is generated by a backend of a security server of the computer network and / or by the network node.
3. The method of claim 1 , further comprising utilizing local behavioral models associated with a plurality of the network nodes to understand behavior of individual network nodes.
4. 10. The method of claim 1, wherein the filtering of the one or more of the input events is further based on one or more of a self-learning rule set, a decision tree, a deep learning neural network, or other machine learning model.
5. The method of claim 1 , wherein the raw data is received by a backend of a security server from multiple network nodes of a computer network or by a network node of a computer network.
6. The method of claim 1 , wherein the filtering of the input events is performed by a backend of a security server and / or by a network node of a computer network.
7. 10. The method of claim 1, wherein processing the input events to generate the security-related decision comprises using at least one of predefined rules, heuristics, machine learning models, fuzzy logic-based models, and statistical inference-based models.
8. performing additional actions to protect the computer network and / or any associated network nodes; The additional action is: preventing one or more of said network nodes from being turned off; turning on a firewall on one or more of said network nodes; throttling or blocking one or more network connections of said network node; Deleting or quarantining suspicious files; collecting logs from said network nodes; executing a set of commands at said network node; alerting one or more users of said network node that an indication of a security breach has been detected; and / or transmitting software updates to one or more of said network nodes; The method of claim 1 , comprising any one or more of:
9. 1. A method for data-efficient threat detection in a computer network, comprising: at a network node of the computer network, the method comprising: receiving raw data associated with the network node, wherein different data types are aligned as input events; generating one or more local behavior models associated with the network node characterizing normal behavior associated with the network node based on the received input events; generating at least one common model of normal behavior based on local behavior models received from a plurality of said network nodes; and filtering one or more of said input events by using a measure that estimates a likelihood that said input event was generated by said generated common model of normal behavior or by said generated one or more local behavior models, such that only input events that have a likelihood generated by said generated common model of normal behavior and / or by said generated one or more local behavior models that is below a predefined threshold are passed through said filtering; and sharing data related to said generated one or more local behavior models with one or more other network nodes so that said input events that are passed through said filtering can be processed to generate security-related decisions. A method for data-efficient threat detection in computer networks.
10. The method of claim 9 , wherein the data related to the generated one or more local behavior models is shared with a security server backend and / or other network nodes of the computer network.
11. The method of claim 10 , further comprising receiving the generated at least one common model of normal behavior from a backend of the security server or another network node of the computer network.
12. receiving data related to one or more local behavior models associated with one or more other network nodes of the computer network; The method of claim 9 , further comprising: generating at least one common model of the normal behavior based on the received local behavior model.
13. 1. A method for data-efficient threat detection, in a back-end of a security server of a computer network, said method comprising: receiving one or more local behavioral models associated with a plurality of network nodes, the one or more local behavioral models associated with the plurality of network nodes characterizing normal behavior associated with each of the network nodes; generating at least one common model of normal behavior based on the generated local behavior models associated with a plurality of the network nodes to be able to filter one or more of the input events by using a measure that estimates the likelihood that an input event was generated by the generated common model of normal behavior or by the received one or more local behavior models, such that only input events that have a likelihood of being generated by the common model of normal behavior below a predefined threshold are passed through the filtering; processing input events passed through said filtering to generate security-related decisions; A data-efficient method for threat detection.
14. 14. The method of claim 13, further comprising sharing the generated at least one common model of normal behavior with one or more network nodes of the computer network to enable the filtering of the input events at the network nodes.
15. The method of claim 14 , further comprising receiving input events from one or more network nodes associated with the one or more network nodes that are passed through the filtering.
16. receiving input events from the one or more network nodes that are passed through the filtering using a metric that estimates the likelihood that the input events will be generated by the one or more generated local behavior models; 14. The method of claim 13, further comprising: filtering the received input events from the network nodes using the generated at least one common model of normal behavior.
17. receiving raw data associated with a network node, where different data types are aligned as input events; generating one or more local behavior models associated with the network node characterizing normal behavior associated with the network node based on the received input events; generating at least one common model of normal behavior based on local behavior models associated with a plurality of said network nodes; filtering one or more of the input events using a measure that estimates the likelihood that the input event will be generated by the generated common model of normal behavior and / or by the generated one or more local behavior models, such that only input events that have a likelihood of being generated by the generated common model of normal behavior and / or by the generated one or more local behavior models below a predefined threshold are passed through the filtering; and configured to process input events passed through the filtering to generate security-related decisions; and one or more processors further configured to share the generated one or more local behavior models associated with the network node with one or more other network nodes in a computer network and / or with a security server backend of the computer network. A device in a computer network system.
18. 18. The apparatus of claim 17, wherein the one or more local behavioral models associated with the network node are generated by the network node, and the at least one common model of normal behavior is generated by a backend of a security server of a computer network and / or by the network node.
19. 20. The apparatus of claim 17, wherein the processor is further configured to utilize local behavioral models associated with a plurality of the network nodes to understand behavior of individual the network nodes.
20. 20. The apparatus of claim 17, wherein the filtering of the one or more of the input events is further based on one or more of a self-learning rule set, a decision tree, a deep learning neural network, or other machine learning model.
21. 20. The apparatus of claim 17, wherein the raw data is received by a backend of a security server from multiple network nodes of a computer network or by a network node of a computer network.
22. The apparatus of claim 17 , wherein the filtering of the input events is performed by a backend of a security server and / or by a network node of a computer network.
23. 20. A computer program comprising computer readable code which, when executed on a computer system or server, causes said computer system or server to operate as a computer system or server according to claim 17.
24. 24. A computer storage medium having the computer program of claim 23 stored on a non-transitory computer readable medium.
Citation Information
Patent Citations
Machine Learning with Model Filtering and Model Mixing for Edge Devices in Heterogeneous Environments
JP2018510399A
Methods and apparatus for distributed use of a machine learning model
US20190042878A1
Detection of malicious network activity
US20190166144A1