Data processing method and data processing device
Identifying the exception path through the tag propagation method solves the high resource occupation problem caused by the construction of a complete origin map in the existing technology, real-time exception detection and path reconstruction are realized, and detection efficiency and accuracy are improved.
Patent Information
- Application Number
- PCT/CN2024/117446
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2024-09-06
- Publication Date
- 2025-07-17
AI Technical Summary
The existing technology needs to build a complete origin map when detecting abnormal behavior, resulting in high memory and computing resources occupancy, and real-time detection and reconstruction of attack paths cannot be achieved.
Through tag propagation, the tag information of the event start point is passed to the end point, and the exception path is identified, without the need to build a complete origin map. Only the local streaming origin map is output when the alarm conditions are met, reducing memory and computing resource usage.
Real-time detection and path reconstruction of abnormal behavior are realized, reducing the use of computing and memory resources, and improving the accuracy and efficiency of detection.
Smart Images

Figure CN2024117446_17072025_PF_FP_ABST
Abstract
Description
Data processing method and data processing device
[0001] This application claims priority to a Chinese patent application filed with the State Intellectual Property Office of China on January 11, 2024, with application number 202410045734.5 and application name “Data Processing Method and Data Processing Device,” the entire contents of which are incorporated herein by reference. Technical Field
[0002] The embodiments of the present application relate to the field of information security technology, and in particular to a data processing method and a data processing device. Background Art
[0003] Abnormal behavior detection is an important part of building a three-dimensional and in-depth intrusion detection system, and is of great value in protecting the security of information systems.
[0004] Currently, related technologies can obtain offline data by collecting offline data flows and control flows in the system after the fact; then, use the offline data to build a complete provenance graph; then, search for critical paths in the provenance graph to identify anomalies.
[0005] However, this solution requires building a complete provenance graph before detecting abnormal behavior, which consumes high memory and computing resources.
[0006] Summary of the Invention
[0007] The present application provides a data processing method and a data processing device, which can propagate the historical status information of the target path indicated by the label of the starting point in an event to the label of the end point in the event through label propagation. When the historical status information indicated by the label of the end point meets the alarm condition, the abnormal path with abnormal historical status information can be identified to identify the attack behavior. This process does not require the complete construction of the origin graph, which can reduce the usage of memory and computing resources.
[0008] In one possible implementation, an embodiment of the present application provides a data processing method. The method includes: determining a target event to be detected based on real-time collected log data, wherein the target event includes a starting point, an end point, and the relationship between the starting point and the end point; determining a second label of the end point based on a first label of the starting point; wherein the label of each node in the target event is used to indicate historical status information of a target path with the node as the end point, wherein the target path is a path in at least one path with the node as the end point, wherein the at least one path is a path formed by at least one event to be detected determined based on the log data; wherein the node in the target event includes a starting point and an end point in the target event; determining second path information related to the second label based on first path information related to the first label; when the second label meets an alarm condition, outputting a local streaming provenance graph related to the end point of the target event based on the second path information.
[0009] The log data is collected in real time from the device to be detected.
[0010] Log data can be converted into a data stream, which may include multiple events, each of which is a target event to be detected.
[0011] Among them, the historical status information of the target path can be at least one of the following: abnormality degree information, historical information between the node and the network (for example, the inflow of network data into the node, etc.), historical information between the node and the file (for example, information about the node reading sensitive files), historical information of calling or being called between the node and certain processes, etc.
[0012] For ease of explanation, the historical status information is taken as an example to illustrate the abnormality degree information. When the historical status information is other types of information, the implementation principle of the method of the present application is the same and will not be repeated here.
[0013] The starting point and end point of an event can each have their own label. The label is initially empty and contains no data. However, as the number of detected abnormal events increases, and at least one path exists that can terminate at the corresponding node, the information indicated by the node's label can be updated or assigned as historical status information for the target path terminating at the node. For example, the target path can be the path with the highest degree of abnormality among the at least one path terminating at the node, and the historical status information is information about the degree of abnormality of the target path. For example, the information indicated by the node's label can be an anomaly score AS indicating the degree of abnormality, or a normality score RS indicating the degree of normality of the path with the highest degree of abnormality, where RS + AS = 1.
[0014] Furthermore, based on differences in historical state information and application scenarios, the target path can be the path with the highest abnormality among the at least one path terminating at the node, a path with a higher abnormality among the at least one path terminating at the node, a path with a higher semantic quality among the at least one path terminating at the node, or a path with a network connection among the at least one path terminating at the node. The specific target path may vary depending on the application scenario and the specific content of the historical state information, and is not limited here.
[0015] Since the label of each node in the target event is used to indicate the historical status information of the target path with the node as the end point, when the node is the starting point of the target event, the target path is the first path information related to the first label; when the node is the end point of the target event, the target path is the second path information related to the second label.
[0016] In addition, regarding the alarm conditions satisfied by the second tag, the alarm conditions also vary based on the content of the historical status information of the target path indicated by the second tag. The alarm conditions can be flexibly configured based on the application scenario and the specific content of the historical status information.
[0017] For example, when the historical status information is abnormality degree information, the alarm condition may be that the abnormality degree exceeds a certain abnormality degree threshold, thereby triggering an alarm for the abnormal path (ie, the second path information related to the second tag).
[0018] For another example, if the historical status information is information related to at least one of a network and a file, the alarm condition may be the condition that the network has been connected and data has been sent out and sensitive files have been read. In this way, the path formed by the nodes that are connected to the network and read sensitive files can be output as the second path information.
[0019] Therefore, the alarm condition is related to the specific content of the historical status information and the scenario, and is not limited here.
[0020] The following implementation manner is described by taking the information stored in the tag as the normal score RS of the target path as an example.
[0021] In the embodiment of the present application, considering that a single event is difficult to support accurate detection of attack behavior, the present application can complete the discovery of abnormal behavior from the granularity of the path. The application can determine the abnormal path (such as the second path information mentioned above) for the log data collected in real time, and when the abnormal path meets the alarm condition, directly output a local streaming provenance graph based on the determined abnormal path. In this process, there is no need to build a complete streaming provenance graph, thereby reducing the memory and computing resources occupied by building the provenance graph. In addition, the method can perform real-time detection of abnormal paths that meet the alarm conditions for real-time log data, and real-time output of the streaming provenance graph composed of the abnormal path, which is conducive to the rapid real-time detection of attack behavior. Moreover, the historical state information of the target path indicated by the label of the starting point of the event can be passed to the label of the end point of the event, so that the target path of each node can be continuously determined in a label propagation manner (see above for explanation and definition), which can improve the accuracy of the determined abnormal path. And the determined abnormal path can be output in the form of a streaming provenance graph, which can realize the reconstruction of the path of the attack behavior, making it easier for operation and maintenance personnel to analyze the attack behavior.
[0022] In one possible embodiment, the method further includes: determining the normality of the target event; determining historical status information of a third path (such as abnormality information, etc.) based on the first label of the starting point and the normality of the target event; wherein the third path is a path formed by the first path information and the target event; and determining the second label of the end point based on the historical status information of the third path.
[0023] The normality of an event can be expressed by any information that can indicate the normality of the event. As the specific content of the historical status information differs, the information indicating the normality of the event may also differ. Because the historical status information is different, the definition of the abnormality of the abnormal behavior that needs to be detected may be different.
[0024] For example, when the historical status information is abnormality degree information, the normality of the event can be the frequency of the event (referring to normal events) mentioned in the embodiment. When the historical status information is other information, the information representing the normality of the event can be other frequency information of the event, which is not limited here.
[0025] For example, the first label of the starting point (eg, file A) in the target event is the normal score AS of the most abnormal path with file A as the end point that has been detected according to the method of the present application.
[0026] For example, the normality of the third path (eg, normal score RS) can be determined by multiplying the normality of the first label and the target event, and then the normality of the third path can be used to determine its abnormality (eg, abnormality score AS=1-RS).
[0027] Of course, when the historical status information is not abnormality degree information, the algorithm between the first label and the normality degree of the target event is not limited to the multiplication operation here, and can also be other calculation methods, which are not limited here.
[0028] In addition, the normality of the target event determined above may be a preset threshold, or may be read from, for example, the following normal behavior model, without limitation, and may be specifically associated with the content of the historical status information, which is not limited here.
[0029] In an embodiment of the present application, a streaming label propagation method is used to propagate local labels along the path, thereby facilitating fast, lightweight, and real-time detection of abnormal paths based on the labels at the endpoints.
[0030] In a possible implementation, the method further includes: determining the normality of the target event based on a pre-constructed information table; wherein the information table includes first information indicating the normality of a preset event.
[0031] Among them, the information table can be a normal behavior model, which can include the normality of multiple preset events (such as the frequency of event occurrence). Then, when a target event is detected using real-time log data, the normality of the target event can be determined through the normal behavior model.
[0032] In a possible implementation, the method further includes: when the target event is the preset event in the information table, determining the normality of the target event based on first information in the information table indicating the normality of the target event.
[0033] Wherein, when the target event belongs to a preset event in the information table, it means that the target event is a known event, and the first information can be used to determine the normality of the target event.
[0034] The first information may be, for example, the normality of a preset event in the normal behavior model (for example, the normality of event e i Frequency of occurrence M ei ), or other information that indicates the normality of the event.
[0035] Then, when determining the target event e that belongs to a known event i When the frequency M eiAs the target event e i Normal level, or, M ei *0.7+0.3 as the target event e i There is no restriction on the normality or other strategies.
[0036] In the embodiment of the present application, the normality of the target event currently being detected can be obtained from a pre-built information table, thereby facilitating rapid detection of abnormal behavior.
[0037] In a possible implementation, the method further includes: when the target event does not belong to the preset event in the information table, configuring a preset normality threshold (eg, a frequency threshold, such as 0.1, not limited) as the normality of the target event.
[0038] In the embodiment of the present application, M of the event that has not appeared in the information table ei The value (UNSEEN_EVENT_SCORE) cannot be too small to avoid insufficient information collected by the information table. At the same time, it needs to be distinguishable from the events in the information table (also expressed as known events). In actual operation, the M of the unseen events can be ei The value is set to 0.1 to indicate how normal this event is.
[0039] In one possible implementation, determining the historical status information of the third path based on the first tag of the starting point and the normality of the target event includes: when the first tag of the starting point is empty data, if the normality of the target event meets a first preset condition, determining the historical status information of the third path based on the normality of the target event.
[0040] When the label of the starting point is empty, the normality of the target event can express the historical status information of the target path (such as the abnormality). Here, the target path is the path of the target event. However, only when the target event is an abnormal event, it is necessary to determine the label of the end point in the target event. Therefore, when determining the normality of the target event (for example, the M of event 1), e1 ) When the first preset condition is met, the target event can be determined to be an abnormal event.
[0041] For example, the target event is event 1, and its frequency of occurrence in the normal behavior model is M. e1 , since the label of the starting point of event 1 is empty, so, M e1 is the historical status information (such as normal score RS1) of the target path (here, the path of event 1), then in the M e1 When it is less than the normality threshold (eg, 0.35), it indicates that the event 1 is an abnormal event.
[0042] Then we can base the normality of the target event (for example, M e1 ) to determine the historical status information of the third path. Here, the label of the starting point of event 1 is empty, so the third path is also the path of event 1. For example, if the historical status information is the abnormality degree, then the historical status information is RS1=M e1 Alternatively, the historical status information is AS=1-M e1 Accordingly, in combination with the above embodiment, when determining the second tag of the destination based on the historical status information of the third path, if the second tag of the destination currently has no information, for example, the value is empty, the second tag of the destination can be initialized to the historical status information of the third path (for example, M e1 ); For example, if the second tag of the destination has an abnormality level less than the historical status information of the third path (eg, abnormality level), the second tag of the destination is updated to the historical status information of the third path.
[0043] In the embodiment of the present application, when the tag of the starting point of the target event is empty, it means that no event with the starting point as the end point has been detected from the real-time log data. In this case, the normality of the target event (for example, M ei ) as the normal score RS of the third path, and when the RS is less than the normal score threshold, it means that the normal degree of the event meets the first preset condition, and the event is an abnormal time, so the AS of the third path is determined by using RS to determine its abnormal degree.
[0044] In one possible embodiment, the historical status information includes abnormality degree information, and determining the historical status information of the third path based on the first label of the starting point and the normality of the target event includes: determining the normality of the target path indicated by the first label based on the first label of the starting point; and determining the abnormality degree information of the third path based on the normality of the target path and the normality of the target event.
[0045] In the embodiment of the present application, when the tag of the starting point of the target event has data, it means that the method of the present application has detected an event with the starting point as the end point from the real-time log data. Then, the normality of the target path indicated by the first tag of the starting point (for example, AS) can be compared with the normality of the target event currently detected (for example, M ei ) performs operations such as multiplication to obtain the normal score RS of the third path, and then uses RS to determine the AS of the third path to determine its abnormality level.
[0046] In a possible implementation, determining the second tag of the endpoint based on the historical status information of the third path includes: when the second tag of the endpoint is empty data, initializing the second tag of the endpoint to indicate the historical status information of the third path.
[0047] In this embodiment, when the tag at the destination of the target event is empty, the tag at the destination of the target node is initialized to indicate the historical status information of the third path (e.g., abnormality level information). For example, the tag may store the normality score RS of the third path or the abnormality score AS of the third path, without limitation.
[0048] In a possible implementation, determining the second path information associated with the second tag based on the first path information associated with the first tag includes: when it is determined that the second path information is empty data, initializing the second path information associated with the second tag of the destination to the information of the third path.
[0049] When the second tag is empty data, the second path information related to the second tag is also empty.
[0050] Then, in the embodiment of the present application, the third path information (here, the path of the current target event) can be directly used as the second path information related to the label of the end point of the target event.
[0051] In one possible implementation, determining the second tag of the endpoint based on historical status information (e.g., degree of abnormality) of the third path includes: when it is determined that the second tag needs to be refreshed based on the historical status information of the third path and historical status information of the target path indicated by the second tag, updating the second tag of the endpoint to indicate the historical status information of the third path.
[0052] In this embodiment, the second tag of the destination is not empty data. Therefore, when determining whether the second tag needs to be refreshed, it can be determined based on the relationship between the historical status information of the third path and the historical status information of the target path indicated by the second tag.
[0053] For example, when the historical status information is abnormality degree information, when it is determined that the abnormality degree of the third path is greater than the abnormality degree of the target path indicated by the second label, it is determined that the second label needs to be refreshed, and the second label of the end point can be updated to information indicating the abnormality degree of the third path.
[0054] In an embodiment of the present application, the endpoint of a target event currently has a label, meaning its second label contains corresponding information. For example, the label of the node stores the RS of the most abnormal path terminating at the node. It is then possible to determine whether the abnormality level of the currently detected third path that passes through the target event and reaches the endpoint is greater than the abnormality level currently indicated by the second label of the target event's endpoint. If so, the second label needs to be updated, and thus updated to contain information about the abnormality level of the third path. This ensures that the label of the detected target event's endpoint always maintains the latest historical status information.
[0055] In a possible implementation, determining the second path information related to the second tag based on the first path information related to the first tag includes: updating the second path information related to the second tag of the destination to the information of the third path.
[0056] In this embodiment of the present application, when the label of the target event's endpoint is updated, the path information associated with the label also needs to be updated synchronously with the corresponding third path information, so that the node label and the path information associated with the label (e.g., the most abnormal path) remain synchronized.
[0057] In a possible implementation manner, the first label and the first path information, and the second label and the second path information are all cached in a memory.
[0058] In the above embodiment, the first label of the starting point of the target event of the present application and its first path information can be cached in the memory, and the label of the end point of the target event and its second path information can also be cached in the memory.
[0059] In this way, for a node in a detected event, the present application can cache information about the degree of abnormality of the most abnormal path with the node as the end point, as well as information about the most abnormal path, so that there is no need to completely construct a streaming origin graph formed by the paths corresponding to all events. Instead, it is only necessary to cache relevant information about the abnormal paths with abnormalities in the streaming origin graph, which can reduce cache occupancy and computing resource occupancy.
[0060] In a possible implementation, the method further includes: when at least one of the first tag at the starting point and the first path information related to the first tag meets a second preset condition, deleting the first tag information and the first path information cached in the memory.
[0061] This embodiment mainly expresses the situation of reducing the labels of nodes. The steps of this method can be performed before the label is propagated, for example, before the step of determining the second label of the end point based on the first label of the starting point.
[0062] In this embodiment of the present application, if the label of the starting point of the currently detected event, or at least one of its corresponding most anomalous paths, meets a second preset condition, the cached label information and path information for that starting point in memory can be deleted to reduce the further spread of the label. This is because attack behavior is not often mixed with normal behavior. When the path is very long but has not yet reached the alarm condition, clearing the label of the corresponding node can reduce the alarm rate for anomalous paths.
[0063] In a possible embodiment, the method further includes: when the historical status information of the target path indicated by the first label at the starting point meets the label reduction condition (for example, the abnormality degree is less than a second preset abnormality degree threshold), deleting the information of the first label and the first path information cached in the memory.
[0064] For example, in event e i If the RS (normality score) stored in the first label of the source node is greater than a preset normality score threshold, the path associated with the first label is not abnormal enough; or if the AS indicated by the first label is less than a preset abnormality score threshold, the path associated with the first label is not abnormal enough. In this case, the label of the source node and the most abnormal path information indicated by the RS in the label can be deleted to reduce memory usage. This prevents the propagation of labels with a normality score greater than the preset normality score threshold.
[0065] In a possible implementation, the method further includes: when a path length of the first path information associated with the first tag is greater than a preset length threshold, deleting the first tag information and the first path information cached in the memory.
[0066] Among them, event e i The path length length of the first path information related to the label of the starting point src (for example, the most abnormal path with the starting point as the end point of the path) can indicate the number of label propagation rounds;
[0067] When the number of label propagation rounds is greater than the preset round threshold, it means that the label has been propagated too many times without triggering an alarm. The label of the node and the historical status information indicated by the label (such as the most abnormal path information) can be deleted to prevent excessive label propagation.
[0068] In a possible implementation, the method further includes: when at least one of the second tag of the destination and the second path information related to the second tag meets a third preset condition, deleting the second tag information and the second path information cached in the memory.
[0069] This embodiment mainly expresses the situation of reducing the labels of nodes. The steps of this method can be performed after the label propagation, for example, after the step of determining the second label of the end point based on the first label of the starting point.
[0070] In this embodiment of the present application, if the label of the endpoint of the currently detected event, or at least one of its corresponding most anomalous paths, meets a third pre-set condition, the cached label and path information for that endpoint in memory can be deleted to reduce the further spread of the label. This is because attack behavior is not often mixed with normal behavior. When the path is very long but has not yet reached the alarm condition, clearing the label of the corresponding node can reduce the alarm rate for anomalous paths.
[0071] In one possible embodiment, the method further includes: when the historical status information of the target path indicated by the second label at the end point meets the label reduction condition (for example, the abnormality level is less than a preset abnormality level threshold, which may be the same as or different from the above-mentioned second preset abnormality level threshold), deleting the second label information and the second path information cached in the memory.
[0072] The implementation principle is the same as the principle of reducing the first label of the starting point, and will not be repeated here.
[0073] In a possible implementation, the method further includes: when the path length of the second path information related to the second tag is greater than a certain length threshold (which may be the same as or different from the above-mentioned preset length threshold), deleting the information of the second tag and the second path information cached in the memory.
[0074] Among them, event e i The path length length of the second path information related to the label of the end point dst (for example, the most abnormal path with the end point as the end point of the path) may indicate the number of label propagation rounds;
[0075] When the number of label propagation rounds is greater than the preset round threshold, it means that the label has been propagated too many times without triggering an alarm. The label of the node and the historical status information indicated by the label (such as the most abnormal path information) can be deleted to prevent excessive label propagation.
[0076] In one possible embodiment, the historical status information includes abnormality degree information; when the second tag meets the alarm condition, the local streaming origin graph related to the end point of the target event is output based on the second path information, including: when the abnormality degree information of the target path indicated by the second tag is greater than a third preset abnormality degree threshold, the local streaming origin graph related to the end point of the target event is output based on the second path information.
[0077] For example, the label of the node stores the normal score RS of the corresponding third path. When determining whether an alarm is needed for the third path, it can be determined whether the normal score RS is less than the preset normal score alarm threshold. When the normal score RS stored in the label is less than the preset normal score alarm threshold, it means that the label of the target node meets the alarm condition, and the second abnormal path in the streaming origin graph can be output.
[0078] For another example, the preset alarm threshold may also be a preset anomaly score alarm threshold. Then, when determining whether an alarm is required for the third path, the normal score RS indicated by the second label of the target node may be calculated as AS according to Formula 5; then, it may be determined whether the AS is greater than the preset anomaly score alarm threshold. When the AS is greater than the preset anomaly score alarm threshold, it indicates that the label of the target node meets the alarm condition, and the second anomaly path in the streaming origin graph may be output. The larger the preset anomaly score alarm threshold, the higher the difficulty in triggering the alarm. The threshold in the alarm condition may be flexibly set according to the need to trigger the alarm, such as a preset normal score alarm threshold, or a preset anomaly score alarm threshold or other threshold indicating the alarm condition.
[0079] In a possible embodiment, the outputting of the local streaming provenance graph related to the endpoint of the target event based on the second path information includes: aggregating the second path information and the local streaming provenance graph of the target node in the second path information to obtain an aggregated local streaming provenance graph; wherein the target node is a node whose label meets the alarm condition; and outputting the aggregated local streaming provenance graph as the local streaming provenance graph related to the endpoint of the target event.
[0080] In a possible implementation, the method further includes: updating the local streaming provenance graph of each node in the aggregated local streaming provenance graph to the aggregated local streaming provenance graph.
[0081] In a possible implementation, determining the second label of the end point based on the first label of the starting point includes: determining the second label of the end point based on the first label of the starting point and a preset normalization coefficient.
[0082] This can be combined with any of the above implementations of using the first label of the starting point to determine the second label of the end point. Each time the second label of the end point is determined, a preset normalization coefficient (eg, α) can be used.
[0083] In the embodiment of the present application, considering that the longer the path, the greater the anomaly score, a normalization coefficient is used to offset the effect of the path length on the accuracy of the label.
[0084] In one possible implementation, an embodiment of the present application provides a data processing device. The device may include: a first determination module for determining a target event to be detected based on real-time collected log data, wherein the target event includes a start point, an end point, and a relationship between the start point and the end point; a second determination module for determining a second label of the end point based on a first label of the start point; wherein the label of each node in the target event is used to indicate historical state information of a target path with the node as the end point; wherein the target path is a path in at least one path with the node as the end point, wherein the at least one path is a path formed by at least one event to be detected determined based on the log data; wherein the node in the target event includes a start point and an end point in the target event; a third determination module for determining second path information related to the second label based on first path information related to the first label; and an alarm module for outputting a local streaming provenance graph related to the end point of the target event based on the second path information when the second label meets an alarm condition.
[0085] In one possible embodiment, the device also includes: a fourth determination module, used to determine the normality of the target event; the second determination module is specifically used to: determine the historical status information of the third path based on the first label of the starting point and the normality of the target event; wherein the third path is a path composed of the first path information and the target event; based on the historical status information of the third path, determine the second label of the end point.
[0086] In a possible implementation, the fourth determination module is specifically configured to determine the normality of the target event based on a pre-constructed information table; wherein the information table includes first information indicating the normality of a preset event.
[0087] In a possible implementation, the fourth determination module is specifically configured to determine the normality of the target event based on first information in the information table indicating the normality of the target event when the target event is the preset event in the information table.
[0088] In a possible implementation manner, the fourth determining module is specifically configured to configure a preset normality threshold as the normality of the target event when the target event does not belong to the preset events in the information table.
[0089] In a possible implementation, the second determination module is specifically configured to determine the historical status information of the third path based on the normality of the target event when the first tag at the starting point is empty data and if the normality of the target event meets a first preset condition.
[0090] In one possible implementation, the historical status information includes abnormality degree information, and the second determination module is specifically used to: determine the normality of the target path indicated by the first label based on the first label of the starting point; and determine the abnormality degree information of the third path based on the normality of the target path and the normality of the target event.
[0091] In a possible implementation manner, the second determining module is specifically configured to initialize the second label of the endpoint to indicate historical status information of the third path when the second label of the endpoint is empty data.
[0092] In a possible implementation, the third determining module is specifically configured to initialize the second path information related to the second tag of the destination as the third path information when determining that the second path information is empty data.
[0093] In a possible implementation, the second determination module is specifically configured to update the second tag of the endpoint to information indicating the degree of abnormality of the third path when the second tag of the endpoint is not empty data and the degree of abnormality of the third path is greater than the degree of abnormality of the target path indicated by the second tag.
[0094] In one possible implementation, the third determination module is specifically configured to update the second label of the destination to indicate the historical status information of the third path when determining that the second label needs to be refreshed based on the historical status information of the third path and the historical status information of the target path indicated by the second label.
[0095] In a possible implementation manner, the first label and the first path information, and the second label and the second path information are all cached in a memory.
[0096] In a possible implementation, the device further includes: a deletion module configured to delete the information of the first label and the first path information cached in the memory when at least one of the first label at the starting point and the first path information related to the first label meets a second preset condition.
[0097] In a possible implementation, the deleting module is specifically configured to delete the first label information and the first path information cached in the memory when the historical status information of the target path indicated by the first label at the starting point meets a label reduction condition.
[0098] In a possible implementation, the deleting module is specifically configured to delete the first tag information and the first path information cached in the memory when a path length of the first path information associated with the first tag is greater than a preset length threshold.
[0099] In one possible embodiment, the historical status information includes abnormality degree information; the alarm module is specifically used to output a local streaming origin graph related to the end point of the target event based on the second path information when the abnormality degree information of the target path indicated by the second tag is greater than a third preset abnormality degree threshold.
[0100] In a possible embodiment, the device also includes: an aggregation module, used to aggregate the second path information and the local streaming provenance graph of the target node in the second path information to obtain an aggregated local streaming provenance graph; wherein the target node is a node whose label meets the alarm condition; and an output module, used to output the aggregated local streaming provenance graph as a local streaming provenance graph related to the end point of the target event.
[0101] In a possible implementation, the apparatus further includes: an updating module configured to update the local streaming provenance graph of each node in the aggregated local streaming provenance graph to the aggregated local streaming provenance graph.
[0102] The effects of the data processing devices of the above embodiments are similar to the effects of the data processing methods of the above embodiments, and will not be described in detail here.
[0103] In one possible implementation, an embodiment of the present application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the data processing method of the first aspect or any possible implementation of the first aspect.
[0104] The effect of the computing device cluster in this embodiment is similar to the effect of the data processing method in the above embodiments, and will not be repeated here.
[0105] In one possible implementation, an embodiment of the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, causes the computing device cluster to execute the data processing method of the first aspect or any possible implementation of the first aspect.
[0106] The effects of the computer program product of this embodiment are similar to the effects of the data processing methods in the above embodiments, and will not be described in detail here.
[0107] In one possible implementation, an embodiment of the present application provides a computer-readable storage medium including computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the data processing method in any one of the above implementations.
[0108] The effects of the computer-readable storage medium of this embodiment are similar to the effects of the data processing methods in the above embodiments, and are not described again here. BRIEF DESCRIPTION OF THE DRAWINGS
[0109] FIG1 is a schematic diagram of an exemplary system architecture;
[0110] FIG2 is a schematic diagram illustrating an exemplary data processing process;
[0111] FIG3a is a schematic diagram of a path between nodes shown as an example;
[0112] FIG3 b is a schematic diagram of paths between nodes shown as an example;
[0113] FIG3c is a schematic diagram showing an exemplary path between nodes;
[0114] FIG3 d is a schematic diagram of paths between nodes shown as an example;
[0115] FIG3e is a schematic diagram of paths between nodes shown as an example;
[0116] FIG4 a is a schematic diagram illustrating an exemplary data processing process;
[0117] FIG4 b is a schematic diagram illustrating an exemplary data processing process;
[0118] FIG4c is a schematic diagram illustrating an exemplary data processing process;
[0119] FIG5 is a schematic diagram illustrating an exemplary local loss origin map;
[0120] FIG6 is a block diagram of an exemplary data processing device;
[0121] FIG7 is a schematic diagram illustrating the structure of a computing device;
[0122] FIG8 is a schematic diagram illustrating the structure of a computing device;
[0123] FIG9 is a schematic diagram showing the structure of an exemplary computing device cluster. DETAILED DESCRIPTION
[0124] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0125] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0126] In the description and claims of the embodiments of this application, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order of objects. For example, the terms "first target object" and "second target object" are used to distinguish different objects, rather than to describe a specific order of objects.
[0127] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0128] In the description of the embodiments of this application, unless otherwise specified, "multiple" means two or more. For example, "multiple processing units" means two or more processing units; "multiple systems" means two or more systems.
[0129] In order to facilitate understanding of the technical solutions of the embodiments of the present application, before describing the technical solutions of the embodiments of the present application, a brief introduction to the technical terms involved in the embodiments of the present application is first given:
[0130] Event triples: Most behaviors in systems can be abstracted as events (s, p, o), consisting of a subject, a predicate, and an object. Here, s represents the performer or subject of the action, and o represents the object or object of the action. p represents the action or predicate, and p represents the relationship or attribute between objects. o represents the time when the event occurs.
[0131] Streaming provenance graphs (SRGs) contain both vertex and edge information, but are not necessarily in graph format. A SRG can construct a provenance graph describing system behavior by treating the subject in the aforementioned event triples as the source node (also known as the starting point), the object as the target node (also known as the end point), and the predicate as the edge. The chronological data stream of these event triples is the SRG. SRGs share most of the characteristics of graphs and support customized graph computing algorithms.
[0132] Path: A path in the provenance graph is formed by connecting several event triplets end to end in chronological order. The chronological order is the order of the events from the earliest to the latest time.
[0133] Tag Propagation: This is a customized graph computation algorithm specifically suited for streaming provenance graph analysis. It uses the labels of labeled nodes to predict the labels of unlabeled nodes. The Tag Propagation algorithm stores the results of local computations in the node's label and propagates the node's label along the event direction (i.e., the edges in the streaming provenance graph) to the next node in the streaming provenance graph, completing the overall computation.
[0134] Anomaly Path: Due to the covert nature of attacks, a single event is difficult to accurately detect. Therefore, this application can detect abnormal behaviors at the granularity of the path, i.e., the anomaly path.
[0135] Abnormal behavior detection is a crucial component of a comprehensive, in-depth intrusion detection system and is extremely valuable for protecting information system security. A good abnormal behavior detection solution must first accurately identify abnormal behavior from a vast array of normal behaviors. Failure to accurately identify abnormal behavior or excessive false positives significantly reduces the solution's value. Secondly, it must provide sufficient information to help analysts understand, confirm, and address attacks.
[0136] Traditional host-side (endpoint) anomaly detection solutions are primarily based on dynamic and static analysis of individual malicious programs, or analysis of specific logs. However, as network attacks evolve into advanced persistent threats (APTs), attack behaviors become more holistic, making localized behaviors difficult to distinguish from normal behavior.
[0137] In the traditional solution 1, anomaly detection can be performed through a process whitelist (as a trusted process). Specifically, if a process is not in the process whitelist or is not a descendant process of a trusted process in the process whitelist, it is determined to be an anomaly.
[0138] However, attackers can easily exploit existing processes using methods like memory injection to launch attacks. This type of attack can effectively bypass deterministic anomaly detection. For example, during memory injection, malicious code can be injected into a whitelisted process running in memory to bypass detection and launch an attack.
[0139] In traditional solution 2, anomaly detection can be performed based on anomaly analysis of process call chains. Specifically, all process chains with normal behavior can be collected and stored as a set as a normal behavior model. Any process call chain not included in the set will trigger an alert.
[0140] While this method can detect attacks based on malicious processes, attackers can also exploit existing processes using methods like memory injection rather than creating new ones. This type of attack can also bypass anomaly detection based on process call chains.
[0141] Traditional solution three involves anomaly detection based on a provenance graph. In this solution, data and control flows within the system are collected post-incident to generate offline data. This offline data is then used to construct a complete provenance graph. Next, critical paths are searched within the provenance graph to identify anomalous paths. For example, algorithms such as graph matching and similarity calculation can be used to capture anomalous behavior within the global attack (a subgraph consisting of interactions between multiple system entities).
[0142] However, traditional solution three is a post-hoc algorithm that identifies abnormal paths from a constructed complete provenance graph. It cannot detect abnormal behavior in real time and cannot reconstruct attack paths in real time, making it difficult for alarm analysts to operate and understand. In addition, the algorithm for constructing the provenance graph has a high computational overhead, so constructing a complete provenance graph will occupy a large amount of computing resources. In addition, the process of constructing the complete provenance graph will occupy a large amount of cache space.
[0143] To this end, the present application provides a data processing method, which can collect log data in real time for determining events, and the event includes a starting point and an end point, and the starting point and the end point can each have their own labels. For example, the label is initially empty (specifically, it can also be non-empty), and the label of a node can be used to indicate the historical status information of the target path with the node as the end point, wherein the target path can be one of at least one path with the node as the end point. Wherein, the at least one path is a path composed of at least one event (also referred to as an event to be detected) determined based on the log data. The present application can determine the label of the end point of the event based on the label of the starting point in the event through label propagation, and determine the label-related path of the end point (that is, the path indicated by the label determined for the end point) based on the path related to the label of the starting point (that is, the target path indicated by the label); then, when the label meets the alarm condition, the label-related path can be shown in the form of a local streaming provenance graph. This method does not require a complete provenance graph to be constructed first. Instead, it only processes the real-time collected log data to determine the abnormal paths that meet the alarm conditions and reach the corresponding nodes. The abnormal paths with alarms are then reconstructed and output to obtain a local provenance graph, thereby achieving the detection of abnormal behavior. In this method, the abnormal paths (e.g., the paths related to the node labels, such as the target paths) can be determined by real-time analysis of log data, which can achieve real-time detection of abnormal behavior. The abnormal paths with alarms can also be reconstructed and output in real time to assist alarm personnel in analyzing and understanding abnormal behaviors that may constitute attacks and making corresponding decisions. In addition, this process does not require a complete provenance graph to be constructed before detecting abnormal behavior. Only a local provenance graph is constructed for the abnormal paths with alarms, which can significantly reduce the computational resources and cache space occupied by the provenance graph algorithm. The historical state information of the target path indicated by the label of the event starting point can be transferred to the label of the event end point, so that the target path of each node can be continuously determined in a label propagation manner (see above for explanation and definition), which can improve the accuracy of the determined abnormal paths. Furthermore, the identified abnormal path can be output in the form of a streaming origin graph, which enables reconstruction of the attack path and facilitates operation and maintenance personnel to analyze the attack behavior.
[0144] Among them, the historical status information of the target path can be at least one of the following: abnormality degree information, historical information between the node and the network (for example, the inflow of network data into the node, etc.), historical information between the node and the file (for example, information about the node reading sensitive files), historical information of calling or being called between the node and certain processes, etc.
[0145] For the sake of convenience, the historical status information is taken as abnormality degree information, and the target path is taken as an example to illustrate the path with the highest abnormality among at least one path with the node as the end point. When the historical status information is other types of information, the implementation principle of the method of the present application is the same and will not be repeated here.
[0146] FIG1 is a schematic diagram showing an exemplary system architecture of the present application.
[0147] As shown in FIG1 , the system may include a streaming platform 100 and terminals 1 to n.
[0148] The streaming processing platform 100 can be implemented as software or hardware.
[0149] In one possible implementation, the stream processing platform 100 may be deployed on a cloud computing node, a computing node and a storage node, or a computing cluster and a storage cluster. The computing cluster may include multiple computing nodes, and the storage cluster may include multiple storage nodes. The computing nodes are computing servers for computing, and the storage nodes are storage servers for storage.
[0150] In one possible implementation, the streaming processing platform 100 can also be deployed on a cloud-based device. For example, the device can be at least one of a terminal device and a cloud-based server. The terminal device can be a personal computer, a laptop, a wearable device, etc. The cloud-based server can be a single server or a server cluster.
[0151] In addition, the number of terminals shown in FIG1 may be one or more, n≥1, and the number of n is not specifically limited.
[0152] The terminal shown in Figure 1 is a device whose system is to be tested for abnormal attacks. The streaming processing platform 100 shown in Figure 1 can be used to perform real-time detection of attack behaviors on any terminal.
[0153] Specifically, as shown in FIG1 , each terminal from terminal 1 to terminal n may be deployed with a collection module 201 .
[0154] The collection module 201 can be used to collect log data on the terminal in real time, and report the collected log data to the streaming processing platform 100 (eg, a computing node deployed on the cloud) in real time.
[0155] The log data may include system control flow and data flow information.
[0156] For example, the log data may include system control classes and data flow information of operations related to the entity.
[0157] The entity may include but is not limited to at least one of the following: a file entity, a process entity, and a network entity.
[0158] In one possible implementation, the log data may include: data on at least one operation, including file-related operations (such as opening a file, creating a new file, etc.), process-related operations, and network-related operations (such as Domain Name System (DNS) queries, Transmission Control Protocol (TCP) connections, etc.).
[0159] As shown in FIG1 , the streaming processing platform 100 can read the log data reported by each terminal in real time, and detect abnormal behaviors on the corresponding terminal and restore the attack path of the abnormal behaviors based on the log data.
[0160] In one possible implementation, as shown in Figure 1, the streaming processing platform 100 may include but is not limited to at least one of the following modules: an event frequency calculation module 101, an event frequency storage module 102, an abnormal path mining module 202, a node data cache module 203, an abnormal path alarm module 204, an abnormal path aggregation module 205, etc.
[0161] Among them, the event frequency calculation module 101 and the event frequency storage module 102 are mainly used to construct an offline normal behavior model.
[0162] The abnormal path mining module 202, the node data cache module 203, and the abnormal path alarm module 204 are mainly used to perform real-time abnormal path detection based on label propagation according to a pre-built normal behavior model.
[0163] The abnormal path alarm module 204 may be configured to output abnormal paths for which an alarm is issued.
[0164] The abnormal path aggregation module 205 is mainly used to aggregate the abnormal paths with alarms detected above, and output the aggregated abnormal paths (a local stream provenance graph) to complete the restoration of the attack path.
[0165] In a possible implementation, the event frequency calculation module 101 can be used to calculate the frequency of occurrence of each single event involved in the offline log data, wherein the offline log data is given log data related to only normal events.
[0166] The frequency indicates the normality of a single event. In this embodiment, the single event involved in the offline log data is a normal event by default, rather than an abnormal event to be detected by this application. In actual applications, the events involved in the real-time collected log data can be divided into normal events and abnormal events. This application is intended to detect abnormal events.
[0167] The event frequency storage module 102 may be used to store the occurrence frequency of each single event calculated by the event frequency calculation module 101 to obtain a normal behavior model, which may be stored in a database (hereinafter referred to as a baseline database).
[0168] The event frequency storage module 102 may store the occurrence frequency of each single event in a matrix format to obtain a normal behavior model (eg, a matrix).
[0169] Before the streaming processing platform 100 of the present application detects and restores abnormal paths, the present application can obtain a normal behavior model through the event frequency calculation module 101 and the event frequency storage module 102, and then use the normal behavior model to detect and restore abnormal paths for the log data collected in real time.
[0170] In some embodiments, since events belonging to normal events may change, the offline log data used to construct the normal behavior model may be updated regularly, and the regularly updated offline log data is log data related to the updated normal events.
[0171] As shown in Figure 1, the event frequency calculation module 101 and the event frequency storage module 102 can run periodically so that the event frequency calculation module 101 can recalculate the frequency of occurrence of each single event related to the updated offline log data (such as user input) based on the updated offline log data, and pass the frequency of occurrence of the event to the event frequency storage module 102.
[0172] The event frequency storage module 102 can update the single events and their occurrence frequencies in the pre-built normal behavior model based on the occurrence frequencies of each single event regularly sent by the event frequency calculation module 101, so that the occurrence frequencies of events in the normal behavior model of this application are consistent with the offline log data after the most recent update.
[0173] In one possible implementation, the event frequency calculation module 101 can count the frequency of a single event occurring within a certain time range in offline log data. For example, the certain time range can be the most recent week, the most recent month, or a natural year, and the specific time range can be set according to needs and is not limited.
[0174] For example, the single event may be process A opening file B. Then, the frequency of occurrence of the single event in the past week can be counted for offline log data.
[0175] In a possible implementation, the event frequency calculation module 101 may further calculate the frequency of a single event occurring on a group of hosts based on offline log data.
[0176] The host may be a physical host, such as a terminal as shown in FIG1 . The host may also be a virtual instance (such as a virtual machine or a container, etc.) running on a physical device, which is not limited here.
[0177] For example, the single event may be process A opening file B. Then, the frequency of occurrence of the single event on 10 hosts can be counted for offline log data.
[0178] In a possible implementation, the event frequency calculation module 101 may also count the frequency of a single event occurring on a group of hosts within a certain time range based on offline log data.
[0179] Example 1
[0180] The following example 1 is used to illustrate the process of the streaming processing platform 100 building a normal behavior model.
[0181] This application can define a standard single event, event e i =(src i ,dst i ,rel i ), where, according to the event triple, src i Indicates event e i The main body (corresponding to the source node of the origin graph), dst i Indicates event e i The object (corresponding to the target node of the origin graph), rel i Indicates event e i The predicate of (corresponding to the edge between the source node and the target node of the origin graph, that is, the relationship).
[0182] The event frequency calculation module 101 can count the number of events e that meet the conditions within a specific time range t on a group of hosts (the number of hosts is h) for offline log data. i Frequency of occurrence M ei .
[0183] Among them, event e i The above condition is that the source node of the event is src i , the target node of the event is dst i , the relationship of the event is rel i .
[0184] Where i is a positive integer, 1≤i≤n, where n is the maximum number of types of single events that appear in the log data.
[0185] For example, event e1 is the first normal event: process A opens file B.
[0186] For another example, event e2 is the second normal event: process 1 opens file 2.
[0187] For example, the event frequency calculation module 101 can calculate the frequency of each single event e in the offline log data by using formula 1, formula 2, and formula 3. i Frequency of occurrence M ei .
[0188] In Formula 1 and Formula 2, j represents host j, 1≤j≤h, and the number of hosts to be counted is h.
[0189] In formula 1, Freq(e i ) represents event e i The number of times the event occurs repeatedly in the log data (also called frequency), where the number is the number of times the event occurs repeatedly in the log data (also called frequency). i The number of occurrences sum(src i ,dst i ,rel i ,j,t). In the following, Freq(e i ) is referred to as event e i Frequency of occurrence.
[0190] Optionally, to avoid an event on a single host, i The number of times is too high, thus affecting the event e i The statistical frequency M ei The accuracy of the normal behavior model is reduced, and the accuracy of the event e i The number of times sum(src i ,dst i ,rel i ,j,t), if the number is greater than or equal to 1, then sum(src i ,dst i ,rel i ,j,t)=1, if the number is 0, then sum(src i ,dst i ,rel i ,j,t)=0.
[0191] For example, event e i If the file B is opened for process A, then the event of process A opening file B occurs 3 times on host 1 within the time range t, then sum(src i ,dst i ,rel i ,1,t)=1, instead of taking the value as 3.
[0192] In formula 2, This means that the event e is ignored. i dst in i Events i ' is the number of times it occurs repeatedly in the log data (also called frequency). In formula 2, "*" indicates that dst is ignored. i The number is the number of times the event e i 'In a specific time range t, the number of occurrences on each host sum(src i ,*,rel i ,j,t). In the following, Referred to as event e i ′The frequency of occurrence.
[0193] Similarly, optionally, to avoid an event e occurring on a single host i ′ is too high, thus affecting the event e i The statistical frequency M ei The accuracy of the normal behavior model is reduced, and the accuracy of the event e i ′The number of times it occurs on a host j within a specific time range t is sum(src i ,*,rel i ,j,t), if the number is greater than or equal to 1, then sum(src i ,*,rel i ,j,t)=1, if the number is 0, then sum(src i ,*,rel i ,j,t)=0.
[0194] Let's take an example to understand, for example, event e i Open file B for process A, then event e i 'Open the file for process A, then Freq(e i ) is the number of times that process A opens file B on a group of hosts within a specific time range t, calculated from offline log data; then The number of times that process A opens a file within a specific time range t on the above set of hosts is calculated based on the offline log data.
[0195] Referring to the above formula 3, the ratio between the absolute value of the result of formula 1 and the absolute value of the result of formula 2 is the frequency M of the event of process A opening file B on a group of hosts within a specific time range t, calculated for offline log data. ei .
[0196] In this way, according to the above formulas 1, 2 and 3, each single event e that meets the conditions in the offline log data can be calculated. i Frequency of occurrence M ei In the following, M ei Referred to as event e i Frequency of occurrence.
[0197] Then, the event frequency calculation module 101 can calculate each single event e based on the offline log data. i The frequency of occurrence M ei The data is transmitted to the event frequency storage module 102 .
[0198] Finally, the event frequency storage module 102 can store each single event e i The frequency of occurrence M ei It is stored as a matrix (also called a transition probability matrix), which can be stored in the baseline database as a normal behavior model.
[0199] Example 2
[0200] The following describes the implementation process of the streaming processing platform 100 processing offline log data to construct a normal behavior model with reference to Example 2.
[0201] Please refer to FIG2 , which exemplarily shows a schematic diagram of a process for constructing a normal behavior model. As shown in FIG2 , the process may include the following steps:
[0202] S101: Determine the data type of each piece of received log data.
[0203] The offline log data can be converted into an offline data stream. In the subsequent embodiment of detecting abnormal paths for real-time log data, the real-time log data can also be converted into a real-time data stream.
[0204] A data stream may include multiple pieces of data arranged in chronological order. Each piece of data may have a data type. The data type may be an entity type, which means that the piece of data is information about an entity. Alternatively, the data type may be an event type, which means that the piece of data is an event and has event information about the event.
[0205] As described above, the log data may include system control flow and data flow information for operations associated with the entity.
[0206] The entity may include but is not limited to at least one of the following: a file entity, a process entity, and a network entity.
[0207] Then the type of event may include but is not limited to: Event type (ie event type) and Entity type, wherein the Entity type may be a file entity, a process entity, or a network entity.
[0208] S102a: When the data type of the received data is Entity type, the entity information of the entity is written into the node information table as node information of a node.
[0209] For example, if the data type of a received piece of data is an entity type, for example, the received piece of data is information about process A, then the data can carry relatively complete entity information about process A. A node A about process A can be created, and the entity information of process A can be stored as the node information of node A in the node information table.
[0210] The node information table may be a data table in a database, and the storage structure of the node information table is not limited here.
[0211] S102b: When the data type of the received data is Event type, extract the corresponding source node src and target node dst for the event.
[0212] For example, if the data type of a received data piece is an event type, it means that the data piece is an event, for example, the event is: process A opens file B. Then, the information that the source node src is process A and the information that the destination node dst is file B can be extracted from the event.
[0213] However, the entity information about the source node and the target node in the event information of the data is incomplete. For example, the information about the source node and the target node in the event information may only include the node ID. Therefore, the following S103 needs to be executed.
[0214] S103: Check in the node information table whether a source node src and a destination node dst exist.
[0215] After S103 , when it is determined that at least one of the source node src and the target node dst does not exist in the node information table, the process ends.
[0216] If the node information table does not contain both the source node src and the destination node dst, for example, the node information table does not contain node information for process A or file B. Alternatively, the node information table contains node information for process A but not for file B. Alternatively, the node information table does not contain node information for process A but contains node information for file B. This indicates that the complete entity information for the source and destination nodes in the current Event type event cannot be fully obtained, and the process ends.
[0217] After S103 , when it is determined that the source node src and the destination node dst exist in the node information table, S104 may be executed.
[0218] S104, obtaining node information of the source node src and the destination node dst.
[0219] Among them, when there is a source node (such as process A) and a target node (such as file B) in the node information table, the node information of the source node (such as process A) and the node information of the target node (such as file B) can be read from the node information table to obtain the entity information of each of the two entities in the event.
[0220] S105: Construct a single event based on the node information of the source node src and the target node dst, and the event information of the above event.
[0221] The event information of the event is the event description information of a piece of data (here, an event) whose data type is the Event type currently received.
[0222] In addition, the complete entity information of the source node and the target node in the event is also obtained through the above S104.
[0223] In this way, the three types of information mentioned above can be used to construct a single event.
[0224] Optionally, S106, generalize the single event to obtain the generalized event e i .
[0225] In order to ensure the normal behavior model (e.g. by all M ei In order to improve the transferability of the transition probability matrix formed by the event triples, and avoid the model being too large, the embodiment of the present application performs a generalization operation on the entity information corresponding to the source node (subject, also called the starting point) and the target node (object, also called the end point) of the event triples, wherein the strategy of the generalization operation may include but is not limited to: abstracting non-deterministic information, removing non-deterministic information, abstracting or removing limiting information about specific scenarios, etc.
[0226] The generalization strategy of entity information for different types of entities is given below:
[0227] Process-type entities: The attribute fields of process-type entities generally only retain: 1. The corresponding executable file path, 2. The command line parameters.
[0228] File entities: Mainly retain the file path, generalize the path, and remove user-related information. For example, / home / user / mediaplayer will be changed to / home / * / mediaplayer.
[0229] Network (socket) entities: Each network connection entity contains two addresses: the source address of the source node and the destination address of the destination node, along with their corresponding port numbers. Generally speaking, the external address is retained while the internal address is removed. For example, for outbound links (such as requests sent to an external server), the source address and port number are removed while the destination address and port number are retained. For inbound links (such as requests received from an external server), the destination address and port number are removed while the source address and port number are retained.
[0230] The generalization configuration for entities can be customized. To make the trained normal behavior model more applicable across different hosts, it's necessary to remove fields related to the host, user, and network environment from the entity fields. For example, " / home / zhenyuan / core.sh" should be generalized to " / home / * / core.sh," and the local IP address and port of the network connection can be ignored. Specific generalization operations require analysis based on specific scenarios.
[0231] S107, determining a group of hosts within a specific time range t, each event e in the event stream i (For example, the generalized event e i ) frequency of occurrence M ei .
[0232] For example, according to Formulas 1 to 3 in Example 1 above, each generalized event e i Calculate its frequency M ei .
[0233] In formula 1, how to determine two events as the same event e i , we can compare whether the node information (such as entity information) of the source nodes of the two events is the same, whether the node information (such as entity information) of the target nodes of the two events is the same, and whether the relationship (also called predicate) between the two events is the same. If the node information of the source node, the node information of the target node and the relationship between the two events are the same, then it means that they are the same single event.i , which can realize the single event e i Calculation of frequency of occurrence.
[0234] In formula 2, how to determine that two events are the same event e i ', we can compare whether the node information (such as entity information) of the source nodes of the two events is the same, and whether the relationship (also called predicate) of the two events is the same. If the node information and relationship of the source nodes of the two events are the same, then it means that they are the same event. i ′, so that the event e can be realized i Calculation of the frequency of occurrence.
[0235] Finally, use the above formula 3 to calculate the generalized event e i The frequency of occurrence M ei .
[0236] S108, based on each event e i Frequency of occurrence M ei , generating a normal behavior model, and optionally writing the normal behavior model into a baseline database.
[0237] The value of i is a positive integer and is not specifically limited. Its specific value depends on the type of single event involved in the currently processed log data.
[0238] In this step, multiple events can be i The respective event frequency M ei Stored as a table as a normal generation model.
[0239] Optionally, the normal behavior model may be written into a database.
[0240] In the embodiment of the present application, when constructing the normal behavior model, each single event e in the offline log data can be counted. i Frequency of occurrence M ei. In this process, events can be first extracted from the log data, and the entity information of the events can be generalized. Then, multiple generalized events obtained by processing the log data can be compared to see if they are the same, so as to determine the frequency of occurrence of each generalized event, and then obtain the frequency of occurrence of the corresponding event. In this process, before calculating the corresponding frequency of the event, the event extracted from the log data can be generalized first. On the one hand, the number of event types involved in the normal behavior model obtained by this application can be reduced, thereby avoiding the problem that the frequency of occurrence of events stored in the normal behavior model is too high, which leads to an overly large normal behavior model. On the other hand, by generalizing the entity information of the event, the single event corresponding to the frequency stored in the normal behavior model can be an event after the entity information is generalized, so that the frequency of occurrence of the event statistics can be applied to other scenarios, ensuring the portability of the normal behavior model. In addition, in the process of Figure 2 above, by writing the entity information in the identified Entity type data into the node information table, when an Event type event is subsequently received, the complete node information (such as entity information) of the source node and target node of the event can be obtained from the node information table, which is convenient for comparing the generalized events (for example, the source node, target node and relationship need to be compared in the above formula 1, while the target node does not need to be compared in formula 2).
[0241] In a possible implementation, the offline log data may be updated periodically. Thus, the process shown in FIG2 may be executed periodically to periodically update the occurrence frequency M of normal events in the normal behavior model. ei .
[0242] Example 3
[0243] The following describes how the stream processing platform 100 uses offline log data to generate the frequency M of each single event involved in the log data, in conjunction with Example 3. ei process.
[0244] In step 1, the log data of the test machine within 3 hours can be collected, and each event in the log data can be analyzed one by one, and the statistical data of each event can be updated.
[0245] S10, when analyzing the event / var / spool / cron / root->FILE_READ->cron, the frequency statistics of the event are updated by combining the above formula 1 (1) Freq( / var / spool / cron / root,FILE_READ,cron)+=1, and the above formula 2 to update the frequency statistics of the event (2) Freq( / var / spool / cron / root,FILE_READ)+=1;
[0246] Among them, the event i The file in the path / var / spool / cron / root is read by the process cron, where the file in the path / var / spool / cron / root is the source node and the process cron is the target node.
[0247] Among them, the event i ′ means the file with the path / var / spool / cron / root is read, where the source node is the file with the path / var / spool / cron / root, the relationship rel is the file read operation, and the target node is omitted.
[0248] S20, when analyzing the event cron->PROCESS_FORK->bash (an event in which a process creates a child process), the frequency statistics of the event are updated by combining the above formula 1 (1) Freq(cron,PROCESS_FORK,bash)+=1, and the above formula 2 to update the frequency statistics of the event (2) Freq(cron,PROCESS_FORK)+=1;
[0249] S30, when analyzing the event bash->PROCESS_FORK->bash (an event in which a process creates a child process), the frequency statistics of the event are updated by combining the above formula 1 (1) Freq(bash,PROCESS_FORK,bash)+=1, and the above formula 2 to update the frequency statistics of the event (2) Freq(bash,PROCESS_FORK)+=1;
[0250] By analogy, all events in the log data can be processed and the corresponding frequency of each event can be counted according to Formula 1 and Formula 2.
[0251] Step 2: Traverse the frequency statistics (1) and frequency statistics (2) of each event in step 1 above, and calculate the occurrence frequency of each single event according to the above formula 3.
[0252] For the event in S10, its occurrence frequency M( / var / spool / cron / root,FILE_READ,cron)=Freq( / var / spool / cron / root,FILE_READ,cron) / Freq( / var / spool / cron / root,FILE_READ);
[0253] For the event in S20, its occurrence frequency M(cron,PROCESS_FORK,bash)=Freq(cron,PROCESS_FORK,bash) / Freq(cron,PROCESS_FORK);
[0254] For the event in S30 above, its occurrence frequency M(bash,PROCESS_FORK,bash)=Freq(bash,PROCESS_FORK,bash) / Freq(bash,PROCESS_FORK);
[0255] By analogy, the frequency of occurrence of all events M can be calculated. ei , to build a model of normal behavior.
[0256] Example 4
[0257] The following describes the implementation principle of the streaming processing platform 100 of the present application using a label propagation method to calculate the anomaly score of a path.
[0258] Formula 3 above is used to calculate the frequency of a single event (deemed a normal event). The result of Formula 3 can reflect the normality of a single event. The stream processing platform 100 can then use the result of Formula 3 to further calculate the normality score of any path using the principle of Formula 4 to reflect the normality of the path, and calculate the abnormality score of the path using Formula 5 to reflect the abnormality of the path. The path is the most abnormal path to the target node in a received single event, calculated during the real-time detection of abnormal behavior in log data.
[0259] Among them, RS is the normal score of path P, length is the length of path P (specifically, the number of single events constituting path P), M ei The single event e calculated by the above formula 3 i The frequency of occurrence M ei Formula 4 represents the M of each single event constituting the path P. ei Perform cumulative multiplication to obtain the normal score of the path P.
[0260] For easier understanding, reference may be made to the path schematic diagram shown in FIG3 a .
[0261] For example, in the log data, events 1 and 2 occur in chronological order. Event 1: File A is read by process B; Event 2: Process B creates child process C. The source node of event 1 is node A (representing file A) as shown in Figure 3a, and its target node is node B (representing process B) as shown in Figure 3a. Path 1: from node A to node B, can represent event 1.
[0262] Similarly, node C shown in Figure 3a can be used as the target node of event 2 to represent process C in event 2, and path 2: from node B to node C, can represent event 2.
[0263] For example, the occurrence frequency of event 1 stored in the normal behavior model is M e1 , the frequency of event 2 is M e2 .
[0264] Path P is path 1 + path 2, that is, the path from node A to node C via node B.
[0265] Then the length of path P is 2 (including the above two events 1 and 2). Then according to formula 4, the normal score of path P is RS(P) = M e1 *M e2 .
[0266] Formula 4 reflects the cumulative multiplication of the normality of multiple events, event e i The more normal, The closer it is to 1, the closer RS is to 1.
[0267] Intuitively, the abnormality of a path is the cumulative abnormality of each event in a single path. i The more abnormal, The smaller the value, the greater the abnormality of the path, and the smaller RS(P). Therefore, in order to intuitively reflect the abnormality of the path, the abnormality score AS(P) of the path P can be calculated using Formula 5 to express the abnormality of the path P:
[0268] AS(P)=1-RS(P), Formula 5;
[0269] In one possible implementation, the streaming processing platform 100 may use the normal score RS(P) of the path P as a label of the target node (also called the end point, such as the node C shown in FIG3 a ) of the path P.
[0270] In a possible implementation, referring to the principle of formula 4, when the stream processing platform 100 calculates the label of the target node in a single event, it can calculate the label of the source node in the single event by comparing the label of the source node in the single event with the frequency M of the single event. eiMultiply them to get the label of the target node.
[0271] Continuing with Figure 3a as an example, after event 1 is detected, for example, node A has no label, that is, no event with node A as the target node (such as an abnormal event) is detected, then RS(1) of path 1 can be calculated according to the above formula 4 as M e1 , where M e1 is the frequency of occurrence of the above event 1. Then, this application can take the value as M e1 RS(1) is used as the label of the target node (here, node B) in path 1.
[0272] Then, the stream processing platform 100 of the present application detects event 2 (occurring after event 1) in the log data, and then obtains the label of the source node in event 2, that is, the label of node B, and obtains the frequency M of event 2 from the normal behavior model. e2 , then when determining the label of the target node in event 2 (here is node C), we can refer to the principle of formula 4 and compare the label of the source node of event 2 (here is node B) with the frequency M of event 2. e2 Multiplying them together to achieve the cumulative multiplication of the normal scores of path 1 and path 2 on path P, thereby obtaining the normal score RS(P)=M of path P (the complete path composed of path 1 and path 2 as shown in FIG3a ). e1 *M e2 and the normal score RS(P)=M e1 *M e2 As the label of the target node (here is node C) in event 2. In this way, the label of the source node (node B) in event 2 and the frequency M of the current event 2 can be e2 Calculate the normal score RS(P) of path P. When the normal score RS(P) is less than the normal score RS in the label of the target node (node C) in event 2, the label propagated from the source node (here, node B) can be used to update the normal score of the most abnormal path (e.g., path 1 + path 2) to the target node in the current event 2.
[0273] Of course, in some embodiments, the information stored in the label of the target node of the event may also be other information that can indicate the anomaly score of the corresponding most abnormal path. For example, the label may store an anomaly score (such as AS(P) calculated by Formula 5), etc., which is not limited here.
[0274] In one possible implementation, for an arbitrary path, wherein the arbitrary path is the most abnormal path to the target node in the single event calculated for a received single event in the process of real-time detection of abnormal behavior in log data. The present application can calculate the abnormal score of the arbitrary path by the principle of formula 4 and in combination with formula 5. However, the longer (the larger the length) the abnormal path, the greater its abnormal score. Therefore, in order to eliminate the interference of the path length on the abnormal score calculated by the present application, so as to improve the accuracy of the abnormal score calculated for the abnormal path, the present application can obtain a normalization coefficient α in advance based on the sampling method, wherein α>1, and the normalization coefficient α can be used to eliminate the negative impact of the path length on the abnormal score of the path calculated by the present application.
[0275] In a possible implementation, the following formula 6 may be used as a possible implementation of calculating the normal score of the path P.
[0276] The parameters in Formula 6 are similar to those in Formula 4 and are not described in detail here. The difference is that a normalization coefficient α is added, where the normalization coefficient α is a pre-calculated constant.
[0277] Continuing with FIG3a as an example, when calculating the normal score RS(1) of the path 1 shown in FIG3a, according to Formula 6, RS(1)=M e1 *α. Among them, the normal score RS(1) can be used as the label of node B.
[0278] Then, upon receiving event 2, RS(P)=RS(1)*M can be calculated. e2 *α, where path P in Figure 3a is path 1 + path 2.
[0279] Similarly, for subsequent events received, the normal score RS of the corresponding path is calculated according to the principle of Formula 6, and then the above Formula 5 is used to obtain the abnormal score of the corresponding path.
[0280] In other words, each time the streaming processing platform 100 of the present application calculates the normal score of the most abnormal path with the target node as the end point for the target node in a single event detected in real time, it can multiply it by a normalization coefficient α according to Formula 6 to obtain the normal score of the most abnormal path, and then use Formula 5 to obtain the abnormal score of the most abnormal path.
[0281] Example 5
[0282] The following describes the process of detecting abnormal paths using the label propagation method by the streaming processing platform 100 of the present application.
[0283] FIG4 a is a schematic diagram illustrating an exemplary detection process.
[0284] Before introducing the process, the following technical terms are explained. The source node and target node of a single event in an embodiment of the present application can have their own labels and abnormal path information related to the label, and the abnormal path information can be cached in the memory. Among them, the label of a node can indicate the abnormal score or normal score of the path, and the path is the most abnormal path (i.e., the path with the highest abnormal score) for the target node to be the node. Optionally, the label of the target node can be cached in the memory by means of a label storage table, without specific limitation.
[0285] Specifically, in the embodiment of the present application, the label of the source node of a single event can be named as the first label, the anomaly score indicated by the first label can be named as the first anomaly score, and the path related to the first label can be named as the first anomaly path; and the label of the target node of the single event can be named as the second label, the anomaly score indicated by the second label can be named as the second anomaly score, and the path related to the second label can be named as the second anomaly path; and the label of the source node and the frequency M of the occurrence of the single event can be used to generate the first anomaly score. ei The calculated anomaly score of the path is named the third anomaly score, and the path with the third anomaly score is named the third path. The third path is the path obtained by connecting the first anomaly path associated with the first label of the source node of the single event and the path of the single event.
[0286] As shown in Figure 4a, the process may include the following steps:
[0287] S201, determine the event e in the real-time collected log data i .
[0288] Taking FIG. 1 as an example, terminal 1 to terminal n can report the real-time collected log data to the streaming processing platform 100 .
[0289] The abnormal path mining module 202 (or other modules not shown in the stream processing platform 100, etc., are not limited thereto) can determine the event e in the real-time collected log data. i .
[0290] In a possible implementation, the abnormal path mining module 202 can determine the event e according to the processing flow and principle of S101 to S106 shown in Figure 2 of Example 2 for the received real-time log data. i , where the event e i This is a generalized single event. The specific implementation refers to the relevant introduction of Example 2 and will not be repeated here.
[0291] In this way, the abnormal path mining module 202 can detect each single event in the real-time event stream according to the time sequence of the events in the log data, and generate a corresponding event for each detected single event. i , the following process is performed to detect abnormal paths (abnormal paths for short). Optionally, the abnormal paths can be output and reconstruction of the abnormal paths can be achieved to facilitate operation and maintenance personnel to detect problems.
[0292] S203, based on event e i The first label of the source node src determines the event e i The second label of the destination node dst and the second abnormal path information related to the second label.
[0293] Among them, by propagating the label of the source node src, the label of the source node src can be used to obtain the label of the target node dst, thereby determining the abnormal path related to the label of the target node dst.
[0294] Now, combined with the current events i The implementation process of S203 is described in combination with Example 5.1 and Example 5.2, based on whether the source node src and the destination node dst have labels and the label sizes.
[0295] Example 5.1
[0296] FIG4 b is a schematic diagram showing an exemplary process of initializing a label and a path for a target node dst.
[0297] In the process of FIG4 b , the label obtained by initializing the target node is the second label in S203 , and the path obtained by initializing the target node is the second abnormal path information in S203 .
[0298] As shown in FIG4b , the process performed by the streaming processing platform 100 may include the following steps:
[0299] S1001, determine event e i The target node does not have a second label.
[0300] For example, the tag storage table in the memory can be queried. If the tag information of the target node is not found, it means that the target node does not have a tag, which further indicates that the current event e i If the path (for example, the path from node A to node B) is the first currently detected path that can reach node B, then the node B needs to be assigned an initial label.
[0301] It should be understood that the source node and target node of an event initially do not have their own labels, or the node is initially assigned a label, but the initial value of the label is empty. Only when a path that can reach the node is detected will the node be assigned a label indicating the historical status information of the corresponding path (such as the degree of abnormality) or its label will be updated.
[0302] S1002: Determine a third anomaly score of a third path based on the first label of the source node src.
[0303] In S1002 , the third anomaly score is determined optionally in combination with the normal behavior model.
[0304] In other embodiments, the occurrence frequency M of the current event can also be calculated in real time. ei , or other information that can indicate the normality of the event.
[0305] S1004: Initialize the second label of the target node dst to information indicating the third anomaly score.
[0306] S1005 : Initialize the second abnormal path information related to the second label of the target node dst as the information of the third path.
[0307] Optionally, S1003, in event e i When the source node does not have the first label, the event e is determined to be i If it is an abnormal event, S1004 to S1005 will be executed.
[0308] For example, the third path is the single event e i The first abnormal path related to the first label of the source node is related to the single event e i Path, the path after connection.
[0309] Event e i Take event 1 from node A to node B as an example, for example, process A opens file B. Process A is the source node and file B is the target node.
[0310] 1, first, the abnormal path mining module 202 can query the event frequency storage module 102 for the event e i Frequency of occurrence M ei For example, event e i The frequency of occurrence is M e1 .
[0311] Event-based i Whether the source node has a label, the implementation of S1002 can be divided into two cases.
[0312] Please refer to Figure 3b(1), Case 1: Node A does not have a label, and node B does not have a label either.
[0313] Then, when calculating the anomaly score of path AB (an example of the third path), M e1 According to the above formula 4 (here length = 1), the normal score of path AB is obtained as RS = M e1 .
[0314] The present application also presets a label threshold (an example of the first threshold mentioned above). In this example, the label threshold is the threshold of the normal score indicated by the label, such as 0.35, or other data.
[0315] Among them, the size of the label threshold can affect the number of labels initialized in the process of real-time anomaly detection by the method of the present application. When the label threshold is the threshold of the indicated normal score, the larger the label threshold, the more labels are initialized, and the computational and memory overhead is also greater. However, setting the label threshold too low may result in missed detection of abnormal events. Therefore, the present application reasonably sets it to a value around 0.35, such as a value within the range of 0.2 to 0.45. The specific value is not limited to this range and can be flexibly set according to the detection accuracy requirements of abnormal behavior and the requirements for system computation and memory overhead.
[0316] In other embodiments, the label threshold may also be a threshold of the abnormal score indicated by the label, and then the calculated AS is compared with the label threshold. When AS is greater than or equal to the label threshold, it indicates that the event e i is an abnormal event. On the contrary, event e i It is a normal event.
[0317] After getting the normal score RS=M of path AB e1 Afterwards, due to the incident i The source node has no label information, so the RS calculated for it needs to be compared with the label threshold (for example, 0.35). When the RS is greater than or equal to 0.35, it means that the event e i If the RS is less than 0.35, it means that the event e i If it is an abnormal event, then the event needs to be i Here, node B is given a label and the associated abnormal path.
[0318] As shown in Figure 3b(2), a tag can be assigned to node B, where Tag_B is M e1In addition, the most abnormal path information related to the label of node B can be initialized as: event 1 recorded in chronological order from earliest to latest. The path corresponding to the most abnormal path information is: A->B, where information indicating the relationship between the two nodes in the path may exist.
[0319] In this embodiment, the node label is taken as the normal score of the corresponding path. In other embodiments, the label may also be information indicating the normal score or information indicating the abnormal score of the path, which is not limited here.
[0320] Please refer to Figure 3c(1), Case 2: Node A has a label, but node B does not have a label.
[0321] As shown in FIG3c(1), the method of the present application has previously detected event 3 consisting of path MA and has initialized the tag of node A of the event, for example, Tag_A is M e3 , where M e3 The frequency of event 3 stored in the normal behavior model.
[0322] Therefore, the currently detected event e i The source node, that is, node A has a label, but node B does not have a label.
[0323] Then the RS of the path MAB (an example of the third path) from path MA to path AB can be calculated according to the above formula 4, which is the label of the source node (node A) and the occurrence frequency M of event e1. e1 The product of RS = M e3 *M e1 , since the source node, that is, node A, has a label, the label of node B can be directly initialized as Tag_B as shown in Figure 3c(2) is M e3 *M 31 . In addition, the second abnormal path information related to the label of node B can be initialized as: third path information, for example: event 3 and event 1 recorded in order from early to late. The third path information records an event list, and the events in the event list can be sorted in order from early to late in order to facilitate reconstruction of the most abnormal path. The path corresponding to the most abnormal path information recorded here is: M->A->B, where there may be information representing the relationship between the two nodes in the path.
[0324] In conjunction with the modules in FIG. 1 , the abnormal path mining module 202 can extract the current event using reported log data and read the frequency of occurrence of the event from the event frequency storage module 102. Based on the label of the source node of the event, the normal score of the third path is calculated, and the abnormal score of the third path is obtained using this normal score. The node data cache module 203 can then store the abnormal score of the third path in memory as the label of the target node of the event, and store the path information of the third path in memory as the path information associated with the label of the target node. In this way, only the partial path in the streaming provenance graph with the abnormality is constructed in memory. This partial path is the path with the abnormality.
[0325] In the above embodiments, the abnormal path mining module 202 searches the normal behavior model for the current event e i The frequency of occurrence of the event e is used as an example to illustrate that i An event that belongs to the normal behavior model, also known as a known event.
[0326] However, in a possible implementation, there is no information about the event e in the normal behavior model in the event frequency storage module 102. i When the frequency of occurrence is i It is an unknown event. This application can set the frequency of the unknown event to a frequency threshold (UNSEEN_EVENT_SCORE), such as 0.1. Of course, this application does not limit this and can be configured according to the situation so that it can be combined with the M of the known event. e Values are separated.
[0327] Among them, M is the number of events that have not appeared in the baseline database (which stores the normal behavior model). e Value (UNSEEN_EVENT_SCORE): This value cannot be too small to avoid insufficient information collected by the baseline database. At the same time, it needs to be distinguishable from the events in the baseline database (also known as known events). In actual operation, the M of the unseen event can be ei Set the value to 0.1, and then use the M ei Directly use the above formula 4 or formula 6 to calculate the normal score RS of the event.
[0328] In a possible implementation, in order to distinguish the frequency of occurrence of events in the baseline database (also known as known events) from the frequency of occurrence of the unknown event, the present application may perform a multi-step process on each known event e in the baseline database. i The frequency of occurrence M ei , processed according to formula 7, using the processed M ei′ is used as M in formula 4 or formula 6 to calculate the normal score RS of the path ei .
[0329] M ei ′=M ei *0.7+0.3, formula 7;
[0330] In this way, the occurrence frequency of the known event can be made into the frequency M after processing according to formula 7. ei ′, the value is in the range of (0.3-1], and has certain differences from unknown events.
[0331] In a possible implementation, there is a normal behavior model in the event frequency storage module 102 about the event e i When the frequency of occurrence is i is a known event, then the event can be read from the normal behavior model (such as the baseline database) i The frequency of occurrence M ei , optionally, the frequency M ei Calculate according to the above formula 7 to get M ei ' as the event e i The frequency of occurrence M ei , used to calculate the event e according to Formula 4 or Formula 6 i The normal score RS of the path AB.
[0332] Example 5.2
[0333] FIG. 4c exemplarily shows a method for i Schematic diagram of the process of determining the latest label and the latest most abnormal path of the target node dst based on the current label of the target node dst.
[0334] In the process of FIG4c , the label of the target node after being updated is the second label in S203 of FIG4a , and the second abnormal path information of the target node after being updated is the second abnormal path information related to the second label in S203 .
[0335] The implementation or principle of most of the contents of this Example 5.2 and the above Example 5.1 are the same, and the main difference is: Difference 1, the target node of the event detected in Example 5.2 currently has a label, so it is necessary to determine whether it is necessary to use the third anomaly score obtained by this calculation to update the label of the target node instead of initializing the label. Specifically: Since the event detected in Example 5.2 currently has a label, then according to the method of label propagation of the source node of this application, for example, according to the above formula 4 and formula 5, or, formula 6 and formula 5 to calculate the third anomaly score of the third path, if the third anomaly score is greater than the second anomaly score indicated by the current label of the target node, it is necessary to update the label of the target node and update the most abnormal path information related to the label, rather than directly initializing the calculated anomaly score and abnormal path information of the third path as the label and abnormal path of the target node in Example 5.1. Thus, with the help of a streaming origin graph, a local origin graph is constructed based on the most abnormal path found step by step, reducing memory usage and the computational overhead of the origin graph.
[0336] As shown in FIG4c , the process performed by the streaming processing platform 100 may include the following steps:
[0337] S2001, determine event e i The target node has a second label.
[0338] For example, the abnormal path mining module 202 can query the label storage table in the memory. When the label information of the target node is found, it means that the target node has a label, which further indicates that the application has detected a path that can reach the current event e. i The abnormal path of the target node, and the event e i The label (second label) of the target node stores the normal score (or abnormal score) of the most abnormal path whose end point is the target node. In addition, the memory also stores the path information (second path information) of the most abnormal path.
[0339] S2002 : Determine a third anomaly score of a third path based on the first label of the source node src.
[0340] In S2002 , the third anomaly score is determined optionally in combination with a normal behavior model.
[0341] S2004: When the third anomaly score is greater than the second anomaly score indicated by the second label of the target node dst, update the second label of the target node dst based on the third anomaly score.
[0342] S2005 : When the third anomaly score is greater than the second anomaly score, the second anomaly path information associated with the second label of the target node dst before the update is updated to the third path information.
[0343] Optionally, S2003, in event e i When the source node does not have the first label, the event e is determined to be i S2004 to S2005 will only be executed if it is an abnormal event.
[0344] For example, the third path is the single event e i The first abnormal path related to the first label of the source node is related to the single event e i Path, the path after connection.
[0345] Event e i Take event 1 from node A to node B as an example, for example, process A opens file B. Process A is the source node and file B is the target node.
[0346] 1, first, the abnormal path mining module 202 can query the event frequency storage module 102 for the event e i Frequency of occurrence M ei For example, event e i The frequency of occurrence is M e1 .
[0347] Event-based i The implementation of S2002 can be divided into two cases depending on whether the source node has a label.
[0348] Please refer to Figure 3d(1), Case 1: Node A does not have a label, but node B does have a label.
[0349] As shown in FIG3d(1), the present application has previously detected event 4 where the path is path CB, causing node B to have a tag Tag_B of M e4 .
[0350] Then, the present application detects event 1 with path AB, and the occurrence frequency of event 1 can be obtained from the normal behavior model as M. e1 , and node A does not have a label, then according to the above formula 4 (here length = 1), the normal score of path AB can be calculated as RS = M e1 .
[0351] The present application also presets a label threshold (an example of the first threshold mentioned above). In this example, the label threshold is the threshold of the normal score indicated by the label, such as 0.35, or other data.
[0352] Among them, the size of the label threshold can affect the number of labels initialized in the process of real-time anomaly detection by the method of the present application. When the label threshold is the threshold of the indicated normal score, the larger the label threshold, the more labels are initialized, and the computational and memory overhead is also greater. However, setting the label threshold too low may result in missed detection of abnormal events. Therefore, the present application reasonably sets it to a value around 0.35, such as a value within the range of 0.2 to 0.45. The specific value is not limited to this range and can be flexibly set according to the detection accuracy requirements of abnormal behavior and the requirements for system computation and memory overhead.
[0353] After getting the normal score RS=M of path AB e1 Afterwards, due to the incident i The source node has no label information, so the RS calculated for it needs to be compared with the label threshold (for example, 0.35). When the RS is greater than or equal to 0.35, it means that the event e i If the RS is less than 0.35, it means that the event e i If it is an abnormal event, then the event needs to be i The target node in , here is node B updater label and related abnormal path.
[0354] Next, we can compare RS (for M e1 ), and Tag_B (for M e4 ) between the size relationship, such as M e1 <M e4 , then the normal score of path AB is less than the normal score of path CB, which means that the abnormal score of path AB is (1-M e1 ) is greater than the anomaly score of path CB (1-M e4 ), so that the path AB is more abnormal than the path indicated by the label of node B. Therefore, the abnormal path mining module 202 shown in FIG1 can replace RS=M e1 , and the path AB (eg event 1) is notified to the data cache module 203 of the node, and the data cache module 203 of the node can cache the tag Tag_B=M about node B in the memory e4 Refresh to Tag_B=M e1 , and the most abnormal path information related to Tag_B cached in the memory is updated from event 4 to event 1, where event 1 may indicate path AB.
[0355] On the contrary, if M e1 ≥M 34, it means that the anomaly score of path AB shown in FIG3 d is less than or equal to the anomaly score of path CB, so there is no need to update the label of node B and the anomaly path information related to the label.
[0356] In this embodiment, the node label is taken as the normal score of the corresponding path. In other embodiments, the label may also be information indicating the normal score, or information indicating the abnormal score of the path, or other historical status information, which is not limited here.
[0357] Please refer to Figure 3e(1), Case 2: Node A has a label, and node B also has a label.
[0358] As shown in FIG3e(1), the method of the present application has previously detected event 3 consisting of path MA and has initialized the tag of node A of the event, for example, Tag_A is M 33 , for example, M e3 The frequency of event 3 stored in the normal behavior model.
[0359] Therefore, the currently detected event e i The source node, node A, has a label.
[0360] In addition, as shown in FIG3e(1), the method of the present application has previously detected event 4 consisting of path CB, and has initialized the tag of node B of event 4, for example, Tag_B is M e4 , for example, M e4 The frequency of event 4 stored in the normal behavior model.
[0361] Therefore, the currently detected event e i The target node, node B, also has a label.
[0362] Then the abnormal path mining module 202 can calculate the path MAB (an example of the third path) from path MA to path AB according to the above formula 4. RS is the label of the source node (node A) and the occurrence frequency M of event 1. e1 The product of the normal score RS of the path MAB can be obtained as RS = M e3 *M e1 .
[0363] Next, we can compare RS=M e3 *M e1 , and Tag_B=M e4 The size relationship between them, such as M e3 *M e1 <M e4, then the normal score of path MAB is less than the normal score of path CB, which means that the abnormal score of path MAB is (1-M e3 *M e1 ) is greater than the anomaly score of path CB (1-M e4 ), so the path MAB is more abnormal than the path indicated by the label of node B. Therefore, the abnormal path mining module 202 shown in FIG3e(2) and FIG1 can convert RS=M e3 *M e1 The data cache module 203 of the node is notified, and the data cache module 203 of the node can cache the tag Tag_B=M about node B in the memory e4 Refresh to Tag_B=M e3 *M e1 .
[0364] Furthermore, as described above, the memory also stores the most anomalous path information associated with node A's tag, Tag_A. This most anomalous path is path MA, as shown in Figure 3e (the recording method may be to record information about event 3). The recorded information about event 3 may include, but is not limited to: the source node is node M, the target node is node A, and the relationship between nodes M and A. The anomalous path mining module 202, as shown in Figure 1, can then instruct the node data cache module 203 to concatenate the most anomalous path information associated with event 1 and node A's tag to serve as the most anomalous path associated with node B's tag. This updates node B's most anomalous path information associated with the tag from path CB to path MAB.
[0365] In short, the updated label-related path information at node B is the result of connecting the label-related path information of the source node of event 1 with the path of event 1. From the perspective of events, this path information is recorded in chronological order: event 3 with path MA and event 1 with path AB.
[0366] Returning to FIG. 4 a , optionally, after S201 and before S203 , the process may further include S202 to reduce the label of the source node of the currently detected event.
[0367] S202, in event e i When the first label of the source node src or the first abnormal path information related to the first label meets a first preset condition, the first label of the source node src and the first abnormal path information related to the first label are removed.
[0368] Among them, event e i The length of the most abnormal path from the source node src to the source node can indicate the number of label propagation rounds;
[0369] If the number of label propagation rounds exceeds a preset threshold, it indicates that the first abnormal path information meets the first preset condition. The source node label and the most abnormal path information indicated by the RS in the label can then be deleted. This can reduce memory usage and prevent excessive false alarms caused by excessive label propagation.
[0370] In addition, in the event i If the RS (normality score) stored in the first label of the source node is greater than a preset normality score threshold, the path associated with the first label is not abnormal enough; or if the AS indicated by the first label is less than a preset abnormality score threshold, the path associated with the first label is not abnormal enough. In both cases, the first label of the source node meets the preset conditions. Then, the label of the source node and the most abnormal path information indicated by the RS in the label can be deleted. This can reduce memory usage and prevent the spread of labels with normality scores RS greater than the preset normality score threshold.
[0371] In some embodiments, the present application may also determine the value of α and M ei , we can infer the average propagation distance of the label.
[0372] In addition, for the labels of each node cached in the memory, if the initialization time of the label of the node (i.e. the time when the label of the node is initialized) exceeds the preset time threshold from the current time, it means that the label of the node has been created too long, which also means that the label of the node meets the third preset condition. Then the content of the label of the node that meets the third preset condition cached in the memory can be cleared, or the label can be deleted. In addition, the abnormal path information related to the label can also be cleared from the memory. Optionally, the initialization time (or creation time) of the label of each node in the cache can be regularly checked to see if the difference from the current time is too large, so as to clean up the labels that have been created too long and clear the abnormal path information related to the label.
[0373] This preset time threshold is the Decay_Time_Threshold. To cache a large number of tags in the tag database and prevent outdated tags from interfering with detection, old tags can be removed. For example, the preset time threshold can be set to 30*60*1000000ms.
[0374] Generally speaking, attack behaviors are not mixed with many normal behaviors. When the path is very long and has not yet reached the alarm condition, clearing the label can prevent the score of the normal process from being contaminated.
[0375] Returning to FIG. 4 a , optionally, after S203 , the process may further include S204 , for reducing the label of the target node of the currently detected event.
[0376] It should be understood that the above-mentioned step S202 of reducing the label of the source node and the step S204 of reducing the label of the target node of the event can be performed selectively as needed, and do not need to be performed both.
[0377] S204, in event e i When the second label of the target node dst or the second abnormal path information related to the second label meets the second preset condition, the second label of the target node dst and the second abnormal path information related to the second label are removed.
[0378] The implementation principle of S204 is the same as that of the above-mentioned S202, except that the second preset condition in S204 may be the same as or different from the first preset condition in S202, and can be flexibly configured according to needs and scenarios.
[0379] Returning to FIG. 4 a , after S203 , the method may further include S205 .
[0380] S205 : When the second label of the target node dst meets the alarm condition, output the second abnormal path in the streaming provenance graph.
[0381] As described above, in the above-mentioned embodiments of Figures 4b and 4c, the second abnormal path indicated by the second label of the target node dst can be initialized or updated to the third path calculated for the current event. Therefore, the second abnormal path here can be the third path. The third path is the path formed by the first abnormal path information associated with the first label of the source node of the event and the path of the event. The second anomaly score indicated by the second label of the target node dst is also updated to the third anomaly score of the third path.
[0382] For example, the label of the target node dst stores the normal score RS of the corresponding third path. When determining whether an alarm is needed for the third path, it is possible to determine whether the normal score RS is less than the preset normal score alarm threshold. When the normal score RS stored in the label is less than the preset normal score alarm threshold, it means that the label of the target node meets the alarm condition, and the second abnormal path in the streaming origin graph can be output.
[0383] For another example, the preset alarm threshold may also be a preset anomaly score alarm threshold. Then, when determining whether an alarm is required for the third path, the normal score RS indicated by the second label of the target node may be calculated as AS according to the above formula 5; then, it may be determined whether the AS is greater than the preset anomaly score alarm threshold. When the AS is greater than the preset anomaly score alarm threshold, it indicates that the label of the target node meets the alarm condition, and the second anomaly path in the streaming origin graph may be output. The larger the preset anomaly score alarm threshold, the higher the difficulty in triggering the alarm. The threshold in the alarm condition may be flexibly set according to the need to trigger the alarm, such as a preset normal score alarm threshold, or a preset anomaly score alarm threshold or other threshold indicating the alarm condition.
[0384] As shown in FIG1 , the abnormal path mining module 202 may notify the abnormal path alarm module 204 of the most abnormal path (ie, second abnormal path information) related to the second label of the target node whose second label meets the alarm condition.
[0385] The abnormal path alarm module 204 can obtain the abnormal path with an alarm from the data cache module 203 of the node.
[0386] The present application caches the most abnormal path information related to the tag (for example, the second abnormal path related to the second tag) in the memory, so that the transmission history information of the tag can be cached in the memory. Then, when an alarm occurs, it means that abnormal behavior has been discovered during the real-time detection process. The most abnormal path information related to the tag with the alarm cached in the memory can be used to quickly restore the most abnormal path, thereby realizing the restoration of the attack path without the need to re-query the massive logs.
[0387] When outputting the second abnormal path, the present application may output all nodes in the second abnormal path and the relationships between the nodes.
[0388] The second abnormal path may be output in the form of text or a stream-based provenance graph, which is not limited here.
[0389] For ease of explanation, the present application may output the output second abnormal path with an alarm in the form of an alarm graph (eg, a graph format of a streaming provenance graph).
[0390] In one possible implementation, the abnormal path alarm module 204 can also notify the abnormal path aggregation module 205 of the abnormal path for which an alarm exists. The abnormal path aggregation module 205 can aggregate the second abnormal path of the current alarm and the local streaming origin graph of the candidate nodes in the two abnormal paths to obtain an aggregated local streaming origin graph; wherein the candidate node is a node whose label meets the alarm condition; and the aggregated local streaming origin graph is output as a local streaming origin graph related to the end point of the target event.
[0391] For example, when an alarm is received, the most abnormal path related to the second tag of the alarm (i.e., the second abnormal path mentioned above) can be obtained. The second abnormal path is the currently detected path with the highest abnormality score AS ending at the target node. The second abnormal path is represented by singleAlertPath.
[0392] Then, the abnormal path aggregation module 205 can traverse the nodes in the traversal path singleAlertPath, for example, traverse to node 1, add the path singleAlertPath to the alert graph 1 (an example of a local streaming provenance graph of a candidate node) where node 1 is located, and obtain the aggregated alert graph 1';
[0393] Then, the abnormal path aggregation module 205 traverses to the node 2 in the path singleAlertPath, and can aggregate the alert graph 2 (an example of a local streaming provenance graph of a candidate node) where the node 2 is located with the above-mentioned alert graph 1' to obtain the alert graph 1".
[0394] Next, the abnormal path aggregation module 205 traverses to node 3 in the path singleAlertPath, and then aggregates the alert graph 3 (an example of a local streaming provenance graph of a candidate node) where node 3 is located with the above-mentioned Figure 1″ to obtain Figure 1″′:
[0395] And so on, until all nodes in the path singleAlertPath are traversed and the corresponding aggregation operations are completed, thereby obtaining the final target aggregation graph M (an example of the local streaming provenance graph after aggregation).
[0396] Optionally, each node in the target aggregate graph M is traversed, and the alarm graph where each node is located is updated to the target aggregate graph M.
[0397] Finally, the target aggregation graph M can be output.
[0398] The alarm graph of each node in the above process represents: an abnormal path related to the label of a target node with an alarm, and a local stream provenance graph composed of paths aggregated with other abnormal paths with alarms.
[0399] The label propagation path detection method of the present application may include a storage and four logics, wherein one storage is a structure of label cache data, and the four logics are: label initialization logic, label propagation logic, label reduction logic and alarm triggering logic. This process will use the association analysis characteristics of the graph, but it does not need to completely construct the origin graph in the memory, thus saving memory overhead. At the same time, the processing results of each event will be cached to avoid repeated calculations, thus saving computing overhead. This type of algorithm can be modified to adapt to different computing scenarios. For the specific problem of abnormal path discovery, the information stored in the label can be the abnormal score of the path or the normal score of the path, and the most abnormal path with the abnormal score or normal score is cached for the target node.
[0400] Using streaming label propagation, local anomaly scores are propagated and aggregated along the path. This allows for fast, lightweight, and real-time anomaly detection. By caching the most anomaly-severe paths in the labels, when an anomaly is detected in real time, the cache allows for fast, lightweight, and efficient attack path restoration.
[0401] FIG1 is a schematic diagram illustrating the structure of a streaming processing platform 100. It should be understood that the streaming processing platform 100 shown in FIG1 is merely an example, and the system of the present application may have more or fewer modules than shown, may combine two or more modules, or may have a different module configuration. The various modules shown in FIG1 may be implemented in hardware, including one or more signal processing and / or application-specific integrated circuits, software, or a combination of hardware and software.
[0402] Example 6
[0403] The implementation process of the method of the present application is described in conjunction with specific events detected in real-time log data. Taking the normalization coefficient α as 2 as an example, other embodiments can flexibly set its value according to the scenario.
[0404] a) Analyze the event redis->FILE_WRITE-> / var / spool / cron / root [A process wrote a file]. Neither the source nor the target node in this event has a label. The offline normal behavior model does not show the frequency of this event, so it is an unknown event. Therefore, the normality score RS for this unknown event is assigned an UNSEEN_EVENT_SCORE = 0.1. The normality score RS for this event is 0.1 < 0.35, where 0.35 is the label threshold for a single event. Therefore, this event is an anomaly. The label and path need to be written to the target node / var / spool / cron / root.
[0405] The source node of the event is the Redis process, which has no label and an anomaly score of 0. The most abnormal path stored in the / var / spool / cron / root file label is redis->FILE_WRITE-> / var / spool / cron / root. The normal score RS of anomalyPath is 0.1, and the anomaly score AS = 1-0.1 = 0.9. This does not exceed the anomaly score alarm threshold ALERT_THRESHOLD (here, 0.95), so no alarm is generated.
[0406] b) Analyze the event / var / spool / cron / root->FILE_READ->cron. Query the offline normal behavior model to find the occurrence frequency of this event M( / var / spool / cron / root,FILE_READ,cron) = 0.6. The normal score of this event RS = 0.6*0.7+0.3 = 0.72 (calculated according to the above formula 7). The label of the source node / var / spool / cron / root (RS = 0.1 in the above step a)) is propagated to the cron process of the target node. The normal score RS of the anomaly path (anomalyPath) is calculated as 0.1*0.72*α = 0.144, and the anomaly score AS = 1-0.144 = 0.856, which does not exceed the ALERT_THRESHOLD (for example, 0.95). No alarm is issued, and the anomalyPath is cached in the label of the target node cron, which is 0.144 here: and the most abnormal path is cached: redis->FILE_WRITE-> / var / spool / cron / root->FILE_READ->cron.
[0407] c) Analyze the event cron->PROCESS_FORK->bash. In the offline normal behavior model, find M(cron,PROCESS_FORK,bash) = 0.2. The normal score for this event, RS, is 0.2 * 0.7 + 0.3 = 0.44. The cron tag is propagated to the bash process. The normal score, RS, of the anomalyPath is calculated to be 0.144 * 0.44 * α = 0.127, and the anomaly score, AS, is 1 - 0.127 = 0.873. This does not exceed the ALERT_THRESHOLD (0.95). No alert is generated. The anomalyPath is cached in the bash tag: redis->FILE_WRITE-> / var / spool / cron / root->FILE_READ->cron->PROCESS_FORK->bash.
[0408] d) Analyze the event bash->PROCESS_EXEC->curl. The offline normal behavior model does not find M(bash,PROCESS_EXEC,curl). The normal score for this event is RS = UNSEEN_EVENT_SCORE = 0.1. The bash tag is propagated to the curl process, and the normal score RS for the anomalyPath is calculated to be 0.127*0.1* α (2) = 0.0254, anomaly score AS = 1-0.0254 = 0.9746, exceeding ALERT_THRESHOLD (0.95), an alarm is issued, and the anomalyPath is cached in the curl process label: redis->FILE_WRITE-> / var / spool / cron / root->FILE_READ->cron->PROCESS_FORK->bash->PROCESS_EXEC->curl.
[0409] Each event is analyzed and processed in this way, and an abnormal alarm is issued in real time for processes whose anomaly score of the anomalyPath in the label exceeds ALERT_THRESHOLD.
[0410] Analyze each event one by one to propagate labels and cache the anomalyPath, detect in real time any processes whose anomaly scores exceed the ALERT_THRESHOLD, and issue an alarm.
[0411] When the anomaly score of the curl process is detected to exceed the alarm threshold, the attack path is restored based on the anomalyPath cached in the curl tag: redis->FILE_WRITE-> / var / spool / cron / root->FILE_READ->cron->PROCESS_FORK->bash->PROCESS_EXEC->curl, and a visual alarm diagram (a local churn origin diagram) is generated as shown in Figure 5.
[0412] In each of the above embodiments, the list of events detected in real time is configurable, for example, at least one of process, network, and file can be used. As an anomaly detection solution, the solution can selectively analyze anomalies in part of the event chain. For example, if only process-related events are considered, the method and system of the present application is an anomaly detection system for the process tree. Generally speaking, the more event types are considered, the stronger the detection capability, but the greater the overhead. In actual operation, it can be recommended to analyze events such as process establishment, file creation and reading and writing, network connection and data sending / receiving.
[0413] This application designs a real-time anomaly path discovery system based on the provenance graph and label transfer technology. It uses the correlation analysis capability of the provenance graph to correlate scattered abnormal events and construct a path of abnormal behavior. The construction of this path has two main values for detection: 1) Real-time attack detection: Correlation analysis can aggregate small local anomalies into large anomalies, while discovering hidden suspicious behavior and avoiding a large number of false positives; 2) Real-time reconstruction of attack paths: The abnormal path itself reflects the path of the attack, which can help analysts understand the attack behavior and make corresponding measures.
[0414] In addition, the present application configures a label propagation framework based on the origin graph, and based on security experience, designs a specific implementation algorithm for abnormal label propagation, propagates and aggregates the local abnormal score as the label of the node along the path of the origin graph, and performs anomaly detection in real time based on the abnormal score. When the abnormal score reaches the set threshold, an alarm is generated in real time. In the process of abnormal label propagation, the abnormal path with the highest abnormal score to each node in the cache is updated in real time. Among them, each node has many paths to reach, one branch is transmitted to a certain node, and another branch is transmitted to the node. Each node has a cache information, which stores the path corresponding to the abnormal score in the label. If another path to the node makes the abnormal score of the node exceed the current score, the cache information of the node is updated to the path with the highest abnormal score. When an anomaly is detected, the attack path (the path with the highest abnormal score) can be restored in real time based on the cache data of the node, and multiple abnormal paths can be fused to generate a complete attack graph.
[0415] An embodiment of the present application provides a data processing device 800. Optionally, the device 800 can be deployed on a cloud management platform, which is used to manage the infrastructure for providing cloud services. The infrastructure includes multiple cloud data centers located in different regions, and each region has at least one cloud data center.
[0416] FIG6 is a schematic diagram showing the structure of an exemplary data processing device 800. Referring to FIG6 , the data processing device 800 includes:
[0417] In one possible implementation, an embodiment of the present application provides a data processing device. The device may include: a first determination module 801, configured to determine a target event to be detected based on real-time collected log data, wherein the target event includes a start point, an end point, and a relationship between the start point and the end point; a second determination module 802, configured to determine a second label of the end point based on the first label of the start point; wherein the label of each node in the target event is used to indicate historical state information of a target path with the node as the end point; wherein the target path is a path in at least one path with the node as the end point, wherein the at least one path is a path formed by at least one event to be detected determined based on the log data; wherein the node in the target event includes a start point and an end point in the target event; a third determination module 803, configured to determine second path information related to the second label based on first path information related to the first label; and an alarm module 804, configured to output a local streaming provenance graph related to the end point of the target event based on the second path information when the second label meets an alarm condition.
[0418] In one possible embodiment, the device also includes: a fourth determination module, used to determine the normality of the target event; the second determination module 802 is specifically used to: determine the historical status information of the third path based on the first label of the starting point and the normality of the target event; wherein the third path is a path composed of the first path information and the target event; based on the historical status information of the third path, determine the second label of the end point.
[0419] In a possible implementation, the fourth determination module is specifically configured to determine the normality of the target event based on a pre-constructed information table; wherein the information table includes first information indicating the normality of a preset event.
[0420] In a possible implementation, the fourth determination module is specifically configured to determine the normality of the target event based on first information in the information table indicating the normality of the target event when the target event is the preset event in the information table.
[0421] In a possible implementation manner, the fourth determining module is specifically configured to configure a preset normality threshold as the normality of the target event when the target event does not belong to the preset events in the information table.
[0422] In a possible implementation, the second determination module 802 is specifically configured to determine the historical status information of the third path based on the normality of the target event when the first tag at the starting point is empty data and if the normality of the target event meets a first preset condition.
[0423] In one possible implementation, the historical status information includes abnormality degree information, and the second determination module 802 is specifically used to: determine the normality of the target path indicated by the first label based on the first label of the starting point; and determine the abnormality degree information of the third path based on the normality of the target path and the normality of the target event.
[0424] In a possible implementation manner, the second determining module 802 is specifically configured to initialize the second label of the endpoint to indicate historical status information of the third path when the second label of the endpoint is empty data.
[0425] In a possible implementation, the third determining module 803 is specifically configured to initialize the second path information associated with the second tag of the destination as the third path information when determining that the first path information is empty data.
[0426] In one possible implementation, the second determination module 802 is specifically configured to update the second tag of the endpoint to information indicating the abnormality level of the third path when the second tag of the endpoint is not empty data and the abnormality level of the third path is greater than the abnormality level of the target path indicated by the second tag.
[0427] In one possible implementation, the third determination module 803 is specifically configured to update the second label of the endpoint to indicate the historical status information of the third path when determining that the second label needs to be refreshed based on the historical status information of the third path and the historical status information of the target path indicated by the second label.
[0428] In a possible implementation manner, the first label and the first path information, and the second label and the second path information are all cached in a memory.
[0429] In a possible implementation, the device further includes: a deletion module configured to delete the information of the first label and the first path information cached in the memory when at least one of the first label at the starting point and the first path information related to the first label meets a second preset condition.
[0430] In a possible implementation, the deleting module is specifically configured to delete the first label information and the first path information cached in the memory when the historical status information of the target path indicated by the first label at the starting point meets a label reduction condition.
[0431] In a possible implementation, the deleting module is specifically configured to delete the first tag information and the first path information cached in the memory when a path length of the first path information associated with the first tag is greater than a preset length threshold.
[0432] In one possible embodiment, the historical status information includes abnormality degree information; the alarm module 804 is specifically used to output a local streaming origin graph related to the end point of the target event based on the second path information when the abnormality degree information of the target path indicated by the second tag is greater than a third preset abnormality degree threshold.
[0433] In a possible embodiment, the device also includes: an aggregation module, used to aggregate the second path information and the local streaming provenance graph of the target node in the second path information to obtain an aggregated local streaming provenance graph; wherein the target node is a node whose label meets the alarm condition; and an output module, used to output the aggregated local streaming provenance graph as a local streaming provenance graph related to the end point of the target event.
[0434] In a possible implementation, the apparatus further includes: an updating module configured to update the local streaming provenance graph of each node in the aggregated local streaming provenance graph to the aggregated local streaming provenance graph.
[0435] The effects of the data processing device 800 in the above-mentioned embodiments are similar to the effects of the data processing methods in the above-mentioned embodiments, and are not described in detail here.
[0436] Among them, the above modules can all be implemented by software, or by hardware. Among them, the module is an example of a software functional unit, and the above module may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the first determination module 801 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Among them, usually a region can include multiple AZs.
[0437] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0438] As an example of a hardware functional unit, a module may include at least one computing device, such as a server. Alternatively, the module may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0439] The multiple computing devices included in the data processing apparatus can be distributed in the same region or in different regions. The multiple computing devices included in the data processing apparatus can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the data processing apparatus can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0440] It should be noted that, in other embodiments, the above modules can be used to execute corresponding steps in the data processing method to realize all functions of the data processing device.
[0441] This application also provides a computing device 900. As shown in Figure 7, computing device 900 includes a bus 902, a processor 904, a memory 906, and a communication interface 909. Processor 904, memory 906, and communication interface 909 communicate with each other via bus 902. Computing device 900 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 900.
[0442] Bus 902 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG7 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 902 may include a path for transmitting information between various components of computing device 900 (e.g., memory 906, processor 904, and communication interface 909).
[0443] The processor 904 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0444] The memory 906 may include volatile memory, such as random access memory (RAM). The processor 904 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0445] The memory 906 stores executable program code, and the processor 904 executes the executable program code to respectively implement the functions of the first determination module, the second determination module, the third determination module, and the alarm module, thereby implementing the data processing method. In other words, the memory 906 stores instructions for executing the data processing method.
[0446] The communication interface 909 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 900 and other devices or a communication network.
[0447] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0448] As shown in Figure 8, the computing device cluster includes at least one computing device 1000. The memory 1006 in one or more computing devices 1000 in the computing device cluster may store the same instructions for executing the data processing method.
[0449] The computing device 1000 includes a bus 1002 , a processor 1004 , a memory 1006 , and a communication interface 1008 . The processor 1004 , the memory 1006 , and the communication interface 1008 communicate with each other via the bus 1002 .
[0450] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also store some instructions for executing the data processing method. In other words, the combination of one or more computing devices 1000 can jointly execute the instructions for executing the data processing method.
[0451] It should be noted that the memory 1006 in different computing devices 1000 in the computing device cluster can store different instructions, each for executing a portion of the functions of the data processing apparatus. In other words, the instructions stored in the memory 1006 in different computing devices 1000 can implement the functions of one or more of the first determination module, the second determination module, the third determination module, and the alarm module.
[0452] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. The network can be a wide area network (WAN), a local area network (LAN), or the like. FIG. 9 illustrates one possible implementation. As shown in FIG. 9 , two computing devices 1100A and 1100B are connected via a network. Specifically, each computing device is connected to the network via a communication interface 1108 within the computing device.
[0453] The computing device 1100A includes a bus 1102 , a processor 1104 , a memory 1106 , and a communication interface 1108 . The processor 1104 , the memory 1106 , and the communication interface 1108 communicate with each other via the bus 1102 .
[0454] The computing device 1100B includes a bus 1102 , a processor 1104 , a memory 1106 , and a communication interface 1108 . The processor 1104 , the memory 1106 , and the communication interface 1108 communicate with each other via the bus 1102 .
[0455] In this possible implementation, the memory 1106 of the computing device 1100A stores instructions for executing the functions of the first determination module, the second determination module, the third determination module, and the alarm module. Simultaneously, the memory 1106 of the computing device 1100B stores instructions for executing the functions of the deletion module, the update module, the aggregation module, the fourth determination module, and the output module.
[0456] It should be understood that the functionality of the computing device 1100A shown in FIG9 may also be implemented by multiple computing devices 1100. Similarly, the functionality of the computing device 1100B may also be implemented by multiple computing devices 1100.
[0457] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection methods of the computing device clusters described in Figures 7 and 9. However, the memory 1106 in one or more computing devices 1100 in this computing device cluster can store the same instructions for executing the data processing method.
[0458] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more computing devices 1100 can jointly execute the instructions for executing the data processing method.
[0459] It should be noted that the memory 1106 in different computing devices 1100 in the computing device cluster may store different instructions for executing partial functions of the data processing apparatus. In other words, the instructions stored in the memory 1106 in different computing devices 1100 may implement the functions of one or more devices in the data processing apparatus.
[0460] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the data processing method described in the above embodiments.
[0461] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the data processing method in the above embodiment.
[0462] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A data processing method, characterized in that, The method includes: Based on the log data collected in real time, determining a target event to be detected, where the target event includes a starting point, an ending point, and the relationship between the starting point and the ending point; Based on the first label of the starting point, determining the second label of the ending point; Wherein, the label of each node in the target event is used to indicate the historical status information of the target path with that node as the ending point; Wherein, the target path is one of at least one path with that node as the ending point, where the at least one path is a path formed by at least one event to be detected determined based on the log data; wherein, the nodes in the target event include the starting point and the ending point in the target event; Based on the first path information related to the first label, determining the second path information related to the second label; When the second label meets the alarm condition, outputting a local streaming origin graph related to the ending point of the target event based on the second path information.
2. The method according to claim 1, wherein The method further includes: Determining the normality level of the target event; Based on the first label of the starting point and the normality level of the target event, determining the historical status information of the third path; Wherein, the third path is the path formed by the first path information and the target event; Based on the historical status information of the third path, determining the second label of the ending point.
3. The method according to claim 2, wherein The method further includes: Based on a pre-constructed information table, determining the normality level of the target event; Wherein, the information table includes first information indicating the normality level of a preset event.
4. The method according to claim 3, wherein The method further includes: When the target event is the preset event in the information table, determining the normality level of the target event based on the first information in the information table indicating the normality level of the target event.
5. The method according to claim 3 or 4, characterized in that, The method further includes: When the target event does not belong to the preset event in the information table, configuring a preset normality level threshold as the normality level of the target event.
6. The method according to any one of claims 2 to 5, characterized in that, The determining the historical status information of the third path based on the first label of the starting point and the normality level of the target event includes: When the first label of the starting point is empty data, if the normality level of the target event meets the first preset condition, determining the historical status information of the third path based on the normality level of the target event.
7. The method according to any one of claims 2 to 5, characterized in that, The historical status information includes abnormality level information, and the determining the historical status information of the third path based on the first label of the starting point and the normality level of the target event includes: Based on the first label of the starting point, determining the normality level of the target path indicated by the first label; Based on the normality level of the target path and the normality level of the target event, determining the abnormality level information of the third path.
8. The method according to any one of claims 2 to 7, characterized in that, The determining the second label of the ending point based on the historical status information of the third path includes: When the second label of the ending point is empty data, initializing and setting the second label of the ending point to indicate the historical status information of the third path.
9. The method according to claim 8, wherein Determining second path information related to the second tag based on the first path information related to the first tag includes: When determining that the second path information is empty data, initializing and setting the second path information related to the second tag at the end point to the information of the third path.
10. The method according to any one of claims 2 to 7, characterized in that Determining the second tag at the end point based on the historical status information of the third path includes: When it is determined that the second tag needs to be refreshed based on the historical status information of the third path and the historical status information of the target path indicated by the second tag, updating the second tag at the end point to indicate the historical status information of the third path.
11. The method according to claim 10, wherein Determining second path information related to the second tag based on the first path information related to the first tag includes: Refreshing the second path information related to the second tag at the end point to the information of the third path.
12. The method according to any one of claims 1 to 11, characterized in that, The first tag and the first path information, and the second tag and the second path information are both cached in the memory.
13. The method according to claim 12, characterized in that, The method further includes: When at least one of the first tag at the start point and the first path information related to the first tag satisfies a second preset condition, deleting the information of the first tag and the first path information cached in the memory.
14. The method according to claim 13, wherein The method further includes: When the historical status information of the target path indicated by the first tag at the start point satisfies the tag reduction condition, deleting the information of the first tag and the first path information cached in the memory.
15. The method according to claim 13 or 14, characterized in that The method further includes: When the path length of the first path information related to the first tag is greater than a preset length threshold, deleting the information of the first tag and the first path information cached in the memory.
16. The method according to any one of claims 1 to 15, characterized in that, The historical status information includes abnormality degree information; When the second tag satisfies the warning condition, outputting a local streaming origin graph of the end point of the target event based on the second path information includes: When the abnormality degree information of the target path indicated by the second tag is greater than a preset abnormality degree threshold, outputting a local streaming origin graph related to the end point of the target event based on the second path information.
17. The method according to any one of claims 1 to 16, characterized in that, Outputting a local streaming origin graph of the end point of the target event based on the second path information includes: Aggregating the second path information and the local streaming origin graph of the target node in the second path information to obtain an aggregated local streaming origin graph; wherein the target node is a node whose tag satisfies the warning condition; Outputting the aggregated local streaming origin graph as a local streaming origin graph related to the end point of the target event.
18. The method according to claim 17, wherein The method further includes: Updating the local streaming origin graph of each node in the aggregated local streaming origin graph to the aggregated local streaming origin graph.
19. The method according to any one of claims 1 to 18, characterized in that, Determining the second tag at the end point based on the first tag at the start point includes: Determining the second tag at the end point based on the first tag at the start point and a preset normalization coefficient.
20. A data processing device, characterized in that, The device includes: A first determination module, configured to determine a target event to be detected based on real-time collected log data, where the target event includes a starting point, an ending point, and the relationship between the starting point and the ending point; A second determination module, configured to determine a second tag of the ending point based on a first tag of the starting point; Wherein, the tag of each node in the target event is used to indicate the historical status information of the target path with that node as the ending point; wherein, the target path is the path with the highest anomaly degree among at least one path with that node as the ending point, and wherein, the at least one path is a path formed by at least one event to be detected determined based on the log data; wherein, the nodes in the target event include the starting point and the ending point in the target event; A third determination module, configured to determine second path information related to the second tag based on first path information related to the first tag; An alarm module, configured to output a local streaming provenance graph related to the ending point of the target event based on the second path information when the second tag meets the alarm condition.
21. A cluster of computing devices, characterized in that, Comprising at least one computing device, each computing device includes a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 19.
22. A computer program product comprising instructions, characterized in that, When the instructions are run by the computing device cluster, the computing device cluster is caused to execute the method according to any one of claims 1 to 19.
23. A computer-readable storage medium, characterized in that, Comprising computer program instructions, when the computer program instructions are executed by the computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 19.
Citation Information
Patent Citations
Webshell real-time detection method based on system audit log and scoring mechanism
CN112506885A
Abnormality tracing method combining system log and origin graph
CN112765603A
Attack tracing method and device, equipment and storage medium
CN115622802A
APT detection system based on data traceability graph label
CN116260627A
Attack chain real-time detection method and device, electronic equipment and storage medium
CN116318990A
Cited By
Operation and maintenance log analysis method and system based on AI intelligent agent
CN122240422A
Detection of anomalous activities in an enterprise network
US12621327B1