Abnormal Event Detection Method, Device, Computer Equipment and Storage Medium
By matching nodes and feature extraction between the event graph to be detected and query graphs, filtering out target node pairs and performing sub-graph matching, the problem of insufficient detection accuracy of abnormal event in the prior art is solved, and higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202210520174.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-05-13
AI Technical Summary
When the prior art detects suspicious data or suspicious behaviors present in the data, it is easy to identify normal data as abnormal data, resulting in poor accuracy.
By obtaining the event graph to be detected and query graphs, node matching and feature extraction are performed, target node pairs that meet the matching conditions are selected, and sub-graph matching is performed based on these node pairs to generate a target sub-graph describing abnormal events.
The accuracy of abnormal event detection is improved, and anomaly events that exist in the event diagram to be detected can be more accurately identified.
Smart Images

Figure CN115114484B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to an abnormal event detection method, apparatus, computer device, storage medium, and computer program product. Background Art
[0002] As a data structure for representing information, a graph structure is often used to describe data with inherent relevance and close connections. In many application fields, problems of information mining can be solved by relevant theories and corresponding technologies of graphs. As a basic operation for efficiently querying on graph data, the subgraph matching technology is also widely applied in various fields such as social network analysis and bioinformatics. For example, it can detect suspicious data or suspicious behaviors existing in the data.
[0003] However, the current methods for detecting suspicious data or suspicious behaviors existing in the data are prone to misidentifying normal data as abnormal data, resulting in poor accuracy. Summary of the Invention
[0004] Based on this, it is necessary to provide an abnormal event detection method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve accuracy for the above technical problems.
[0005] On the one hand, the present application provides an abnormal event detection method. The method includes:
[0006] Obtain a to-be-detected event graph and a query graph; both the to-be-detected event graph and the query graph include a plurality of nodes and edges, the edges of the to-be-detected event graph represent events occurring between corresponding nodes, and the edges of the query graph represent abnormal events occurring between corresponding nodes;
[0007] Perform matching processing on the nodes of the query graph and the nodes of the to-be-detected event graph to obtain a plurality of candidate node pairs;
[0008] Extract features for each node in the plurality of candidate node pairs respectively to obtain a node representation corresponding to each node;
[0009] For each of the candidate node pairs, according to the node representations of the nodes in the corresponding candidate node pair, screen out target node pairs that meet the matching conditions from the candidate node pairs;
[0010] Based on each of the target node pairs, perform subgraph matching on the to-be-detected event graph to obtain a target subgraph in the to-be-detected event graph that matches the query graph; the target subgraph is used to describe abnormal events existing in the to-be-detected event graph.
[0011] On the other hand, the present application also provides an abnormal event detection device. The device includes:
[0012] An acquisition module, configured to acquire an event graph to be detected and a query graph; both the event graph to be detected and the query graph include a plurality of nodes and edges, the edges of the event graph to be detected represent events occurring between corresponding nodes, and the edges of the query graph represent abnormal events occurring between corresponding nodes;
[0013] A node matching module, configured to perform matching processing on the nodes of the query graph and the nodes of the event graph to be detected to obtain a plurality of candidate node pairs;
[0014] An extraction module, configured to perform feature extraction on each node in the plurality of candidate node pairs respectively to obtain a node representation corresponding to each node;
[0015] A screening module, configured to, for each of the candidate node pairs, screen out target node pairs that meet the matching conditions from the candidate node pairs according to the node representations of the nodes in the corresponding candidate node pairs;
[0016] A subgraph matching module, configured to perform subgraph matching on the event graph to be detected based on each of the target node pairs to obtain a target subgraph in the event graph to be detected that matches the query graph; the target subgraph is used to describe the abnormal events existing in the event graph to be detected.
[0017] In one embodiment, the multiple nodes of the query graph include a plurality of first process nodes, the multiple nodes of the event graph to be detected include a plurality of second process nodes, and the candidate node pairs include candidate process node pairs obtained by performing matching processing on the first process nodes and the second process nodes.
[0018] In one embodiment, the multiple nodes of the query graph include first neighbor nodes of each of the first process nodes, the multiple nodes of the event graph to be detected include second neighbor nodes of each of the second process nodes, and the target node pairs include target process node pairs and target neighbor node pairs; the screening module is further configured to, for each of the target process node pairs, perform matching processing on the first neighbor nodes of the first process nodes in the corresponding target process node pairs and the second neighbor nodes of the second process nodes to obtain target neighbor node pairs formed by each of the first neighbor nodes and the second neighbor nodes that match; the subgraph matching module is further configured to perform subgraph matching on the event graph to be detected based on each of the target process node pairs and each of the target neighbor node pairs to obtain a target subgraph in the event graph to be detected that matches the query graph.
[0019] In one embodiment, the screening module is further configured to, for each of the candidate process node pairs, determine the similarity between the node representation of the first process node and the node representation of the second process node in the corresponding candidate process node pair, so as to obtain the similarity corresponding to each of the candidate process node pairs; for each of the first process nodes, based on the similarities corresponding to the respective candidate process node pairs to which the corresponding first process node belongs, screen out the target process node pairs that meet the matching conditions from the respective candidate process node pairs to which the corresponding first process node belongs.
[0020] In one embodiment, the screening module is further configured to, for each of the target process node pairs, determine the similarity between each first neighbor node of the first process node and each second neighbor node of the second process node in the corresponding target process node pair; for each of the first neighbor nodes, determine the second neighbor node that matches the corresponding first neighbor node according to the similarity between the corresponding first neighbor node and each of the second neighbor nodes, so as to obtain the target neighbor node pairs formed by each of the first neighbor nodes and the matching second neighbor nodes respectively.
[0021] In one embodiment, the screening module is further configured to, for each of the target process node pairs, determine the first neighbor nodes corresponding to the first process node in each node distance and the second neighbor nodes corresponding to the second process node in each node distance in the corresponding target process node pair; determine the similarity between each first neighbor node in each node distance and each second neighbor node in the corresponding node distance; for each node distance, determine the second neighbor node that matches each of the first neighbor nodes according to the similarity between each of the first neighbor nodes and each second neighbor node in the corresponding node distance, so as to obtain the target neighbor node pairs formed by each of the first neighbor nodes and the matching second neighbor nodes in each node distance respectively.
[0022] In one embodiment, the subgraph matching module is further configured to, for each of the target process node pairs, determine the target neighbor node pairs corresponding to the second process node in each node distance according to the node distance between the second process node and the matching second neighbor nodes in the corresponding target process node pair; for each of the target neighbor node pairs in each node distance, select the second neighbor nodes in the target neighbor node pairs that meet the association conditions from the target neighbor node pairs corresponding to the corresponding node distance; based on the association relationship between the selected second neighbor nodes and the matching second process nodes in the event graph to be detected in each node distance, generate the target subgraphs corresponding to each of the second process nodes respectively.
[0023] In one embodiment, the sub-graph matching module is further configured to select the pair of target process nodes with the largest similarity from each of the pairs of target process nodes; determine, according to the node distance between the second process node and the matched second neighbor node in the selected pair of target process nodes, the corresponding target neighbor node pairs of the second process node in the selected pair of target process nodes at each node distance; generate a sub-graph corresponding to the second process node in the selected pair of target process nodes based on the association relationship between the selected second process node and the selected second neighbor node at each node distance in the event graph to be detected; select the pair of target process nodes with the largest similarity from each of the unselected pairs of target process nodes, and return the step of determining, according to the node distance between the second process node and the matched second neighbor node in the selected pair of target process nodes, the corresponding target neighbor node pairs of the second process node in the selected pair of target process nodes at each node distance and continue to execute, until the generation of the sub-graph corresponding to the second process node in each pair of target process nodes stops, so as to obtain the target sub-graph in the event graph to be detected that matches the query graph.
[0024] In one embodiment, the screening module is further configured to determine the similarity between the node representations of the nodes in each corresponding candidate node pair for each of the candidate node pairs, and obtain the similarity corresponding to each of the candidate node pairs; screen out the target node pairs that meet the matching conditions from each of the candidate node pairs based on the respective similarities.
[0025] In one embodiment, the apparatus further includes a preprocessing module; the preprocessing module is configured to collect a plurality of nodes and edges from the event graph to be detected and the query graph to construct a plurality of triples, where each triple includes a node with an initialized node representation, a node with a node representation to be solved, and an edge with an initialized feature representation; construct a target loss function according to the initialized node representation, the node representation to be solved, and the initialized feature representation corresponding to the plurality of triples; perform iterative solution based on the target loss function until the iteration stops to obtain the node representation corresponding to each of the nodes.
[0026] In one embodiment, the preprocessing module is further configured to construct a first loss function according to the initialized node representation, the node representation to be solved, and the initialized feature representation corresponding to the plurality of triples; determine the initialized context representation corresponding to each process node in the event graph to be detected and the query graph; construct a second loss function based on the node representation to be solved and the corresponding initialized context representation corresponding to each process node; construct a target loss function according to the first loss function and the second loss function.
[0027] On the other hand, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0028] Obtain a to-be-detected event graph and a query graph; both the to-be-detected event graph and the query graph include a plurality of nodes and edges. The edges of the to-be-detected event graph represent events occurring between corresponding nodes, and the edges of the query graph represent abnormal events occurring between corresponding nodes;
[0029] Perform matching processing on the nodes of the query graph and the nodes of the to-be-detected event graph to obtain a plurality of candidate node pairs;
[0030] Extract features for each node in the plurality of candidate node pairs respectively to obtain a node representation corresponding to each node;
[0031] For each of the candidate node pairs, based on the node representations of the nodes in the corresponding candidate node pair, screen out target node pairs that meet the matching conditions from the candidate node pairs;
[0032] Based on each of the target node pairs, perform subgraph matching on the to-be-detected event graph to obtain a target subgraph in the to-be-detected event graph that matches the query graph; the target subgraph is used to describe abnormal events existing in the to-be-detected event graph.
[0033] On the other hand, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0034] Obtain a to-be-detected event graph and a query graph; both the to-be-detected event graph and the query graph include a plurality of nodes and edges. The edges of the to-be-detected event graph represent events occurring between corresponding nodes, and the edges of the query graph represent abnormal events occurring between corresponding nodes;
[0035] Perform matching processing on the nodes of the query graph and the nodes of the to-be-detected event graph to obtain a plurality of candidate node pairs;
[0036] Extract features for each node in the plurality of candidate node pairs respectively to obtain a node representation corresponding to each node;
[0037] For each of the candidate node pairs, based on the node representations of the nodes in the corresponding candidate node pair, screen out target node pairs that meet the matching conditions from the candidate node pairs;
[0038] Based on each of the target node pairs, perform subgraph matching on the event graph to be detected to obtain a target subgraph in the event graph to be detected that matches the query graph; the target subgraph is used to describe the abnormal events existing in the event graph to be detected.
[0039] On the other hand, the present application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the following steps:
[0040] Obtain an event graph to be detected and a query graph; both the event graph to be detected and the query graph include a plurality of nodes and edges, the edges of the event graph to be detected represent events occurring between corresponding nodes, and the edges of the query graph represent abnormal events occurring between corresponding nodes;
[0041] Perform matching processing on the nodes of the query graph and the nodes of the event graph to be detected to obtain a plurality of candidate node pairs;
[0042] Extract features for each node in a plurality of the candidate node pairs to obtain a node representation corresponding to each node;
[0043] For each of the candidate node pairs, based on the node representations of the nodes in the corresponding candidate node pair, screen out target node pairs that meet the matching conditions from the candidate node pairs;
[0044] Based on each of the target node pairs, perform subgraph matching on the event graph to be detected to obtain a target subgraph in the event graph to be detected that matches the query graph; the target subgraph is used to describe the abnormal events existing in the event graph to be detected.
[0045] In the above abnormal event detection method, device, computer device, storage medium, and computer program product, the query graph includes a plurality of nodes and edges, and the edges of the query graph represent abnormal events occurring between corresponding nodes, so the query graph can be used to describe the abnormal events occurring between nodes. The event graph to be detected includes a plurality of nodes and edges, and the edges of the event graph to be detected represent events occurring between corresponding nodes. Then, by matching the query graph and the event graph to be detected and detecting whether there is a subgraph in the event graph to be detected that matches the query graph, it can be accurately determined whether there are abnormal events in the event graph to be detected.
[0046] The matching between the query graph and the event graph to be detected can be refined into node matching. The nodes of the query graph and the nodes of the event graph to be detected are matched to preliminarily match the nodes in these two graphs, and the candidate node pairs for matching can be roughly determined. Feature extraction is performed on each node in multiple candidate node pairs to obtain the node representation corresponding to each node. The node representation contains the key information of the node and the relevant topological structure information of the node in the graph. For each candidate node pair, the target node pairs that meet the matching conditions are screened out from the candidate node pairs according to the node representations of the nodes in the corresponding candidate node pair. Using the node representation as the condition for screening the target node pairs can make full use of the information contained in the graph to further perform fine matching on the nodes, so as to improve the accuracy of screening. Based on each target node pair, subgraph matching is performed on the event graph to be detected to determine the target subgraph in the event graph to be detected that matches the query graph, so that the abnormal events existing in the event graph to be detected and the nodes where the abnormal events exist can be accurately described by the target subgraph. Description of the Drawings
[0047] Figure 1 It is a diagram of the application environment of the abnormal event detection method in an embodiment;
[0048] Figure 2 It is a schematic flowchart of the abnormal event detection method in an embodiment;
[0049] Figure 3 It is a schematic flowchart of the process of determining the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node in an embodiment;
[0050] Figure 4 It is a schematic flowchart of the steps of performing subgraph matching on the event graph to be detected based on each target process node pair and each target neighbor node pair to obtain the target subgraph in the event graph to be detected that matches the query graph in an embodiment;
[0051] Figure 5 It is a schematic flowchart of the process of generating the target subgraph corresponding to each second process node in an embodiment;
[0052] Figure 6 It is a schematic diagram of the query graph and the event graph to be detected in an embodiment;
[0053] Figure 7 It is a schematic diagram of the intermediate graph obtained by matching neighbor nodes with a node distance of 1 in an embodiment;
[0054] Figure 8 It is a schematic diagram of the target subgraph matched from the event graph to be detected in an embodiment;
[0055] Figure 9A schematic flow chart of constructing a target loss function according to the initialization node representation, the node representation to be solved, and the initialization feature representation corresponding to multiple triples in an embodiment;
[0056] Figure 10 A schematic flow chart of an abnormal event detection method in an embodiment;
[0057] Figure 11 A structural block diagram of an abnormal event detection device in an embodiment;
[0058] Figure 12 An internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0059] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0060] The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, data mining, etc. For example, it is applied to the field of artificial intelligence (AI) technology. Among them, artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning, and decision-making. The solution provided by the embodiments of the present application relates to an abnormal event detection method in artificial intelligence, which will be specifically described through the following embodiments.
[0061] The abnormal event detection method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed on the cloud or other network servers. Both the terminal 102 and the server 104 can separately execute the abnormal event detection method provided in the embodiments of the present application. The terminal 102 and the server 104 can also be used in cooperation to execute the abnormal event detection method provided in the embodiments of the present application. When the terminal 102 and the server 104 are used in cooperation to execute the abnormal event detection method provided in the embodiments of the present application, the terminal 102 obtains a to-be-detected event graph, which includes multiple nodes and edges. The edges of the to-be-detected event graph represent the events occurring between the corresponding nodes. The terminal 102 sends the to-be-detected event graph to the server 104. The server 104 obtains a corresponding query graph, which also includes multiple nodes and edges. The edges of the query graph represent the abnormal events occurring between the corresponding nodes. The server 104 performs a matching process on the nodes of the query graph and the nodes of the to-be-detected event graph to obtain multiple candidate node pairs. The server 104 respectively extracts features for each node in the multiple candidate node pairs to obtain a node representation corresponding to each node. For each candidate node pair, the server 104 filters out the target node pairs that meet the matching conditions from the candidate node pairs according to the node representations of the nodes in the corresponding candidate node pairs. The server 104 performs subgraph matching on the to-be-detected event graph based on the target node pairs to obtain a target subgraph in the to-be-detected event graph that matches the query graph; this target subgraph is used to describe the abnormal events existing in the to-be-detected event graph. The server 104 returns this target subgraph to the terminal 102, and the terminal 102 can discover the abnormal events existing in the to-be-detected event graph through the target subgraph and perform corresponding processing on the abnormal events.
[0062] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, portable wearable devices, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 102 and the server 104 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.
[0063] It should be noted that the quantity referred to by "multiple" and the like mentioned in the embodiments of the present application all refer to the quantity of "at least two".
[0064] In one embodiment, as Figure 2As shown, an abnormal event detection method is provided. Taking the application of this method to a computer device (the computer device can be a Figure 1 terminal or server in
[0065] Step S202, obtain the event graph to be detected and the query graph; both the event graph to be detected and the query graph include multiple nodes and edges. The edges of the event graph to be detected represent the events occurring between the corresponding nodes, and the edges of the query graph represent the abnormal events occurring between the corresponding nodes.
[0066] Among them, the graph structure is a data structure formed by multiple nodes connected to each other through edges. The entities after mathematical abstraction are called nodes, such as roles, processes, resources. The relationships between nodes form edges. The edges in a graph can be directed or undirected. The graph composed of directed edges is called a directed graph, and the graph composed of undirected edges is called an undirected graph. At the same time, the nodes and edges in the graph can also have multiple types. When there is only one type of node and one type of edge in the graph, the graph is called a single graph; otherwise, it is called a heterogeneous graph.
[0067] The multiple nodes in the event graph to be detected and the query graph include process nodes and the neighbor nodes of each process node; the neighbor nodes are resource nodes related to the process nodes, such as files, network addresses, monitors, etc.
[0068] The event graph to be detected is a directed heterogeneous graph, which represents the events occurring between nodes through the graph structure. The nodes represent entities, such as roles, processes, resources, and the edges represent the events occurring between entities, such as creation, access, establishment of connections, etc. The event graph to be processed can include at least one of a system event graph, a risk control event graph, and an object operation event graph. The nodes in the system event graph represent entities in the system, such as processes, files, network addresses, monitors, etc., but not limited to this; the edges in the system event graph represent the events between entities in the system, such as creation, access, establishment of connections, etc., but not limited to this.
[0069] A process refers to the running process of a program on a data set. A program running on different data sets or multiple runs of a program on the same data set are different processes. A monitor defines a data structure and a set of operations that can be executed by concurrent processes on this data structure. This set of operations can synchronize processes and change the data in the monitor.
[0070] The query graph includes at least one of a graph structure related to system abnormal events, a graph structure related to risk events, and a graph structure related to object abnormal behaviors.
[0071] It can be understood that the system event graph, the risk control event graph, and the object operation event graph each have their corresponding query graphs. For example, the query graph corresponding to the system event graph refers to a directed heterogeneous graph used to describe the time sequence and mutual relationships of abnormal events involving threats to the system. The query graph corresponding to the risk control event graph refers to a directed heterogeneous graph used to describe the time sequence and mutual relationships of events with risks. The query graph corresponding to the object operation event graph refers to a directed heterogeneous graph used to describe the time sequence and mutual relationships of abnormal behaviors of objects.
[0072] An abnormal event can refer to any event with abnormalities, including events with threats, events with risks, abnormal operation events, etc., but not limited to these.
[0073] Specifically, the computer device can determine each entity and the events occurring between the entities from the log, use each entity as a node, and use the events occurring between the entities as the edges connecting the nodes, so as to abstract the event graph to be detected from the log.
[0074] The computer device can obtain the query graph corresponding to the event graph to be detected. The query graph can be a pre-constructed graph structure used to describe the time sequence and relationships of abnormal events. The computer device can pre-determine each object and the abnormal events occurring between the objects from the abnormal event report, use each object as a node, and use the abnormal events occurring between the objects as the edges connecting the nodes to generate the query graph.
[0075] In one embodiment, the computer device can determine the type labels corresponding to each node in the event graph to be detected. The type label represents the type of the node. The type label can be, for example, a file, a process, a monitor, etc. The computer device can classify the nodes in the event graph to be detected according to the type label to obtain the nodes under each classification. The computer device can also determine the neighbor nodes corresponding to each node in the event graph to be detected, and can divide the neighbor nodes of each node into their respective neighbor node sets.
[0076] Similarly, the computer device can determine the type labels corresponding to each node in the query graph. The type label represents the type of the node. The type label can be, for example, a file type, a process type, a monitor type, etc.
[0077] The computer device can classify each node in the query graph according to the type label to obtain the nodes under each classification. The computer device can also determine the neighbor nodes corresponding to each node in the query graph, and can divide the neighbor nodes of each node into their respective neighbor node sets. The event graph to be detected and the query graph correspond to the same type label.
[0078] It can be understood that a node belonging to the process type is a process node, and a node belonging to the monitor type is a monitor node.
[0079] In one embodiment, after constructing a query graph, a computer device may pre-classify each node in the query graph according to type tags. It may also pre-determine the neighbor nodes corresponding to each node in the query graph and divide each neighbor node into the neighbor node set corresponding to the corresponding node, so as to improve the efficiency of subsequent node matching.
[0080] In one embodiment, a computer device may select a target type tag from multiple type tags, filter out the nodes belonging to the target type tag from the event graph to be detected, and filter out the nodes belonging to the target type tag from the query graph. For example, if the target type tag is the process type, the computer device may filter out the nodes belonging to the process type from the event graph to be detected and filter out the nodes belonging to the process type from the query graph. That is, filter out the process nodes in the event graph to be detected and the process nodes in the query graph.
[0081] In this embodiment, the computer device may determine the neighbor nodes corresponding to each process node and divide the neighbor nodes into the neighbor node set corresponding to the corresponding process node. The neighbor nodes are resource nodes directly or indirectly operated by the process node. For example, the neighbor node may be at least one of a file, a network address, and a monitor operated by the process node.
[0082] The neighbor nodes corresponding to a process node may be at least one of the direct neighbor nodes and the indirect neighbor nodes corresponding to the process node. A direct neighbor node refers to a neighbor node directly connected to the process node, and an indirect neighbor node refers to a neighbor node indirectly connected to the process node. There is at least one neighbor node between the indirect neighbor node and the process node, and they are indirectly connected through another neighbor node. For example, if process node A is directly connected to neighbor node B through an edge, and neighbor node B is directly connected to neighbor node C through an edge, and process node A is not connected to neighbor node C, then neighbor node B is the direct neighbor node of process node A, and neighbor node C is the indirect neighbor node of process node A.
[0083] The node distances between a process node and its direct neighbor node and indirectly connected node are different. For example, the node distance between a process node and its direct neighbor node is 1, and the node distance between a process node and an indirectly connected node can be 2, 3, 4, etc., but not limited thereto. The computer device may determine the node distance between a process node and each neighbor node to determine the neighbor nodes corresponding to the process node at each node distance, such as the neighbor nodes corresponding to a node distance of 1, the neighbor nodes corresponding to a node distance of 2, and the neighbor nodes corresponding to a node distance of 3. It may also divide the various neighbor nodes of the process node into the neighbor node sets corresponding to the corresponding node distances according to the node distance.
[0084] Step S204: Match the nodes of the query graph with the nodes of the event graph to be detected to obtain multiple candidate node pairs.
[0085] Among them, a candidate node pair refers to a node pair formed by the matching of the nodes of the query graph and the nodes of the event graph to be detected.
[0086] Specifically, the computer device can match each node in the query graph with the nodes in the event graph to be detected respectively, so as to obtain candidate node pairs formed by the nodes in the query graph and the nodes in the event graph to be detected.
[0087] In this embodiment, the computer device can determine the nodes with the same type of labels in the query graph and the event graph to be detected, and match the nodes with the same type of labels in the query graph and the event graph to be detected, so as to obtain multiple candidate node pairs corresponding to the corresponding type of labels. Further, for the nodes with the same type of labels in the query graph and the event graph to be detected, the computer device can match one node in the query graph with each node in the event graph to be detected respectively, so that one node in the query graph and each node in the event graph to be detected form a candidate node pair, that is, one node in the query graph corresponds to multiple candidate node pairs. According to the same processing method, the candidate node pairs formed by each node in the query graph and each node in the event graph to be detected can be obtained.
[0088] For example, the computer device can match each node belonging to the process type in the query graph with each node belonging to the process type in the event graph to be detected, so as to obtain multiple candidate node pairs under the process type.
[0089] Step S206: Extract features for each node in multiple candidate node pairs respectively to obtain a node representation corresponding to each node.
[0090] Among them, the node representation represents the fusion of the key information of the node and the relevant topological structure information of the node in the graph.
[0091] Specifically, for multiple candidate node pairs, the computer device can extract features for each node in each candidate node pair respectively to obtain a node representation corresponding to each node respectively.
[0092] In one embodiment, the computer device can preprocess the query graph and the event graph to be detected to obtain a node representation corresponding to each node in the query graph respectively, and obtain a node representation corresponding to each node in the event graph to be detected respectively.
[0093] In this embodiment, the computer device can preprocess the query graph. In the preprocessing, initial node representations are defined for some or all of the nodes in the query graph respectively, initial feature representations are defined for the edges of each node respectively, and by iteratively solving the initial node representations and the initial feature representations, the node representations corresponding to each node in the query graph are obtained when the iteration stops.
[0094] Furthermore, in the preprocessing, an initial context representation can also be defined for the context of the process nodes in the query graph. The context of the process node consists of the previous neighbor node and the next neighbor node adjacent to the process node, the edge between the process node and the previous neighbor node, and the edge between the process node and the next neighbor node. The computer device iteratively solves the initial node representations, the initial feature representations, and the initial context representation, so as to obtain the node representations corresponding to each node in the query graph when the iteration stops.
[0095] The preprocessing process of the event graph to be detected is the same as that of the query graph. Through preprocessing, the node representations corresponding to each node in the event graph to be detected can be obtained.
[0096] In one embodiment, the computer device can input the query graph and the event graph to be detected into a generation model. The generation model extracts features for each node of the query graph to obtain the node representations corresponding to each node of the query graph. And, the generation model extracts features for each node of the event graph to be detected to obtain the node representations corresponding to each node of the event graph to be detected.
[0097] Step S208, for each candidate node pair, according to the node representations of the nodes in the corresponding candidate node pair, filter out the target node pairs that meet the matching conditions from each candidate node pair.
[0098] Among them, the matching condition refers to a preset condition for filtering target node pairs. The matching condition can specifically be that the similarity is greater than the similarity threshold, or a preset number of candidate node pairs can be selected according to the similarity.
[0099] When a node in the query graph forms candidate node pairs with each node in the event graph to be detected respectively, or a node in the event graph to be detected forms candidate node pairs with each node in the query graph respectively, the matching condition can also be to select the candidate node pair with the highest similarity among the candidate node pairs belonging to the same node.
[0100] Specifically, the computer device calculates the similarity between the nodes in the candidate node pair according to the node representations corresponding to each node in the candidate node pair, and takes the calculated similarity as the similarity corresponding to the candidate node pair. By the same processing, the similarities corresponding to each candidate node pair can be obtained.
[0101] The computer device can obtain the matching conditions, match the similarities corresponding to each candidate node pair with the matching conditions to determine whether each candidate node pair meets the matching conditions. The computer device screens out the candidate node pairs that meet the matching conditions from each candidate node pair as the target node pairs.
[0102] In this embodiment, one node in the query graph forms candidate node pairs with each node in the event graph to be detected respectively. For the same node in the query graph, determine the candidate node pairs to which the same node belongs. According to the similarities corresponding to the candidate node pairs to which the same node belongs respectively, screen out the target node pairs that meet the matching conditions from the candidate node pairs to which the same node belongs. According to the same processing, the candidate node pairs to which each node in the query graph belongs respectively can be screened to obtain the target node pairs corresponding to each node in the query graph respectively.
[0103] In this embodiment, one node in the event graph to be detected forms candidate node pairs with each node in the query graph respectively. For the same node in the event graph to be detected, determine the candidate node pairs to which the same node belongs. According to the similarities corresponding to the candidate node pairs to which the same node belongs respectively, screen out the target node pairs that meet the matching conditions from the candidate node pairs to which the same node belongs. According to the same processing, the candidate node pairs to which each node in the event graph to be detected belongs respectively can be screened to obtain the target node pairs corresponding to each node in the event graph to be detected respectively.
[0104] Furthermore, the matching condition can be to select the candidate node pair with the highest similarity among the candidate node pairs to which the same node belongs. The computer device can screen out the candidate node pair with the highest similarity as the target node pair from the candidate node pairs to which the same node belongs according to the similarities corresponding to the candidate node pairs to which the same node belongs respectively.
[0105] Step S210, perform subgraph matching on the event graph to be detected based on each target node pair to obtain a target subgraph in the event graph to be detected that matches the query graph; the target subgraph is used to describe the abnormal events existing in the event graph to be detected.
[0106] Among them, the target subgraph includes multiple nodes and edges, and the target subgraph is a subgraph in the event graph to be detected that matches the query graph. The edges of the target subgraph represent the abnormal events occurring between the corresponding nodes. The target subgraph, as a subgraph of the event graph to be detected, is also a subgraph of the query graph.
[0107] For two graphs G and H, when the node set of H is a subset of the node set of G and the edge set of H is a subset of the edge set of G, then H is called a subgraph of G.
[0108] Specifically, the computer device performs subgraph matching on the event graph to be detected based on each selected target node pair, so as to match a subgraph in the event graph to be detected that matches the query graph. Further, the computer device performs subgraph matching on the event graph to be detected based on each selected target node pair, so as to match a subgraph in the event graph to be detected that is similar to the query graph.
[0109] In one embodiment, performing subgraph matching on the event graph to be detected based on each target node pair to obtain a target subgraph in the event graph to be detected that matches the query graph includes: sequentially selecting target node pairs according to the similarity of each target node pair, and generating a corresponding target subgraph according to the nodes in the selected target node pairs that belong to the event graph to be detected and the association relationships of the selected nodes in the event graph to be detected.
[0110] In the above abnormal event detection method, the query graph includes multiple nodes and edges, and the edges of the query graph represent abnormal events occurring between the corresponding nodes. Then the query graph can be used to describe the abnormal events occurring between the nodes. The event graph to be detected includes multiple nodes and edges, and the edges of the event graph to be detected represent events occurring between the corresponding nodes. By matching the query graph and the event graph to be detected and detecting whether there is a subgraph in the event graph to be detected that matches the query graph, it can be accurately determined whether there is an abnormal event in the event graph to be detected.
[0111] The matching of the query graph and the event graph to be detected can be refined into node matching. The nodes of the query graph and the nodes of the event graph to be detected are matched to perform preliminary matching on the nodes in these two graphs, and the candidate node pairs for matching can be roughly determined. Feature extraction is respectively performed on each node in multiple candidate node pairs to obtain a node representation corresponding to each node. The node representation includes the key information of the node and the relevant topological structure information of the node in the graph. For each candidate node pair, target node pairs that meet the matching conditions are selected from the candidate node pairs according to the node representations of the nodes in the corresponding candidate node pairs. By using the node representation as the condition for screening the target node pairs, the information contained in the graph can be fully utilized to further perform fine matching on the nodes, so as to improve the accuracy of screening. Subgraph matching is performed on the event graph to be detected based on each target node pair to determine a target subgraph in the event graph to be detected that matches the query graph, so that the abnormal events existing in the event graph to be detected and the nodes where the abnormal events occur can be accurately described by the target subgraph.
[0112] In one embodiment, the multiple nodes of the query graph include multiple first process nodes, the multiple nodes of the event graph to be detected include multiple second process nodes, and the candidate node pairs include candidate process node pairs obtained by matching the first process nodes and the second process nodes.
[0113] Specifically, the nodes belonging to the process type in the query graph and the event graph to be detected are called process nodes. The process nodes in the query graph are called first process nodes, and the process nodes in the event graph to be detected are called second process nodes. The candidate node pairs include candidate process node pairs.
[0114] The computer device can match each first process node with each second process node respectively, so that one first process node corresponds to multiple second process nodes, and each first process node forms a corresponding candidate process node pair with each second process node.
[0115] The computer device extracts features from each process node in multiple candidate process node pairs respectively to obtain the node representation corresponding to each process node. For each candidate node pair, the computer device filters out the target process node pairs that meet the matching conditions from each candidate process node pair according to the node representations of the process nodes in the corresponding candidate process node pair. Based on each target process node pair, subgraph matching is performed on the event graph to be detected to obtain the target subgraph in the event graph to be detected that matches the query graph.
[0116] In one embodiment, the multiple nodes of the query graph include multiple first process nodes, the multiple nodes of the event graph to be detected include multiple second process nodes, and the candidate node pairs include candidate process node pairs; the first process nodes of the query graph and the second process nodes of the event graph to be detected are matched to obtain multiple candidate process node pairs, including:
[0117] Extract each first process node from the query graph and extract each second process node from the event graph to be detected; match each first process node with each second process node respectively to obtain the candidate process node pairs formed by each first process node and each second process node respectively.
[0118] Specifically, the computer device can extract each first process node from the query graph and extract each second process node from the event graph to be detected, and match the extracted first process nodes and the extracted second process nodes respectively to obtain the candidate process node pairs formed by each first process node and each second process node respectively.
[0119] In one embodiment, the candidate node pairs can also include candidate neighbor node pairs. The process nodes in the query graph and the event graph to be detected respectively correspond to their respective neighbor nodes. The neighbor nodes in the query graph are called first neighbor nodes, and the neighbor nodes in the event graph to be detected are called second neighbor nodes. The computer device matches each first neighbor node of the first process node in the candidate process node pair with each second neighbor node of the second process node in the candidate process node pair respectively to obtain the candidate neighbor node pairs formed by each first neighbor node and each second neighbor node matching respectively.
[0120] In this embodiment, the process node serves as the subject of the event. By matching the first process node in the query graph with the second process node in the graph to be detected, the approximate position of the process node in the graph to be detected that may match the first process node can be determined through preliminary matching, facilitating subsequent fine matching of the process nodes. For each pair of candidate process nodes, target process node pairs that meet the matching conditions are selected from each pair of candidate process nodes according to the node representations of the process nodes in the corresponding pair of candidate process nodes. Using the node representation as the condition for screening the target process node pairs can make full use of the information contained in the graph to further perform fine matching on the process nodes, thereby improving the accuracy of screening. Based on each target process node pair, subgraph matching is performed on the graph to be detected to determine the target subgraph in the graph to be detected that matches the query graph, so that the abnormal events and the nodes where the abnormal events exist in the graph to be detected can be accurately described by the target subgraph.
[0121] In one embodiment, the multiple nodes of the query graph include the first neighbor nodes of each first process node, the multiple nodes of the graph to be detected include the second neighbor nodes of each second process node, and the target node pairs include target process node pairs and target neighbor node pairs. Based on each target node pair, subgraph matching is performed on the graph to be detected to obtain the target subgraph in the graph to be detected that matches the query graph, including:
[0122] For each target process node pair, the first neighbor nodes of the first process node in the corresponding target process node pair are matched with the second neighbor nodes of the second process node to obtain target neighbor node pairs formed by each first neighbor node and the matched second neighbor node. Based on each target process node pair and each target neighbor node pair, subgraph matching is performed on the graph to be detected to obtain the target subgraph in the graph to be detected that matches the query graph.
[0123] Specifically, the multiple nodes of the query graph include the first neighbor nodes corresponding to each first process node respectively, and the multiple nodes of the graph to be detected include the second neighbor nodes corresponding to each second process node respectively.
[0124] The computer device can match each first process node with each second process node respectively, so that one first process node corresponds to multiple second process nodes, and each first process node and each second process node form corresponding candidate process node pairs respectively. The computer device extracts features from each process node in the multiple candidate process node pairs respectively to obtain the node representation corresponding to each process node. For each pair of candidate process nodes, the computer device selects target process node pairs that meet the matching conditions from each pair of candidate process nodes according to the node representations of the process nodes in the corresponding pair of candidate process nodes.
[0125] The computer device can determine each first neighbor node corresponding to the first process node in the target process node pair, and determine each second neighbor node corresponding to the second process node in the corresponding process node pair. Each of the first neighbor nodes is respectively matched with the second neighbor nodes to obtain a target neighbor node pair formed by each first neighbor node and the matched second neighbor node.
[0126] In this embodiment, for each target process node pair, matching the first neighbor node of the first process node with the second neighbor node of the second process node in the corresponding target process node pair to obtain a target neighbor node pair formed by each first neighbor node and the matched second neighbor node includes:
[0127] For each target process node pair, each first neighbor node corresponding to the first process node in the corresponding target process node pair is respectively matched with each second neighbor node of the second process node to obtain candidate neighbor node pairs formed by each first neighbor node and each second neighbor node; target neighbor node pairs that meet the matching conditions are selected from the respective candidate neighbor node pairs.
[0128] Selecting target neighbor node pairs that meet the matching conditions from the candidate neighbor node pairs includes: for each candidate neighbor node pair, determining the similarity between the node representation of the first neighbor node and the node representation of the second neighbor node in the corresponding candidate neighbor node pair; based on the similarity, selecting target neighbor node pairs that meet the matching conditions from the respective candidate neighbor node pairs.
[0129] It can be understood that the matching condition for screening the target process node pair can be referred to as the first matching condition, and the matching condition for screening the target neighbor node pair can be referred to as the second matching condition. The second matching condition may be the same as or different from the first matching condition.
[0130] In this embodiment, for each target process node pair, matching the first neighbor node of the first process node with the second neighbor node of the second process node in the corresponding target process node pair to obtain a target neighbor node pair formed by each first neighbor node and the matched second neighbor node includes:
[0131] For each target process node pair, determining the similarity between each first neighbor node of the first process node in the corresponding target process node pair and each second neighbor node of the second process node; for each first neighbor node, determining the second neighbor node that matches the corresponding first neighbor node according to the similarity between the corresponding first neighbor node and each second neighbor node, so as to obtain a target neighbor node pair formed by each first neighbor node and the matched second neighbor node respectively.
[0132] The computer device selects target process node pairs from the target process node pairs based on the corresponding similarities of each target process node pair, and determines each target neighbor node pair corresponding to the selected target process node pairs. Based on the similarities of each target neighbor node pair corresponding to the selected target process node pairs, it selects target neighbor node pairs from the target neighbor node pairs. According to the association relationship between the second process node in the selected target process node pair and the second neighbor node in the selected target neighbor node pair in the event graph to be detected, it generates a corresponding target subgraph.
[0133] In this embodiment, the multiple nodes of the query graph include a first process node and the first neighbor node of each first process node, and the multiple nodes of the event graph to be detected include a second process node and the second neighbor node of each second process node. For each target process node pair, the first neighbor node of the first process node in the corresponding target process node pair is matched with the second neighbor node of the second process node to obtain a target neighbor node pair formed by each first neighbor node and the matched second neighbor node; based on each target process node pair and each target neighbor node pair, subgraph matching is performed on the event graph to be detected to obtain a target subgraph in the event graph to be detected that matches the query graph. First, a preliminary match is made between the process nodes of the query graph and the process nodes of the event graph to be detected, and based on the node representations of the process nodes in each candidate process node pair, a fine match is made on the multiple candidate process node pairs generated by the preliminary match to filter out target process node pairs that meet the matching conditions.
[0134] The first neighbor node of the first process node in the target process node pair is matched with the second neighbor node of the second process node to obtain a target neighbor node pair formed by each first neighbor node and the matched second neighbor node, so that the matching of neighbor nodes can be performed after the screening of the target process node pair is completed, which can reduce the amount of data to be matched and improve the processing efficiency. The target process node pair is composed of a first process node and a second process node with a higher matching degree, and the target neighbor node pair is composed of a first neighbor node and a second neighbor node with a higher matching degree. Then, based on each target process node pair and each target neighbor node pair with a higher matching degree, subgraph matching is performed on the event graph to be detected, and the target subgraph that matches the query graph can be more accurately determined from the event graph to be detected.
[0135] In one embodiment, the multiple nodes of the query graph include multiple first process nodes and the first neighbor nodes of each first process node, the multiple nodes of the event graph to be detected include multiple second process nodes and the second neighbor nodes of each second process node, and the candidate node pairs include candidate process node pairs and candidate neighbor node pairs; performing a matching process on the nodes of the query graph and the nodes of the event graph to be detected to obtain multiple candidate node pairs, including: performing a matching process on the first process nodes of the query graph and the second process nodes of the event graph to be detected to obtain multiple candidate process node pairs; performing a matching process on the first neighbor nodes corresponding to the first process nodes and the second neighbor nodes corresponding to the corresponding second process nodes to obtain multiple candidate neighbor node pairs.
[0136] The matching conditions include a first matching condition and a second matching condition, and the target node pairs include target process node pairs and target neighbor node pairs; for each candidate node pair, screening out the target node pairs that meet the matching conditions from each candidate node pair according to the node representations of the nodes in the corresponding candidate node pair, including: for each candidate process node pair, screening out the target process node pairs that meet the first matching condition from each candidate process node pair according to the node representations of the first process node and the second process node in the corresponding candidate process node pair; for each candidate neighbor node pair, screening out the target neighbor node pairs that meet the second matching condition from each candidate neighbor node pair according to the node representations of the first neighbor node and the second neighbor node in the corresponding candidate neighbor node pair. The first matching condition and the first matching condition may be the same or different.
[0137] In one embodiment, for each candidate node pair, screening out the target node pairs that meet the matching conditions from each candidate node pair according to the node representations of the nodes in the corresponding candidate node pair, including:
[0138] For each candidate process node pair, determining the similarity between the node representation of the first process node and the node representation of the second process node in the corresponding candidate process node pair to obtain the similarity corresponding to each candidate process node pair respectively; for each first process node, screening out the target process node pairs that meet the matching conditions from the candidate process node pairs to which the corresponding first process node belongs based on the similarities corresponding to the candidate process node pairs to which the corresponding first process node belongs.
[0139] Specifically, a candidate process node pair includes a first process node and a second process node, and the computer device calculates the similarity between the first process node and the second process node according to the node representation corresponding to the first process node and the node representation corresponding to the second process node in the candidate process node pair, and takes the calculated similarity as the similarity corresponding to the candidate process node pair. By the same process, the similarity corresponding to each candidate process node pair can be obtained.
[0140] The same first process node can form different candidate process node pairs with different second process nodes. Then, the computer device can determine each candidate process node pair to which the same first process node belongs, and determine the similarities respectively corresponding to each candidate process node pair to which the same first process node belongs, so as to obtain multiple similarities corresponding to the same first process node. Based on the multiple similarities corresponding to the same first process node, the computer device filters out the target process node pairs that meet the matching conditions from each candidate process node pair to which the first process node belongs, and obtains the target process node pairs corresponding to the first process node. According to the same processing, the computer device can filter out the target process node pairs corresponding to each first process node respectively.
[0141] In this embodiment, the same second process node can form different candidate process node pairs with different first process nodes. Then, the computer device can determine each candidate process node pair to which the same second process node belongs, and determine the similarities respectively corresponding to each candidate process node pair to which the same second process node belongs, so as to obtain multiple similarities corresponding to the same second process node. Based on the multiple similarities corresponding to the same second process node, the computer device filters out the target process node pairs that meet the matching conditions from each candidate process node pair to which the second process node belongs, and obtains the target process node pairs corresponding to the second process node. According to the same processing, the computer device can filter out the target process node pairs corresponding to each second process node respectively.
[0142] In this embodiment, meeting the matching conditions can specifically be that the similarity of the candidate process node pair is greater than the similarity threshold, or it can be to select a preset number of candidate process node pairs according to the similarity, or it can be to select the candidate process node pair with the highest similarity among each candidate process node pair to which the same process node belongs.
[0143] For example, when the matching condition is to select the candidate process node pair with the highest similarity among each candidate process node pair to which the same process node belongs, the computer device can filter out the candidate process node pair with the highest similarity from each candidate process node pair to which the same process node belongs as the target process node pair according to the similarities respectively corresponding to each candidate process node pair to which the same process node belongs. This process node can be the first process node or the second process node.
[0144] In this embodiment, the similarity can specifically be the sine similarity, the cosine similarity, etc., but is not limited thereto, and the similarity threshold corresponds to the sine similarity threshold and the cosine similarity threshold.
[0145] In this embodiment, for each candidate process node pair, the similarity between the node representation of the first process node and the node representation of the second process node in the corresponding candidate process node pair is determined, and the similarity corresponding to each candidate process node pair is obtained. Thus, the similarity between the node representations of the nodes can be used as a screening condition to further screen the candidate process node pairs. For each first process node, based on the similarities corresponding to the respective candidate process node pairs to which the corresponding first process node belongs, the target process node pairs that meet the matching conditions are accurately screened out from the respective candidate process node pairs to which the corresponding first process node belongs.
[0146] In one embodiment, for each target process node pair, the first neighbor node of the first process node in the corresponding target process node pair is matched with the second neighbor node of the second process node to obtain a target neighbor node pair formed by each first neighbor node and the matched second neighbor node, including:
[0147] For each target process node pair, the similarity between each first neighbor node of the first process node and each second neighbor node of the second process node in the corresponding target process node pair is determined; for each first neighbor node, according to the similarities between the corresponding first neighbor node and each second neighbor node, the second neighbor node that matches the corresponding first neighbor node is determined, so as to obtain a target neighbor node pair formed by each first neighbor node and the matched second neighbor node respectively.
[0148] Specifically, the computer device can determine each first neighbor node corresponding to the first process node in the target process node pair, and determine each second neighbor node corresponding to the second process node in the corresponding process node pair. Then, the computer device can calculate the similarity between the same first neighbor node and each second neighbor node respectively, and obtain multiple similarities corresponding to the same first neighbor node. According to the multiple similarities corresponding to the same first neighbor node, the second neighbor node that matches the first neighbor node is screened out from the second neighbor nodes corresponding to each similarity. The first neighbor node and the matched second neighbor node are used as the target neighbor node pair.
[0149] Furthermore, the computer device can determine the node representation of each first neighbor node and the node representation of each second neighbor node, and calculate the similarity between the node representation of the same first neighbor node and the node representation of each second neighbor node respectively, so as to obtain the similarity between the same first neighbor node and each second neighbor node respectively.
[0150] For example, from the second neighbor nodes corresponding to each similarity, second neighbor nodes with a similarity greater than the similarity threshold are filtered out as the second neighbor nodes that match the first neighbor node. Alternatively, from the second neighbor nodes corresponding to each similarity, the second neighbor node with the highest similarity is filtered out as the second neighbor node that matches the first neighbor node.
[0151] According to the same process, the computer device can determine the second neighbor nodes that match each first neighbor node, thereby obtaining multiple target neighbor node pairs.
[0152] In this embodiment, the computer device can determine each first neighbor node corresponding to the first process node in the target process node pair, and determine each second neighbor node corresponding to the second process node in the corresponding process node pair. Then, the computer device can calculate the similarity between the same second neighbor node and each first neighbor node respectively, to obtain multiple similarities corresponding to the same second neighbor node. According to the multiple similarities corresponding to the same second neighbor node, the first neighbor nodes that match the second neighbor node are filtered out from the first neighbor nodes corresponding to each similarity. The second neighbor node and the matching first neighbor node form a target neighbor node pair. According to the same process, the first neighbor nodes that match each second neighbor node can be determined, thereby obtaining multiple target neighbor node pairs.
[0153] In this embodiment, for each target process node pair, determining the similarity between each first neighbor node of the first process node and each second neighbor node of the second process node in the corresponding target process node pair can filter out the most matching second neighbor node for each first neighbor node based on the similarity. For each first neighbor node, according to the similarity between the corresponding first neighbor node and each second neighbor node respectively, the second neighbor node that matches the corresponding first neighbor node is determined, so as to obtain the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node respectively, such that the matching degree between the neighbor nodes in the generated target neighbor node pairs is high, that is, the matching degree of the neighbor node pairs used for subgraph matching is high, which can effectively improve the accuracy of subsequent subgraph matching.
[0154] In one embodiment, for each pair of target process nodes, determining the similarity between each first neighbor node of the first process node in the corresponding pair of target process nodes and each second neighbor node of the second process node includes: for each pair of target process nodes, matching each first neighbor node of the first process node in the corresponding pair of target process nodes with each second neighbor node of the second process node to obtain candidate neighbor node pairs formed by each first neighbor node and each second neighbor node respectively; for each candidate neighbor node pair, determining the similarity between the node representation of the first neighbor node and the node representation of the second neighbor node in the corresponding candidate neighbor node pair to obtain the similarity corresponding to each candidate neighbor node pair respectively.
[0155] For each first neighbor node, determining the second neighbor node that matches the corresponding first neighbor node according to the similarity between the corresponding first neighbor node and each second neighbor node respectively, so as to obtain target neighbor node pairs formed by each first neighbor node and the matching second neighbor node respectively, includes: for each first neighbor node, based on the similarities corresponding to each candidate neighbor node pair to which the corresponding first neighbor node belongs, screening out target neighbor node pairs that meet the matching conditions from each candidate neighbor node pair to which the corresponding first neighbor node belongs.
[0156] In one embodiment, the method further includes: for each pair of target process nodes, determining the node distance between the second process node in the corresponding pair of target process nodes and each matching second neighbor node; for each target neighbor node pair, dividing according to the node distance between the second neighbor node in the corresponding target neighbor node pair and the corresponding second process node to obtain the target neighbor node pairs corresponding to each second process node under each node distance respectively.
[0157] In one embodiment, as Figure 3 shown, for each pair of target process nodes, determining the similarity between each first neighbor node of the first process node in the corresponding pair of target process nodes and each second neighbor node of the second process node includes steps S302 - S304:
[0158] Step S302, for each pair of target process nodes, determining the first neighbor nodes corresponding to the first process node in the corresponding pair of target process nodes under each node distance respectively, and the second neighbor nodes corresponding to the second process node under each node distance respectively.
[0159] Specifically, the node distances between a first process node in the query graph and multiple first neighbor nodes may be different. For example, a certain first process node in the query graph has 3 first neighbor nodes with a node distance of 1 and 2 first neighbor nodes with a node distance of 2.
[0160] After the computer device filters out each target process node pair, it can determine the first neighbor nodes corresponding to the first process node in each target process node pair under each node distance, and determine the second neighbor nodes corresponding to the second process node in each target process node under each node distance.
[0161] In the same processing manner, the first neighbor nodes corresponding to the first process node in each target process node under each node distance, and the second neighbor nodes corresponding to the second process node in each target process node under each node distance can be obtained.
[0162] Step S304, determine the similarity between each first neighbor node under each node distance and each second neighbor node under the corresponding node distance.
[0163] Specifically, for the first process node and the second process node in a single target process node pair, calculate the similarity between each corresponding first neighbor node and each second neighbor node under the same node distance, to obtain the similarity between each first neighbor node and each second neighbor node under the same node distance. The same processing is performed for the first neighbor nodes and the second neighbor nodes under each node distance, and the similarity between each first neighbor node and each second neighbor node under each node distance can be obtained. Thus, the similarity calculation between the first neighbor nodes and the second neighbor nodes corresponding to different node distances of a single target process node pair is completed.
[0164] The same processing is performed for each target process node pair, and the similarity between the neighbor nodes of the two process nodes in each target process node pair under the same node distance can be obtained.
[0165] For each first neighbor node, according to the similarity between the corresponding first neighbor node and each second neighbor node, determine the second neighbor node that matches the corresponding first neighbor node, so as to obtain the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node respectively, including step S306:
[0166] Step S306, for each node distance, according to the similarity between each first neighbor node and each second neighbor node under the corresponding node distance, determine the second neighbor node that matches each first neighbor node, so as to obtain the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node under each node distance respectively.
[0167] Specifically, the computer device determines the second neighbor nodes that match the first neighbor nodes according to the similarities between the first neighbor nodes at the same node distance and each second neighbor node. The first neighbor node and the matching second neighbor node form a target neighbor node pair at the same node distance. By performing the same processing on each pair of first neighbor nodes, the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node at the same node distance can be obtained.
[0168] By performing the same processing on the first neighbor nodes at each node distance, the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node at each node distance can be obtained.
[0169] For example, in the target process node pair A, the first process node has two first neighbor nodes a1 and a2 at a node distance of 1, and two first neighbor nodes a3 and a4 at a node distance of 2. The second process node in the target process node pair A has three second neighbor nodes b1, b2, and b3 at a node distance of 1, and two second neighbor nodes b4 and b5 at a node distance of 2. The computer device calculates the similarities between the first neighbor nodes a1 and a2 and the second neighbor nodes b1, b2, and b3 respectively at a node distance of 1, that is, the similarities between a1 and b1, a1 and b2, a1 and b3, a2 and b1, a2 and b2, and a2 and b3. For the three similarities corresponding to a1, the second neighbor node with the highest similarity is selected as the second neighbor node that matches a1. If the second neighbor node with the highest similarity is b2, then b2 is used as the second neighbor node that matches a1, and the target neighbor node pair <a1, b2> formed by a1 and b2 is obtained. The node distance corresponding to the target neighbor node pair <a1, b2> is 1. The same applies to the three similarities corresponding to a2, and the second neighbor node that matches a2 among b1, b2, and b3 can be determined to form the target neighbor node pair corresponding to a2 at a node distance of 1. After the two first neighbor nodes a1 and a2 at a node distance of 1 both form their respective target neighbor node pairs, the target neighbor node pairs corresponding to the first process node at a node distance of 1 can be obtained, which are also the target neighbor node pairs corresponding to the second process node at a node distance of 1.
[0170] The computer device calculates the similarities between the first neighbor nodes a3 and a4 and the second neighbor nodes b4 and b5 respectively at a node distance of 2, that is, the similarities between a3 and b4, a3 and b5, a4 and b4, and a4 and b5. The processing method is the same as that at a node distance of 1, and the target neighbor node pairs corresponding to the first process node at a node distance of 2 can be obtained.
[0171] In this embodiment, for each target process node pair, the first neighbor nodes respectively corresponding to the first process node in each node distance in the corresponding target process node pair and the second neighbor nodes respectively corresponding to the second process node in each node distance are determined, so that the neighbor nodes of the process nodes can be divided according to different node distances. The similarity between each first neighbor node in each node distance and each second neighbor node in the corresponding node distance is determined to divide the similarity between neighbor nodes according to the node distance as well. For each node distance, according to the similarity between each first neighbor node and each second neighbor node in the corresponding node distance, the second neighbor node that matches each first neighbor node is determined, so as to obtain, for each node distance, the target neighbor node pair formed by each first neighbor node and the matching second neighbor node, enabling the screening of the matching first neighbor node and second neighbor node at the same node distance, and the formed target neighbor node pair has a higher matching degree, which helps to improve the accuracy of subsequent subgraph matching.
[0172] In one embodiment, based on each target process node pair and each target neighbor node pair, subgraph matching is performed on the event graph to be detected to obtain the target subgraph in the event graph to be detected that matches the query graph, including:
[0173] For each target process node pair, according to the node distance between the second process node in the corresponding target process node pair and each matching second neighbor node, the target neighbor node pair respectively corresponding to the second process node in each node distance is determined; for each target neighbor node pair in each node distance, the second neighbor node is selected from each target neighbor node pair corresponding to the corresponding node distance; based on the association relationship between the second neighbor node selected in each node distance and the matching second process node in the event graph to be detected, the target subgraph respectively corresponding to each second process node is generated.
[0174] In one embodiment, for each target neighbor node pair in each node distance, selecting the second neighbor node from each target neighbor node pair corresponding to the corresponding node distance includes: for each target neighbor node pair in each node distance, selecting the second neighbor node from each target neighbor node pair corresponding to the corresponding node distance.
[0175] In one embodiment, for each target neighbor node pair in each node distance, selecting the second neighbor node from each target neighbor node pair corresponding to the corresponding node distance includes: for each target neighbor node pair in each node distance, selecting the second neighbor node from the target neighbor node pairs that meet the association condition corresponding to the corresponding node distance.
[0176] In one embodiment, as Figure 4 shown, based on each target process node pair and each target neighbor node pair, perform subgraph matching on the event graph to be detected to obtain a target subgraph in the event graph to be detected that matches the query graph, including:
[0177] Step S402, for each target process node pair, determine the target neighbor node pair corresponding to the second process node in each node distance according to the node distance between the second process node in the corresponding target process node pair and each matching second neighbor node.
[0178] Specifically, for each target process node pair, determine the first neighbor node corresponding to the first process node in the corresponding target process node pair in each node distance, and the second neighbor node corresponding to the second process node in each node distance; determine the similarity between each first neighbor node in each node distance and each second neighbor node in the corresponding node distance; for each node distance, according to the similarity between each first neighbor node and each second neighbor node in the corresponding node distance, determine the second neighbor node that matches each first neighbor node, so as to obtain, in each node distance, the target neighbor node pair formed by each first neighbor node and the matching second neighbor node, that is, the target neighbor node pair corresponding to the first process node in each target process node pair in each node distance, and also the target neighbor node pair corresponding to the second process node in each target process node pair in each node distance. Thus, the target neighbor node pair corresponding to the second process node in each target process node pair in each node distance can be obtained.
[0179] Step S404, for each target neighbor node pair in each node distance, select the second neighbor node in the target neighbor node pair that meets the association condition from the target neighbor node pairs corresponding to the corresponding node distance.
[0180] Among them, the association condition refers to a preset condition for selecting the second neighbor node in the target neighbor node pair, which can specifically be that the similarity of the target neighbor node pair is the largest, or the similarity of the target neighbor node pair is greater than the similarity threshold.
[0181] Specifically, the computer device can determine the similarity corresponding to each pair of target neighbor nodes. For each pair of target neighbor nodes corresponding to the same pair of target process nodes under each node distance, the computer device selects, from the pairs of target neighbor nodes under the same node distance, the pairs of target neighbor nodes whose similarity meets the association condition, and selects the second neighbor nodes from the pairs of target neighbor nodes that meet the association condition. According to the same processing, for the same target process node, pairs of target neighbor nodes whose similarity meets the association condition can be selected from the pairs of target neighbor nodes under each node distance, so as to select the second neighbor nodes.
[0182] According to the same processing for each target process node, the second neighbor nodes corresponding to each target process node under each node distance can be selected.
[0183] In one embodiment, the computer device can select at least one pair of target neighbor nodes from the pairs of target neighbor nodes corresponding to each node distance in ascending order of node distance.
[0184] Step S406: Based on the association relationship between the second neighbor nodes selected under each node distance and the matching second process nodes in the event graph to be detected, generate a target subgraph corresponding to each second process node.
[0185] Specifically, after the computer device selects the second neighbor nodes corresponding to the same target process node under each node distance, it can generate a corresponding target subgraph according to the association relationship between the second process node in the same target process node and the selected second neighbor nodes in the event graph to be detected. The target subgraph corresponds to the second process node in the same target process node.
[0186] This association relationship can be based on the edge representation of the second process node and each second neighbor node in the event graph to be detected. According to the edges connecting the second process node and each second neighbor node, and the edges connecting each second neighbor node in the event graph to be detected, the second process node and the selected second neighbor nodes are connected to generate the target subgraph.
[0187] The same processing is performed for each target process node, that is, a target subgraph corresponding to the second process node in each target process node is generated.
[0188] In this embodiment, the multiple nodes of the target subgraph include at least one process node and at least one neighbor node corresponding to the at least one process node, and the edges of the subgraph are used to describe the abnormal events occurring between the corresponding nodes.
[0189] In this embodiment, for each target process node pair, according to the node distances between the second process node in the corresponding target process node pair and each matching second neighbor node, the target neighbor node pairs respectively corresponding to the second process node at each node distance are determined. For each target neighbor node pair at each node distance, the second neighbor node in the target neighbor node pair that meets the association condition is selected from the target neighbor node pairs corresponding to the corresponding node distance, so that the neighbor node most closely associated with the second process node can be selected at each node distance. Therefore, based on the association relationship between the second neighbor node selected at each node distance and the matching second process node in the event graph to be detected, the target subgraph respectively corresponding to each second process node can be accurately constructed.
[0190] In one embodiment, as Figure 5 shown, for each target process node pair, according to the node distances between the second process node in the corresponding target process node pair and each matching second neighbor node, the target neighbor node pairs respectively corresponding to the second process node at each node distance are determined, including steps S502 - S504:
[0191] Step S502: Select the target process node pair with the largest similarity from each target process node pair.
[0192] Specifically, the computer device selects the target process node pair with the largest similarity from each target process node pair to generate a corresponding subgraph for the second process node in the selected target process node pair.
[0193] In one embodiment, the target process node pair with the largest chi - square value is selected from each target process node pair, and the chi - square value of the target process node pair is determined based on the similarity of the target process node pair.
[0194] Step S504: According to the node distances between the second process node in the selected target process node pair and the matching second neighbor nodes, determine the target neighbor node pairs respectively corresponding to the second process node in the selected target process node pair at each node distance.
[0195] Specifically, the computer device determines the second process node in the selected target process node pair, and determines the node distances between the second process node and the matching second neighbor nodes, so as to determine the second neighbor nodes respectively corresponding to the second process node at each node distance, and obtain the target neighbor node pairs formed by the second process node and the second neighbor nodes at each node distance, that is, obtain the target neighbor node pairs respectively corresponding to the second process node at each node distance.
[0196] For each pair of target neighbor nodes of the second process node at each node distance, select a second neighbor node from the pairs of target neighbor nodes corresponding to the respective node distance.
[0197] In one embodiment, for each pair of target neighbor nodes of the second process node at each node distance, select the second neighbor node from the pair of target neighbor nodes with the greatest similarity among the pairs of target neighbor nodes corresponding to the respective node distance.
[0198] In one embodiment, for each pair of target neighbor nodes of the second process node at each node distance, select the second neighbor node from the pair of target neighbor nodes with the largest chi-square value among the pairs of target neighbor nodes corresponding to the respective node distance.
[0199] Based on the association relationship between the selected second neighbor node at each node distance and the matching second process node in the event graph to be detected, generate a target subgraph corresponding to each second process node, including steps S506 - S508:
[0200] Step S506, based on the association relationship between the selected second process node and the selected second neighbor node at each node distance in the event graph to be detected, generate a subgraph corresponding to the second process node in the selected pair of target process nodes.
[0201] Specifically, after obtaining the second process node and the selected second neighbor node at each node distance, a subgraph corresponding to the second process node can be generated according to the association relationship between the second process node and each selected second neighbor node in the event graph to be detected. This target subgraph corresponds to the second process node in the same target process node. This association relationship can be represented by the edges between the second process node and each second neighbor node in the event graph to be detected, that is, by connecting the second process node and each second neighbor node with edges in the event graph to be detected to obtain a subgraph, which is the subgraph corresponding to the selected second process node.
[0202] In this embodiment, multiple nodes of this subgraph include a process node and at least one neighbor node, and the edges of this subgraph are used to describe abnormal events occurring between the corresponding nodes.
[0203] Step S508, check whether there are unselected target process nodes. If so, execute step S510; otherwise, end.
[0204] Step S510, select the pair of target process nodes with the greatest similarity from the unselected pairs of target process nodes, and return to step S504 to continue execution until stopping when generating subgraphs corresponding to the second process nodes in each pair of target process nodes, so as to obtain a target subgraph in the event graph to be detected that matches the query graph.
[0205] Specifically, the computer device selects the target process node pair with the greatest similarity from each of the unselected target process node pairs, and returns the step of determining the target neighbor node pairs corresponding to the second process node in each selected target process node pair according to the node distance between the second process node and the matching second neighbor node, and continues to execute, so as to generate the subgraph corresponding to the second process node selected this time. According to the same processing, it stops until the subgraphs corresponding to the second process nodes in each target process node pair are generated. The node set in this subgraph is a subset of the node set of the event graph to be detected, and the edge set in this subgraph is a subset of the edge set of the event graph to be detected. The node set in this subgraph is a subset of the node set of the query graph, and the edge set in this subgraph is a subset of the edge set of the query graph.
[0206] The computer device can use each subgraph as the target subgraph, or can screen out the target subgraph from each subgraph. For example, screen out a specified number of subgraphs from each subgraph as the target subgraph.
[0207] As Figure 6 shown, it is a schematic diagram of the query graph and the event graph to be detected in an embodiment. The query graph includes the first process node M1, and the first neighbor nodes m1, m2, and m3 whose node distance from the first process node M1 is 1, and the first neighbor node m4 whose node distance from the first process node M1 is 2.
[0208] The event graph to be detected includes the second process nodes N1, N2, and the second neighbor nodes n1, n2, and n3 whose node distance from the second process node N1 is 1, and the second neighbor nodes n4, n5 whose node distance from the second process node N1 is 2, and the second neighbor node n6 whose node distance from the second process node N1 is 3.
[0209] The first process node M1 and the second process node N1 form a target process node pair <M1, N1>.
[0210] <m1, n1>, <m2, n2>, and <m3, n3> are used as the target neighbor node pairs when the node distance is 1.
[0211] <m4, n4> is used as the target neighbor node pair when the node distance is 2.
[0212] For the second process node N1, select the target neighbor node pair with the largest similarity from the target neighbor node pairs <m1, n1>, <m2, n2>, and <m3, n3> when the node distance is 1. Connect the second neighbor node in the selected target neighbor node pair to the second process node N1 through the edge between them in the event graph to be detected. Then, select the target neighbor node pair with the largest similarity from the remaining 2 pairs of target neighbor node pairs corresponding to a node distance of 1. Connect the second neighbor node in the selected target neighbor node pair to the second process node N1 through the edge between them in the event graph to be detected. Then, select the target neighbor node pair with the largest similarity from the remaining 1 pair of target neighbor node pairs corresponding to a node distance of 1. Connect the second neighbor node in the selected target neighbor node pair to the second process node N1 through the edge between them in the event graph to be detected, thereby completing the matching of neighbor nodes with a node distance of 1 and obtaining the intermediate graph as shown by the dashed lines in Figure 7 the figure.
[0213] Alternatively, select the target neighbor node pairs with a similarity greater than the similarity threshold, and connect the second neighbor node in the selected target neighbor node pairs to the second process node N1 through the edge in the event graph to be detected, obtaining the intermediate graph as shown by the dashed lines in Figure 7 the figure.
[0214] When the target neighbor node pair with a node distance of 2 from the second process node N1 is only <m4, n4>, select the second neighbor node n4 from <m4, n4> and add the second neighbor node n4 to the intermediate graph to obtain a new intermediate graph, as shown by the dashed lines in Figure 8 the figure. Since there are no target neighbor node pairs when the node distance from the second process node N1 is 3, then Figure 8 the intermediate graph shown in the figure is the target subgraph corresponding to the second process node N1.
[0215] In this embodiment, the target process node pair with the largest similarity is selected from each pair of target process nodes, and the second process node in the selected target process node pair is used as the starting node, so as to sequentially match each neighbor node of the starting node at different node distances. According to the node distance between the second process node and the matched second neighbor node in the selected target process node pair, the target neighbor node pair corresponding to the second process node in the selected target process node pair at each node distance is determined. For each target neighbor node pair at each node distance, the second neighbor node in the target neighbor node pair that meets the association condition is selected from the target neighbor node pairs corresponding to the corresponding node distance. Based on the association relationship between the selected second process node and the selected second neighbor node at each node distance in the event graph to be detected, the second process node can be accurately matched with each neighbor node at different node distances, so as to accurately generate the corresponding subgraph. The subgraph includes multiple nodes and edges, and the edges of the subgraph represent the abnormal events occurring between the corresponding nodes, so that the abnormal events occurring between the selected second process node and the corresponding neighbor nodes can be described by the subgraph, effectively realizing the detection of abnormal events in the event graph to be detected.
[0216] The target process node pair with the largest similarity is selected from the unselected target process node pairs, and the step of determining the target neighbor node pair corresponding to the second process node in the selected target process node pair at each node distance according to the node distance between the second process node and the matched second neighbor node in the selected target process node pair is returned and continued to be executed until the generation of the subgraph corresponding to the second process node in each target process node pair stops, so as to obtain the target subgraph in the event graph to be detected that matches the query graph, so that based on the second process node, the abnormal events occurring between each second process node and the corresponding neighbor nodes can be detected more precisely, making the abnormal event detection more refined and the detection accuracy higher, thus effectively reducing the detection error rate.
[0217] In one embodiment, for each candidate node pair, according to the node representations of the nodes in the corresponding candidate node pair, the target node pairs that meet the matching conditions are screened out from each candidate node pair, including:
[0218] For each candidate node pair, the similarity between the node representations of the nodes in the corresponding candidate node pair is determined to obtain the similarity corresponding to each candidate node pair; based on each similarity, the target node pairs that meet the matching conditions are screened out from each candidate node pair.
[0219] Specifically, the computer device calculates the similarity between the node representations corresponding to each node in the candidate node pairs, and uses the calculated similarity as the similarity corresponding to the candidate node pair. By the same processing, the similarity corresponding to each candidate node pair can be obtained.
[0220] The computer device can obtain a matching condition, and match the similarity corresponding to each candidate node pair with the matching condition, so as to screen out the candidate node pairs that meet the matching condition from each candidate node pair as target node pairs.
[0221] In this embodiment, meeting the matching condition may specifically be that the similarity of the candidate node pair is greater than the similarity threshold. Then, the computer device can compare the similarity corresponding to each candidate node pair with the similarity threshold respectively, and use the candidate node pairs with similarity greater than the similarity threshold as target node pairs.
[0222] In this embodiment, meeting the matching condition may specifically be to select a preset number of candidate node pairs according to the similarity. The computer device can screen out a preset number of candidate node pairs with high similarity from each candidate node pair as target node pairs. For example, sort each candidate node pair according to the similarity, and select a preset number of candidate node pairs as target node pairs in the order from high to low similarity.
[0223] In this embodiment, the similarity may specifically be sine similarity, cosine similarity, etc., but is not limited thereto. The similarity threshold corresponds to the sine similarity threshold and the cosine similarity threshold. For example, the computer device can calculate the sine similarity between the node representations of each node in the candidate node pair <v, q>.
[0224] In this embodiment, for each candidate node pair, the similarity between the node representations of the nodes in the corresponding candidate node pair is determined, and the similarity corresponding to each candidate node pair is obtained. Thus, the similarity between the node representations of the nodes can be used as a screening condition to accurately screen out the target node pairs that meet the matching condition from each candidate node pair based on each similarity.
[0225] In one embodiment, based on each target node pair, subgraph matching is performed on the graph of the event to be detected to obtain a target subgraph in the graph of the event to be detected that matches the query graph, including: sequentially selecting target node pairs based on the similarity of each target node pair, and generating a corresponding target subgraph according to the nodes in the selected target node pairs that belong to the graph of the event to be detected and the association relationship of the selected nodes in the graph of the event to be detected.
[0226] In one embodiment, the node representation is obtained through a preprocessing step, and the preprocessing step includes:
[0227] Collect multiple nodes and edges from the event graph to be detected and the query graph to construct multiple triples. Each triple includes a node represented by an initialized node, a node represented by a node to be solved, and an edge represented by an initialized feature. Based on the initialized node representations, the node representations to be solved, and the initialized feature representations corresponding to the multiple triples, construct an objective loss function. Perform iterative solution based on the objective loss function until the iterative process stops, and obtain the node representations corresponding to each node respectively.
[0228] Specifically, the computer device can define an initialized node representation for each node in the event graph to be detected, and define an initialized feature representation for each edge. The computer device can also define an initialized node representation for each node in the query graph, and define an initialized feature representation for each edge. The computer device can collect multiple nodes and edges from the event graph to be detected and the query graph to construct multiple triples. Each triple consists of two nodes and the association relationship between the two nodes. The two nodes are connected through the association relationship, and the association relationship can be represented by the edge between the two nodes.
[0229] When constructing a triple, define the corresponding initialized node representation for any one of the two collected nodes, set the node representation of the other node as the node to be solved, that is, the node representation to be solved, and define the initialized feature representation for the association relationship between the two nodes, so that the triple consists of the node represented by the initialized node representation, the node represented by the node to be solved, and the edge represented by the initialized feature representation.
[0230] In one embodiment, the computer device can connect the event graph to be detected and the query graph to generate a sample graph, and collect multiple nodes and edges from the sample graph to form multiple triples. The constructed multiple triples can be divided into positive samples and negative samples. In the triple as a positive sample, the two nodes are adjacent; in the triple as a negative sample, the two nodes are not adjacent. That is, the computer device can collect two adjacent nodes from the sample graph, obtain the initialized node representation of any one of the two adjacent nodes, set the node representation of the other node as the node to be solved, so as to form a triple of positive samples with the two adjacent nodes and the association relationship between the two nodes. The computer device can collect two non - adjacent nodes from the sample graph, obtain the initialized node representation of any one of the two non - adjacent nodes, set the node representation of the other node as the node to be solved, and form a triple of negative samples with the two non - adjacent nodes and the association relationship between the two nodes.
[0231] In one embodiment, the computer device defines an initialized node representation for each node in the sample graph generated by connecting the event graph to be detected and the query graph, and defines an initialized feature representation for each edge in the sample graph.
[0232] The computer device can construct an objective loss function according to the initialization node representation, the node representation to be solved, and the initialization feature representation corresponding to each triple. The computer device iteratively solves the node representation of each node and the feature representation of the edge based on the objective loss function until the node representation corresponding to each node is obtained when the iteration stops. Further, the computer device iteratively solves the node representation of each node and the feature representation of the edge based on the objective loss function, updates the node representation corresponding to each node and each feature representation after each iterative solution, and performs the next iteration after updating the node representation and the feature representation until the node representation corresponding to each node and the feature representation corresponding to each edge are obtained when the iteration stops.
[0233] In one embodiment, the computer device can take the derivative of the objective loss function to obtain the gradient corresponding to the node of the node representation to be solved, and update the current node representation and the current feature representation in the way of gradient descent. After the update, continue to perform the next iterative solution through the objective loss function until it stops when the iteration stop condition is met, and obtain the node representation corresponding to the node of the node representation to be solved, as well as the node representations and feature representations corresponding to the other respective nodes.
[0234] In calculus, taking the partial derivative of the parameters of a multivariate function and writing out the partial derivatives of each obtained parameter in the form of a vector is the gradient vector, which can be simply referred to as the gradient.
[0235] The gradient is the place where the objective loss function changes the fastest. Along the direction of the gradient vector, the maximum value of the objective loss function can be solved. On the contrary, along the direction opposite to the gradient vector, the gradient decreases the fastest, and the minimum value of the objective loss function can be solved. Therefore, through the way of gradient descent, the loss value of the objective loss function can be minimized.
[0236] In this embodiment, the way of gradient descent can be any one of batch gradient descent (Batch Gradient Descent, abbreviated as BGD), stochastic gradient descent (Stochastic Gradient Descent, abbreviated as SGD), and mini-batch gradient descent (Mini-batch Gradient Descent, abbreviated as MBGD).
[0237] Batch gradient descent means using all samples to perform gradient solution and update, that is, all positive samples and all negative samples participate in each iteration. Using batch gradient descent to perform each iterative solution can determine the direction of the gradient through a large number of samples, so that the objective loss function can converge to the local optimal solution faster, and the obtained result is more accurate and reliable, that is, the node representation and feature representation determined by the iterative solution of a large number of samples are more accurate.
[0238] The principle of the stochastic gradient descent method is similar to that of the batch gradient descent method. However, when solving the gradient, the stochastic gradient descent method only selects one sample to participate, that is, only any one sample is selected for each iteration. Since the stochastic gradient descent method only uses one sample for iterative solution each time, the computational complexity of each iteration is small, and the iteration speed is very fast, which can more quickly determine the node representations corresponding to each node and the feature representations corresponding to each edge.
[0239] The mini-batch gradient descent method is a compromise between the batch gradient descent method and the stochastic gradient descent method, and it selects a preset number of samples from all samples to participate in each iteration. Selecting some samples to participate in the iterative solution each time can improve the processing efficiency while ensuring the accuracy of the results.
[0240] In this embodiment, the iteration stop condition can be to minimize the objective loss function, that is, to solve the minimum value of the objective loss function. The iteration stop condition can also be that the number of iterations reaches a preset number. By using the gradient descent method to solve the minimum value of the objective loss function, the node representations of each node and the feature representations of each edge corresponding to the minimum loss value can be determined, effectively eliminating the influence of inappropriate definitions of the node representations of each node on subgraph matching. Moreover, the nodes and topological structures in the event graph to be detected and the query graph can be transformed into computable and reusable representations, and the information contained in the heterogeneous graph can be fully utilized through the node representations, thereby effectively improving the accuracy of subgraph matching.
[0241] It can be understood that during the multiple iterative solution processes, the node representations and feature representations are updated in each iteration. Then, the initial node representation used in the next iteration is the node representation obtained by the update in the previous iteration, and the initial feature representation used in the next iteration is the feature representation obtained by the update in the previous iteration.
[0242] In this embodiment, multiple nodes and edges are collected from the event graph to be detected and the query graph to construct multiple triples. The triple includes a node with an initial node representation, a node whose node representation is to be solved, and an edge with an initial feature representation. By constructing the triples, the association relationship between a node and its directly adjacent nodes can be established, so as to predict the node representation of another node through a certain node and the association relationship in the triples. According to the initial node representations, the node representations to be solved, and the initial feature representations corresponding to the multiple triples, an objective loss function is constructed, so that the node representation to be solved can be used as the parameter to be solved, and the parameter to be solved is iteratively solved step by step based on the objective loss function, and the node representations and feature representations of each node are updated based on the results of each iterative solution until the final node representation of each node is accurately obtained when the iteration stops.
[0243] Moreover, in this embodiment, the node representation of another node is predicted through the node representation of a certain node in the triple and the feature representation corresponding to the association relationship. This is to predict the node representation of the node through the unsupervised graph pre-training technology, and more abundant topological structure information of the graph structure is combined during the prediction process, making the iteration more accurate. At the same time, it can effectively predict the node representation for all nodes of the graph structure without labels, and is applicable to the prediction of the node representation of all graph structures.
[0244] In one embodiment, the preprocessing step can be executed by a generative model. The computer device inputs the to-be-detected event graph and the query graph into the to-be-trained generative model. The to-be-trained generative model collects multiple nodes and edges from the to-be-detected event graph and the query graph to construct multiple triples. The triple includes a node for initializing the node representation, a node for solving the node representation to be solved, and an edge for initializing the feature representation. A target loss function is constructed according to the initialized node representation, the node representation to be solved, and the initialized feature representation corresponding to the multiple triples. The to-be-trained generative model is trained based on the target loss function, and the parameters of the generative model are adjusted during the training process until the iteration stop condition is met and then stopped. The trained generative model is obtained, and the node representations corresponding to each node output by the trained generative model are obtained.
[0245] In this embodiment, the iteration stop condition can be to minimize the target loss function. The gradient descent method can be used to perform iterative solution step by step to obtain the model parameters corresponding to the generative model when the target loss function is minimized, so as to obtain the node representations corresponding to each node in the query graph output by the generative model, and the node representations corresponding to each node in the to-be-detected event graph. Moreover, in this embodiment, the node representation of another node is predicted through the node representation of a certain node in the triple and the feature representation corresponding to the association relationship, which can use the method of unsupervised graph representation learning to predict the node representation of the node, avoiding the problem of large training errors caused by inaccurate labels. And it can also effectively predict the node representation for all nodes of the graph structure without labels, and is applicable to the prediction of the node representation of all graph structures. Unsupervised graph representation learning refers to the process of converting the original data into a form that can be directly used by machine learning through unsupervised learning methods. Unsupervised learning is a type of machine learning, and the model parameters are updated according to the training samples without pre-labeled labels.
[0246] In one embodiment, as Figure 9 shown, constructing the target loss function according to the initialized node representation, the node representation to be solved, and the initialized feature representation corresponding to the multiple triples includes:
[0247] Step S902: Construct a first loss function based on the initialization node representations, nodes to be solved representations, and initialization feature representations corresponding to multiple triples.
[0248] Specifically, the computer device can determine the product of the initialization node representation, the node to be solved representation, and the initialization feature representation corresponding to a single triple. In the same way, the computer device can determine the products corresponding to each triple respectively, and construct the first loss function according to the products corresponding to each triple respectively.
[0249] Furthermore, the computer device can determine the first ratio of the product corresponding to a single triple to a preset parameter, and determine the exponent of the first ratio. In the same way, the exponents of the first ratios corresponding to each triple can be obtained. The computer device sums up the exponents of the first ratios corresponding to each triple to obtain the sum of exponents, and constructs the first loss function according to the exponents of the first ratios and the sum of exponents. Furthermore, the computer device determines the ratio between the exponent of the first ratio and the sum of exponents to obtain a second ratio, and uses the function of taking the negative logarithm of the second ratio as the first loss function.
[0250] For example, the first loss function constructed by the computer device is as follows:
[0251]
[0252] where u and v are nodes in the same triple, and R is the association relationship between node u and node v. is the node representation corresponding to node u, and the node representation is the representation vector corresponding to the node; W R represents the feature representation of the association relationship R between two nodes in the triple, and this feature representation can be characterized by a representation matrix; h v is the node representation corresponding to node v, and h i is the node representation corresponding to the i-th node v; τ represents a parameter, and i represents the number of nodes v.
[0253] In one embodiment, during the construction process of the first loss function, the node representation of one node in the triple corresponding to the numerator of the first loss function is set to the node to be solved representation, and for the remaining triples in the denominator except the triple corresponding to the numerator, the node representation of each node in the remaining triples uses the initialization node representation.
[0254] Step S904: Determine the initialization context representations corresponding to each process node in the event graph to be detected and the query graph respectively.
[0255] Specifically, the computer device can determine each process node in the event graph to be detected and determine the corresponding context of each process node in the event graph to be detected. The computer device can pre-define the initialization context representation corresponding to the context of each process node.
[0256] The context of a process node consists of the previous neighbor node and the next neighbor node adjacent to the process node, the edge between the process node and the previous neighbor node, and the edge between the process node and the next neighbor node. The computer device can define the initialization context representation corresponding to the context of the process node according to the initialization node representations corresponding to the previous neighbor node and the next neighbor node respectively, the initialization feature representation corresponding to the edge between the process node and the previous neighbor node, and the initialization feature representation corresponding to the edge between the process node and the next neighbor node.
[0257] Furthermore, the computer device can also pre-define the initialization node representation corresponding to each process node respectively.
[0258] Step S906, construct a second loss function based on the node representation to be solved corresponding to each process node and the corresponding initialization context representation.
[0259] Specifically, the computer device can construct a second objective loss function according to the node representation to be solved corresponding to each process node and the initialization context representation corresponding to the context of each process node. The context is used to describe various property information of the process node. Furthermore, the computer device can also pre-define the initialization node representation corresponding to each process node respectively, and can set the node representations corresponding to some process nodes as the nodes to be solved according to the requirements of each iteration, so as to be used as the parameters to be solved in the iteration. The computer device constructs a second objective loss function according to some node representations to be solved, some initialization node representations, and each initialization context representation.
[0260] It can be understood that the samples used in the process of constructing the first loss function can be called the first samples, and the samples used in the process of constructing the second loss function can be called the second samples.
[0261] The computer device may use a process node and the context corresponding to the process node as a second sample. Based on each process node and its corresponding context, multiple second samples can be constructed. For each second sample, the node representation corresponding to the process node in the second sample can be an initialization node representation, or can be adjusted to a node representation to be solved in different iterations as the parameter to be solved in the iteration. For example, if the computer device only uses one second sample in an iterative solution, the node representation corresponding to the process node in the used second sample can be set as the node representation to be solved, and the node representations corresponding to the remaining second samples use the initialization node representation. If the computer device uses multiple second samples in an iterative solution, the node representations corresponding to the process nodes in the used multiple second samples are set as the node representations to be solved, and the node representations corresponding to the remaining second samples use the initialization node representation.
[0262] In one embodiment, the computer device determines the product of the node representation to be solved of the process node in a single second sample and the initialization context representation corresponding to the process node. According to the same process, the products corresponding to each second sample can be calculated. A second loss function is constructed based on the products corresponding to each second sample.
[0263] Further, the computer device may determine the third ratio of the product corresponding to a single second sample to a preset parameter, and determine the exponent of the third ratio. According to the same process, the exponents of the third ratios corresponding to each second sample can be obtained. The computer device sums the exponents of the third ratios corresponding to each second sample to obtain the sum of exponents, and constructs a second loss function based on the exponent of the third ratio and the sum of exponents. Further, the computer device determines the ratio between the exponent of the third ratio and the sum of exponents to obtain a fourth ratio, and uses the function representing the negative logarithm of the fourth ratio as the second loss function.
[0264] For example, taking a process node u as an example for description, in the event graph to be detected, the direct neighbor nodes of a process node are the resources directly manipulated by the process node. The computer device may determine the direct neighbor node v of the process node, and denote the set of direct neighbor nodes of the process node u as:
[0265] PC u ={v i |(u,v i )∈A}
[0266] The second loss function constructed by the computer device is as follows:
[0267]
[0268]
[0269]
[0270] Among them, u is the process node, and c u is the context representation corresponding to the process node, and v i represents the adjacent neighbor node of the process node u, that is, the direct neighbor node; is the node representation corresponding to the i-th neighbor node v adjacent to the process node u. c i is the context representation corresponding to the i-th process node. By performing high-level feature extraction on the neighbor nodes involved in the context, and then through aggregating the high-level features, the corresponding context representation c u is obtained.
[0271] In one embodiment, during the construction of the second loss function, the node representation of the process node in the second sample corresponding to the numerator of the second loss function is set to the node representation to be solved, and in the denominator, except for the second sample corresponding to the numerator, the node representations of the remaining second samples in the denominator use the initialized node representation.
[0272] It can be understood that during the multiple iterative solution processes, the node representation and the context representation are updated in each iteration. Then, the initialized node representation used in the next iteration is the node representation obtained by the update in the previous iteration, and the initialized context node representation used in the next iteration is the context representation obtained by the update in the previous iteration.
[0273] Step S908, constructing a target loss function according to the first loss function and the second loss function.
[0274] Specifically, the computer device can sum the first loss function and the second loss function to obtain the target loss function.
[0275] In one embodiment, the computer device can obtain the weights corresponding to the first loss function and the second loss function respectively, multiply the first loss function and the second loss function by their respective weights respectively, and then sum them to obtain the target loss function.
[0276] The computer device iteratively solves the node representations of each node, the feature representations of the edges, the node representations of the process nodes, and the context representations of the process nodes based on the target loss function until the node representations corresponding to each node are obtained when the iteration stops. Further, the computer device iteratively solves the node representations of each node, the feature representations of the edges, the node representations of the process nodes, and the context representations of the process nodes based on the target loss function, and updates the node representations corresponding to each node, each feature representation, the node representations of each process node, and each context representation after each iterative solution, and then performs the next iteration after the update until the node representations corresponding to each node, the feature representations corresponding to each edge, the node representations of each process node, and the context representations corresponding to each process node are obtained when the iteration stops.
[0277] It can be understood that the nodes include process nodes. When a node is a process node, the node representation of this node is the node representation of the process node.
[0278] In this embodiment, a triple is composed of two nodes and the association relationship between these two nodes. By constructing the first loss function through the triple, the self-supervised task at the node level can be realized, that is, the prediction of the node representation of a single node can be realized. The process nodes in the graph structure are used as the event subjects, and to a large extent, the properties of the neighbor nodes related to them can be inferred. Then, by constructing the second loss function through the process nodes and the context representations of the process nodes, richer topological structure information is used, and the self-supervised task at the regional level can be realized, that is, the neighborhood matching prediction can be realized. By performing pre-training through the self-supervised task at the node level and the neighborhood matching prediction task, the node representations can be iteratively solved from the node level and the neighborhood level of the nodes, so that the result of the iterative solution not only considers the characteristics of the nodes themselves, but also considers the characteristics of the neighborhoods of the nodes, making the determined node representations more accurate. Moreover, both the self-supervised task at the node level and the self-supervised task at the regional level are unsupervised learning tasks, and the prediction of node representations can be performed without using labels, with a wider application range and higher processing efficiency.
[0279] In one embodiment, an abnormal event detection method is provided, which is applied to a computer device and includes:
[0280] Obtain a graph of events to be detected and a query graph; both the graph of events to be detected and the query graph include multiple nodes and edges. The multiple nodes of the query graph include multiple first process nodes and the first neighbor nodes of each first process node. The edges of the graph of events to be detected represent the events occurring between the corresponding nodes; the multiple nodes of the graph of events to be detected include multiple second process nodes and the second neighbor nodes of each second process node. The edges of the query graph represent the abnormal events occurring between the corresponding nodes.
[0281] Collect multiple nodes and edges from the event graph to be detected and the query graph to construct multiple triples, where a triple includes a node represented by an initialization node, a node represented by a node to be solved, and an edge represented by an initialization feature; construct a first loss function according to the initialization node representation, the node representation to be solved, and the initialization feature representation corresponding to the multiple triples.
[0282] Determine the initialization context representation corresponding to each process node in the event graph to be detected and the query graph; construct a second loss function based on the node representation to be solved and the corresponding initialization context representation corresponding to each process node.
[0283] Construct an objective loss function according to the first loss function and the second loss function; perform iterative solution based on the objective loss function until the node representations corresponding to each node in the event graph to be detected and the node representations corresponding to each node in the query graph are obtained when the iteration stops.
[0284] Extract each first process node from the query graph and extract each second process node from the event graph to be detected; match each first process node with each second process node to obtain candidate process node pairs formed by each first process node and each second process node.
[0285] For each candidate process node pair, determine the similarity between the node representation of the first process node and the node representation of the second process node in the corresponding candidate process node pair to obtain the similarity corresponding to each candidate process node pair.
[0286] For each first process node, based on the similarities corresponding to the candidate process node pairs to which the corresponding first process node belongs, screen out the target process node pairs that meet the matching conditions from the candidate process node pairs to which the corresponding first process node belongs.
[0287] For each target process node pair, determine the first neighbor nodes corresponding to the first process node in each node distance and the second neighbor nodes corresponding to the second process node in each node distance;
[0288] Determine the similarity between each first neighbor node in each node distance and each second neighbor node in the corresponding node distance;
[0289] For each node distance, according to the similarities between each first neighbor node and each second neighbor node in the corresponding node distance, determine the second neighbor nodes that match each first neighbor node to obtain, respectively, target neighbor node pairs formed by each first neighbor node and the matching second neighbor nodes in each node distance.
[0290] Select the pair of target process nodes with the maximum similarity from each pair of target process nodes, and determine the corresponding pair of target neighbor nodes for the selected pair of target process nodes at each node distance;
[0291] For each pair of target neighbor nodes at each node distance, select the second neighbor node from the corresponding pairs of target neighbor nodes corresponding to the respective node distances; based on the association relationship between the selected second process node and the second neighbor node selected at each node distance in the event graph to be detected, generate the subgraph corresponding to the second process node in the selected pair of target process nodes.
[0292] Select the pair of target process nodes with the maximum similarity from the unselected pairs of target process nodes, and return to the step of determining the corresponding pair of target neighbor nodes for the selected pair of target process nodes at each node distance and continue to execute until the subgraph corresponding to the second process node in each pair of target process nodes is generated and then stop, so as to obtain the target subgraph in the event graph to be detected that matches the query graph.
[0293] In this embodiment, a triple is composed of two nodes and the association relationship between these two nodes. By constructing the first loss function through the triple, the node-level self-supervised task can be realized, that is, the prediction of the node representation of a single node can be achieved. The process nodes in the graph structure are used as the event subjects, and to a large extent, the properties of the neighbor nodes related to them can be inferred. Then, by constructing the second loss function through the process nodes and the context representation of the process nodes, more abundant topological structure information is used, and the regional-level self-supervised task can be realized, that is, the neighborhood matching prediction can be achieved. Through the node-level self-supervised task and the neighborhood matching prediction task for pre-training, the representation of the nodes can be iteratively solved from the node level and the neighborhood level of the nodes, so that the result of the iterative solution not only takes into account the characteristics of the nodes themselves, but also takes into account the characteristics of the neighborhoods of the nodes, making the determined node representation more accurate. Moreover, both the node-level self-supervised task and the regional-level self-supervised task belong to unsupervised learning tasks, and the prediction of the node representation can be carried out without using labels, with a wider application range and higher processing efficiency.
[0294] Match the first process node in the query graph with the second process node in the event graph to be detected, and the rough position of the process node in the event graph to be detected that may match the first process node can be determined through the preliminary match. First, preliminarily match the process nodes of the query graph and the process nodes of the event graph to be detected, and based on the similarity between the node representations of the process nodes in each candidate process node pair, perform fine matching on the multiple candidate process node pairs generated by the preliminary match to screen out the target process node pairs that meet the matching conditions.
[0295] Perform matching processing on the first neighbor nodes of the first process node and the second neighbor nodes of the second process node in the target process node pair to obtain target neighbor node pairs formed by each first neighbor node and the matching second neighbor node, so that the matching of neighbor nodes can be performed after the screening of the target process node pair is completed, which can reduce the amount of data to be matched and improve the processing efficiency. The target process node pair is composed of the first process node and the second process node with a relatively high matching degree, and the target neighbor node pair is composed of the first neighbor node and the second neighbor node with a relatively high matching degree.
[0296] Select the target process node pair with the largest similarity from each target process node pair, and use the second process node in the selected target process node pair as the starting node to sequentially match each neighbor node at different node distances for this starting node.
[0297] For each target neighbor node pair at each node distance, select the second neighbor node in the target neighbor node pair that meets the association condition from each target neighbor node pair corresponding to the corresponding node distance, which can select the neighbor node most closely associated with the second process node at each node distance. Thus, based on the association relationship between the second neighbor node selected at each node distance and the matching second process node in the event graph to be detected, the second process node can be accurately matched with each neighbor node at different node distances, and the corresponding subgraph can be accurately generated. The subgraph includes multiple nodes and edges, and the edges of the subgraph represent the abnormal events occurring between the corresponding nodes. Thus, the abnormal events occurring between the selected second process node and the corresponding neighbor nodes can be described through the subgraph, effectively realizing the detection of abnormal events in the event graph to be detected.
[0298] Select the target process node pair with the largest similarity from the unselected target process node pairs, and return to the step of determining the target neighbor node pairs corresponding to the selected target process node pair at each node distance and continue to execute until the generation of the subgraph corresponding to the second process node in each target process node pair stops, so as to obtain the target subgraph matching the query graph in the event graph to be detected. Thus, based on the second process node, the abnormal events occurring between each second process node and the corresponding neighbor nodes can be detected more precisely, making the abnormal event detection more refined and the detection accuracy higher, thereby effectively reducing the detection error rate.
[0299] As Figure 10 shown, it is an application scenario of the abnormal event detection method in an embodiment. Specifically, the application scenario is to detect the advanced persistent threat (APT) existing in the system through the abnormal event detection method. An advanced persistent threat refers to a stealthy and persistent computer intrusion process, usually carefully planned, targeting specific targets, and maintaining high concealment for a long time, which is an intrusion with a relatively high degree of complexity.
[0300] 1. Caching and Indexing of Relevant Information in System Event Graphs
[0301] The computer device abstracts a directed heterogeneous graph from the system log to describe the timing sequence of system events and the relationships between system events, which is the system event graph. It also abstracts a directed heterogeneous graph from the digital threat report to describe the timing sequence of system events related to APT and the relationships between them, which is the query graph.
[0302] For the query graph and the system event graph, the following data can be extracted and temporarily stored respectively to improve the matching efficiency:
[0303] Extract the nodes in the query graph and the system event graph, as well as the list of entity labels corresponding to the nodes; the list of entity labels refers to the type labels of the nodes, including file types, process types, tube process types, etc.
[0304] For the process node q in the query graph, aggregate the neighbor nodes with a node distance of 1 from each process node into the neighbor set and aggregate the neighbor nodes with a node distance of 2 from each process node into the neighbor set N 2 (q) to obtain the neighbor set with a node distance of m from each process node The process nodes in the query graph are called first process nodes, and the neighbor set can be a neighbor list.
[0305] For the process node v in the system event graph, aggregate the neighbor nodes with a node distance of 1 from each process node into the neighbor set and aggregate the neighbor nodes with a node distance of 2 from each process node into the neighbor set to obtain the neighbor set with a node distance of m from each process node The process nodes in the query graph are called second process nodes.
[0306] Determine the list of labels for each neighbor node. The list of labels for neighbor nodes includes file types, network address types, tube process types, etc.
[0307] 2. Entity Feature Pre-training
[0308] The self-supervised pre-training of the heterogeneous graph in this embodiment is based on the idea of contrastive learning. The purpose is to use the large-scale system event graph to learn an unsupervised generation model without supervision. This generation model can be a graph neural network GNN model. This generation model can be directly used to generate the node representations corresponding to each node in the system event graph and the query graph, or to sort and adjust multiple subsequent generated subgraphs.
[0309] The present embodiment defines unsupervised tasks for pre-training, including node-level self-supervised tasks and region-level self-supervised tasks. Moreover, a contrastive learning approach is used to contrast positive and negative samples in the feature space for representation learning.
[0310] (1) Node-level self-supervised task - node representation prediction: The system event graph and the query graph can be connected to form a heterogeneous graph G, and this heterogeneous graph G serves as the sample graph.
[0311] In the heterogeneous graph G, an observed <u, R, v> triple represents that nodes u ∈ V and v ∈ V are connected by a specific association relationship R. The node adjacency matrix corresponding to the heterogeneous graph G is denoted as A. The node-level self-supervised task is to predict the node representation of another node based on a certain node u (or v) and the association relationship R in the triple, that is: <u, R,?> or <?, R, v>. <u, R,?> means that the node representation of node u is known, the feature representation corresponding to the association relationship R is known, and the node representation of node v is to be solved, that is,? refers to the node representation to be solved. <?, R, v> means that the feature representation corresponding to the association relationship R is known, the node representation of node v is known, and the node representation of node u is to be solved, that is,? refers to the node representation to be solved.
[0312] The actually observed triples in the heterogeneous graph G can be used as positive samples, and negative samples are sampled from the following set:
[0313]
[0314] where, v - is a node that is not adjacent to node u in the heterogeneous graph G, indicates that v has the same type label as v - . sim(v, v - ) < δ limits the node distance between v and v - to ensure the discrimination difficulty between negative and positive samples.
[0315] Construct the following first loss function for the node-level self-supervised task
[0316]
[0317] This pre-training task only models the association relationship between a node and its directly adjacent nodes.
[0318] (2) Regional-level self-supervised task - Neighborhood matching prediction of process nodes: Compared with general heterogeneous graphs, the remarkable feature of the system event graph is that process nodes, as the main bodies of events, can largely be used to infer the properties of related resource nodes (files, network addresses, monitors, etc.). Therefore, pre-training can be carried out through the neighborhood matching prediction task.
[0319] Here, taking a process node u as an example for description, in the system event graph, the direct neighbor nodes of a process node are the resources directly manipulated by the process node. Then, the set of direct neighbor nodes of the process node u is denoted as:
[0320] PC u ={v i |(u,v i )∈A}
[0321] In this way, and its related topological structure and labels constitute the context of the process node u. The context describes various property information of the process node u. This regional-level self-supervised task is to predict the node representation of the process node corresponding to the context based on the context. Similar to the node-level self-supervised task, positive and negative samples also need to be constructed for contrastive learning. The process nodes actually observed from the heterogeneous graph G and the context corresponding to the process nodes are used as positive samples, and the negative samples are:
[0322]
[0323] Here, u - is an arbitrary process node with the same type label as u in the heterogeneous graph G, indicating that x u has the same type label as . For the regional-level self-supervised task, the following second loss function is constructed
[0324]
[0325] where c u is the context representation of the process node u. The calculation of c u is as follows:
[0326] Perform high-level feature extraction on the nodes involved in the context;
[0327] Aggregate the high-level features.
[0328] The first loss function, the second loss function, and their respective weights are weighted and summed to obtain the target loss function. Based on the target loss function, multiple iterative solutions are performed. When the iterative solution stops, a trained generation model is obtained, and the node representations corresponding to each node in the system event graph and the node representations corresponding to each node in the query graph are obtained when the iterative solution stops.
[0329] III. Preliminary Matching: In a computer system scenario, a process is usually the subject of a system event. And the type label of an entity (such as file, process, monitor, etc.) is one of the most important pieces of information for describing an entity. Therefore, each process node in the query graph can be preliminarily matched in the system event graph according to the type label to obtain the matching candidate set VP of the process nodes in the query graph in the system event graph:
[0330]
[0331] Here, v i , q i respectively represent the process nodes in the system event graph and the query graph. means that v i , q i have the same type label.
[0332] In the preliminary matching process, based on the process nodes in the query graph, the rough positions where matches may exist in the system event graph can be found. Then, the preliminary matching results in the matching candidate set VP need to be further sorted to determine the most likely matching node pairs. Traditional subgraph matching methods generally compare the similarity based on the label information corresponding to the neighbor nodes of the nodes, ignoring the richer topological structure information, and the intermediate results are difficult to reuse. In this embodiment, the node representations corresponding to each node are obtained through the generation model, and the similarity between the nodes is calculated through the node representations, that is, entering the fine matching part.
[0333] IV. Fine Matching of Process Nodes (Sensitivity Analysis)
[0334] Through the generation model, the features of the nodes in the system event graph and the query graph can be extracted to obtain the node representations of each process node and the node representations of the neighbor nodes of each process node.
[0335] Similarity scoring between process nodes: For the candidate process node pairs <v ∈ G, q ∈ Q> in the matching candidate set VP, the similarity score sim(v, q) between the node representations corresponding to the process nodes in the candidate process node pairs can be calculated to screen out multiple process node pairs with high similarity scores to obtain a high-quality node matching subset VP * , the node matching subset Then, VP *The process nodes in the system event graph matched are extended to the neighborhood of the process nodes, that is, extended to the neighbor nodes of the process nodes, to complete the subgraph matching between the query graph and the system event graph.
[0336] V. Neighborhood Matching
[0337] The similarity score between the neighbor nodes of process nodes v and q, that is, calculate the neighbor sets and neighbor set The similarity score sim(v' i , q' i ) of each pair of neighbor nodes within. The pair of process nodes composed of process node v and process node q belongs to the node matching subset VP * . represents the set of neighbor nodes at a distance of 1 from process node v, represents the set of neighbor nodes at a distance of 1 from process node q.
[0338]
[0339] When the similarity between the neighbor nodes at a distance of 1 from process node v (or q) is completed, the and The nodes with the best matching degree in are paired and recorded in , and the involved neighbor nodes are marked as matched. For example, calculate the similarity between neighbor node and neighbor node , the similarity between neighbor node and neighbor node , the similarity between neighbor node and neighbor node to screen out the neighbor node with the best matching degree with neighbor node from neighbor nodes , and write the screened neighbor node and neighbor node as the target neighbor node pair into
[0340] Next, the similarity calculation is extended to the neighbor nodes at a distance of 2, 3... from process node v (or q) until all nodes in the query graph are marked as matched, that is, the and The nodes with the best matching degree in are paired and recorded in ... until
[0341] VI. Generate Fine Matching, that is, Subgraph Matching
[0342] Store all candidate process node pairs in VP and their χ2 (i.e., similarity) in the form of <v, q, χ2<v, q>> into the max-heap max-heap-1. Put all node pairs involved in 2 and their χ
[0343] values into the max-heap max-heap-(m + 1). Subsequently, repeat the following process until K subgraph matches are completed: First, select the process node pair <v, q> with the maximum similarity from VP, and then select the list of multiple neighbor nodes corresponding to the process node v in the process node pair <v, q> with the maximum similarity Start from the neighbor set with a node distance of 1, and select the neighbor node with the maximum similarity from to connect with the process node v, and then select the neighbor node with the maximum similarity from to connect... Select the neighbor node with the maximum similarity from
[0344] to connect until max-heap-(m + 1) is empty, and obtain the subgraph corresponding to the process node v. After the generation of the subgraph corresponding to this process node is completed, enter the generation of the next subgraph, that is, select the <v, q> with the maximum similarity in VP at this time... Continue to execute according to the above processing until all nodes in the query graph are in the subgraph, that is, until the subgraphs corresponding to each process node are generated and then stop, and obtain the subgraphs corresponding to each process node respectively. In this embodiment, under the unsupervised condition, using the graph pre-training method, the labels and topological structures of the nodes in the system behavior graph are transformed into computable and reusable representations, thereby avoiding the repeated use of high-complexity graph algorithms during the matching process and improving the real-time performance of the matching method. At the same time, the pre-training method can make full use of the information contained in the heterogeneous graph to improve the matching accuracy. For the queried subgraph, a corresponding relationship between the subgraph nodes and the query graph nodes is also established through an unsupervised node matching method.
[0345] The abnormal event detection method provided in this embodiment has good real-time performance, accuracy, and can conveniently integrate additional information and domain knowledge. This method can also be widely applied to fields such as financial risk control and user spamming behavior detection, and has generality and flexibility.
[0346] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0347] Based on the same inventive concept, an embodiment of the present application also provides an abnormal event detection device for implementing the abnormal event detection method described above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the abnormal event detection device provided below can refer to the limitations on the abnormal event detection method in the above text, and will not be repeated here.
[0348] In one embodiment, as Figure 11 shown, an abnormal event detection device 1100 is provided, including: an acquisition module 1102, a node matching module 1104, an extraction module 1106, a screening module 1108, and a sub-graph matching module 1110, where:
[0349] The acquisition module 1102 is configured to acquire a graph of events to be detected and a query graph; both the graph of events to be detected and the query graph include multiple nodes and edges. The edges of the graph of events to be detected represent the events that occur between the corresponding nodes, and the edges of the query graph represent the abnormal events that occur between the corresponding nodes.
[0350] The node matching module 1104 is configured to perform matching processing on the nodes of the query graph and the nodes of the graph of events to be detected to obtain multiple candidate node pairs.
[0351] The extraction module 1106 is configured to perform feature extraction on each node in multiple candidate node pairs to obtain a node representation corresponding to each node.
[0352] The screening module 1108 is configured to, for each candidate node pair, screen out the target node pairs that meet the matching conditions from each candidate node pair according to the node representations of the nodes in the corresponding candidate node pair.
[0353] The sub-graph matching module 1110 is configured to perform sub-graph matching on the graph of events to be detected based on each target node pair to obtain a target sub-graph in the graph of events to be detected that matches the query graph; this target sub-graph is used to describe the abnormal events existing in the graph of events to be detected.
[0354] In this embodiment, the query graph includes multiple nodes and edges. The edges of the query graph represent abnormal events occurring between the corresponding nodes. Then, the query graph can be used to describe the abnormal events occurring between the nodes. The event graph to be detected includes multiple nodes and edges. The edges of the event graph to be detected represent the events occurring between the corresponding nodes. By matching the query graph with the event graph to be detected and detecting whether there is a subgraph in the event graph to be detected that matches the query graph, it can be accurately determined whether there is an abnormal event in the event graph to be detected.
[0355] The matching of the query graph and the event graph to be detected can be refined to node matching. The nodes of the query graph and the nodes of the event graph to be detected are matched to preliminarily match the nodes in these two graphs, and the candidate node pairs for matching can be roughly determined. Feature extraction is performed on each node in multiple candidate node pairs to obtain the node representation corresponding to each node. The node representation includes the key information of the node and the relevant topological structure information of the node in the graph. For each candidate node pair, the target node pair that meets the matching conditions is selected from the candidate node pairs according to the node representations of the nodes in the corresponding candidate node pair. By using the node representation as the condition for screening the target node pair, the information contained in the graph can be fully utilized to further perform fine matching on the nodes, so as to improve the accuracy of screening. Based on each target node pair, subgraph matching is performed on the event graph to be detected to determine the target subgraph in the event graph to be detected that matches the query graph, so that the abnormal events existing in the event graph to be detected and the nodes where the abnormal events occur can be accurately described by the target subgraph.
[0356] In one embodiment, the multiple nodes of the query graph include multiple first process nodes, the multiple nodes of the event graph to be detected include multiple second process nodes, and the candidate node pairs include the candidate process node pairs obtained by matching the first process nodes and the second process nodes.
[0357] In this embodiment, the process node serves as the subject of the event. By matching the first process node in the query graph with the second process node in the graph of events to be detected, it is possible to preliminarily determine the approximate location of the process node in the graph of events to be detected that may match the first process node, facilitating subsequent further fine matching of the process nodes. For each pair of candidate process nodes, based on the node representations of the process nodes in the corresponding pair of candidate process nodes, the target process node pairs that meet the matching conditions are screened out from each pair of candidate process nodes. By using the node representation as the condition for screening the target process node pairs, the information contained in the graph can be fully utilized to further perform fine matching on the process nodes, thereby improving the accuracy of screening. Based on each target process node pair, subgraph matching is performed on the graph of events to be detected to determine the target subgraph in the graph of events to be detected that matches the query graph, so that the abnormal events existing in the graph of events to be detected and the nodes where the abnormal events occur can be accurately described by the target subgraph.
[0358] In one embodiment, the multiple nodes of the query graph include the first neighbor nodes of each first process node, the multiple nodes of the graph of events to be detected include the second neighbor nodes of each second process node, and the target node pairs include target process node pairs and target neighbor node pairs; the screening module 1108 is further configured to, for each target process node pair, perform matching processing on the first neighbor nodes of the first process node in the corresponding target process node pair and the second neighbor nodes of the second process node to obtain the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node; the subgraph matching module is further configured to perform subgraph matching on the graph of events to be detected based on each target process node pair and each target neighbor node pair to obtain the target subgraph in the graph of events to be detected that matches the query graph.
[0359] In this embodiment, the multiple nodes of the query graph include the first process node and the first neighbor nodes of each first process node, and the multiple nodes of the graph of events to be detected include the second process node and the second neighbor nodes of each second process node. For each target process node pair, perform matching processing on the first neighbor nodes of the first process node in the corresponding target process node pair and the second neighbor nodes of the second process node to obtain the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node; perform subgraph matching on the graph of events to be detected based on each target process node pair and each target neighbor node pair to obtain the target subgraph in the graph of events to be detected that matches the query graph. First, perform preliminary matching on the process nodes of the query graph and the process nodes of the graph of events to be detected, and based on the node representations of the process nodes in each pair of candidate process nodes, perform fine matching on the multiple pairs of candidate process nodes generated by the preliminary matching to screen out the target process node pairs that meet the matching conditions.
[0360] Perform matching processing on the first neighbor node of the first process node and the second neighbor node of the second process node in the target process node pair, so as to obtain a target neighbor node pair formed by each first neighbor node and the matching second neighbor node. Enabling neighbor node matching after completing the screening of the target process node pair can reduce the amount of data to be matched and improve processing efficiency. The target process node pair is composed of a first process node and a second process node with a relatively high matching degree, and the target neighbor node pair is composed of a first neighbor node and a second neighbor node with a relatively high matching degree. Then, based on each target process node pair and each target neighbor node pair with a relatively high matching degree, perform subgraph matching on the event graph to be detected, and the target subgraph matching the query graph can be more accurately determined from the event graph to be detected.
[0361] In one embodiment, the screening module 1108 is further configured to, for each candidate process node pair, determine the similarity between the node representation of the first process node and the node representation of the second process node in the corresponding candidate process node pair, so as to obtain the similarity corresponding to each candidate process node pair; for each first process node, based on the similarities corresponding to the respective candidate process node pairs to which the corresponding first process node belongs, screen out the target process node pairs that meet the matching conditions from the respective candidate process node pairs to which the corresponding first process node belongs.
[0362] In this embodiment, for each candidate process node pair, determine the similarity between the node representation of the first process node and the node representation of the second process node in the corresponding candidate process node pair, so as to obtain the similarity corresponding to each candidate process node pair, so that the similarity between the node representations of the nodes can be used as a screening condition to further screen the candidate process node pairs. For each first process node, based on the similarities corresponding to the respective candidate process node pairs to which the corresponding first process node belongs, accurately screen out the target process node pairs that meet the matching conditions from the respective candidate process node pairs to which the corresponding first process node belongs.
[0363] In one embodiment, the screening module 1108 is further configured to, for each target process node pair, determine the similarity between each first neighbor node of the first process node and each second neighbor node of the second process node in the corresponding target process node pair; for each first neighbor node, determine the second neighbor node that matches the corresponding first neighbor node according to the similarities between the corresponding first neighbor node and each second neighbor node, so as to obtain a target neighbor node pair formed by each first neighbor node and the matching second neighbor node respectively.
[0364] In this embodiment, for each pair of target process nodes, by determining the similarity between each first neighbor node of the first process node in the corresponding pair of target process nodes and each second neighbor node of the second process node, the most matching second neighbor node can be selected for each first neighbor node based on the similarity. For each first neighbor node, according to the similarity between the corresponding first neighbor node and each second neighbor node, the second neighbor node that matches the corresponding first neighbor node is determined, so as to obtain the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node respectively. In this way, the matching degree between the neighbor nodes in the generated target neighbor node pairs is high, that is, the matching degree of the neighbor node pairs used for subgraph matching is high, which can effectively improve the accuracy of subsequent subgraph matching.
[0365] In one embodiment, the screening module 1108 is further configured to, for each pair of target process nodes, determine the first neighbor nodes corresponding to the first process node in each node distance in the corresponding pair of target process nodes, and the second neighbor nodes corresponding to the second process node in each node distance; determine the similarity between each first neighbor node in each node distance and each second neighbor node in the corresponding node distance; for each node distance, according to the similarity between each first neighbor node in the corresponding node distance and each second neighbor node, determine the second neighbor node that matches each first neighbor node, so as to obtain the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node in each node distance respectively.
[0366] In this embodiment, for each pair of target process nodes, by determining the first neighbor nodes corresponding to the first process node in each node distance in the corresponding pair of target process nodes, and the second neighbor nodes corresponding to the second process node in each node distance, the neighbor nodes of the process nodes can be divided according to different node distances. By determining the similarity between each first neighbor node in each node distance and each second neighbor node in the corresponding node distance, the similarity between neighbor nodes is also divided according to the node distance. For each node distance, according to the similarity between each first neighbor node in the corresponding node distance and each second neighbor node, the second neighbor node that matches each first neighbor node is determined, so as to obtain the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node in each node distance respectively, so that the matching first neighbor node and second neighbor node in the same node distance can be selected, and the matching degree of the formed target neighbor node pairs is higher, which helps to improve the accuracy of subsequent subgraph matching.
[0367] In one embodiment, the sub - graph matching module 1110 is further configured to, for each pair of target process nodes, determine, according to the node distances between the second process node in the corresponding pair of target process nodes and each of the matched second neighbor nodes, the target neighbor node pairs respectively corresponding to the second process node at each node distance; for each of the target neighbor node pairs at each node distance, select the second neighbor nodes in the target neighbor node pairs that meet the association condition from among the target neighbor node pairs corresponding to the corresponding node distance; and generate, based on the association relationships in the event graph to be detected between the second neighbor nodes selected at each node distance and the matched second process nodes, the target sub - graphs respectively corresponding to each second process node.
[0368] In this embodiment, for each pair of target process nodes, determining, according to the node distances between the second process node in the corresponding pair of target process nodes and each of the matched second neighbor nodes, the target neighbor node pairs respectively corresponding to the second process node at each node distance, and for each of the target neighbor node pairs at each node distance, selecting the second neighbor nodes in the target neighbor node pairs that meet the association condition from among the target neighbor node pairs corresponding to the corresponding node distance can select, at each node distance, the neighbor nodes most closely associated with the second process node. Thus, based on the association relationships in the event graph to be detected between the second neighbor nodes selected at each node distance and the matched second process nodes, the target sub - graphs respectively corresponding to each second process node can be accurately constructed.
[0369] In one embodiment, the sub - graph matching module 1110 is further configured to select the pair of target process nodes with the maximum similarity from among the pairs of target process nodes; determine, according to the node distance between the second process node and the matched second neighbor node in the selected pair of target process nodes, the target neighbor node pairs respectively corresponding to the second process node in the selected pair of target process nodes at each node distance; generate, based on the association relationships in the event graph to be detected between the selected second process node and the second neighbor nodes selected at each node distance, the sub - graph corresponding to the second process node in the selected pair of target process nodes; select the pair of target process nodes with the maximum similarity from among the unselected pairs of target process nodes, and return and continue to execute the step of determining, according to the node distance between the second process node and the matched second neighbor node in the selected pair of target process nodes, the target neighbor node pairs respectively corresponding to the second process node in the selected pair of target process nodes at each node distance, until the generation of the sub - graphs corresponding to the second process nodes in each pair of target process nodes stops, so as to obtain the target sub - graphs in the event graph to be detected that match the query graph.
[0370] In this embodiment, the target process node pair with the maximum similarity is selected from each pair of target process nodes, and the second process node in the selected target process node pair is used as the starting node, so as to sequentially match each neighbor node of the starting node at different node distances. According to the node distance between the second process node and the matched second neighbor node in the selected target process node pair, the target neighbor node pairs corresponding to the second process node in the selected target process node pair at each node distance are determined. For each target neighbor node pair at each node distance, the second neighbor node in the target neighbor node pair that meets the association condition is selected from the target neighbor node pairs corresponding to the corresponding node distance. Based on the association relationship between the selected second process node and the selected second neighbor node at each node distance in the event graph to be detected, the second process node can be accurately matched with each neighbor node at different node distances, so as to accurately generate the corresponding subgraph. The subgraph includes multiple nodes and edges, and the edges of the subgraph represent the abnormal events occurring between the corresponding nodes, so that the abnormal events occurring between the selected second process node and the corresponding neighbor nodes can be described by the subgraph, effectively realizing the detection of abnormal events in the event graph to be detected.
[0371] Select the target process node pair with the maximum similarity from the unselected target process node pairs, and return the step of determining the target neighbor node pairs corresponding to the second process node in the selected target process node pair at each node distance according to the node distance between the second process node and the matched second neighbor node in the selected target process node pair, and continue to execute until the subgraph corresponding to the second process node in each target process node pair is generated and then stop, so as to obtain the target subgraph in the event graph to be detected that matches the query graph, so that based on the second process node, the abnormal events occurring between each second process node and the corresponding neighbor nodes can be detected more precisely, making the abnormal event detection more refined and the detection accuracy higher, thus effectively reducing the detection error rate.
[0372] In one embodiment, the screening module 1108 is further configured to determine the similarity between the node representations of the nodes in each candidate node pair for each candidate node pair, and obtain the similarity corresponding to each candidate node pair; and screen out the target node pairs that meet the matching conditions from each candidate node pair based on the respective similarities.
[0373] In this embodiment, for each candidate node pair, the similarity between the node representations of the nodes in the corresponding candidate node pair is determined, and the similarity corresponding to each candidate node pair is obtained, so that the similarity between the node representations of the nodes can be used as a screening condition to accurately screen out the target node pairs that meet the matching conditions from each candidate node pair based on the respective similarities.
[0374] In one embodiment, the device further includes a preprocessing module; the preprocessing module is configured to collect a plurality of nodes and edges from the event graph to be detected and the query graph to construct a plurality of triples, where the triples include the nodes represented by the initialization nodes, the nodes represented by the nodes to be solved, and the edges represented by the initialization features; construct a target loss function according to the initialization node representations, the node representations to be solved, and the initialization feature representations corresponding to the plurality of triples; perform iterative solution based on the target loss function until the node representations corresponding to each node are obtained when the iteration stops.
[0375] In this embodiment, a plurality of nodes and edges are collected from the event graph to be detected and the query graph to construct a plurality of triples. The triples include the nodes represented by the initialization nodes, the nodes represented by the nodes to be solved, and the edges represented by the initialization features. By constructing the triples, the association relationship between a node and its directly adjacent nodes can be established, so as to predict the node representation of another node through a certain node and the association relationship in the triples. According to the initialization node representations, the node representations to be solved, and the initialization feature representations corresponding to the plurality of triples, a target loss function is constructed, so that the node representation to be solved can be used as the parameter to be solved, and the parameter to be solved can be iteratively solved step by step based on the target loss function, and the node representations and feature representations of each node are updated based on the results of each iterative solution until the final node representations of each node are accurately obtained when the iteration stops.
[0376] Moreover, in this embodiment, the node representation of another node is predicted through the node representation of a certain node in the triple and the feature representation corresponding to the association relationship. The prediction of the node representation of the node is performed by an unsupervised graph pre-training technique, and more abundant topological structure information of the graph structure is combined in the prediction process, making the iteration more accurate. At the same time, it can effectively predict the node representations of all nodes of the graph structure without labels, and is applicable to the prediction of the node representations of all graph structures.
[0377] In one embodiment, the preprocessing module is further configured to construct a first loss function according to the initialization node representations, the node representations to be solved, and the initialization feature representations corresponding to the plurality of triples; determine the initialization context representations corresponding to each process node in the event graph to be detected and the query graph; construct a second loss function based on the node representations to be solved and the corresponding initialization context representations corresponding to each process node; construct a target loss function according to the first loss function and the second loss function.
[0378] In this embodiment, a triple is composed of two nodes and the association relationship between these two nodes. By constructing a first loss function with the triple, a node-level self-supervised task can be achieved, that is, the prediction of the node representation of a single node can be realized. The process node in the graph structure serves as the event subject, and to a great extent, it can be used to infer the properties of the neighbor nodes related to it. Then, a second loss function is constructed through the process node and the context representation of the process node, using richer topological structure information, and a region-level self-supervised task can be achieved, that is, the neighborhood matching prediction can be realized. Through the node-level self-supervised task and the neighborhood matching prediction task for pre-training, the representation of the node can be iteratively solved from the node level and the neighborhood level of the node, so that the result of the iterative solution not only takes into account the characteristics of the node itself, but also takes into account the characteristics of the neighborhood of the node, making the determined node representation more accurate. Moreover, both the node-level self-supervised task and the region-level self-supervised task belong to unsupervised learning tasks, and the prediction of the node representation can be carried out without using labels, with a wider application range and higher processing efficiency.
[0379] Each module in the above abnormal event detection device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor in the computer device in the form of hardware or be independent of the processor, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0380] In one embodiment, a computer device is provided. The computer device can be a terminal or a server. Taking the computer device as a server as an example, its internal structure diagram can be as Figure 12 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store relevant detection data of abnormal events. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an abnormal event detection method is implemented.
[0381] Those skilled in the art can understand, Figure 12The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0382] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0383] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0384] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0385] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0386] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0387] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0388] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. An abnormal event detection method, characterized in that, the method includes: Obtain a to-be-detected event graph and a query graph; both the to-be-detected event graph and the query graph include a plurality of nodes and edges, the edges of the to-be-detected event graph represent events occurring between corresponding nodes, and the edges of the query graph represent abnormal events occurring between corresponding nodes; the plurality of nodes of the query graph include a plurality of first process nodes and first neighbor nodes of each of the first process nodes, and the plurality of nodes of the to-be-detected event graph include a plurality of second process nodes and second neighbor nodes of each of the second process nodes, and the second process nodes represent processes in a computer system; the first process nodes represent processes where abnormal events occur; Perform a matching process on the nodes of the query graph and the nodes of the to-be-detected event graph to obtain a plurality of candidate node pairs; Extract features for each node in the plurality of candidate node pairs to obtain a node representation corresponding to each node; For each of the candidate node pairs, according to the node representations of the nodes in the corresponding candidate node pair, screen out target node pairs that meet the matching conditions from the candidate node pairs; the target node pairs include target process node pairs and target neighbor node pairs; the target neighbor node pairs include a first process node and a second process node that match; the target neighbor node pairs include a first neighbor node and a second neighbor node that match; the first neighbor node represents a resource node associated with the process where the abnormal event occurs in the computer system, and the resource node includes a file, a network address, or a monitor; For each of the target process node pairs, determine the target neighbor node pairs corresponding to the second process node at each node distance according to the node distance between the second process node in the target process node pair and the corresponding second neighbor nodes; For each of the target neighbor node pairs at each node distance, select the second neighbor nodes in the target neighbor node pairs that meet the association conditions from the target neighbor node pairs corresponding to the respective node distances; Based on the association relationship between the second neighbor nodes selected at each node distance and the corresponding second process nodes in the to-be-detected event graph, generate a target subgraph corresponding to each second process node; the target subgraph is used to describe the abnormal events existing in the to-be-detected event graph.
2. The method according to claim 1, characterized in that, the candidate node pairs include candidate process node pairs obtained by performing a matching process on the first process nodes and the second process nodes.
3. The method according to claim 2, characterized in that, the subgraph matching of the to-be-detected event graph based on each of the target node pairs to obtain a target subgraph in the to-be-detected event graph that matches the query graph includes: For each of the target process node pairs, perform a matching process on the first neighbor nodes of the first process nodes in the corresponding target process node pair and the second neighbor nodes of the second process nodes to obtain target neighbor node pairs formed by each first neighbor node and the corresponding second neighbor node; Based on each of the target process node pairs and each of the target neighbor node pairs, perform subgraph matching on the event graph to be detected, and obtain a target subgraph in the event graph to be detected that matches the query graph.
4. The method according to claim 2, wherein, for each of the candidate node pairs, screening out target node pairs that meet the matching conditions from each of the candidate node pairs according to the node representations of the nodes in the corresponding candidate node pairs, includes: For each of the candidate process node pairs, determine the similarity between the node representation of the first process node and the node representation of the second process node in the corresponding candidate process node pair, and obtain the similarity corresponding to each of the candidate process node pairs. For each of the first process nodes, based on the similarities corresponding to the respective candidate process node pairs to which the corresponding first process node belongs, screen out target process node pairs that meet the matching conditions from the respective candidate process node pairs to which the corresponding first process node belongs.
5. The method according to claim 3, wherein, for each of the target process node pairs, performing matching processing on the first neighbor nodes of the first process node and the second neighbor nodes of the second process node in the corresponding target process node pair, and obtaining a target neighbor node pair formed by each of the first neighbor nodes and the second neighbor nodes that match, includes: For each of the target process node pairs, determine the similarity between each of the first neighbor nodes of the first process node and each of the second neighbor nodes of the second process node in the corresponding target process node pair. For each of the first neighbor nodes, determine the second neighbor node that matches the corresponding first neighbor node according to the similarity between the corresponding first neighbor node and each of the second neighbor nodes, so as to obtain a target neighbor node pair formed by each of the first neighbor nodes and the second neighbor nodes that match.
6. The method according to claim 5, wherein, for each of the target process node pairs, determining the similarity between each of the first neighbor nodes of the first process node and each of the second neighbor nodes of the second process node in the corresponding target process node pair, includes: For each of the target process node pairs, determine the first neighbor nodes corresponding to the first process node in each node distance in the corresponding target process node pair, and the second neighbor nodes corresponding to the second process node in each node distance. Determine the similarity between each of the first neighbor nodes at each node distance and each of the second neighbor nodes at the corresponding node distance. For each of the first neighbor nodes, determine the second neighbor node that matches the corresponding first neighbor node according to the similarity between the corresponding first neighbor node and each of the second neighbor nodes, so as to obtain a target neighbor node pair formed by each of the first neighbor nodes and the second neighbor nodes that match. For each node distance, determine the second neighbor node that matches each of the first neighbor nodes according to the similarity between each of the first neighbor nodes at the corresponding node distance and each second neighbor node, so as to obtain, respectively for each node distance, a target neighbor node pair formed by each of the first neighbor nodes and the matching second neighbor node.
7. The method according to claim 1, wherein, for each of the target process node pairs, determining the target neighbor node pair corresponding to the second process node at each node distance according to the node distance between the second process node in the target process node pair and each matching second neighbor node includes: selecting the target process node pair with the largest similarity from each of the target process node pairs; determining the target neighbor node pair corresponding to the second process node in the selected target process node pair at each node distance according to the node distance between the second process node and the matching second neighbor node in the selected target process node pair; generating a target subgraph corresponding to each of the second process nodes based on the association relationship between the selected second neighbor nodes and the matching second process nodes in the event graph to be detected at each node distance includes: generating a subgraph corresponding to the second process node in the selected target process node pair based on the association relationship between the selected second process node and the selected second neighbor nodes in the event graph to be detected at each node distance; selecting the target process node pair with the largest similarity from each of the unselected target process node pairs, and returning to the step of determining the target neighbor node pair corresponding to the second process node in the selected target process node pair at each node distance and continuing to execute until the subgraph corresponding to the second process node in each target process node pair is generated and then stopping, so as to obtain the target subgraph in the event graph to be detected that matches the query graph.
8. The method according to claim 1, wherein, for each of the candidate node pairs, screening out the target node pairs that meet the matching conditions from each of the candidate node pairs according to the node representations of the nodes in the corresponding candidate node pairs includes: for each of the candidate node pairs, determining the similarity between the node representations of the nodes in the corresponding candidate node pair to obtain the similarity corresponding to each of the candidate node pairs; screening out the target node pairs that meet the matching conditions from each of the candidate node pairs based on each similarity.
9. The method according to any one of claims 1 to 8, wherein, the node representation is obtained through a preprocessing step, and the preprocessing step includes: collecting a plurality of nodes and edges from the event graph to be detected and the query graph to construct a plurality of triples, where the triple includes the node for initializing the node representation, the node for solving the node representation, and the edge for initializing the feature representation. Construct a target loss function based on the initialization node representations, nodes to be solved representations, and initialization feature representations corresponding to the multiple triples. Perform iterative solution based on the target loss function until node representations corresponding to each of the nodes are obtained when the iteration stops.
10. The method according to claim 9, wherein, the constructing a target loss function based on the initialization node representations, nodes to be solved representations, and initialization feature representations corresponding to the multiple triples includes: Construct a first loss function based on the initialization node representations, nodes to be solved representations, and initialization feature representations corresponding to the multiple triples; Determine the initialization context representations corresponding to each process node in the event graph to be detected and the query graph; Construct a second loss function based on the nodes to be solved representations and the corresponding initialization context representations corresponding to each process node; Construct a target loss function according to the first loss function and the second loss function.
11. An abnormal event detection device, wherein, the device includes: An acquisition module, configured to acquire an event graph to be detected and a query graph; both the event graph to be detected and the query graph include a plurality of nodes and edges, the edges of the event graph to be detected represent events occurring between corresponding nodes, and the edges of the query graph represent abnormal events occurring between corresponding nodes; the plurality of nodes of the query graph include a plurality of first process nodes and first neighbor nodes of each of the first process nodes, the plurality of nodes of the event graph to be detected include a plurality of second process nodes and second neighbor nodes of each of the second process nodes, and the second process nodes represent processes in a computer system; the first process nodes represent processes where abnormal events occur; A node matching module, configured to perform matching processing on the nodes of the query graph and the nodes of the event graph to be detected to obtain a plurality of candidate node pairs; An extraction module, configured to perform feature extraction on each of the nodes in the plurality of candidate node pairs to obtain a node representation corresponding to each of the nodes; A screening module, configured to, for each of the candidate node pairs, screen out target node pairs that meet the matching conditions from the candidate node pairs according to the node representations of the nodes in the corresponding candidate node pairs; the target node pairs include target process node pairs and target neighbor node pairs; the target neighbor node pairs include matching first process nodes and second process nodes; the target neighbor node pairs include matching first neighbor nodes and second neighbor nodes; the first neighbor nodes represent resource nodes associated with the processes where the abnormal events occur in the computer system, and the resource nodes include files, network addresses, or monitors; The sub - graph matching module is used to, for each of the target process node pairs, determine the target neighbor node pairs corresponding to the second process node under each node distance according to the node distances between the second process node in the target process node pair and each of the matched second neighbor nodes; for each of the target neighbor node pairs under each node distance, select the second neighbor nodes in the target neighbor node pairs that meet the association condition from the target neighbor node pairs corresponding to the corresponding node distance; based on the association relationship between the selected second neighbor nodes under each node distance and the matched second process nodes in the event graph to be detected, generate the target sub - graphs corresponding to each of the second process nodes; the target sub - graphs are used to describe the abnormal events existing in the event graph to be detected.
12. The abnormal event detection device according to claim 11, wherein, the candidate node pairs include the candidate process node pairs obtained by matching the first process node and the second process node.
13. The abnormal event detection device according to claim 12, wherein, the screening module is further used to, for each of the target process node pairs, match the first neighbor nodes of the first process node in the corresponding target process node pair with the second neighbor nodes of the second process node to obtain the target neighbor node pairs formed by each of the first neighbor nodes and the matched second neighbor nodes; the sub - graph matching module is further used to, based on each of the target process node pairs and each of the target neighbor node pairs, perform sub - graph matching on the event graph to be detected to obtain the target sub - graph in the event graph to be detected that matches the query graph.
14. The abnormal event detection device according to claim 12, wherein, the screening module is further used to, for each of the candidate process node pairs, determine the similarity between the node representation of the first process node and the node representation of the second process node in the corresponding candidate process node pair to obtain the similarity corresponding to each of the candidate process node pairs; for each of the first process nodes, based on the similarities corresponding to the respective candidate process node pairs to which the corresponding first process node belongs, screen out the target process node pairs that meet the matching conditions from the respective candidate process node pairs to which the corresponding first process node belongs.
15. The abnormal event detection device according to claim 13, wherein, the screening module is further used to, for each of the target process node pairs, determine the similarity between each of the first neighbor nodes of the first process node in the corresponding target process node pair and each of the second neighbor nodes of the second process node; for each of the first neighbor nodes, determine the second neighbor node that matches the corresponding first neighbor node according to the similarities between the corresponding first neighbor node and each of the second neighbor nodes, so as to obtain the target neighbor node pairs formed by each of the first neighbor nodes and the matched second neighbor nodes respectively.
16. The abnormal event detection device according to claim 15, wherein, The screening module is further configured to, for each of the target process node pairs, determine the first neighbor nodes respectively corresponding to the first process node in each node distance in the corresponding target process node pair, and the second neighbor nodes respectively corresponding to the second process node in each node distance; determine the similarity between each first neighbor node in each node distance and each second neighbor node in the corresponding node distance. For each node distance, according to the similarity between each of the first neighbor nodes and each second neighbor node in the corresponding node distance, determine the second neighbor node that matches each first neighbor node, so as to obtain, for each node distance, the target neighbor node pairs formed by each first neighbor node and the matching second neighbor node.
17. The abnormal event detection device according to claim 11, wherein, the screening module is further configured to, for each of the target process node pairs, determine the target neighbor node pairs respectively corresponding to the second process node in each node distance according to the node distance between the second process node in the target process node pair and the matching second neighbor nodes, including: selecting the target process node pair with the largest similarity from each of the target process node pairs; determining the target neighbor node pairs respectively corresponding to the second process node in each node distance in the selected target process node pair according to the node distance between the second process node and the matching second neighbor node in the selected target process node pair; the sub-graph matching module is further configured to generate a sub-graph corresponding to the second process node in the selected target process node pair based on the association relationship between the selected second process node and the selected second neighbor nodes in each node distance in the event graph to be detected; selecting the target process node pair with the largest similarity from each of the unselected target process node pairs, and returning the step of determining the target neighbor node pairs respectively corresponding to the second process node in each node distance in the selected target process node pair according to the node distance between the second process node and the matching second neighbor node in the selected target process node pair and continuing to execute, until the sub-graph corresponding to the second process node in each target process node pair is generated and then stopping, so as to obtain the target sub-graph matching the query graph in the event graph to be detected.
18. The abnormal event detection device according to claim 11, wherein, the screening module is further configured to, for each of the candidate node pairs, determine the similarity between the node representations of the nodes in the corresponding candidate node pair, so as to obtain the similarity respectively corresponding to each candidate node pair; Based on each similarity, screen out the target node pairs that meet the matching conditions from each of the candidate node pairs.
19. The abnormal event detection device according to any one of claims 11 to 18, wherein, The device further includes a preprocessing module; the preprocessing module is configured to collect a plurality of nodes and edges from the event graph to be detected and the query graph to construct a plurality of triples, where each triple includes a node representing an initialized node, a node representing a node to be solved, and an edge representing an initialized feature; construct an objective loss function according to the initialized node representations, the nodes to be solved representations, and the initialized feature representations corresponding to the plurality of triples; perform iterative solution based on the objective loss function until node representations corresponding to each of the nodes are obtained when the iteration stops.
20. The abnormal event detection device according to claim 19, wherein, the preprocessing module is further configured to construct a first loss function according to the initialized node representations, the nodes to be solved representations, and the initialized feature representations corresponding to the plurality of triples; determine the initialized context representations corresponding to each process node in the event graph to be detected and the query graph; construct a second loss function based on the nodes to be solved representations and the corresponding initialized context representations corresponding to each process node; construct an objective loss function according to the first loss function and the second loss function.
21. A computer device, including a memory and a processor, where the memory stores a computer program, wherein, when the processor executes the computer program, the method according to any one of claims 1 to 10 is implemented.
22. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
23. A computer program product, including a computer program, wherein, when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Method and device for obtaining isomorphic subgraph, computer equipment and readable storage medium
CN113779085A
Rapid anomaly detection method for multi-attribute network
CN114401136A